Compare commits
55
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9803d3b768 | ||
|
|
d4635d28f0 | ||
|
|
dea96f0195 | ||
|
|
557a0d791e | ||
|
|
120c05cd60 | ||
|
|
cb952c9bf3 | ||
|
|
7e5e58aa1b | ||
|
|
844c199fb8 | ||
|
|
bb28ba3006 | ||
|
|
0857ea1c23 | ||
|
|
2eb086f400 | ||
|
|
5f9793f9c0 | ||
|
|
0bc5a541ac | ||
|
|
8739d29550 | ||
|
|
dcfe9d16ff | ||
|
|
5e74b991ac | ||
|
|
05a2a42983 | ||
|
|
2d166f567b | ||
|
|
27be01af93 | ||
|
|
9949f914df | ||
|
|
537a2170fb | ||
|
|
255c8c493e | ||
|
|
b7c7dbe2a1 | ||
|
|
c583762352 | ||
|
|
1641dab505 | ||
|
|
e5f1b9d825 | ||
|
|
ff85258d03 | ||
|
|
e4636e6a44 | ||
|
|
e461f01c97 | ||
|
|
0e7921a02a | ||
|
|
2310f6cbd8 | ||
|
|
1037e68091 | ||
|
|
3184818c08 | ||
|
|
51cb552715 | ||
|
|
cab493206c | ||
|
|
82fec45bc6 | ||
|
|
510553d800 | ||
|
|
62ed009281 | ||
|
|
230345a867 | ||
|
|
e5c7cc1ca2 | ||
|
|
ec73440c19 | ||
|
|
562284faf4 | ||
|
|
62270b03ca | ||
|
|
54fd112fde | ||
|
|
45f3336201 | ||
|
|
ba7f8e1bad | ||
|
|
f853dafaf1 | ||
|
|
04573e1d86 | ||
|
|
f2f9c0326c | ||
|
|
29134789a9 | ||
|
|
66ab9d7d93 | ||
|
|
f5f52c61e8 | ||
|
|
5796d1c22f | ||
|
|
89f05850f2 | ||
|
|
e77268c951 |
No files matched your search
@@ -5,6 +5,29 @@ auto-loads this file as instructions, overriding the built-in defaults for
|
||||
this project. Custom agents in `.opencode/agents/` inherit these rules
|
||||
unless their own prompts say otherwise.
|
||||
|
||||
## Session Continuity (keep the agent loop alive)
|
||||
|
||||
opencode ends the turn whenever an assistant message contains no tool
|
||||
call — including messages that are pure analysis. Long reasoning bursts
|
||||
are welcome in this repo (they pre-catch errors and self-correct), but a
|
||||
burst that ends as analysis-only text silently stops the session
|
||||
mid-task. Past sessions documented this repeatedly ("the session
|
||||
stalled"; the working fix discovered there: "call tools frequently, keep
|
||||
thinking bursts short"). Keep the depth; change where the burst ends:
|
||||
|
||||
1. **Never end a turn with analysis-only text.** Every visible message
|
||||
must either issue a tool call or be a final report for a genuinely
|
||||
completed phase/task. When a thinking burst converges on a decision,
|
||||
act on it (read, edit, bash) in the same turn.
|
||||
2. **Land work incrementally.** Once a design decision is settled, write
|
||||
the code before analyzing the next one. Do not emit full-design
|
||||
essays in a single chat message; reasoning belongs in thinking tokens
|
||||
or committed docs, not in the transcript.
|
||||
3. **On a silent turn end, resume without re-deriving.** If the turn
|
||||
ended after an analysis-only message and the task is incomplete,
|
||||
pick up from the last settled decision — do not redo the analysis
|
||||
and do not ask the user whether to continue.
|
||||
|
||||
## Git Workflow
|
||||
|
||||
**Commit and push when reasonable.** When a change is complete and
|
||||
|
||||
+227
@@ -0,0 +1,227 @@
|
||||
# Changelog
|
||||
|
||||
All notable changes to this crate are documented here. The format is
|
||||
based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
|
||||
this crate adheres to [Semantic Versioning](https://semver.org/).
|
||||
|
||||
## [0.3.0] - 2026-09-07
|
||||
|
||||
The compiled-forms release. The packed read path — the hot path for
|
||||
stream parsing — is driven by a compile-once `ReadPlan` instead of a
|
||||
per-call walk of the BAST typed tree; byte validation runs a compiled
|
||||
`ValidationPlan`; plans and offset maps fingerprint to a stable hash
|
||||
(ADR-011/ADR-012). Reads of SFTP-shaped packet streams went from
|
||||
~189× hand-rolled Rust to ~74× (~2.4× faster), fixed-stride chunk
|
||||
reads from ~18× to ~11×, and `SequentialReader::read_next_borrowed`
|
||||
makes the per-field hot loop allocation-free.
|
||||
|
||||
### Breaking changes
|
||||
|
||||
- **`Bast*` types are owned.** `BastDoc`/`BastStruct`/`BastField`/… no
|
||||
longer borrow from the source `serde_json::Value`; all v0.2.0
|
||||
lifetimes are gone. `BastDoc::new` parses the root eagerly; `$ref`s
|
||||
resolve lazily.
|
||||
- **`OffsetMap::get` / `PackedLayout::get`** return `&OffsetEntry`
|
||||
(was `Option<OffsetEntry>` by value), backed by an O(log n)
|
||||
`BTreeMap` path→index (first-occurrence-wins for duplicate names).
|
||||
- **`SequentialReader::new`** takes the compiled plan; construct via
|
||||
`AlkTypeEngine::sequential_reader()` (packed mode only).
|
||||
- **`materialize_packed` / `materialize_aligned`** take the compiled
|
||||
plan / `(&BastDoc, &OffsetMap)` pair respectively.
|
||||
- **Field-name-discriminator union wire convention** (ADR-011
|
||||
addendum): the builder lays out the union's declared `fields`
|
||||
(shared) first, then the variant's own fields. Variants must not
|
||||
re-declare the discriminator or any shared field, and the
|
||||
discriminator field must be the first entry in `fields` — all
|
||||
enforced at parse with clean `Schema` errors. Schemas relying on
|
||||
0.2.0's variant-only layout are rejected (they produced
|
||||
reader↔builder-disagreeing bytes).
|
||||
- **`maxLength` is string/bytes-only** — rejected at parse on every
|
||||
other kind (it was silently unenforced there).
|
||||
- **Aligned-mode `Record` fields reject `offset-indirect`** (the
|
||||
materializer always walks the inline count-prefixed form — the
|
||||
annotated shape was never readable).
|
||||
- **Schema input bounds** (untrusted-schema hardening, AGENTS.md §3):
|
||||
array `count` ≤ 2^16 and `count × stride` ≤ 2^26 bytes; `align` ≤
|
||||
4096; `maxLength` ≤ 2^26; cyclic `$ref` graphs and >128-deep nesting
|
||||
are rejected by every public walker (`OffsetMap::compute`,
|
||||
`LayoutBuilder::new`, `materialize_aligned` included), not just the
|
||||
engine.
|
||||
|
||||
### Additions
|
||||
|
||||
- **`ReadPlan`** (ADR-011) — the compiled packed-read plan, re-exported
|
||||
with `CompositePlan`/`FieldPlan`/`ReadKind`/`DiscriminatorPlan`.
|
||||
`ReadPlan::compile` is untrusted-input-safe standalone (depth cap +
|
||||
cycle set). `fixed_size()` exposes the compile-time-known byte size
|
||||
for fixed structs.
|
||||
- **`ValidationPlan`** (ADR-012 §3) — the compiled `validate_bytes`
|
||||
walker, with `ValidNode`/`ValidVariant` sub-types.
|
||||
- **`fingerprint()`** on `ReadPlan`/`OffsetMap`/`ValidationPlan` +
|
||||
`Hash`/`Eq` derives on the plan types (ADR-012 §1/§4) — plan
|
||||
identity for cache-keying across processes.
|
||||
- **`OffsetMap` `LeafMeta`** — each entry records whether it is
|
||||
fixed/length-prefixed/offset-indirect so `read_field`/`write_field`
|
||||
dispatch without re-walking the schema; `OffsetEntry` type re-exported.
|
||||
- **`SequentialReader::read_next_borrowed`** — zero-allocation variant
|
||||
of `read_next` (field name borrowed from the plan).
|
||||
- **`AlkTypeEngine::validate_bytes`** now runs the compiled
|
||||
`ValidationPlan` (was an interpretive BAST walk in 0.2.0).
|
||||
|
||||
### Fixes (post-release-commit hardening — reviews #006, #007, #008)
|
||||
|
||||
All found and fixed before the first crates.io publish of 0.3.0, so
|
||||
no published version ever exhibited them.
|
||||
|
||||
- **Untrusted-input crashes removed.** A huge declared array count
|
||||
OOM-aborted the process (`Vec::with_capacity(count)` before reading
|
||||
a byte) — now compile-capped and walked with push-only growth.
|
||||
Cyclic `$ref` graphs stack-overflowed the three standalone layout
|
||||
walkers — now guarded by a shared reference-graph check. Deeply
|
||||
nested stride-0 arrays briefly allowed ~477 MB of simultaneous
|
||||
allocation from a ~1 KB schema — restored to incremental growth.
|
||||
- **Cross-consumer divergences closed.** Builder, reader,
|
||||
materializer, tunion, and the validation plan now agree on
|
||||
field-disc union layout (shared-then-variant), on the discriminator
|
||||
field's position (must be first), and on union mapping-key matching
|
||||
(numeric fast-path dispatch only for canonical keys like `"2"`;
|
||||
`"01"`/`"+1"` fall back to the string comparison all consumers
|
||||
share). The legacy BAST walker's field-disc union arm walks shared
|
||||
fields before the variant (it previously materialized variant fields
|
||||
from shared fields' bytes).
|
||||
- **Silently-corrupt layouts rejected.** Aligned record fields with
|
||||
`maxLength`/`offset-indirect`; non-final inline length-prefixed
|
||||
fields (records included — the ADR-006 check now sees them);
|
||||
aligned-mode `maxLength`/`offset-indirect` on records; unions in
|
||||
aligned mode (pre-existing, now tested).
|
||||
- **Coverage**: 90.67% lines / 86.32% functions at review #007's
|
||||
audit, 91.66% after its fixes; every uncovered region outside test
|
||||
modules read and classified in-tree (docs/reviews/007).
|
||||
|
||||
### Non-breaking improvements
|
||||
|
||||
- Engine compile is one-shot and allocation-tidy; plans are
|
||||
`Send + Sync` (statically asserted) and fingerprintable.
|
||||
- Zero-progress array-element guard on all three array walkers (a
|
||||
zero-size element makes the declared count unbounded on the wire).
|
||||
- WASM-clean unchanged: two dependencies (`jsonschema`
|
||||
default-features off, `serde_json` with `preserve_order`), no
|
||||
`async`, no `unsafe`, no feature flags.
|
||||
- Benches (`benches/wire_vs_bast.rs`): read/write chunk streams, an
|
||||
SFTP-shaped union packet stream, and `validate_bytes` per buffer —
|
||||
the numbers quoted above and in ADR-007/ADR-011.
|
||||
|
||||
## [0.2.0] - 2026-08-17
|
||||
|
||||
A breaking release that replaces the v0.1.0 `AlkType:*` custom-keyword
|
||||
JSON Schema format with BAST (Binary Abstract Syntax Tree) — a JSON
|
||||
document that describes binary layouts using a `kind`-based vocabulary
|
||||
with `$defs`/`$ref` for composition. BAST is itself a valid JSON Schema
|
||||
instance (it has a meta-schema), making it self-validating,
|
||||
editor-friendly, and trivially consumable from any language with a JSON
|
||||
parser. The engine works the same way as before: compile a document
|
||||
once into an `AlkTypeEngine`, then read/write fields at computed offsets
|
||||
and validate bytes/JSON. The pivot was made now because v0.1.0 has no
|
||||
real consumers (≈15 crates.io downloads, mostly bots/scanners), so the
|
||||
custom-keyword wart could be removed cleanly.
|
||||
|
||||
### Breaking changes
|
||||
|
||||
- **Schema format.** The v0.1.0 `AlkType:*` custom-keyword JSON Schema
|
||||
format (`{ "AlkType:Struct": true, "fields": [...] }`) is removed.
|
||||
Schemas are now BAST documents:
|
||||
`{ "$defs": { "<TypeName>": { "kind": "struct", "fields": [...] } } }`.
|
||||
The `kind`-based vocabulary covers 18 binary kinds (integers, floats,
|
||||
bytes, string, struct, union, enum, array, etc.).
|
||||
- **`AlkTypeEngine::compile` signature.** Now takes
|
||||
`(bast_doc: &Value, root_name: &str, mode: LayoutMode, json_schema: Option<&Value>)`.
|
||||
The root type name is a required parameter — it selects which `$defs`
|
||||
entry is the top-level type (previously the root was implicit from the
|
||||
single top-level schema object).
|
||||
- **Builder API output.** `Definitions`/`Schema`/`Discriminator` now
|
||||
produce BAST JSON via `Definitions::build_doc(name, schema)`. The
|
||||
builder method names are unchanged; only the emitted JSON shape
|
||||
changed. `Schema::struct_()` produces a BAST struct;
|
||||
`Schema::object()` produces a standard JSON Schema (for the
|
||||
`validate_json` path).
|
||||
- **Validation split.** v0.1.0 used a single `jsonschema` validator
|
||||
with 19 custom `AlkType:*` keywords for both bytes and JSON
|
||||
validation. 0.2.0 splits this into two independent paths:
|
||||
- `validate_bytes` uses a new BAST-native validator
|
||||
(`bast_validation`) — a recursive walker over the BAST type tree.
|
||||
- `validate_json` / `is_valid_json` use a standard
|
||||
`jsonschema::Validator` built from a consumer-provided JSON Schema
|
||||
(passed to `compile` as the `json_schema` parameter). No custom
|
||||
keywords; BAST is not involved — BAST describes bytes, not JSON
|
||||
shape.
|
||||
Both paths return `AlkTypeError::Validation` with a uniform
|
||||
`jsonschema::ValidationError<'static>` payload.
|
||||
- **Removed.** The v0.1.0 custom-keyword accessor layer
|
||||
(`AlkTypeKind::FromStr`, `parse_*`, `resolve_ref*`,
|
||||
`DiscriminatorKind`) is removed. The BAST parser (`bast` module)
|
||||
exposes a cleaner typed surface (`BastDoc`/`BastDef`/`BastStruct`/
|
||||
`BastField`/`BastType`/etc.) that borrows from the source
|
||||
`serde_json::Value` without cloning field data.
|
||||
- **Public module surface.** New public modules: `bast`, `bast_meta`,
|
||||
`bast_validation`, `builder`, `materialize`. The `schema` module is
|
||||
retained but now holds only `Endian`/`AlkTypeKind`/`VariableEncoding`
|
||||
(the binary-kind vocabulary); the v0.1.0 custom-keyword machinery is
|
||||
gone.
|
||||
|
||||
### Additions
|
||||
|
||||
- **BAST meta-schema.** Embedded in the crate as
|
||||
`BAST_META_SCHEMA` (re-exported from the crate root) and published at
|
||||
`https://alk.dev/bast/v1/schema`. BAST documents are validated against
|
||||
it at compile time (`AlkTypeEngine::compile` calls
|
||||
`validate_bast_doc` before parsing).
|
||||
- **`materialize` module.** Materializes a `serde_json::Value` tree from
|
||||
a binary buffer by walking the BAST typed tree. Used by
|
||||
`AlkTypeEngine::validate_bytes` (ADR-010).
|
||||
- **Builder for JSON Schemas.** `Schema::object()` produces a standard
|
||||
JSON Schema object (for the `validate_json` path), complementing
|
||||
`Schema::struct_()` which produces a BAST struct (for the bytes path).
|
||||
One builder, two output shapes — the method name selects which.
|
||||
|
||||
### Bug fixes vs v0.1.0
|
||||
|
||||
- **Enum index bounds are now checked.** The v0.1.0 validator had a
|
||||
dead constraint: enum variant indices were never bounds-checked
|
||||
against `values.len()`. The BAST-native validator enforces it
|
||||
(`validate_enum` checks `idx < values.len()`).
|
||||
- **Offset-indirect, field-level endian, and aligned materialization**
|
||||
bugs found during review #003 are fixed.
|
||||
|
||||
### Non-breaking improvements
|
||||
|
||||
- `$ref` is restricted to `#/$defs/<name>` — one hash lookup, no
|
||||
`normalize_refs` pass (the v0.1.0 engine needed one).
|
||||
- Schemas remain untrusted input: every engine path that walks a BAST
|
||||
document returns `Err` on a malformed document, never `panic!`/
|
||||
`unreachable!`. Overflow-safe arithmetic (`checked_add`,
|
||||
`usize::try_from`) on all offset/count casts.
|
||||
- Still two dependencies (`jsonschema` with `default-features = false`,
|
||||
`serde_json` with `preserve_order`), no `async`, no `unsafe`, no
|
||||
platform deps, no feature flags. Compiles to
|
||||
`wasm32-unknown-unknown`.
|
||||
|
||||
### Upgrade notes
|
||||
|
||||
There is no migration path from v0.1.0 `AlkType:*` schemas — the format
|
||||
is incompatible. Rewrite schemas as BAST documents (the `builder` API
|
||||
produces them; see the README usage example) and update `compile` calls
|
||||
to pass the root type name and the optional JSON Schema. The
|
||||
read/write/validate API surface (`read_field`, `write_field`,
|
||||
`sequential_reader`, `validate_bytes`, `validate_json`,
|
||||
`is_valid_json`) is unchanged.
|
||||
|
||||
## [0.1.0] - 2025-11-10
|
||||
|
||||
Initial crates.io release. Custom-keyword JSON Schema format
|
||||
(`AlkType:*`), single `jsonschema` validator for both bytes and JSON,
|
||||
`AlkTypeEngine` with packed/aligned layout modes, builder API producing
|
||||
`serde_json::Value`.
|
||||
|
||||
[0.3.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.3.0
|
||||
[0.2.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.2.0
|
||||
[0.1.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.1.0
|
||||
Generated
+188
-1
@@ -27,8 +27,9 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "alktype"
|
||||
version = "0.1.0"
|
||||
version = "0.3.0"
|
||||
dependencies = [
|
||||
"criterion",
|
||||
"jsonschema",
|
||||
"serde_json",
|
||||
]
|
||||
@@ -39,6 +40,18 @@ version = "0.2.21"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "683d7910e743518b0e34f1186f92494becacb047c7b6bf616c96772180fef923"
|
||||
|
||||
[[package]]
|
||||
name = "anes"
|
||||
version = "0.1.6"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "4b46cbb362ab8752921c97e041f5e366ee6297bd428a31275b9fcf1e380f7299"
|
||||
|
||||
[[package]]
|
||||
name = "anstyle"
|
||||
version = "1.0.14"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "940b3a0ca603d1eade50a4846a2afffd5ef57a9feac2c0e2ec2e14f9ead76000"
|
||||
|
||||
[[package]]
|
||||
name = "autocfg"
|
||||
version = "1.5.1"
|
||||
@@ -84,12 +97,107 @@ version = "0.6.9"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "175812e0be2bccb6abe50bb8d566126198344f707e304f45c648fd8f2cc0365e"
|
||||
|
||||
[[package]]
|
||||
name = "cast"
|
||||
version = "0.3.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "37b2a672a2cb129a2e41c10b1224bb368f9f37a2b16b612598138befd7b37eb5"
|
||||
|
||||
[[package]]
|
||||
name = "cfg-if"
|
||||
version = "1.0.4"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801"
|
||||
|
||||
[[package]]
|
||||
name = "ciborium"
|
||||
version = "0.2.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "42e69ffd6f0917f5c029256a24d0161db17cea3997d185db0d35926308770f0e"
|
||||
dependencies = [
|
||||
"ciborium-io",
|
||||
"ciborium-ll",
|
||||
"serde",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "ciborium-io"
|
||||
version = "0.2.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "05afea1e0a06c9be33d539b876f1ce3692f4afea2cb41f740e7743225ed1c757"
|
||||
|
||||
[[package]]
|
||||
name = "ciborium-ll"
|
||||
version = "0.2.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "57663b653d948a338bfb3eeba9bb2fd5fcfaecb9e199e87e1eda4d9e8b240fd9"
|
||||
dependencies = [
|
||||
"ciborium-io",
|
||||
"half",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "clap"
|
||||
version = "4.6.6"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "473c7e07f409a8d772161724aa8db6a765a2532a70f9667eeb7b49d3d02fbdca"
|
||||
dependencies = [
|
||||
"clap_builder",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "clap_builder"
|
||||
version = "4.6.6"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "7b48fea5a88e9ae728a2dcbedbfc0e730f7d60da42e1cb049a83c9fb8b789889"
|
||||
dependencies = [
|
||||
"anstyle",
|
||||
"clap_lex",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "clap_lex"
|
||||
version = "1.1.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9"
|
||||
|
||||
[[package]]
|
||||
name = "criterion"
|
||||
version = "0.7.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "e1c047a62b0cc3e145fa84415a3191f628e980b194c2755aa12300a4e6cbd928"
|
||||
dependencies = [
|
||||
"anes",
|
||||
"cast",
|
||||
"ciborium",
|
||||
"clap",
|
||||
"criterion-plot",
|
||||
"itertools",
|
||||
"num-traits",
|
||||
"oorandom",
|
||||
"regex",
|
||||
"serde",
|
||||
"serde_json",
|
||||
"tinytemplate",
|
||||
"walkdir",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "criterion-plot"
|
||||
version = "0.6.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "9b1bcc0dc7dfae599d84ad0b1a55f80cde8af3725da8313b528da95ef783e338"
|
||||
dependencies = [
|
||||
"cast",
|
||||
"itertools",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "crunchy"
|
||||
version = "0.2.4"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "460fbee9c2c2f33933d720630a6a0bac33ba7053db5344fac858d4b8952d77d5"
|
||||
|
||||
[[package]]
|
||||
name = "data-encoding"
|
||||
version = "2.11.0"
|
||||
@@ -107,6 +215,12 @@ dependencies = [
|
||||
"syn 2.0.119",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "either"
|
||||
version = "1.18.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "252afb9ae5eaa683babdc6a068b3f5726eb19e05070c731f9b2a23a7c3e8ed34"
|
||||
|
||||
[[package]]
|
||||
name = "email_address"
|
||||
version = "0.2.9"
|
||||
@@ -174,6 +288,17 @@ dependencies = [
|
||||
"wasm-bindgen",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "half"
|
||||
version = "2.7.1"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "6ea2d84b969582b4b1864a92dc5d27cd2b77b622a8d79306834f1be5ba20d84b"
|
||||
dependencies = [
|
||||
"cfg-if",
|
||||
"crunchy",
|
||||
"zerocopy",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "hashbrown"
|
||||
version = "0.16.1"
|
||||
@@ -304,6 +429,15 @@ dependencies = [
|
||||
"hashbrown 0.17.1",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "itertools"
|
||||
version = "0.13.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "413ee7dfc52ee1a4949ceeb7dbc8a33f2d6c088194d9f922fb8318faf1f01186"
|
||||
dependencies = [
|
||||
"either",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "itoa"
|
||||
version = "1.0.18"
|
||||
@@ -479,6 +613,12 @@ version = "1.21.4"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50"
|
||||
|
||||
[[package]]
|
||||
name = "oorandom"
|
||||
version = "11.1.5"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
|
||||
|
||||
[[package]]
|
||||
name = "outref"
|
||||
version = "0.5.2"
|
||||
@@ -628,6 +768,15 @@ version = "1.0.23"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "cf54715a573b99ac80df0bc206da022bcd442c974952c7b9720069370852e21f"
|
||||
|
||||
[[package]]
|
||||
name = "same-file"
|
||||
version = "1.0.6"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "93fc1dc3aaa9bfed95e02e6eadabb4baf7e3078b0bd1b4d7b6b0b68378900502"
|
||||
dependencies = [
|
||||
"winapi-util",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "scopeguard"
|
||||
version = "1.2.0"
|
||||
@@ -733,6 +882,16 @@ dependencies = [
|
||||
"zerovec",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "tinytemplate"
|
||||
version = "1.2.1"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "be4d6b5f19ff7664e8c98d03e2139cb510db9b0a60b55f8e8709b689d939b6bc"
|
||||
dependencies = [
|
||||
"serde",
|
||||
"serde_json",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "unicode-general-category"
|
||||
version = "1.1.0"
|
||||
@@ -773,6 +932,16 @@ version = "0.8.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "5c3082ca00d5a5ef149bb8b555a72ae84c9c59f7250f013ac822ac2e49b19c64"
|
||||
|
||||
[[package]]
|
||||
name = "walkdir"
|
||||
version = "2.5.0"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "29790946404f91d9c5d06f9874efddea1dc06c5efe94541a7d6863108e3a5e4b"
|
||||
dependencies = [
|
||||
"same-file",
|
||||
"winapi-util",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "wasip2"
|
||||
version = "1.0.4+wasi-0.2.12"
|
||||
@@ -827,12 +996,30 @@ dependencies = [
|
||||
"unicode-ident",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "winapi-util"
|
||||
version = "0.1.11"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "c2a7b1c03c876122aa43f3020e6c3c3ee5c05081c9a00739faf7503aeba10d22"
|
||||
dependencies = [
|
||||
"windows-sys",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "windows-link"
|
||||
version = "0.2.1"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5"
|
||||
|
||||
[[package]]
|
||||
name = "windows-sys"
|
||||
version = "0.61.2"
|
||||
source = "registry+https://github.com/rust-lang/crates.io-index"
|
||||
checksum = "ae137229bcbd6cdf0f7b80a31df61766145077ddf49416a728b02cb3921ff3fc"
|
||||
dependencies = [
|
||||
"windows-link",
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "wit-bindgen"
|
||||
version = "0.57.1"
|
||||
|
||||
+12
-4
@@ -1,15 +1,15 @@
|
||||
[package]
|
||||
name = "alktype"
|
||||
version = "0.1.0"
|
||||
version = "0.3.0"
|
||||
edition = "2021"
|
||||
rust-version = "1.85"
|
||||
license = "MIT OR Apache-2.0"
|
||||
description = "Binary struct engine: takes a JSON Schema with AlkType:* custom keywords and produces an offset map, read/write functions, and validation"
|
||||
description = "Binary struct engine: takes a BAST (Binary Abstract Syntax Tree) document and produces an offset map, read/write functions, and validation"
|
||||
repository = "https://git.alk.dev/alkdev/alktype"
|
||||
readme = "README.md"
|
||||
keywords = ["binary", "jsonschema", "wire-format", "serialization", "layout"]
|
||||
categories = ["encoding", "data-structures", "parsing"]
|
||||
exclude = [".opencode/", "docs/reviews/", "docs/research/", "docs/sdd_process.md", "Cargo.lock"]
|
||||
exclude = [".opencode/", "docs/reviews/", "docs/research/", "docs/sdd_process.md", "Cargo.lock", "AGENTS.md"]
|
||||
|
||||
[lib]
|
||||
name = "alktype"
|
||||
@@ -19,4 +19,12 @@ default = []
|
||||
|
||||
[dependencies]
|
||||
jsonschema = { version = "0.46", default-features = false }
|
||||
serde_json = { version = "1", features = ["preserve_order"] }
|
||||
serde_json = { version = "1", features = ["preserve_order"] }
|
||||
|
||||
[dev-dependencies]
|
||||
serde_json = "1"
|
||||
criterion = { version = "0.7", default-features = false }
|
||||
|
||||
[[bench]]
|
||||
name = "wire_vs_bast"
|
||||
harness = false
|
||||
@@ -1,51 +1,62 @@
|
||||
# alktype
|
||||
|
||||
The binary struct engine: a small Rust crate that takes a JSON Schema
|
||||
with `AlkType:*` custom keywords and produces an offset map, read/write
|
||||
The binary struct engine: a small Rust crate that takes a BAST (Binary
|
||||
Abstract Syntax Tree) document and produces an offset map, read/write
|
||||
functions, and validation — all driven by the schema. The schema is the
|
||||
format definition; the engine is generic.
|
||||
|
||||
`alktype` is a standalone crate with **two dependencies**: `jsonschema`
|
||||
(for validation) and `serde_json` (for schema parsing). No tokio, no
|
||||
platform deps, no `unsafe`. Compiles to `wasm32-unknown-unknown`.
|
||||
(for JSON validation and BAST meta-schema validation) and `serde_json`
|
||||
(for BAST document parsing). No tokio, no platform deps, no `unsafe`.
|
||||
Compiles to `wasm32-unknown-unknown`.
|
||||
|
||||
## What it is
|
||||
|
||||
A JSON Schema annotated with `AlkType:*` custom keywords serves three
|
||||
roles simultaneously:
|
||||
BAST is a JSON document that describes binary data layouts using a
|
||||
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
|
||||
itself a valid JSON Schema instance (it has a meta-schema), making it
|
||||
self-validating, editor-friendly, and trivially consumable from any
|
||||
language with a JSON parser. See
|
||||
[`docs/architecture/bast-format.md`](docs/architecture/bast-format.md)
|
||||
for the normative format spec.
|
||||
|
||||
A BAST document serves three roles simultaneously:
|
||||
|
||||
| Role | Mechanism | When |
|
||||
|------|-----------|------|
|
||||
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
|
||||
| **Validation spec (bytes)** | Compiled `ValidationPlan` walk over the materialized `Value` (ADR-012) | Access time (`validate_bytes`) |
|
||||
| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) |
|
||||
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map / packed layout) |
|
||||
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
|
||||
| **Wire access (packed)** | Compiled `ReadPlan` (ADR-011) — compile-once, no per-read schema walk | Access time (`SequentialReader`) |
|
||||
|
||||
No separate format definition, no separate parser, no separate
|
||||
validator. The schema is the single source of truth for the binary
|
||||
format. Adding a new field to a protocol is adding a property to the
|
||||
schema JSON — the engine computes the new offsets automatically.
|
||||
validator. The BAST document is the single source of truth for the
|
||||
binary format. Adding a new field to a protocol is adding an entry to
|
||||
the BAST `fields` array — the engine computes the new offsets
|
||||
automatically.
|
||||
|
||||
This is the same principle as `#[repr(C)]` struct field access, but at
|
||||
runtime from a portable JSON Schema instead of at compile time from
|
||||
language-specific annotations. The schema is the ABI contract.
|
||||
runtime from a portable JSON document instead of at compile time from
|
||||
language-specific annotations. The BAST document is the ABI contract.
|
||||
|
||||
## Usage
|
||||
|
||||
Build the schema with the fluent Rust builder (ADR-009), compile it
|
||||
once into an [`AlkTypeEngine`], then read/write fields at computed
|
||||
Build the BAST document with the fluent Rust builder (ADR-009), compile
|
||||
it once into an [`AlkTypeEngine`], then read/write fields at computed
|
||||
offsets:
|
||||
|
||||
```rust
|
||||
use alktype::{AlkTypeEngine, Endian, LayoutMode, Schema, FieldValue};
|
||||
use alktype::{AlkTypeEngine, Definitions, Endian, LayoutMode, Schema, FieldValue};
|
||||
|
||||
// Channels' 8-byte chunk header: big-endian, packed mode.
|
||||
let mut schema = Schema::struct_()
|
||||
let doc = Definitions::new().build_doc("ChunkHeader", Schema::struct_()
|
||||
.endian(Endian::Big)
|
||||
.field("channel_id", Schema::uint32())
|
||||
.field("length", Schema::uint32())
|
||||
.build();
|
||||
.field("length", Schema::uint32()));
|
||||
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
// `json_schema: None` — no JSON-validation path needed for a binary-only schema.
|
||||
let engine = AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)?;
|
||||
|
||||
// Write a frame into a buffer. For fixed-size structs, the byte
|
||||
// positions are a direct read off the layout — channel_id at 0,
|
||||
@@ -55,8 +66,8 @@ let mut buf = vec![0u8; 8];
|
||||
alktype::data_access::write_u32(&mut buf, 0, 42, "channel_id", Endian::Big)?;
|
||||
alktype::data_access::write_u32(&mut buf, 4, 7, "length", Endian::Big)?;
|
||||
|
||||
// Validate the bytes against the schema in one call.
|
||||
engine.validate_bytes(&buf)?; // materializes a Value, then validates
|
||||
// Validate the bytes against the BAST document in one call.
|
||||
engine.validate_bytes(&buf)?; // materializes a Value, then runs the BAST-native validator
|
||||
|
||||
// Read the frame back sequentially (packed mode is sequential by
|
||||
// construction — variable-length fields shift subsequent fields).
|
||||
@@ -67,43 +78,64 @@ assert_eq!(value, FieldValue::U32(42));
|
||||
# Ok::<(), alktype::AlkTypeError>(())
|
||||
```
|
||||
|
||||
Schemas may also be authored as plain `serde_json::json!{...}` literals
|
||||
and passed directly to `AlkTypeEngine::compile` — the builder is a
|
||||
construction convenience, not a requirement.
|
||||
BAST documents may also be authored as plain `serde_json::json!{...}`
|
||||
literals and passed directly to `AlkTypeEngine::compile` — the builder
|
||||
is a construction convenience, not a requirement.
|
||||
|
||||
## The 19 `AlkType:*` kinds
|
||||
```rust
|
||||
use alktype::{AlkTypeEngine, LayoutMode};
|
||||
use serde_json::json;
|
||||
|
||||
| Kind | Rust type | Size | Notes |
|
||||
|------|-----------|-----:|-------|
|
||||
| `AlkType:Int8` | `i8` | 1 | |
|
||||
| `AlkType:Int16` | `i16` | 2 | endian-sensitive |
|
||||
| `AlkType:Int32` | `i32` | 4 | endian-sensitive |
|
||||
| `AlkType:Int64` | `i64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
|
||||
| `AlkType:Uint8` | `u8` | 1 | |
|
||||
| `AlkType:Uint16` | `u16` | 2 | endian-sensitive |
|
||||
| `AlkType:Uint32` | `u32` | 4 | endian-sensitive; also the enum/string/bytes length-prefix width |
|
||||
| `AlkType:Uint64` | `u64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
|
||||
| `AlkType:Float32` | `f32` | 4 | endian-sensitive; NaN/inf rejected by validator |
|
||||
| `AlkType:Float64` | `f64` | 8 | endian-sensitive; NaN/inf rejected by validator |
|
||||
| `AlkType:Boolean` | `bool` | 1 | |
|
||||
| `AlkType:Enum` | `u32` index | 4 | index into the schema's `"enum"` array |
|
||||
| `AlkType:String` | length-prefixed UTF-8 | 4 + N | `[length: u32][bytes]` by default |
|
||||
| `AlkType:Bytes` | length-prefixed raw bytes | 4 + N | `[length: u32][bytes]` by default |
|
||||
| `AlkType:Timestamp` | length-prefixed RFC 3339 | 4 + N | non-strict string check (see inline docs) |
|
||||
| `AlkType:Struct` | record of fields | composite | nested; field paths are dotted (`"header.version"`) |
|
||||
| `AlkType:Union` | tagged union | composite | byte-offset or field-name discriminator |
|
||||
| `AlkType:Array` | repeated element | composite | fixed-size elements with stride, or variable count |
|
||||
| `AlkType:Record` | string-keyed map | composite | `[count: u32][key, value]...` |
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"ChunkHeader": {
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [
|
||||
{ "name": "channel_id", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)?;
|
||||
# Ok::<(), alktype::AlkTypeError>(())
|
||||
```
|
||||
|
||||
The engine recognizes a kind when the schema object has a key starting
|
||||
with `AlkType:` whose value is `true` (the boolean shorthand) or an
|
||||
annotation object (e.g. `{ "AlkType:String": { "encoding": "offset-indirect" } }`).
|
||||
## The 18 BAST kinds
|
||||
|
||||
| `kind` | Rust type | Size | Notes |
|
||||
|--------|-----------|-----:|-------|
|
||||
| `int8` | `i8` | 1 | |
|
||||
| `int16` | `i16` | 2 | endian-sensitive |
|
||||
| `int32` | `i32` | 4 | endian-sensitive |
|
||||
| `int64` | `i64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
|
||||
| `uint8` | `u8` | 1 | |
|
||||
| `uint16` | `u16` | 2 | endian-sensitive |
|
||||
| `uint32` | `u32` | 4 | endian-sensitive; also the enum/string/bytes length-prefix width |
|
||||
| `uint64` | `u64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
|
||||
| `float32` | `f32` | 4 | endian-sensitive; NaN/inf rejected by validator |
|
||||
| `float64` | `f64` | 8 | endian-sensitive; NaN/inf rejected by validator |
|
||||
| `bool` | `bool` | 1 | `0x00`=false, `0x01`=true |
|
||||
| `enum` | `u32` index | 4 | index into the `values` array; bounds-checked by the BAST-native validator |
|
||||
| `string` | length-prefixed UTF-8 | 4 + N | `[length: u32][bytes]` by default |
|
||||
| `bytes` | length-prefixed raw bytes | 4 + N | `[length: u32][bytes]` by default |
|
||||
| `struct` | record of fields | composite | nested; field paths are dotted (`"header.version"`) |
|
||||
| `union` | tagged union | composite | byte-offset or field-name discriminator |
|
||||
| `array` | repeated element | composite | fixed-size elements with stride; `count` required in v1 (D-BAST-004) |
|
||||
| `record` | string-keyed map | composite | `[count: u32][key, value]...` |
|
||||
|
||||
The 18 kinds map to the `AlkTypeKind` Rust enum. `AlkTypeKind::to_bast_str`/
|
||||
`from_bast_str` convert between the enum and the lowercase BAST strings
|
||||
(D-BAST-002). Only `struct`, `union`, and `enum` can appear as named
|
||||
`$defs` entries; primitives, arrays, and records appear as field/element/
|
||||
value types via [TypeRef](docs/architecture/bast-format.md#typeref).
|
||||
|
||||
## Two layout modes
|
||||
|
||||
The consumer selects the layout mode at engine construction time via
|
||||
`AlkTypeEngine::compile(schema, mode)`. The same schema can be compiled
|
||||
in either mode. Decided in ADR-002.
|
||||
`AlkTypeEngine::compile(bast_doc, root_name, mode, json_schema)`. The
|
||||
same BAST document can be compiled in either mode. Decided in ADR-002.
|
||||
|
||||
| Mode | Use case | Read API | Write API |
|
||||
|------|----------|----------|-----------|
|
||||
@@ -118,60 +150,114 @@ in either mode. Decided in ADR-002.
|
||||
- **Aligned mode**: a 4-byte length prefix sits at a known offset; the
|
||||
variable data is not part of the static layout. Offset indirection
|
||||
(the metatensor blob pattern: `{offset, length}` pointing into a
|
||||
separate data region) is opt-in via the `encoding` annotation.
|
||||
separate data region) is opt-in via the field-level `encoding`
|
||||
annotation. Fixed-size reservation via `maxLength` is also supported.
|
||||
|
||||
### TUnion discriminators
|
||||
### Union discriminators
|
||||
|
||||
`AlkType:Union` supports two discriminator kinds (ADR-003):
|
||||
`kind: "union"` supports two discriminator kinds (ADR-003):
|
||||
|
||||
- **Byte-offset** — a fixed-size integer at a known byte offset. The
|
||||
SFTP `Packet` pattern: byte 0 is the type byte, bytes 1..N are the
|
||||
variant struct. Mapping keys are stringified integers.
|
||||
- **Field-name** — a named field within the struct. The TypeBox
|
||||
- **Byte-offset** — a fixed-size integer (`uint8`/`uint16`/`uint32`) at
|
||||
a known byte offset. The SFTP `Packet` pattern: byte 0 is the type
|
||||
byte, bytes 1..N are the variant struct. Mapping keys are stringified
|
||||
integers. With all-canonical numeric keys the compiled reader
|
||||
dispatches on the raw integer (no per-read stringification).
|
||||
- **Field-name** — a named field within the union. The TypeBox
|
||||
`typedef.ts` pattern. Mapping keys are string values matching the
|
||||
discriminator field's value.
|
||||
discriminator field's value. The `fields` array declares the
|
||||
discriminator field (D-BAST-005), which must be its first entry; the
|
||||
variant must not re-declare it or any shared field. The builder lays
|
||||
out the declared `fields` first, then the variant's own fields
|
||||
(ADR-011 addendum) — builder, reader, materializer, and validator all
|
||||
agree on that convention.
|
||||
|
||||
Variant `$ref`s are resolved lazily — no compile-time inlining step.
|
||||
|
||||
## Endianness
|
||||
|
||||
Per-schema, default little-endian. Set `"endian": "big"` on the
|
||||
top-level schema (or via `Schema::endian(Endian::Big)`) and the engine
|
||||
byte-swaps every multi-byte read/write accordingly. SFTP consumers
|
||||
specify big-endian; channels' chunk header is big-endian.
|
||||
Per-schema, default little-endian. Set `"endian": "big"` on the root
|
||||
struct (or via `Schema::endian(Endian::Big)`) and the engine byte-swaps
|
||||
every multi-byte read/write accordingly. Field-level `endian` overrides
|
||||
the struct default. SFTP consumers specify big-endian; channels' chunk
|
||||
header is big-endian.
|
||||
|
||||
## Validation
|
||||
|
||||
Two entry points on [`AlkTypeEngine`], one underlying `jsonschema`
|
||||
validator (ADR-010):
|
||||
Two entry points on [`AlkTypeEngine`], two validators for two input
|
||||
types (ADR-VAL-SPLIT):
|
||||
|
||||
- `validate_bytes(&[u8])` — for raw byte buffers (channels' chunk
|
||||
header, SFTP packets). Materializes a `Value` tree from the bytes via
|
||||
the layout engine, then runs the compiled **`ValidationPlan`** (0.2.0
|
||||
used an interpretive BAST walker; 0.3.0 compiles the value-domain
|
||||
constraints — integer ranges, `maxLength`, enum index bounds, union
|
||||
variant dispatch — once at compile time). No
|
||||
`jsonschema` involvement; the BAST document is the complete
|
||||
validation spec for bytes (D-BAST-006).
|
||||
- `validate_json(&Value)` / `is_valid_json(&Value)` — for already-parsed
|
||||
JSON (call's payload schemas).
|
||||
- `validate_bytes(&[u8])` — materializes a `Value` tree from the bytes
|
||||
via the layout engine, then validates that `Value`. Single-call binary
|
||||
buffer validation.
|
||||
JSON (call's `OperationSpec.input_schema` payloads). Validates
|
||||
against a **standard `jsonschema::Validator`** compiled at
|
||||
`AlkTypeEngine::compile` time from a consumer-provided JSON Schema
|
||||
(the `json_schema: Option<&Value>` parameter). BAST is not involved —
|
||||
BAST describes bytes, not JSON shape (D-BAST-007).
|
||||
|
||||
The validator is compiled once at load time; access-time validation is
|
||||
a fast `is_valid()` check. High-throughput paths can skip validation;
|
||||
security-sensitive paths can validate every frame.
|
||||
Both paths return `AlkTypeError::Validation(jsonschema::ValidationError<'static>)`
|
||||
— one uniform payload, one match arm (D-BAST-009).
|
||||
|
||||
Validation is opt-in per operation. High-throughput paths can skip it;
|
||||
security-sensitive paths can validate every frame. The BAST-native
|
||||
validator also fixes a v0.1.0 dead constraint: enum index bounds are
|
||||
now checked (the materializer emits a numeric index; the validator
|
||||
checks it against `values.len()`).
|
||||
|
||||
## BAST document shape
|
||||
|
||||
Every BAST document has the same top-level shape:
|
||||
|
||||
```json
|
||||
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
|
||||
```
|
||||
|
||||
- The `$defs` block is **required** (D-BAST-003).
|
||||
- The **root type name** is a required parameter to
|
||||
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001)
|
||||
— it selects which `$defs` entry is the top-level type.
|
||||
- `$ref` is restricted to `#/$defs/<name>` — one hash lookup, no
|
||||
normalization pass.
|
||||
|
||||
The BAST meta-schema is embedded in the crate as `BAST_META_SCHEMA`
|
||||
(re-exported from the crate root) and published at
|
||||
`https://alk.dev/bast/v1/schema`. Consumers can validate a BAST
|
||||
document's structure with any JSON Schema validator. See
|
||||
[`docs/architecture/bast-format.md`](docs/architecture/bast-format.md)
|
||||
for the full spec.
|
||||
|
||||
## Crate independence
|
||||
|
||||
`alktype` does **not** depend on any application or networking crate.
|
||||
It defines its own types (`AlkTypeError`, `AlkTypeEngine`, `FieldValue`,
|
||||
etc.) and is usable in contexts where networking doesn't exist — CLI
|
||||
tools, test harnesses, schema-building utilities, and WASM targets. The
|
||||
upcoming `alkcall` crate (the `alknet-call` + `alknet-channels`
|
||||
unification) depends on `alktype` for both binary layout and JSON
|
||||
payload schemas; `alktype` knows nothing about `alkcall`.
|
||||
tools, test harnesses, schema-building utilities, and WASM targets.
|
||||
|
||||
## Schemas as untrusted input
|
||||
|
||||
The crate treats schemas as untrusted input. A malformed schema
|
||||
returns `AlkTypeError::Schema` / `AlkTypeError::Offset` from any
|
||||
engine path — never a panic. This matters for hub/spoke topologies
|
||||
where the remote peer provides the schema (e.g. `alkcall` accepting an
|
||||
`OperationSpec` from an arbitrary internet peer). All `unreachable!()`
|
||||
sites in production code were converted to `Err` ahead of v0.1.0
|
||||
(review #002, L2).
|
||||
The crate treats BAST documents as untrusted input. A malformed
|
||||
document returns `AlkTypeError::Schema` from any engine path — never a
|
||||
panic. This matters for hub/spoke topologies where the remote peer
|
||||
provides the schema (e.g. `alkcall` accepting an `OperationSpec` from
|
||||
an arbitrary internet peer). Every `unreachable!()` site in production
|
||||
code was converted to `Err` ahead of v0.1.0 (review #002, L2); the
|
||||
BAST parser preserves this invariant — overflow-safe arithmetic
|
||||
(`checked_add`, `usize::try_from`) on all offset/count casts.
|
||||
|
||||
0.3.0 adds compile-time bounds for adversarial schemas: array counts
|
||||
≤ 2^16 elements, computed array sizes ≤ 2^26 bytes, `align` ≤ 4096,
|
||||
`maxLength` ≤ 2^26, and a shared reference-graph guard that rejects
|
||||
cyclic `$ref`s and >128-deep nesting in every public schema walker.
|
||||
Adversarial buffers fail with `Access` errors at read time — the
|
||||
materializers never preallocate from declared counts. Reviews #006,
|
||||
#007, and #008 document the audit trail
|
||||
([docs/reviews/](docs/reviews/)).
|
||||
|
||||
## Documentation
|
||||
|
||||
@@ -179,20 +265,27 @@ Architecture documentation lives under [`docs/architecture/`](docs/architecture/
|
||||
|
||||
- [Overview](docs/architecture/overview.md) — purpose, "schema is the
|
||||
format" principle, dependencies, consumers, scope boundaries
|
||||
- [Schema layer](docs/architecture/schema-layer.md) — the 19 kinds,
|
||||
jsonschema custom keyword integration, schema annotations
|
||||
- [BAST format](docs/architecture/bast-format.md) — **normative format
|
||||
spec**: meta-schema, TypeRef, TypeDef shapes, validation model
|
||||
- [Schema layer](docs/architecture/schema-layer.md) — the BAST parser
|
||||
(`BastDoc`/`BastDef`/`BastType` typed tree), the 18 kinds, the
|
||||
`AlkTypeKind` enum
|
||||
- [Layout engine](docs/architecture/layout-engine.md) — offset
|
||||
computation, the two layout modes, alignment, endianness
|
||||
- [Data access](docs/architecture/data-access.md) — read/write
|
||||
functions, TUnion dispatch, field paths, zero-copy access
|
||||
- [Validation](docs/architecture/validation.md) — custom keyword
|
||||
validators, `AlkTypeError`, load-time vs access-time validation
|
||||
- [Validation](docs/architecture/validation.md) — the two-validator
|
||||
model, `AlkTypeError`, load-time vs access-time validation
|
||||
- [Builder](docs/architecture/builder.md) — fluent Rust API for
|
||||
constructing alktype JSON Schemas at runtime
|
||||
constructing BAST documents and standard JSON Schemas at runtime
|
||||
- [Architecture decisions (ADRs)](docs/architecture/decisions/) —
|
||||
purpose/scope, two layout modes, schema annotations, error handling,
|
||||
int64/uint64 kinds, packed-mode read factory, TUnion in aligned mode,
|
||||
builder API, `validate_bytes`
|
||||
purpose/scope (ADR-001), BAST format (ADR-BAST), two-validator model
|
||||
(ADR-VAL-SPLIT), two layout modes (ADR-002), schema annotations
|
||||
(ADR-003), error handling (ADR-004), int64/uint64 kinds (ADR-005),
|
||||
non-final inline variable fields (ADR-006), packed-mode read factory
|
||||
(ADR-007), TUnion in aligned mode (ADR-008), builder API (ADR-009),
|
||||
`validate_bytes` (ADR-010), compiled read plan (ADR-011), plan
|
||||
fingerprinting + `ValidationPlan` (ADR-012)
|
||||
|
||||
## License
|
||||
|
||||
|
||||
@@ -0,0 +1,645 @@
|
||||
//! Informal speed comparison: hand-rolled codec logic vs alktype-driven
|
||||
//! codec over the same wire shapes.
|
||||
//!
|
||||
//! History: this bench originated (uncommitted) in `alktty` as the
|
||||
//! curiosity probe that surfaced review #004's 400x read gap — the
|
||||
//! finding that drove the 0.3.0 compiled-forms release (ADR-011/012).
|
||||
//! It now lives here so alktype owns its perf story. The alktty-only
|
||||
//! async roundtrip group (tokio `ChunkReader`/`ChunkWriter` over a
|
||||
//! duplex pipe) was dropped — that measures alktty's I/O stack, not
|
||||
//! this engine.
|
||||
//!
|
||||
//! The alktype engine / layout / plans are built **once outside** the
|
||||
//! measured routine, per the "build cost is paid once" framing.
|
||||
//!
|
||||
//! Shapes:
|
||||
//!
|
||||
//! - **ChunkHeader** — a 5-byte header (`stream_type: uint8`,
|
||||
//! `length: uint32` big-endian). The original shape, kept so numbers
|
||||
//! stay comparable with the historical series (review #004: 400x →
|
||||
//! 0.3.0: ~18x on read p64).
|
||||
//! - **Read** — the hand-rolled path mirrors
|
||||
//! `ChunkReader::read_chunk_after_peek` minus the tokio I/O
|
||||
//! (identical overhead on both sides): validate `stream_type <= 4`,
|
||||
//! parse `u32::from_be_bytes`, slice the payload. The alktype path
|
||||
//! drives `SequentialReader::read_next` over the `ChunkHeader`
|
||||
//! struct, then slices the payload at the parsed length. Both
|
||||
//! return a `&[u8]` payload view — no allocation in either measured
|
||||
//! path.
|
||||
//! - **Write** — serialize the 5-byte header. Hand-rolled mirrors
|
||||
//! `ChunkWriter::write_chunk`'s header writes; alktype uses
|
||||
//! `PackedLayout` offsets (built once) and
|
||||
//! `data_access::write_u8`/`write_u32` at those offsets.
|
||||
//! - **Packet** — a byte-offset-discriminator union
|
||||
//! (`Read {handle, length}` / `Write {handle, length, data: bytes}`),
|
||||
//! the SFTP-shaped case ADR-011's framing argument was about:
|
||||
//! exercises `CompositePlan::Union` dispatch, variant walks, and
|
||||
//! length-prefixed variable reads. The alktype consumer pattern is
|
||||
//! the documented one: `read_next` on the root yields
|
||||
//! `FieldValue::Union { discriminator, variant_start }`, the consumer
|
||||
//! selects the pre-built reader for that variant and walks it over
|
||||
//! `&buf[variant_start..]`.
|
||||
//! - **validate_bytes** — `engine.validate_bytes` per buffer
|
||||
//! (materialize + `ValidationPlan` walk, ADR-010/ADR-012 §3): the
|
||||
//! read+validate-on-untrusted-stream shape `alkcall` cares about. No
|
||||
//! hand comparator: a hand-rolled codec validates inline during the
|
||||
//! (already measured) parse, while `validate_bytes` additionally
|
||||
//! materializes a `Value` tree per buffer — the honest reading is the
|
||||
//! absolute per-chunk cost.
|
||||
//!
|
||||
//! One-shots (paid once at startup, not per chunk):
|
||||
//! `alktype_engine_compile` (dominated by BAST meta-schema
|
||||
//! validation), `alktype_sequential_reader_new` (an `Arc::clone`),
|
||||
//! `alktype_layout_build`.
|
||||
//!
|
||||
//! Two payload sizes (64 B, 4 KiB) so per-chunk fixed overhead is
|
||||
//! visible separately from payload-copy cost.
|
||||
//!
|
||||
//! Run: `cargo bench --bench wire_vs_bast`
|
||||
|
||||
use criterion::{criterion_group, criterion_main, BenchmarkId, Criterion};
|
||||
use std::hint::black_box;
|
||||
|
||||
use alktype::{
|
||||
data_access, AlkTypeEngine, Endian, FieldValue, LayoutBuilder, LayoutMode, PackedLayout,
|
||||
ReadPlan, SequentialReader,
|
||||
};
|
||||
|
||||
/// Mirrors `alktty::wire::MAX_CHUNK_LEN` — the hand-rolled comparator
|
||||
/// validates against the same cap the real codec enforces.
|
||||
const MAX_CHUNK_LEN: u32 = 16 * 1024 * 1024;
|
||||
|
||||
/// The `ChunkHeader` BAST definition. The `StreamType` enum is
|
||||
/// intentionally NOT used — BAST enums encode as `u32`, but the
|
||||
/// on-wire `stream_type` is a `uint8`; both sides read it as `uint8`.
|
||||
const CHUNK_HEADER_BAST: &str = r#"{
|
||||
"$schema": "https://alk.dev/bast/v1/schema",
|
||||
"$defs": {
|
||||
"ChunkHeader": {
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [
|
||||
{ "name": "stream_type", "kind": "uint8" },
|
||||
{ "name": "length", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
}"#;
|
||||
|
||||
/// SFTP-shaped byte-discriminator union: one byte selects the variant,
|
||||
/// `Write` carries a trailing length-prefixed `bytes` field. The root
|
||||
/// struct wraps the union (`AlkTypeEngine::compile` requires a struct
|
||||
/// root); mapping keys are the stringified `uint8` discriminator
|
||||
/// values.
|
||||
const PACKET_BAST: &str = r##"{
|
||||
"$schema": "https://alk.dev/bast/v1/schema",
|
||||
"$defs": {
|
||||
"Packet": {
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [
|
||||
{ "name": "event", "kind": { "$ref": "#/$defs/Event" } }
|
||||
]
|
||||
},
|
||||
"Event": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": {
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
},
|
||||
"Read": {
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" }
|
||||
]
|
||||
},
|
||||
"Write": {
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" },
|
||||
{ "name": "data", "kind": "bytes" }
|
||||
]
|
||||
}
|
||||
}
|
||||
}"##;
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// ChunkHeader fixtures
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// One chunk's worth of bytes on the wire: 5-byte header + payload.
|
||||
fn make_chunk_bytes(stream_type: u8, payload: &[u8]) -> Vec<u8> {
|
||||
let mut buf = Vec::with_capacity(5 + payload.len());
|
||||
buf.push(stream_type);
|
||||
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
|
||||
buf.extend_from_slice(payload);
|
||||
buf
|
||||
}
|
||||
|
||||
/// Concatenate `n` chunks into one buffer, each with `payload_len` bytes.
|
||||
fn make_chunk_stream(n: usize, payload_len: usize) -> Vec<u8> {
|
||||
let payload = vec![0xA5u8; payload_len];
|
||||
let mut buf = Vec::with_capacity(n * (5 + payload_len));
|
||||
for i in 0..n {
|
||||
let st = (i % 5) as u8;
|
||||
buf.extend_from_slice(&make_chunk_bytes(st, &payload));
|
||||
}
|
||||
buf
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Hand-rolled chunk read: mirrors ChunkReader::read_chunk_after_peek minus
|
||||
// the tokio I/O. Returns (stream_type, payload) so the compiler can't
|
||||
// elide the work. Validates stream_type <= 4 and length <= MAX_CHUNK_LEN.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
#[inline]
|
||||
fn hand_read_header(buf: &[u8]) -> Option<(u8, u32)> {
|
||||
if buf.len() < 5 {
|
||||
return None;
|
||||
}
|
||||
let stream_type = buf[0];
|
||||
if stream_type > 4 {
|
||||
return None;
|
||||
}
|
||||
let length = u32::from_be_bytes([buf[1], buf[2], buf[3], buf[4]]);
|
||||
if length > MAX_CHUNK_LEN {
|
||||
return None;
|
||||
}
|
||||
Some((stream_type, length))
|
||||
}
|
||||
|
||||
#[inline]
|
||||
fn hand_read_chunk(buf: &[u8]) -> Option<(u8, &[u8])> {
|
||||
let (st, len) = hand_read_header(buf)?;
|
||||
let end = 5usize.checked_add(len as usize)?;
|
||||
if buf.len() < end {
|
||||
return None;
|
||||
}
|
||||
Some((st, &buf[5..end]))
|
||||
}
|
||||
|
||||
/// Drive `hand_read_chunk` across `n` contiguous chunks in `buf`.
|
||||
/// Returns the total payload bytes consumed (so the loop body is
|
||||
/// meaningfully used and not optimized away).
|
||||
fn hand_read_stream(buf: &[u8], n: usize) -> usize {
|
||||
let mut pos = 0usize;
|
||||
let mut total = 0usize;
|
||||
for _ in 0..n {
|
||||
let (st, payload) = match hand_read_chunk(&buf[pos..]) {
|
||||
Some(v) => v,
|
||||
None => break,
|
||||
};
|
||||
total += payload.len();
|
||||
pos += 5 + payload.len();
|
||||
black_box(st);
|
||||
}
|
||||
black_box(total)
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// alktype chunk read: SequentialReader over ChunkHeader. The reader is
|
||||
// constructed once per benchmark group and reset() between chunks. After
|
||||
// the header read, the payload is sliced at the parsed length — same as
|
||||
// the hand-rolled path. We do NOT re-read a length prefix for the payload
|
||||
// (that would be the double-prefix problem).
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn alktype_read_stream(buf: &[u8], n: usize, reader: &mut SequentialReader) -> usize {
|
||||
let mut pos = 0usize;
|
||||
let mut total = 0usize;
|
||||
for _ in 0..n {
|
||||
reader.reset();
|
||||
let st = match reader.read_next_borrowed(&buf[pos..]) {
|
||||
Ok(Some((_, FieldValue::U8(v)))) => v,
|
||||
_ => break,
|
||||
};
|
||||
let len = match reader.read_next_borrowed(&buf[pos..]) {
|
||||
Ok(Some((_, FieldValue::U32(v)))) => v,
|
||||
_ => break,
|
||||
};
|
||||
if len > MAX_CHUNK_LEN {
|
||||
break;
|
||||
}
|
||||
let end = match 5usize.checked_add(len as usize) {
|
||||
Some(e) if e <= buf.len() - pos => e,
|
||||
_ => break,
|
||||
};
|
||||
let payload = &buf[pos + 5..pos + end];
|
||||
total += payload.len();
|
||||
pos += end;
|
||||
black_box(st);
|
||||
black_box(payload.as_ptr());
|
||||
}
|
||||
black_box(total)
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Hand-rolled chunk write: mirrors ChunkWriter::write_chunk's header
|
||||
// writes into a caller-provided buffer. Writes `n` contiguous chunks.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn hand_write_stream(out: &mut Vec<u8>, n: usize, payload_len: usize) {
|
||||
let payload = vec![0xA5u8; payload_len];
|
||||
for i in 0..n {
|
||||
let st = (i % 5) as u8;
|
||||
let start = out.len();
|
||||
out.resize(start + 5 + payload_len, 0);
|
||||
out[start] = st;
|
||||
out[start + 1..start + 5].copy_from_slice(&(payload_len as u32).to_be_bytes());
|
||||
out[start + 5..start + 5 + payload_len].copy_from_slice(&payload);
|
||||
}
|
||||
black_box(out.len());
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// alktype chunk write: data_access::write_u8 / write_u32 at the
|
||||
// PackedLayout offsets. The layout is built once per group and reused.
|
||||
// Payload bytes are copied with the same slice copy as the hand-rolled
|
||||
// path so the comparison isolates the header-encoding overhead.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn alktype_write_stream(out: &mut Vec<u8>, n: usize, payload_len: usize, layout: &PackedLayout) {
|
||||
let payload = vec![0xA5u8; payload_len];
|
||||
let st_pos = layout.get("stream_type").expect("stream_type field").offset;
|
||||
let len_pos = layout.get("length").expect("length field").offset;
|
||||
for i in 0..n {
|
||||
let start = out.len();
|
||||
out.resize(start + 5 + payload_len, 0);
|
||||
let _ = data_access::write_u8(out, start + st_pos, (i % 5) as u8, "stream_type");
|
||||
let _ = data_access::write_u32(
|
||||
out,
|
||||
start + len_pos,
|
||||
payload_len as u32,
|
||||
"length",
|
||||
Endian::Big,
|
||||
);
|
||||
out[start + 5..start + 5 + payload_len].copy_from_slice(&payload);
|
||||
}
|
||||
black_box(out.len());
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Packet fixtures: byte-disc union stream, alternating Read/Write
|
||||
// variants. Wire layout per packet (packed, big-endian):
|
||||
// Read: disc(1) + handle(4) + length(4) = 9 bytes
|
||||
// Write: disc(1) + handle(4) + length(4) + len(4)+data = 13 + payload
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn make_packet_bytes(disc: u8, payload: &[u8]) -> Vec<u8> {
|
||||
let mut buf = Vec::with_capacity(13 + 4 + payload.len());
|
||||
buf.push(disc);
|
||||
buf.extend_from_slice(&0x0102_0304u32.to_be_bytes());
|
||||
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
|
||||
if disc == 6 {
|
||||
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
|
||||
buf.extend_from_slice(payload);
|
||||
}
|
||||
buf
|
||||
}
|
||||
|
||||
fn make_packet_stream(n: usize, payload_len: usize) -> Vec<u8> {
|
||||
let payload = vec![0xA5u8; payload_len];
|
||||
let mut buf = Vec::new();
|
||||
for i in 0..n {
|
||||
let disc = if i % 2 == 0 { 5u8 } else { 6u8 };
|
||||
buf.extend_from_slice(&make_packet_bytes(disc, &payload));
|
||||
}
|
||||
buf
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Hand-rolled packet read: read the discriminator byte, match the
|
||||
// variant, parse its fields directly. Returns bytes consumed.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn hand_read_packet(buf: &[u8]) -> Option<usize> {
|
||||
let disc = *buf.first()?;
|
||||
match disc {
|
||||
5 => {
|
||||
if buf.len() < 9 {
|
||||
return None;
|
||||
}
|
||||
let handle = u32::from_be_bytes(buf[1..5].try_into().ok()?);
|
||||
let length = u32::from_be_bytes(buf[5..9].try_into().ok()?);
|
||||
black_box((handle, length));
|
||||
Some(9)
|
||||
}
|
||||
6 => {
|
||||
if buf.len() < 13 {
|
||||
return None;
|
||||
}
|
||||
let handle = u32::from_be_bytes(buf[1..5].try_into().ok()?);
|
||||
let length = u32::from_be_bytes(buf[5..9].try_into().ok()?);
|
||||
let data_len = u32::from_be_bytes(buf[9..13].try_into().ok()?);
|
||||
let end = 13usize.checked_add(data_len as usize)?;
|
||||
if buf.len() < end {
|
||||
return None;
|
||||
}
|
||||
black_box((handle, length));
|
||||
black_box(&buf[13..end].as_ptr());
|
||||
Some(end)
|
||||
}
|
||||
_ => None,
|
||||
}
|
||||
}
|
||||
|
||||
fn hand_read_packet_stream(buf: &[u8], n: usize) -> usize {
|
||||
let mut pos = 0usize;
|
||||
let mut total = 0usize;
|
||||
for _ in 0..n {
|
||||
let Some(consumed) = hand_read_packet(&buf[pos..]) else {
|
||||
break;
|
||||
};
|
||||
total += consumed;
|
||||
pos += consumed;
|
||||
}
|
||||
black_box(total)
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// alktype packet read: the documented union consumer contract. The root
|
||||
// reader walks the wrapping struct; `read_next` returns
|
||||
// `FieldValue::Union { discriminator, variant_start }`; the consumer
|
||||
// selects the pre-built reader for that variant and walks it over
|
||||
// `&buf[pos + variant_start..]` until exhausted.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Walk one variant's fields to exhaustion; returns bytes consumed.
|
||||
/// Uses `read_next_borrowed` — the zero-alloc hot-loop pattern for
|
||||
/// consumers that match or discard the field name.
|
||||
fn alktype_walk_variant(reader: &mut SequentialReader, buf: &[u8]) -> Option<usize> {
|
||||
reader.reset();
|
||||
loop {
|
||||
match reader.read_next_borrowed(buf) {
|
||||
Ok(Some((name, value))) => {
|
||||
black_box(name);
|
||||
black_box(&value);
|
||||
}
|
||||
Ok(None) => return Some(reader.position()),
|
||||
Err(_) => return None,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
fn alktype_read_packet_stream(
|
||||
buf: &[u8],
|
||||
n: usize,
|
||||
packet: &mut SequentialReader,
|
||||
read: &mut SequentialReader,
|
||||
write: &mut SequentialReader,
|
||||
) -> usize {
|
||||
let mut pos = 0usize;
|
||||
let mut total = 0usize;
|
||||
for _ in 0..n {
|
||||
packet.reset();
|
||||
let disc = match packet.read_next_borrowed(&buf[pos..]) {
|
||||
Ok(Some((_, FieldValue::Union {
|
||||
discriminator,
|
||||
variant_start,
|
||||
}))) => {
|
||||
pos += variant_start;
|
||||
discriminator
|
||||
}
|
||||
_ => break,
|
||||
};
|
||||
let vbuf = &buf[pos..];
|
||||
let consumed = match disc.as_str() {
|
||||
"5" => alktype_walk_variant(read, vbuf),
|
||||
"6" => alktype_walk_variant(write, vbuf),
|
||||
_ => break,
|
||||
};
|
||||
let Some(consumed) = consumed else {
|
||||
break;
|
||||
};
|
||||
total += consumed;
|
||||
pos += consumed;
|
||||
}
|
||||
black_box(total)
|
||||
}
|
||||
|
||||
/// One-time sanity check (outside the measured loops): the union
|
||||
/// consumer pattern the stream loop relies on — root reader reports the
|
||||
/// mapping key and the variant start; the variant reader's walk to
|
||||
/// exhaustion reports exactly the variant's byte size, so
|
||||
/// `variant_start + consumed` lands on the next packet.
|
||||
fn assert_packet_reader_parity(
|
||||
payload_len: usize,
|
||||
packet: &mut SequentialReader,
|
||||
read: &mut SequentialReader,
|
||||
write: &mut SequentialReader,
|
||||
) {
|
||||
let payload = vec![0u8; payload_len];
|
||||
|
||||
let read_pkt = make_packet_bytes(5, &payload);
|
||||
packet.reset();
|
||||
match packet.read_next_borrowed(&read_pkt) {
|
||||
Ok(Some((_, FieldValue::Union {
|
||||
discriminator,
|
||||
variant_start,
|
||||
}))) => {
|
||||
assert_eq!(discriminator, "5");
|
||||
assert_eq!(variant_start, 1, "variant starts after the 1-byte disc");
|
||||
}
|
||||
_ => panic!("expected union value for Read packet"),
|
||||
}
|
||||
let consumed = alktype_walk_variant(read, &read_pkt[1..]).expect("read variant walk");
|
||||
assert_eq!(consumed, 8, "Read = handle(4) + length(4)");
|
||||
assert_eq!(1 + consumed, read_pkt.len(), "Read packet fully consumed");
|
||||
|
||||
let write_pkt = make_packet_bytes(6, &payload);
|
||||
packet.reset();
|
||||
match packet.read_next_borrowed(&write_pkt) {
|
||||
Ok(Some((_, FieldValue::Union {
|
||||
discriminator,
|
||||
variant_start,
|
||||
}))) => {
|
||||
assert_eq!(discriminator, "6");
|
||||
assert_eq!(variant_start, 1);
|
||||
}
|
||||
_ => panic!("expected union value for Write packet"),
|
||||
}
|
||||
let consumed = alktype_walk_variant(write, &write_pkt[1..]).expect("write variant walk");
|
||||
assert_eq!(
|
||||
consumed,
|
||||
12 + payload_len,
|
||||
"Write = handle(4) + length(4) + len-prefix(4) + data"
|
||||
);
|
||||
assert_eq!(1 + consumed, write_pkt.len(), "Write packet fully consumed");
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Benchmarks
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn bench_read(c: &mut Criterion) {
|
||||
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
|
||||
let engine =
|
||||
AlkTypeEngine::compile(&bast, "ChunkHeader", LayoutMode::Packed, None).expect("compile");
|
||||
let mut reader = engine.sequential_reader().expect("packed reader");
|
||||
|
||||
let mut group = c.benchmark_group("read_chunk_stream");
|
||||
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
|
||||
let n = 1024;
|
||||
let buf = make_chunk_stream(n, payload_len);
|
||||
|
||||
group.bench_with_input(BenchmarkId::new("hand_rolled", label), &n, |b, &n| {
|
||||
b.iter(|| hand_read_stream(black_box(&buf), n));
|
||||
});
|
||||
group.bench_with_input(BenchmarkId::new("alktype", label), &n, |b, &n| {
|
||||
b.iter(|| alktype_read_stream(black_box(&buf), n, &mut reader));
|
||||
});
|
||||
}
|
||||
group.finish();
|
||||
}
|
||||
|
||||
fn bench_write(c: &mut Criterion) {
|
||||
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
|
||||
let builder = LayoutBuilder::new(&bast, "ChunkHeader").expect("builder");
|
||||
let layout = builder
|
||||
.build(&std::collections::HashMap::new())
|
||||
.expect("layout");
|
||||
|
||||
let mut group = c.benchmark_group("write_chunk_stream");
|
||||
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
|
||||
let n = 1024;
|
||||
|
||||
group.bench_with_input(
|
||||
BenchmarkId::new("hand_rolled", label),
|
||||
&(n, payload_len),
|
||||
|b, &(n, pl)| {
|
||||
b.iter(|| {
|
||||
let mut out = Vec::with_capacity(n * (5 + pl));
|
||||
hand_write_stream(&mut out, n, pl);
|
||||
});
|
||||
},
|
||||
);
|
||||
group.bench_with_input(
|
||||
BenchmarkId::new("alktype", label),
|
||||
&(n, payload_len),
|
||||
|b, &(n, pl)| {
|
||||
b.iter(|| {
|
||||
let mut out = Vec::with_capacity(n * (5 + pl));
|
||||
alktype_write_stream(&mut out, n, pl, &layout);
|
||||
});
|
||||
},
|
||||
);
|
||||
}
|
||||
group.finish();
|
||||
}
|
||||
|
||||
fn bench_packet_read(c: &mut Criterion) {
|
||||
let bast: serde_json::Value = serde_json::from_str(PACKET_BAST).expect("bast json");
|
||||
let engine =
|
||||
AlkTypeEngine::compile(&bast, "Packet", LayoutMode::Packed, None).expect("compile");
|
||||
let mut packet_reader = engine.sequential_reader().expect("packed reader");
|
||||
let read_plan = std::sync::Arc::new(ReadPlan::compile(&bast, "Read").expect("read plan"));
|
||||
let write_plan = std::sync::Arc::new(ReadPlan::compile(&bast, "Write").expect("write plan"));
|
||||
let mut read_reader = SequentialReader::new(read_plan);
|
||||
let mut write_reader = SequentialReader::new(write_plan);
|
||||
|
||||
// One-time parity check of the union consumer pattern (not measured).
|
||||
assert_packet_reader_parity(64, &mut packet_reader, &mut read_reader, &mut write_reader);
|
||||
|
||||
let mut group = c.benchmark_group("read_packet_stream");
|
||||
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
|
||||
let n = 1024;
|
||||
let buf = make_packet_stream(n, payload_len);
|
||||
|
||||
group.bench_with_input(BenchmarkId::new("hand_rolled", label), &n, |b, &n| {
|
||||
b.iter(|| hand_read_packet_stream(black_box(&buf), n));
|
||||
});
|
||||
group.bench_with_input(BenchmarkId::new("alktype", label), &n, |b, &n| {
|
||||
b.iter(|| {
|
||||
alktype_read_packet_stream(
|
||||
black_box(&buf),
|
||||
n,
|
||||
&mut packet_reader,
|
||||
&mut read_reader,
|
||||
&mut write_reader,
|
||||
)
|
||||
});
|
||||
});
|
||||
}
|
||||
group.finish();
|
||||
}
|
||||
|
||||
fn bench_validate(c: &mut Criterion) {
|
||||
let header_bast: serde_json::Value =
|
||||
serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
|
||||
let header_engine = AlkTypeEngine::compile(&header_bast, "ChunkHeader", LayoutMode::Packed, None)
|
||||
.expect("compile");
|
||||
let packet_bast: serde_json::Value = serde_json::from_str(PACKET_BAST).expect("bast json");
|
||||
let packet_engine =
|
||||
AlkTypeEngine::compile(&packet_bast, "Packet", LayoutMode::Packed, None).expect("compile");
|
||||
|
||||
let mut group = c.benchmark_group("validate_stream");
|
||||
let n = 1024;
|
||||
|
||||
let headers: Vec<Vec<u8>> = (0..n)
|
||||
.map(|i| make_chunk_bytes((i % 5) as u8, &[0xA5u8; 64]))
|
||||
.collect();
|
||||
group.bench_function("alktype_chunk_header", |b| {
|
||||
b.iter(|| {
|
||||
for h in &headers {
|
||||
header_engine.validate_bytes(black_box(h)).expect("validate");
|
||||
}
|
||||
})
|
||||
});
|
||||
|
||||
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
|
||||
let payload = vec![0xA5u8; payload_len];
|
||||
let packets: Vec<Vec<u8>> = (0..n)
|
||||
.map(|i| make_packet_bytes(if i % 2 == 0 { 5 } else { 6 }, &payload))
|
||||
.collect();
|
||||
group.bench_with_input(
|
||||
BenchmarkId::new("alktype_packet", label),
|
||||
&packets,
|
||||
|b, packets| {
|
||||
b.iter(|| {
|
||||
for p in packets {
|
||||
packet_engine.validate_bytes(black_box(p)).expect("validate");
|
||||
}
|
||||
})
|
||||
},
|
||||
);
|
||||
}
|
||||
group.finish();
|
||||
}
|
||||
|
||||
/// One-shot costs paid once at startup, not per chunk.
|
||||
fn bench_oneshot(c: &mut Criterion) {
|
||||
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
|
||||
c.bench_function("alktype_engine_compile", |b| {
|
||||
b.iter(|| {
|
||||
let _ =
|
||||
AlkTypeEngine::compile(black_box(&bast), "ChunkHeader", LayoutMode::Packed, None)
|
||||
.expect("compile");
|
||||
});
|
||||
});
|
||||
c.bench_function("alktype_sequential_reader_new", |b| {
|
||||
let engine = AlkTypeEngine::compile(&bast, "ChunkHeader", LayoutMode::Packed, None)
|
||||
.expect("compile");
|
||||
b.iter(|| engine.sequential_reader());
|
||||
});
|
||||
c.bench_function("alktype_layout_build", |b| {
|
||||
let builder = LayoutBuilder::new(&bast, "ChunkHeader").expect("builder");
|
||||
b.iter(|| builder.build(&std::collections::HashMap::new()));
|
||||
});
|
||||
}
|
||||
|
||||
criterion_group!(
|
||||
benches,
|
||||
bench_read,
|
||||
bench_write,
|
||||
bench_packet_read,
|
||||
bench_validate,
|
||||
bench_oneshot
|
||||
);
|
||||
criterion_main!(benches);
|
||||
+75
-57
@@ -1,12 +1,12 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-08-11
|
||||
status: accepted
|
||||
last_updated: 2026-08-15
|
||||
---
|
||||
|
||||
# alktype
|
||||
|
||||
The binary struct engine: a small Rust crate that takes a JSON Schema
|
||||
with `AlkType:*` custom keywords and produces an offset map, read/write
|
||||
The binary struct engine: a small Rust crate that takes a BAST (Binary
|
||||
Abstract Syntax Tree) document and produces an offset map, read/write
|
||||
functions, and validation — all driven by the schema. The schema is the
|
||||
format definition; the engine is generic.
|
||||
|
||||
@@ -14,60 +14,75 @@ format definition; the engine is generic.
|
||||
|
||||
| Document | Status | Description |
|
||||
|----------|--------|-------------|
|
||||
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
|
||||
| [schema-layer.md](schema-layer.md) | draft | The 19 `AlkType:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
|
||||
| [overview.md](overview.md) | accepted | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
|
||||
| [`bast-format.md`](bast-format.md) | accepted | **Normative BAST format specification.** Meta-schema, TypeRef, TypeDef shapes (Struct/Union/Enum/FieldDef), examples, validation model. The format the engine consumes. |
|
||||
| [schema-layer.md](schema-layer.md) | accepted | The BAST parser (`src/bast.rs`) — the typed tree (`BastDoc`/`BastDef`/`BastType`/…) every engine module walks, the 18 BAST kinds, the `AlkTypeKind` enum, and the foundational annotation types. |
|
||||
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
|
||||
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
|
||||
| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine`; `validate_bytes` for binary buffers (ADR-010) |
|
||||
| [builder.md](builder.md) | draft | Fluent Rust API for constructing alktype JSON Schemas at runtime, producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema (ADR-009) |
|
||||
| [validation.md](validation.md) | accepted | The two-validator model (BAST-native for `validate_bytes`, standard `jsonschema` for `validate_json`), `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine` as the compiled form of a BAST document (ADR-010, ADR-VAL-SPLIT). |
|
||||
| [builder.md](builder.md) | accepted | Fluent Rust API for constructing BAST documents (`struct_()`) and standard JSON Schemas (`object()`) at runtime, producing `serde_json::Value` (ADR-009, D-BAST-008). |
|
||||
|
||||
### In-progress work
|
||||
|
||||
| Document | Status | Description |
|
||||
|----------|--------|-------------|
|
||||
| [BAST pivot — research record](../research/bast-pivot.md) | accepted | Motivation, POC scope and result, decisions D-BAST-001..009, risks for the BAST format pivot. Implemented in steps 1–10. |
|
||||
| [BAST pivot — implementation plan](../plans/bast-implementation.md) | accepted | Ordered implementation steps, the public-API semver contract, and the ADR-sync checklist for the BAST pivot. Steps 1–10 complete. |
|
||||
|
||||
## Applicable ADRs
|
||||
|
||||
| ADR | Title | Relevance |
|
||||
|-----|-------|-----------|
|
||||
| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
|
||||
| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` |
|
||||
| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations |
|
||||
| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors |
|
||||
| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries. *Format-specific content superseded by ADR-BAST; purpose/scope retained.* |
|
||||
| [BAST](decisions/bast-bast-format.md) | BAST (Binary Abstract Syntax Tree) as the Schema Format | The BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary. Supersedes ADR-001's format-specific content; records D-BAST-001..009. |
|
||||
| [VAL-SPLIT](decisions/val-split-two-validator-model.md) | Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON | `validate_bytes` uses the BAST-native validator; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema. Records D-BAST-006/007/009. |
|
||||
| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` (format-agnostic — input format changed, modes didn't) |
|
||||
| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Annotation *semantics* (carry forward unchanged); annotation *location* moved to BAST type-level properties under the pivot |
|
||||
| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum (shape unchanged, D-BAST-009); load-time build, access-time check; field-path-carrying errors. *Validation-strategy section refined by ADR-VAL-SPLIT.* |
|
||||
| [005](decisions/005-int64-uint64-first-class-kinds.md) | Int64/Uint64 as First-Class Kinds | 64-bit integers (SFTP offsets, metatensor data_offsets); JSON precision caveat |
|
||||
| [006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode | Prevents silent data corruption (inline variable data clobbering subsequent fields) |
|
||||
| [007](decisions/007-packed-mode-read-factory.md) | Packed-Mode Read API — Engine as SequentialReader Factory | `engine.sequential_reader()` returns an owned reader, not a reference |
|
||||
| [008](decisions/008-reject-tunion-in-aligned-mode.md) | Reject TUnion in Aligned Mode for v1 | Unions are the protocol pattern; aligned-mode union semantics were broken |
|
||||
| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
|
||||
| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate; two methods on one struct, not a trait |
|
||||
| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003. *Output format amended to BAST / standard JSON Schema by ADR-BAST.* |
|
||||
| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate. *Validation step amended to the BAST-native validator by ADR-VAL-SPLIT.* |
|
||||
| [011](decisions/011-compiled-read-plan-for-packed-mode.md) | Compiled Read Plan for Packed Mode | `ReadPlan` — the packed read-side compiled form, symmetric to `OffsetMap` (aligned) and `PackedLayout` (packed write). Closes review #004's 400x read-path gap; retires ADR-007's "re-parse on demand" framing. *Accepted — implemented in 0.3.0 (phases 1–2).* |
|
||||
| [012](decisions/012-plan-fingerprinting-and-m1-closure.md) | Plan Fingerprinting, ValidationPlan, and Closing the Deferred M1 Sites in 0.3.0 | `ReadPlan`/`OffsetMap`/`ValidationPlan` `Hash + Eq` + `fingerprint()`; owned `BastDoc` (lifetime removal); `OffsetMap` carries `LeafMeta` to close the aligned-side M1 sites; `ValidationPlan` retires the interpretive `bast_validation` walk (review #005 M3 reversed the original deferral). Bundles with ADR-011 into one 0.3.0 breaking release. *Accepted — fully implemented in 0.3.0 (fingerprinting, owned `BastDoc`, `LeafMeta`, `ValidationPlan`).* |
|
||||
|
||||
## Relevant Open Questions
|
||||
|
||||
| OQ | Title | Status | Relevance |
|
||||
|----|-------|--------|-----------|
|
||||
| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it |
|
||||
| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it. BAST arrays require `count` in v1 (D-BAST-004), aligning with this deferral. |
|
||||
| OQ-002 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case |
|
||||
| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Resolved in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) |
|
||||
| OQ-004 | `Discriminator::Field` name — `&str` or `String` | open | Builder API ownership question; resolve before the SFTP Packet POC's field-name discriminator path |
|
||||
| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | open | Blocks the SFTP Packet `validate_bytes` POC (next round) |
|
||||
| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | open | Documentation fix in builder.md; the engine requires `AlkType:Struct` at the top level |
|
||||
| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | open | Blocks the SFTP use case for `validate_bytes` (binary `handle`/`data` fields) |
|
||||
| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Shipped in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) |
|
||||
| OQ-004 | `Discriminator::Field` name — `&str` or `String` | resolved | `String`, for ownership simplicity |
|
||||
| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | resolved | Both kinds return `{ "__discriminator": <value>, ...variant-fields }` |
|
||||
| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | resolved | [builder.md](builder.md) Example 3 wraps the union in a `Schema::struct_().field("payload", ...)` |
|
||||
| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | resolved | Array of u8: materializer produces `Value::Array` of `Value::Number`; BAST-native validator accepts both `Value::String` and `Value::Array` |
|
||||
| OQ-008 | `UnionValidator` variant dispatch | resolved | BAST-native validator recurses into the selected variant's BAST definition on `__discriminator` lookup — no custom keywords, no `inline_union_variant_refs` |
|
||||
|
||||
## Key Design Principles
|
||||
|
||||
1. **The schema is the format.** A JSON Schema with `AlkType:*` custom
|
||||
keywords is both the validation spec and the layout spec. No separate
|
||||
format definition, no separate parser, no separate validator. One
|
||||
schema, three uses: validate, compute offsets, access data. See
|
||||
[overview.md](overview.md) and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
|
||||
1. **The schema is the format.** A BAST document is both the layout
|
||||
spec and the validation spec for bytes. No separate format
|
||||
definition, no separate parser, no separate validator. One schema,
|
||||
three uses: validate, compute offsets, access data. See
|
||||
[overview.md](overview.md), [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md),
|
||||
and [ADR-BAST](decisions/bast-bast-format.md).
|
||||
|
||||
2. **jsonschema is the validation engine, not a custom engine.** The
|
||||
`jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with
|
||||
custom keyword support. The novel code is the offset computation, not
|
||||
the validation. This eliminates ~14,000 lines of hand-rolled schema
|
||||
engines (typebox-rs, the @alkdev/alktype prototype). See [schema-layer.md](schema-layer.md)
|
||||
and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
|
||||
2. **BAST is a JSON Schema dialect, not a custom format.** A BAST
|
||||
document is valid JSON conforming to the BAST meta-schema (a
|
||||
standard Draft 2020-12 JSON Schema). Any JSON Schema validator can
|
||||
check whether a BAST document is well-formed; editors with JSON
|
||||
Schema support provide autocomplete for free. See
|
||||
[`bast-format.md`](bast-format.md) and
|
||||
[ADR-BAST](decisions/bast-bast-format.md).
|
||||
|
||||
3. **Two layout modes for two use cases.** Packed sequential
|
||||
(`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP,
|
||||
channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats
|
||||
(metatensor). The consumer selects the mode; the schema is the same.
|
||||
See [layout-engine.md](layout-engine.md) and
|
||||
(metatensor). The consumer selects the mode; the BAST document is the
|
||||
same. See [layout-engine.md](layout-engine.md) and
|
||||
[ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md).
|
||||
|
||||
4. **Variable-length types default to inline length-prefixing.**
|
||||
@@ -84,38 +99,41 @@ format definition; the engine is generic.
|
||||
[ADR-003](decisions/003-schema-annotations.md).
|
||||
|
||||
6. **Endianness is per-schema, default little-endian.** The engine reads
|
||||
the `"endian"` annotation and byte-swaps accordingly. SFTP consumers
|
||||
specify `"endian": "big"`. See [layout-engine.md](layout-engine.md)
|
||||
and [ADR-003](decisions/003-schema-annotations.md).
|
||||
the struct-level `"endian"` annotation and byte-swaps accordingly.
|
||||
SFTP consumers specify `"endian": "big"`. See
|
||||
[layout-engine.md](layout-engine.md) and
|
||||
[ADR-003](decisions/003-schema-annotations.md).
|
||||
|
||||
7. **Validation is opt-in, built once at load time.** The jsonschema
|
||||
validator is compiled once at schema load time. Access-time validation
|
||||
is a fast `is_valid()` check. High-throughput paths can skip
|
||||
validation; security-sensitive paths can validate every frame. See
|
||||
7. **Two validators for two input types.** `validate_bytes(&[u8])` uses
|
||||
the BAST-native validator (a recursive walker over the BAST type
|
||||
tree — no `jsonschema` involvement). `validate_json(&Value)` uses a
|
||||
standard `jsonschema::Validator` from a consumer-provided JSON Schema
|
||||
(BAST is not involved — BAST describes bytes, not JSON shape). One
|
||||
`AlkTypeError::Validation` variant covers both (D-BAST-009). See
|
||||
[validation.md](validation.md) and
|
||||
[ADR-004](decisions/004-error-handling-validation-strategy.md).
|
||||
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
|
||||
|
||||
8. **Not a serialization framework.** The alktype engine is not a
|
||||
general-purpose serde replacement. It operates on raw byte buffers at
|
||||
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||
with a known schema, use alktype. See [overview.md](overview.md) and
|
||||
computed offsets — no reflection, no dynamic dispatch per field. For
|
||||
JSON data, use serde. For binary data with a known BAST document, use
|
||||
alktype. See [overview.md](overview.md) and
|
||||
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
9. **Schemas can be built at runtime from Rust (v0.1.0).** A fluent
|
||||
builder API produces `serde_json::Value` for both AlkType-kind
|
||||
schemas and standard JSON Schema, covering alkcall's two roles
|
||||
9. **Schemas can be built at runtime from Rust.** A fluent builder API
|
||||
produces `serde_json::Value` for both BAST documents (`struct_()`) and
|
||||
standard JSON Schemas (`object()`), covering alkcall's two roles
|
||||
(binary layout + JSON payloads) from one module. The builder is
|
||||
additive — consumers with static schemas continue to load JSON.
|
||||
See [builder.md](builder.md) and [ADR-009](decisions/009-builder-api.md).
|
||||
additive — consumers with static BAST documents continue to load
|
||||
JSON. See [builder.md](builder.md) and
|
||||
[ADR-009](decisions/009-builder-api.md).
|
||||
|
||||
10. **Two validation entry points, one engine (v0.1.0).**
|
||||
`validate_json(&Value)` for already-parsed JSON (call's payloads);
|
||||
`validate_bytes(&[u8])` for binary buffers (channels' chunk header).
|
||||
Same underlying `jsonschema` validator; the bytes path materializes
|
||||
a `Value` tree via the layout engine, then validates. See
|
||||
[validation.md](validation.md) and
|
||||
[ADR-010](decisions/010-generalized-validation-validate-bytes.md).
|
||||
10. **Two validation entry points, one engine.** `validate_json(&Value)`
|
||||
for already-parsed JSON (call's payloads); `validate_bytes(&[u8])`
|
||||
for binary buffers (channels' chunk header). Different validators,
|
||||
one `AlkTypeError::Validation` variant. See [validation.md](validation.md),
|
||||
[ADR-010](decisions/010-generalized-validation-validate-bytes.md),
|
||||
and [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
|
||||
|
||||
## References
|
||||
|
||||
@@ -135,4 +153,4 @@ format definition; the engine is generic.
|
||||
> **Note**: The research findings, POC code, and prior-attempt paths above
|
||||
> refer to the parent `@alkdev/alknet` workspace where this crate originated.
|
||||
> They are preserved here as historical context for the architectural
|
||||
> decisions; the artifacts themselves are not part of this standalone repo.
|
||||
> decisions; the artifacts themselves are not part of this standalone repo.
|
||||
@@ -0,0 +1,701 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-08-15
|
||||
---
|
||||
|
||||
# alktype — BAST Format
|
||||
|
||||
**BAST** (Binary Abstract Syntax Tree) is alktype's schema format: a
|
||||
JSON document that describes binary data layouts using a `kind`-based
|
||||
vocabulary with `$defs`/`$ref` for composition. BAST replaces the
|
||||
v0.1.0 `AlkType:*` custom-keyword JSON Schema format.
|
||||
|
||||
This document is the **normative format specification**. It is grounded
|
||||
in the POC on branch `bast-validator-poc` (commit `f371fe4`) and the
|
||||
decisions D-BAST-001 through D-BAST-009 in
|
||||
[the pivot research record](../research/bast-pivot.md#decisions). The
|
||||
implementation plan is
|
||||
[`docs/plans/bast-implementation.md`](../plans/bast-implementation.md).
|
||||
|
||||
Until the BAST pivot lands in code, [`schema-layer.md`](schema-layer.md)
|
||||
describes the *current* (custom-keyword) schema layer. This document
|
||||
describes the *target* (BAST) schema layer. They coexist temporarily;
|
||||
the implementation plan's final step retires `schema-layer.md`'s
|
||||
custom-keyword content.
|
||||
|
||||
## Design Principles
|
||||
|
||||
1. **BAST is a JSON Schema instance.** A BAST document is valid JSON
|
||||
that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON
|
||||
Schema). Any JSON Schema validator can check whether a BAST document
|
||||
is well-formed; editors with JSON Schema support provide autocomplete
|
||||
and inline validation for free.
|
||||
|
||||
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
|
||||
top-level `$defs` block. `$ref` handles cross-references and union
|
||||
variant references — the same pattern as JSON Schema's own `$defs`
|
||||
and TypeBox's `Type.Module`. No custom reference resolution mechanism.
|
||||
|
||||
3. **`kind`-based vocabulary.** Every type has a `kind` field whose
|
||||
value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.).
|
||||
This replaces the `AlkType:*` custom-keyword pattern with a flat,
|
||||
easily-matched string. The 18 `AlkTypeKind` enum variants are
|
||||
unchanged; `AlkTypeKind::from_str`/`to_str` map between the enum and
|
||||
the lowercase BAST strings (D-BAST-002).
|
||||
|
||||
4. **Order is explicit.** Struct fields are an ordered array, not an
|
||||
object with `properties`. Field order is unambiguous — no reliance on
|
||||
`serde_json`'s `preserve_order` for correctness — and matches the
|
||||
mental model of binary layouts.
|
||||
|
||||
5. **Annotations are type-level properties.** Endianness, alignment,
|
||||
encoding, and discriminators are properties of the type definition
|
||||
or field, not custom keywords on a separate schema object. Their
|
||||
*semantics* carry forward unchanged from ADR-003; only their
|
||||
*location* moves.
|
||||
|
||||
## Document Shape
|
||||
|
||||
Every BAST document has the same top-level shape:
|
||||
|
||||
```json
|
||||
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
|
||||
```
|
||||
|
||||
- The `$defs` block is **required** (D-BAST-003). Single-type documents
|
||||
are a special case with one entry. A bare struct at the top level
|
||||
would be a different shape with different parsing logic and no home
|
||||
for additional definitions — rejected.
|
||||
- The **root type name** is a required parameter to
|
||||
`AlkTypeEngine::compile(bast_doc, root_name, mode)` (D-BAST-001). It
|
||||
selects which `$defs` entry is the top-level type. Convention (first
|
||||
entry) is fragile and depends on JSON key order; a `$root` marker is
|
||||
redundant with an explicit parameter.
|
||||
|
||||
## The Meta-Schema
|
||||
|
||||
The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that
|
||||
validates the *structure* of BAST documents (is it well-formed?). It
|
||||
lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded
|
||||
in the crate for offline use. A different validator — the BAST-native
|
||||
validator (see [Validation Model](#validation-model) below) — validates
|
||||
*binary data* against a BAST document (are the bytes a valid instance?).
|
||||
These are different validators for different inputs.
|
||||
|
||||
```json
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://alk.dev/bast/v1/schema",
|
||||
"title": "Binary Abstract Syntax Tree (BAST) v1",
|
||||
"description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"$defs": {
|
||||
"type": "object",
|
||||
"additionalProperties": { "$ref": "#/$defs/TypeDef" }
|
||||
}
|
||||
},
|
||||
"required": ["$defs"],
|
||||
"$defs": {
|
||||
"TypeDef": {
|
||||
"oneOf": [
|
||||
{ "$ref": "#/$defs/StructDef" },
|
||||
{ "$ref": "#/$defs/UnionDef" },
|
||||
{ "$ref": "#/$defs/EnumDef" }
|
||||
]
|
||||
},
|
||||
"StructDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "struct" },
|
||||
"endian": { "enum": ["little", "big"] },
|
||||
"align": { "type": "integer", "minimum": 1 },
|
||||
"fields": {
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/FieldDef" }
|
||||
}
|
||||
},
|
||||
"required": ["kind", "fields"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"FieldDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
|
||||
"kind": { "$ref": "#/$defs/TypeRef" },
|
||||
"endian": { "enum": ["little", "big"] },
|
||||
"align": { "type": "integer", "minimum": 1 },
|
||||
"encoding": { "enum": ["length-prefixed", "offset-indirect"] },
|
||||
"maxLength": { "type": "integer", "minimum": 0 }
|
||||
},
|
||||
"if": {
|
||||
"properties": {
|
||||
"kind": { "enum": ["string", "bytes"] }
|
||||
}
|
||||
},
|
||||
"else": { "properties": { "maxLength": false } },
|
||||
"required": ["name", "kind"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"TypeRef": {
|
||||
"oneOf": [
|
||||
{
|
||||
"description": "Primitive type",
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"int8", "int16", "int32", "int64",
|
||||
"uint8", "uint16", "uint32", "uint64",
|
||||
"float32", "float64",
|
||||
"bool", "string", "bytes"
|
||||
]
|
||||
},
|
||||
{
|
||||
"description": "Reference to a named $defs entry",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
|
||||
},
|
||||
"required": ["$ref"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"description": "Array type (fixed-size only in v1 — count is required)",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "array" },
|
||||
"element": { "$ref": "#/$defs/TypeRef" },
|
||||
"count": { "type": "integer", "minimum": 0 }
|
||||
},
|
||||
"required": ["kind", "element", "count"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"description": "Record (string-keyed map) type",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "record" },
|
||||
"values": { "$ref": "#/$defs/TypeRef" }
|
||||
},
|
||||
"required": ["kind", "values"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
]
|
||||
},
|
||||
"UnionDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "union" },
|
||||
"endian": { "enum": ["little", "big"] },
|
||||
"fields": {
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/FieldDef" }
|
||||
},
|
||||
"discriminator": {
|
||||
"oneOf": [
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "byte" },
|
||||
"offset": { "type": "integer", "minimum": 0 },
|
||||
"type": { "enum": ["uint8", "uint16", "uint32"] }
|
||||
},
|
||||
"required": ["kind", "offset", "type"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "field" },
|
||||
"name": { "type": "string" }
|
||||
},
|
||||
"required": ["kind", "name"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
]
|
||||
},
|
||||
"mapping": {
|
||||
"type": "object",
|
||||
"additionalProperties": { "$ref": "#/$defs/TypeRef" }
|
||||
}
|
||||
},
|
||||
"required": ["kind", "discriminator", "mapping"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"EnumDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "enum" },
|
||||
"values": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"minItems": 1
|
||||
}
|
||||
},
|
||||
"required": ["kind", "values"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Type Definitions
|
||||
|
||||
### Struct
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"align": 256,
|
||||
"fields": [
|
||||
{ "name": "channel_id", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
- `kind` (required): `"struct"`.
|
||||
- `endian` (optional): `"little"` (default) or `"big"`. Sets the default
|
||||
for all fields; field-level `endian` overrides.
|
||||
- `align` (optional): struct-level alignment, only meaningful in aligned
|
||||
static mode (ADR-002/003).
|
||||
- `fields` (required): ordered array of [FieldDef](#fielddef). Array
|
||||
position is field order — no reliance on JSON key order.
|
||||
|
||||
### FieldDef
|
||||
|
||||
```json
|
||||
{ "name": "handle", "kind": "string", "encoding": "offset-indirect", "maxLength": 256 }
|
||||
```
|
||||
|
||||
- `name` (required): identifier, `^[a-zA-Z_][a-zA-Z0-9_]*$`.
|
||||
- `kind` (required): a [TypeRef](#typeref) — primitive string, `$ref`
|
||||
object, array object, or record object.
|
||||
- `endian` (optional): overrides the struct/union default for this field.
|
||||
- `align` (optional): field-level alignment (aligned mode only).
|
||||
- `encoding` (optional): `"length-prefixed"` (default) or
|
||||
`"offset-indirect"`. See [Variable-length encoding](#variable-length-encoding).
|
||||
- `maxLength` (optional, `string`/`bytes` fields only): byte-length
|
||||
cap. See [Variable-length encoding](#variable-length-encoding).
|
||||
Rejected at parse on any other kind (review #006 N3: the annotation
|
||||
was silently unenforced there — the validation plan bakes `maxLength`
|
||||
into string/bytes leaves only).
|
||||
|
||||
### TypeRef
|
||||
|
||||
`TypeRef` is the central mechanism for referencing types. Seven forms:
|
||||
|
||||
| Form | Example | Meaning |
|
||||
|------|---------|---------|
|
||||
| Primitive string | `"uint32"` | A built-in primitive (see [Primitives](#primitives)) |
|
||||
| `$ref` object | `{ "$ref": "#/$defs/Read" }` | Reference to a named `$defs` entry |
|
||||
| Array object | `{ "kind": "array", "element": "uint32", "count": 3 }` | Fixed-size array |
|
||||
| Record object | `{ "kind": "record", "values": "string" }` | String-keyed map |
|
||||
| Inline struct | `{ "kind": "struct", "fields": [...] }` | Anonymous struct |
|
||||
| Inline union | `{ "kind": "union", ... }` | Anonymous union |
|
||||
| Inline enum | `{ "kind": "enum", "values": [...] }` | Anonymous enum |
|
||||
|
||||
The `$ref` form uses standard JSON Pointer syntax **restricted to
|
||||
`#/$defs/<name>`** — no external references, no fragment-only pointers,
|
||||
no bare names. The restriction keeps resolution a single hash lookup
|
||||
and eliminates the `normalize_refs` step the v0.1.0 engine needed for
|
||||
TypeBox's bare-name refs.
|
||||
|
||||
Arrays and records are inline type constructors, not top-level `$defs`
|
||||
entries. Complex element types use nested `$ref`:
|
||||
|
||||
```json
|
||||
{ "kind": "array", "element": { "$ref": "#/$defs/ComplexElement" }, "count": 4 }
|
||||
```
|
||||
|
||||
Inline `struct`/`union`/`enum` TypeRefs are anonymous composites — a
|
||||
field, array element, record value, or union variant whose type is
|
||||
declared inline rather than named in `$defs`. They are structurally
|
||||
identical to their named counterparts (same `StructDef`/`UnionDef`/
|
||||
`EnumDef` shape); only the reference mechanism differs. Named composites
|
||||
are preferred for reuse and for `$ref`-based dispatch; inline composites
|
||||
are convenient for one-off nested types.
|
||||
|
||||
### Primitives
|
||||
|
||||
The 13 primitive `kind` strings map to the unchanged `AlkTypeKind`
|
||||
variants (D-BAST-002 — lowercase strings, PascalCase enum variants):
|
||||
|
||||
| BAST kind | `AlkTypeKind` | Rust type | Size | Category |
|
||||
|-----------|---------------|-----------|------|----------|
|
||||
| `int8` | `Int8` | `i8` | 1 | fixed |
|
||||
| `int16` | `Int16` | `i16` | 2 | fixed |
|
||||
| `int32` | `Int32` | `i32` | 4 | fixed |
|
||||
| `int64` | `Int64` | `i64` | 8 | fixed |
|
||||
| `uint8` | `Uint8` | `u8` | 1 | fixed |
|
||||
| `uint16` | `Uint16` | `u16` | 2 | fixed |
|
||||
| `uint32` | `Uint32` | `u32` | 4 | fixed |
|
||||
| `uint64` | `Uint64` | `u64` | 8 | fixed |
|
||||
| `float32` | `Float32` | `f32` | 4 | fixed |
|
||||
| `float64` | `Float64` | `f64` | 8 | fixed |
|
||||
| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed |
|
||||
| `string` | `String` | length-prefixed UTF-8 | variable | variable |
|
||||
| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable |
|
||||
|
||||
`int64`/`uint64` are alktype additions (not in TypeBox's `typedef.ts`),
|
||||
required by SFTP `offset: u64` and metatensor `data_offsets`. JSON
|
||||
precision caveat per ADR-005 applies: integers beyond `2^53` lose
|
||||
precision in `serde_json::Value::Number`; the binary path is exact.
|
||||
|
||||
### Enum
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "enum",
|
||||
"values": ["Ok", "PermissionDenied", "NoSuchFile", "Failure"]
|
||||
}
|
||||
```
|
||||
|
||||
- `kind` (required): `"enum"`.
|
||||
- `values` (required): non-empty array of strings, in declaration order.
|
||||
- Binary representation: a `u32` index into `values` (0-based), encoded
|
||||
per the struct's endianness. This is a deliberate deviation from
|
||||
TypeBox's string enum in favor of binary efficiency — a `u32` index is
|
||||
compact, fixed-size, and sufficient for any realistic enum.
|
||||
|
||||
**Bug fix vs v0.1.0:** The v0.1.0 engine has a dead constraint on the
|
||||
bytes path — the built-in `enum` keyword checks string membership, but
|
||||
the materializer emits `Value::Number(index)`, which never matches. The
|
||||
BAST-native validator (see [Validation Model](#validation-model)) checks
|
||||
the materialized index against `values.len()` bounds, fixing this.
|
||||
|
||||
### Union
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "union",
|
||||
"endian": "big",
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"1": { "$ref": "#/$defs/Init" },
|
||||
"3": { "$ref": "#/$defs/Open" },
|
||||
"5": { "$ref": "#/$defs/Read" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `kind` (required): `"union"`.
|
||||
- `endian` (optional): default endianness for variant fields.
|
||||
- `discriminator` (required): one of:
|
||||
- **Byte-offset**: `{ "kind": "byte", "offset": <N>, "type": "uint8"|"uint16"|"uint32" }`.
|
||||
The discriminator byte is at `offset`; the variant struct starts at
|
||||
`offset + discriminator_size`. Mapping keys are stringified integers.
|
||||
- **Field-name**: `{ "kind": "field", "name": "<field>" }`. The
|
||||
discriminator is a length-prefixed string field; mapping keys are
|
||||
string values matching the field's value. The union's `fields` array
|
||||
(optional, only valid with field-name discriminators per D-BAST-005)
|
||||
provides the field definitions including the discriminator field.
|
||||
- `mapping` (required): object mapping discriminator values to
|
||||
[TypeRef](#typeref) entries (typically `$ref` to `$defs` variants).
|
||||
|
||||
Variant `$ref`s are resolved **lazily** by the materializer and
|
||||
validator — no `inline_union_variant_refs` compile step (removed under
|
||||
BAST). The validator recurses into the selected variant's BAST
|
||||
definition on `__discriminator` lookup, recovering the OQ-008
|
||||
per-variant constraint enforcement (e.g., `maxLength` on a `bytes`
|
||||
field inside a variant) without custom keywords.
|
||||
|
||||
### Array
|
||||
|
||||
```json
|
||||
{ "kind": "array", "element": "float32", "count": 3 }
|
||||
```
|
||||
|
||||
- `kind` (required): `"array"`.
|
||||
- `element` (required): a [TypeRef](#typeref).
|
||||
- `count` (required in v1): the fixed element count. **Arrays of
|
||||
variable-length elements without a `count` are not supported in v1**
|
||||
(D-BAST-004, aligning with OQ-001). The meta-schema enforces this:
|
||||
`count` is in `required`. Variable-length collections are available
|
||||
via `record` instead.
|
||||
|
||||
For fixed-size elements with a known count, the array size is
|
||||
`element_size × count`. For variable-length elements (e.g.,
|
||||
`"element": "string"`) with a known count, each element carries its own
|
||||
length prefix — the array is count-prefixed in the sense that the count
|
||||
is known at schema time, but the total byte size is not.
|
||||
|
||||
### Record
|
||||
|
||||
```json
|
||||
{ "kind": "record", "values": "string" }
|
||||
```
|
||||
|
||||
- `kind` (required): `"record"`.
|
||||
- `values` (required): a [TypeRef](#typeref) for the value type.
|
||||
|
||||
Binary layout: a count-prefixed sequence of `(key, value)` pairs —
|
||||
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
|
||||
Each key is a length-prefixed UTF-8 string. Each value is encoded per
|
||||
its `values` type. There is no separate `value_len` prefix — the value's
|
||||
size is determined by its kind (fixed-size kinds have a known size;
|
||||
variable-length kinds carry their own length prefix). The count and
|
||||
key-length prefixes respect the struct's endianness.
|
||||
|
||||
## Variable-Length Encoding
|
||||
|
||||
The three strategies from ADR-003 carry forward with the same semantics,
|
||||
expressed as field-level properties instead of keyword-value objects:
|
||||
|
||||
| Strategy | BAST syntax | Behavior |
|
||||
|----------|------------|----------|
|
||||
| Inline length-prefixed (default) | `{ "name": "handle", "kind": "string" }` | `[u32 length][data]` |
|
||||
| Fixed-size reservation | `{ "name": "name", "kind": "string", "maxLength": 256 }` | Reserve `maxLength` bytes (aligned mode); validation constraint (packed mode) |
|
||||
| Offset indirection | `{ "name": "blob", "kind": "bytes", "encoding": "offset-indirect" }` | `{offset: u32, length: u32}` pointing to separate data region |
|
||||
|
||||
**Default strategy selection (unchanged from v0.1.0):**
|
||||
- **Packed sequential mode:** always inline length-prefixing.
|
||||
`maxLength` is a validation constraint only.
|
||||
- **Aligned static mode:** fixed-size reservation if `maxLength` is
|
||||
declared; offset indirection if `"encoding": "offset-indirect"` is
|
||||
declared; inline length-prefixing otherwise.
|
||||
|
||||
**Length prefix endianness:** The 4-byte length prefix (strategies 1
|
||||
and 3) respects the effective endianness (struct default or field
|
||||
override). In little-endian mode, `u32::from_le_bytes`; in big-endian
|
||||
mode, `u32::from_be_bytes`. Ensures SFTP consumers (big-endian) have
|
||||
consistent byte order for field values and length prefixes.
|
||||
|
||||
Applies to variable-length primitive types only: `string` and
|
||||
`bytes`. The parser rejects `maxLength` (and the meta-schema forbids
|
||||
it) on every other kind — including `record` (review #006 N3/M5: no
|
||||
consumer honored it there, so the annotation was either silently
|
||||
unenforced or, in aligned mode, silently corrupt).
|
||||
|
||||
## Endianness
|
||||
|
||||
Struct-level or union-level property with per-field override (same
|
||||
semantics as ADR-003):
|
||||
|
||||
- Struct/union-level `"endian"` sets the default for all fields.
|
||||
- Field-level `"endian"` overrides the struct/union default.
|
||||
- Default is `"little"` when neither is specified.
|
||||
- The length prefix for variable-length fields respects the effective
|
||||
endianness.
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [
|
||||
{ "name": "id", "kind": "uint32" },
|
||||
{ "name": "handle", "kind": "string" },
|
||||
{ "name": "crc", "kind": "uint32", "endian": "little" }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Alignment
|
||||
|
||||
Struct-level or field-level property, only meaningful in aligned static
|
||||
mode (same as ADR-003):
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "struct",
|
||||
"align": 256,
|
||||
"fields": [
|
||||
{ "name": "header", "kind": { "$ref": "#/$defs/Header" } },
|
||||
{ "name": "weight", "kind": "float32", "align": 16 }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
- Struct-level `"align"` sets the default for all fields.
|
||||
- Field-level `"align"` overrides the struct default.
|
||||
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/
|
||||
f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length
|
||||
prefix), 1 for struct/union/array. Unchanged from v0.1.0.
|
||||
- Ignored in packed sequential mode.
|
||||
|
||||
## Validation Model
|
||||
|
||||
BAST separates two concerns that the v0.1.0 format conflates, and in
|
||||
doing so reveals that the engine has **two distinct validation paths**
|
||||
with different inputs and guarantees. This is the validator split,
|
||||
decided in D-BAST-006, D-BAST-007, and D-BAST-009. See
|
||||
[`validation.md`](validation.md) for the current (pre-pivot) validation
|
||||
layer; this section specifies the target model.
|
||||
|
||||
### Two validators, two inputs
|
||||
|
||||
| Path | Input | Validator | Schema source |
|
||||
|------|-------|-----------|---------------|
|
||||
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator | The BAST document (binary layout + value constraints) |
|
||||
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
|
||||
|
||||
**`validate_bytes` — bytes in, BAST is the validator.** The materializer
|
||||
produces a `Value` tree from bytes. By construction, this `Value` is
|
||||
*structurally correct*: all declared fields are present (the
|
||||
materializer iterates the field list), types are correct (`read_u32`
|
||||
produces `Value::Number`), bounds are checked (via `data_access::
|
||||
check_bounds`), UTF-8 is valid (via `from_utf8`), the discriminator is
|
||||
in the mapping, and the boolean byte is 0 or 1. What the materializer
|
||||
does NOT check — and what the validation half checks afterward — are
|
||||
**value-domain constraints expressed in the BAST document**. Under
|
||||
ADR-012 §3 these constraints are compiled once into a `ValidationPlan`
|
||||
at engine-compile time (eager `$ref` resolution, cyclic-graph
|
||||
rejection); each `validate_bytes` call walks the compiled constraint
|
||||
tree against the `Value`. The plan's nodes enforce exactly these
|
||||
constraints (the set is normative; the walker that enforced it
|
||||
interpretively in 0.2.0 is retired):
|
||||
|
||||
| Constraint | `ValidNode` arm |
|
||||
|------------|-----------------|
|
||||
| Integer range (Int8..Uint32) | `Int { min, max }` / `Uint { max }` with `as_i64`/`as_u64` + range check |
|
||||
| Int64/Uint64 (full range) | `I64` / `U64` (JSON precision caveat per ADR-005) |
|
||||
| Float finiteness (Float32/64) | `Float` with `as_f64().is_finite()` |
|
||||
| String `maxLength` (byte length) | `Str { max_len }` — `maxLength` baked in from the owning field at compile time |
|
||||
| Bytes `maxLength` (array length) | `Bytes { max_len }` — accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
|
||||
| Enum index bounds | `Enum { count }` checks `idx < count` — **fixes the v0.1.0 dead constraint** |
|
||||
| Union variant dispatch | `Union { variants }` reads `__discriminator`, dispatches on the compiled variant nodes |
|
||||
| Struct fields | `Struct { fields }` requires each declared field present, recurses |
|
||||
| Array count | `Array { count, element }` checks `arr.len() == count` and recurses per element |
|
||||
| Record values | `Record { values }` recurses into each value |
|
||||
| Boolean | `Bool` (materializer already rejects non-0/1 bytes) |
|
||||
|
||||
No external JSON Schema is required for `validate_bytes`. The BAST
|
||||
document is the complete specification of the binary format — it
|
||||
describes both the layout (how to read) and the constraints (what
|
||||
values are valid). This is the "schema is the format" principle from
|
||||
ADR-001, now fully realized.
|
||||
|
||||
An optional external JSON Schema can be layered on top for constraints
|
||||
BAST doesn't express (cross-field consistency, regex patterns on string
|
||||
content). This is additive, not load-bearing.
|
||||
|
||||
**`validate_json` — JSON in, JSON Schema is the validator.** The
|
||||
consumer provides a JSON `Value` (e.g., an incoming JSON-RPC request).
|
||||
The BAST document is irrelevant — BAST describes bytes, not JSON shape.
|
||||
The right validator for a JSON value is a standard
|
||||
`jsonschema::Validator` built from a standard JSON Schema document the
|
||||
consumer provides. This is the path alkcall uses for its `OperationSpec`
|
||||
JSON validation. No custom keywords; BAST is not involved.
|
||||
|
||||
### What is removed
|
||||
|
||||
Under the BAST pivot, the v0.1.0 validation machinery is removed from
|
||||
the `validate_bytes` path:
|
||||
|
||||
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
|
||||
factories) — replaced by the BAST-native validator (~250 lines, a
|
||||
flat match with no factories, no trait objects, no sub-validator
|
||||
pre-computation).
|
||||
- `inline_union_variant_refs()` — union variant refs are resolved lazily
|
||||
by the validator and materializer.
|
||||
- `build_validator()`'s custom-keyword path — repurposed or removed (see
|
||||
the implementation plan's step 6 for the decision on its fate).
|
||||
|
||||
The `jsonschema` crate **remains a direct dependency** for
|
||||
`validate_json` and for validating BAST documents against the BAST
|
||||
meta-schema. The only thing removed is the custom keyword integration
|
||||
path. The `validate_bytes` path no longer touches `jsonschema` — a
|
||||
small wasm binary-size win in addition to the architecture
|
||||
simplification. (Since ADR-012 §3, the interpretation step itself is
|
||||
also compiled away: see the `ValidationPlan` above.)
|
||||
|
||||
### `AlkTypeError::Validation` payload shape
|
||||
|
||||
**Decided (D-BAST-009):** Keep
|
||||
`Validation(jsonschema::ValidationError<'static>)`.
|
||||
|
||||
The `validate_bytes` path no longer uses `jsonschema`, so its error
|
||||
payload is constructed via `jsonschema::ValidationError::custom` purely
|
||||
to keep the variant's type unchanged. The rationale is consumer
|
||||
ergonomics on the *combined* path: consumers like alkcall use both
|
||||
`validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary
|
||||
channels) and handle `AlkTypeError::Validation` in one place. A single
|
||||
uniform payload type means one match arm covers both sources. The
|
||||
alternative (`Validation(String)`) would force `validate_json` to
|
||||
flatten its structured errors (instance path, schema path, keyword) to
|
||||
a `String` via `Display` — the more information-rich path loses data to
|
||||
accommodate the less rich one. That is the wrong direction.
|
||||
|
||||
The `no_std`/minimal-build angle (OQ-002) that the alternative was
|
||||
meant to enable is moot: `validate_json` requires `jsonschema`
|
||||
regardless, so a bytes-only `no_std` build already has to give up
|
||||
`validate_json` as a separate, larger decision. The right place to
|
||||
revisit is when/if OQ-002 is actually pursued.
|
||||
|
||||
## Relationship to JSON Schema and TypeBox
|
||||
|
||||
### BAST is a JSON Schema dialect
|
||||
|
||||
BAST is a specific JSON Schema instance format — like how JSON Schema
|
||||
itself is a JSON document conforming to the JSON Schema meta-schema.
|
||||
BAST documents conform to the BAST meta-schema. The entire JSON Schema
|
||||
tooling ecosystem works with BAST:
|
||||
|
||||
- **Validation:** `jsonschema::options().build(&bast_meta_schema)?.validate(&bast_doc)`
|
||||
- **Editors:** VSCode with `$schema` pointing to the BAST meta-schema URL
|
||||
- **Documentation:** JSON Schema generators produce human-readable docs
|
||||
from the meta-schema
|
||||
|
||||
### TypeBox interop
|
||||
|
||||
TypeBox's `Type.Module({...})` pattern maps naturally to BAST's `$defs`
|
||||
structure. A TypeBox module defining binary types can serialize to BAST
|
||||
JSON. The relationship:
|
||||
|
||||
- TypeBox → BAST JSON → alktype engine (binary layout)
|
||||
- TypeBox → standard JSON Schema → jsonschema (JSON validation)
|
||||
|
||||
Same TypeBox source, two output formats, two validators.
|
||||
|
||||
### Not a replacement for JSON Schema
|
||||
|
||||
BAST does not replace JSON Schema for JSON data validation. A BAST
|
||||
document cannot validate a JSON payload — it describes binary data
|
||||
layouts and value-domain constraints for bytes. For JSON validation,
|
||||
consumers use standard JSON Schema documents (which may be derived from
|
||||
BAST via future codegen, or authored separately). The `validate_json`
|
||||
path accepts a consumer-provided JSON Schema and uses a standard
|
||||
`jsonschema::Validator` — BAST is not involved.
|
||||
|
||||
This is the split: `validate_bytes` is BAST-native (the BAST document
|
||||
is both the layout spec and the validation spec for bytes);
|
||||
`validate_json` is JSON-Schema-native (a standard JSON Schema is the
|
||||
validation spec for JSON values). One crate, two validators, two input
|
||||
types.
|
||||
|
||||
## Decisions
|
||||
|
||||
The BAST format is grounded in decisions D-BAST-001 through D-BAST-009,
|
||||
recorded in [the pivot research record](../research/bast-pivot.md#decisions).
|
||||
The implementation-relevant summary:
|
||||
|
||||
| Decision | Summary |
|
||||
|----------|---------|
|
||||
| [D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention |
|
||||
| [D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase |
|
||||
| [D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape |
|
||||
| [D-BAST-004](../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) |
|
||||
| [D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` |
|
||||
| [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed |
|
||||
| [D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema |
|
||||
| [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema |
|
||||
| [D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths |
|
||||
|
||||
## References
|
||||
|
||||
- [Pivot research record](../research/bast-pivot.md) — motivation, POC
|
||||
scope and result, decisions D-BAST-001..009, risks
|
||||
- [Implementation plan](../plans/bast-implementation.md) — ordered
|
||||
steps, public-API semver contract, ADR-sync checklist
|
||||
- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics
|
||||
(carry forward unchanged; only location moves)
|
||||
- [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) — Int64/
|
||||
Uint64 as first-class kinds; JSON precision caveat
|
||||
- [`schema-layer.md`](schema-layer.md) — the current (v0.1.0) schema
|
||||
layer; superseded by this document when the pivot lands
|
||||
- [`validation.md`](validation.md) — the current (v0.1.0) validation
|
||||
layer; rewritten for the validator split when the pivot lands
|
||||
+240
-169
@@ -1,29 +1,41 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-08-11
|
||||
status: accepted
|
||||
last_updated: 2026-08-15
|
||||
---
|
||||
|
||||
# alktype — Builder API
|
||||
|
||||
The builder layer: a fluent Rust API for constructing alktype JSON
|
||||
Schemas (both `AlkType:*`-bearing binary-layout schemas and plain
|
||||
JSON-Schema-only operation payload schemas) at runtime, producing
|
||||
`serde_json::Value`. Decided in [ADR-009](decisions/009-builder-api.md);
|
||||
resolves [OQ-003](questions/003-builder-api-for-schema-construction.md).
|
||||
The builder layer: a fluent Rust API for constructing BAST documents
|
||||
(binary-layout schemas) and standard JSON Schemas (JSON-validation
|
||||
schemas) at runtime, producing `serde_json::Value`. Decided in
|
||||
[ADR-009](decisions/009-builder-api.md); resolves
|
||||
[OQ-003](questions/003-builder-api-for-schema-construction.md). The
|
||||
two-output-format split is D-BAST-008, recorded in
|
||||
[ADR-BAST](decisions/bast-bast-format.md).
|
||||
|
||||
## What
|
||||
|
||||
The `builder` module provides a single `Schema` builder type and a
|
||||
`Definitions` helper for named `$defs`. The builder's `.build()` method
|
||||
returns a `serde_json::Value` — the same form alktype already consumes
|
||||
via `AlkTypeEngine::compile` (for `AlkType:*` schemas) and the same form
|
||||
`OperationSpec.input_schema` / `output_schema` / `error_schemas` hold
|
||||
(for plain JSON Schema, no `AlkType:*` kinds).
|
||||
returns a `serde_json::Value` — one of two forms depending on the
|
||||
constructor used (D-BAST-008):
|
||||
|
||||
- **BAST JSON** (binary layout) — `Schema::struct_().field(...).build()`
|
||||
produces a BAST TypeDef (`{ "kind": "struct", "fields": [...] }`).
|
||||
Primitive constructors produce bare TypeRef strings (`"uint32"`).
|
||||
Feed to [`AlkTypeEngine::compile`](validation.md) (the binary-layout
|
||||
path) → `validate_bytes`.
|
||||
- **Standard JSON Schema** (JSON validation) — `Schema::object().field(...)`
|
||||
produces `{ "type": "object", "properties": {...}, "required": [...] }`.
|
||||
No BAST `kind`, no custom keywords — a plain JSON Schema. Feed to a
|
||||
standard `jsonschema::Validator` (or `AlkTypeEngine::compile` with a
|
||||
JSON Schema for the `validate_json` path, D-BAST-007).
|
||||
|
||||
The builder covers:
|
||||
|
||||
- All 19 `AlkType:*` kinds (binary-layout schemas) — see
|
||||
[schema-layer.md](schema-layer.md) for the kinds.
|
||||
- All 18 BAST kinds (binary-layout schemas) — see
|
||||
[schema-layer.md](schema-layer.md) for the kinds and
|
||||
[`bast-format.md`](bast-format.md) for the format.
|
||||
- All standard JSON Schema keywords needed for operation payload
|
||||
schemas: `type`, `properties`, `required`, `items`, `enum`, `format`,
|
||||
`additionalProperties`, `minimum`, `maximum`, `minItems`,
|
||||
@@ -39,16 +51,20 @@ alktype's first consumer and needs to build schemas at runtime from
|
||||
Rust code, for two roles:
|
||||
|
||||
1. **Binary layout schemas** (channels' 8-byte chunk header, future
|
||||
binary call frames) — `AlkType:*` schemas, fed to
|
||||
`AlkTypeEngine::compile` (packed mode, big-endian).
|
||||
binary call frames) — BAST documents, fed to
|
||||
`AlkTypeEngine::compile` (packed mode, big-endian) →
|
||||
`validate_bytes`.
|
||||
2. **JSON payload schemas** (call's `OperationSpec.input_schema` /
|
||||
`output_schema` / `error_schemas`) — plain JSON Schema, no
|
||||
`AlkType:*` kinds, validated via the standard `jsonschema` validator.
|
||||
`output_schema` / `error_schemas`) — plain JSON Schema, no BAST
|
||||
`kind`, validated via the standard `jsonschema` validator
|
||||
(`AlkTypeEngine::compile` with a JSON Schema → `validate_json`).
|
||||
|
||||
A single builder serving both roles means alkcall imports one module
|
||||
for schema construction. See [ADR-009](decisions/009-builder-api.md)
|
||||
for the decision rationale (why `Value` not a typed `Schema` enum, why
|
||||
both AlkType and standard JSON Schema in one builder).
|
||||
both BAST and standard JSON Schema in one builder) and
|
||||
[ADR-BAST](decisions/bast-bast-format.md) for the two-output-format
|
||||
decision (D-BAST-008).
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -62,35 +78,42 @@ duplicating the JSON form that `AlkTypeEngine::compile`,
|
||||
### Module placement
|
||||
|
||||
`src/builder.rs`, re-exported from the crate root. The builder is a
|
||||
peer of `schema.rs` (which parses schemas) and `engine.rs` (which
|
||||
peer of `bast.rs` (which parses BAST documents) and `engine.rs` (which
|
||||
compiles them). The builder constructs; it does not parse or compile.
|
||||
|
||||
```rust
|
||||
// src/lib.rs (additions)
|
||||
pub mod builder;
|
||||
pub use builder::{Schema, Definitions};
|
||||
pub use builder::{Schema, Definitions, Discriminator};
|
||||
```
|
||||
|
||||
### Field order is load-bearing
|
||||
### Field order is explicit
|
||||
|
||||
`serde_json` with `preserve_order` is already a dependency (ADR-001).
|
||||
The builder's `Value` output uses `serde_json::Map` (which preserves
|
||||
insertion order under `preserve_order`), so field declaration order in
|
||||
the builder is the field order in the binary layout. This is critical
|
||||
for packed mode (ADR-002) where field order determines offsets.
|
||||
BAST struct fields are an ordered array (BAST design principle #4 —
|
||||
see [`bast-format.md`](bast-format.md#design-principles)). The builder's
|
||||
`struct_()`/`union_()` accumulates fields in call order and emits them
|
||||
as the `fields` array on `.build()`. Field order in the builder is the
|
||||
field order in the binary layout. This is critical for packed mode
|
||||
(ADR-002) where field order determines offsets.
|
||||
|
||||
(`serde_json`'s `preserve_order` feature remains a dependency, but
|
||||
layout correctness no longer depends on it — the `fields` array makes
|
||||
order explicit. `preserve_order` is still load-bearing for the
|
||||
`mapping` object's iteration order and for `Definitions`' `$defs`
|
||||
block, which the parser walks in document order.)
|
||||
|
||||
## Public API
|
||||
|
||||
### `Schema` builder
|
||||
|
||||
`Schema` is the single entry point. Constructors for each AlkType kind
|
||||
and each standard JSON Schema type; setters for annotations and
|
||||
`Schema` is the single entry point. Constructors for each BAST kind and
|
||||
each standard JSON Schema type; setters for annotations and
|
||||
constraints; `.build()` produces `Value`.
|
||||
|
||||
#### AlkType kind constructors
|
||||
#### BAST kind constructors
|
||||
|
||||
One constructor per `AlkTypeKind` variant (see [schema-layer.md](schema-layer.md)
|
||||
§"The 19 AlkType Kinds"):
|
||||
§"The 18 BAST Kinds"):
|
||||
|
||||
```rust
|
||||
impl Schema {
|
||||
@@ -112,43 +135,44 @@ impl Schema {
|
||||
// Variable-length kinds
|
||||
pub fn string() -> Self;
|
||||
pub fn bytes() -> Self;
|
||||
pub fn timestamp() -> Self;
|
||||
// Composite kinds
|
||||
pub fn struct_() -> Self; // fields added via .field()
|
||||
pub fn union_(disc: Discriminator) -> Self; // variants via .mapping()
|
||||
pub fn array_of(element: Schema) -> Self;
|
||||
pub fn array_of(element: Schema) -> Self; // .count() required for valid BAST (D-BAST-004)
|
||||
pub fn record_of(value: Schema) -> Self;
|
||||
}
|
||||
```
|
||||
|
||||
Each constructor sets the corresponding `"AlkType:<Kind>": true` key.
|
||||
For example, `Schema::uint32()` produces `{"AlkType:Uint32": true}`.
|
||||
Primitive constructors produce the bare BAST TypeRef string on
|
||||
`.build()`. For example, `Schema::uint32().build()` produces `"uint32"`.
|
||||
Composite constructors produce the BAST object form.
|
||||
|
||||
**`enum_of`** sets both `"AlkType:Enum": true` and the standard
|
||||
`"enum"` keyword with the provided values (declaration order is the
|
||||
index order — see [schema-layer.md](schema-layer.md) §"TEnum binary
|
||||
representation"):
|
||||
**`enum_of`** produces a BAST enum TypeDef (`{ "kind": "enum", "values":
|
||||
[...] }`); declaration order is the index order — see
|
||||
[schema-layer.md](schema-layer.md) §"The 18 BAST Kinds"):
|
||||
|
||||
```rust
|
||||
Schema::enum_of(&["read", "write", "execute"])
|
||||
// -> { "AlkType:Enum": true, "enum": ["read", "write", "execute"] }
|
||||
Schema::enum_of(&["read", "write", "execute"]).build()
|
||||
// -> { "kind": "enum", "values": ["read", "write", "execute"] }
|
||||
```
|
||||
|
||||
**`array_of`** and **`record_of`** take the element/value schema as a
|
||||
nested `Schema`:
|
||||
nested `Schema`. `array_of` requires `.count(N)` for valid BAST
|
||||
(D-BAST-004 — arrays of variable-length elements without a count are
|
||||
deferred, aligning with OQ-001):
|
||||
|
||||
```rust
|
||||
Schema::array_of(Schema::uint32())
|
||||
// -> { "AlkType:Array": true, "items": { "AlkType:Uint32": true } }
|
||||
Schema::array_of(Schema::uint32()).count(3).build()
|
||||
// -> { "kind": "array", "element": "uint32", "count": 3 }
|
||||
|
||||
Schema::record_of(Schema::float32())
|
||||
// -> { "AlkType:Record": true, "values": { "AlkType:Float32": true } }
|
||||
Schema::record_of(Schema::float32()).build()
|
||||
// -> { "kind": "record", "values": "float32" }
|
||||
```
|
||||
|
||||
#### Standard JSON Schema type constructors
|
||||
|
||||
For plain JSON Schema (no `AlkType:*` kinds) — call's
|
||||
`input_schema` / `output_schema` / `error_schemas`:
|
||||
For plain JSON Schema (no BAST `kind`) — call's `input_schema` /
|
||||
`output_schema` / `error_schemas`:
|
||||
|
||||
```rust
|
||||
impl Schema {
|
||||
@@ -163,22 +187,24 @@ impl Schema {
|
||||
}
|
||||
```
|
||||
|
||||
The `_` suffix disambiguates standard JSON Schema types from AlkType
|
||||
kinds (`string` is the AlkType kind; `string_` is the standard JSON
|
||||
Schema type — the AlkType kind constructor sets `"AlkType:String":
|
||||
true`, the standard constructor sets `"type": "string"`). This is
|
||||
deliberate: the two are distinct schema forms and the builder makes
|
||||
the distinction visible at the call site.
|
||||
The `_` suffix disambiguates standard JSON Schema types from BAST kinds
|
||||
(`string` is the BAST primitive; `string_` is the standard JSON Schema
|
||||
type — `string()` would produce `"string"` as a BAST TypeRef,
|
||||
`string_()` produces `{ "type": "string" }` as a standard JSON Schema).
|
||||
This is deliberate: the two are distinct schema forms and the builder
|
||||
makes the distinction visible at the call site.
|
||||
|
||||
#### Annotation setters
|
||||
|
||||
Annotation setters mirror ADR-003. Each setter is named after the
|
||||
annotation it produces; calling the setter sets the corresponding JSON
|
||||
key. Setters return `Self` for chaining.
|
||||
Annotation setters mirror ADR-003 (semantics unchanged; location moved
|
||||
to BAST type-level properties under the pivot — see
|
||||
[ADR-BAST](decisions/bast-bast-format.md)). Each setter is named after
|
||||
the annotation it produces; calling the setter sets the corresponding
|
||||
JSON key. Setters return `Self` for chaining.
|
||||
|
||||
```rust
|
||||
impl Schema {
|
||||
/// Schema-level endianness (ADR-003 §1). Default little.
|
||||
/// Struct/union-level endianness (ADR-003 §1). Default little.
|
||||
pub fn endian(mut self, endian: Endian) -> Self;
|
||||
|
||||
/// Struct or field alignment (ADR-003 §2). Struct-level sets the
|
||||
@@ -193,12 +219,19 @@ impl Schema {
|
||||
/// a variable-length type, reserves this many bytes (strategy 2).
|
||||
/// In packed mode, validation constraint only.
|
||||
pub fn max_length(mut self, max: usize) -> Self;
|
||||
|
||||
/// Array count (D-BAST-004 — required for valid BAST arrays in v1).
|
||||
pub fn count(mut self, count: usize) -> Self;
|
||||
}
|
||||
```
|
||||
|
||||
`Endian` and `VariableEncoding` are re-exported from `schema.rs` (no
|
||||
new types — the builder uses the existing enums). The setters produce
|
||||
the exact JSON shapes from ADR-003:
|
||||
new types — the builder uses the existing enums). When applied to a
|
||||
struct, `endian`/`align` are struct-level; when the `Schema` is used as
|
||||
a `.field()` argument, the builder extracts `endian`/`align`/`encoding`/
|
||||
`maxLength` and places them on the *field* object (BAST field-level
|
||||
properties). The setters produce the exact BAST JSON shapes from
|
||||
[`bast-format.md`](bast-format.md):
|
||||
|
||||
```rust
|
||||
Schema::struct_()
|
||||
@@ -207,31 +240,34 @@ Schema::struct_()
|
||||
.field("length", Schema::uint32())
|
||||
.build()
|
||||
// -> {
|
||||
// "AlkType:Struct": true,
|
||||
// "kind": "struct",
|
||||
// "endian": "big",
|
||||
// "properties": {
|
||||
// "channel_id": { "AlkType:Uint32": true },
|
||||
// "length": { "AlkType:Uint32": true }
|
||||
// }
|
||||
// "fields": [
|
||||
// { "name": "channel_id", "kind": "uint32" },
|
||||
// { "name": "length", "kind": "uint32" }
|
||||
// ]
|
||||
// }
|
||||
```
|
||||
|
||||
#### Composite builders
|
||||
|
||||
`struct_()`, `union_()`, `array_of()`, `record_of()` are the
|
||||
composite constructors. `struct_()` and `union_()` need additional
|
||||
setters to populate their children:
|
||||
`struct_()`, `union_()`, `array_of()`, `record_of()` are the composite
|
||||
constructors. `struct_()` and `union_()` need additional setters to
|
||||
populate their children:
|
||||
|
||||
```rust
|
||||
impl Schema {
|
||||
/// Add a field to a struct (or object). Field order is load-bearing
|
||||
/// for binary layouts (packed mode field order = byte order).
|
||||
/// Add a field to a struct (or a field-name-discriminator union).
|
||||
/// Field order is load-bearing for binary layouts (packed mode
|
||||
/// field order = byte order — the `fields` array is ordered).
|
||||
/// The field's schema is built from the passed `Schema`.
|
||||
pub fn field(mut self, name: &str, field: Schema) -> Self;
|
||||
|
||||
/// Mark fields as required (standard JSON Schema `required` keyword).
|
||||
/// Can be called multiple times; required names accumulate.
|
||||
/// Field names must have been added via `.field()`.
|
||||
/// Only meaningful for `object()` (standard JSON Schema) — BAST
|
||||
/// structs require all declared fields present (the validator
|
||||
/// enforces this). Can be called multiple times; required names
|
||||
/// accumulate.
|
||||
pub fn required(mut self, names: &[&str]) -> Self;
|
||||
|
||||
/// Set the items schema for a standard `array` type.
|
||||
@@ -247,9 +283,11 @@ impl Schema {
|
||||
}
|
||||
```
|
||||
|
||||
**`field`** sets `properties[name] = field.build()`. Repeated calls
|
||||
append. Field order in the built `Value` is the call order (because
|
||||
`serde_json::Map` preserves insertion order under `preserve_order`).
|
||||
**`field`** appends a `{ "name": ..., "kind": <field.build()>, ... }`
|
||||
entry to the struct/union's `fields` array, extracting field-level
|
||||
annotations (`endian`, `align`, `encoding`, `maxLength`) from the
|
||||
passed `Schema`. Repeated calls append in order. Field order in the
|
||||
built `Value` is the call order.
|
||||
|
||||
**`required`** sets the standard JSON Schema `"required"` array. The
|
||||
builder does not check that the named fields exist (that's a
|
||||
@@ -260,11 +298,12 @@ accumulates names:
|
||||
|
||||
```rust
|
||||
Schema::object()
|
||||
.field("path", Schema::string_())
|
||||
.field("path", Schema::string_().max_length(4096))
|
||||
.field("offset", Schema::integer().minimum(0))
|
||||
.field("length", Schema::integer().minimum(0))
|
||||
.required(["path"])
|
||||
.required(["offset", "length"])
|
||||
.build()
|
||||
// -> {
|
||||
// "type": "object",
|
||||
// "properties": { "path": {...}, "offset": {...}, "length": {...} },
|
||||
@@ -280,35 +319,31 @@ For operation payload schemas (call's `input_schema` etc.):
|
||||
impl Schema {
|
||||
/// `minimum` (inclusive lower bound for numbers/integers).
|
||||
pub fn minimum(mut self, min: f64) -> Self;
|
||||
|
||||
/// `maximum` (inclusive upper bound for numbers/integers).
|
||||
pub fn maximum(mut self, max: f64) -> Self;
|
||||
|
||||
/// `minLength` (minimum string length).
|
||||
pub fn min_length(mut self, min: usize) -> Self;
|
||||
|
||||
/// `minItems` (minimum array length).
|
||||
pub fn min_items(mut self, min: usize) -> Self;
|
||||
|
||||
/// `maxItems` (maximum array length).
|
||||
pub fn max_items(mut self, max: usize) -> Self;
|
||||
|
||||
/// `format` (e.g. "date-time", "uri", "email").
|
||||
pub fn format(mut self, fmt: &str) -> Self;
|
||||
|
||||
/// `title` (human-readable description).
|
||||
pub fn title(mut self, t: &str) -> Self;
|
||||
|
||||
/// `description` (human-readable description).
|
||||
pub fn description(mut self, d: &str) -> Self;
|
||||
}
|
||||
```
|
||||
|
||||
These set the corresponding standard JSON Schema keywords. They apply
|
||||
to both AlkType-kind schemas and standard JSON Schema type schemas
|
||||
(e.g., `Schema::string().max_length(4096)` sets `maxLength`, which
|
||||
serves as both a validation constraint and, in aligned mode, a
|
||||
fixed-size reservation — ADR-003 §3).
|
||||
to standard JSON Schema type schemas (e.g.,
|
||||
`Schema::string_().max_length(4096)` sets `maxLength`, which on the
|
||||
`validate_json` path is a JSON-Schema validation constraint). On a BAST
|
||||
schema, `max_length` also serves as the aligned-mode fixed-size
|
||||
reservation (ADR-003 §3) and the packed-mode validation constraint
|
||||
(enforced by the BAST-native validator — see
|
||||
[validation.md](validation.md)).
|
||||
|
||||
#### `.build()`
|
||||
|
||||
@@ -343,21 +378,22 @@ to compose them. `from_value` wraps the `Value` so it can be passed to
|
||||
### `Discriminator` for `union_()`
|
||||
|
||||
`union_()` takes a `Discriminator` describing the union's dispatch
|
||||
mechanism. This mirrors `schema.rs::DiscriminatorKind` but with a
|
||||
builder-friendly shape (the kind enum is re-exported from `schema.rs`,
|
||||
not duplicated):
|
||||
mechanism. This mirrors `bast::BastDiscriminator` (the parser's typed
|
||||
view) but with a builder-friendly shape:
|
||||
|
||||
```rust
|
||||
pub enum Discriminator {
|
||||
/// Byte-offset discriminator (ADR-003 §4 Kind A).
|
||||
/// `offset` is the byte position; `disc_type` is the AlkType kind
|
||||
/// `offset` is the byte position; `disc_type` is the BAST kind
|
||||
/// of the discriminator (Uint8/Uint16/Uint32).
|
||||
Byte {
|
||||
offset: usize,
|
||||
disc_type: AlkTypeKind, // restricted to Uint8/Uint16/Uint32
|
||||
},
|
||||
/// Field-name discriminator (ADR-003 §4 Kind B).
|
||||
/// `name` is the field holding the discriminator value.
|
||||
/// `name` is the field holding the discriminator value. The
|
||||
/// discriminator field and any shared fields are declared via
|
||||
/// `.field()` on the union builder.
|
||||
Field {
|
||||
name: String,
|
||||
},
|
||||
@@ -376,9 +412,13 @@ let packet = Schema::union_(Discriminator::Byte {
|
||||
.mapping("101", Schema::ref_def("Status"))
|
||||
.build();
|
||||
// -> {
|
||||
// "AlkType:Union": true,
|
||||
// "discriminator": { "kind": "byte", "offset": 0, "type": "AlkType:Uint8" },
|
||||
// "mapping": { "5": {"$ref":"#/$defs/Read"}, "6": {...}, "101": {...} }
|
||||
// "kind": "union",
|
||||
// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
// "mapping": {
|
||||
// "5": { "$ref": "#/$defs/Read" },
|
||||
// "6": { "$ref": "#/$defs/Write" },
|
||||
// "101": { "$ref": "#/$defs/Status" }
|
||||
// }
|
||||
// }
|
||||
```
|
||||
|
||||
@@ -386,22 +426,29 @@ let packet = Schema::union_(Discriminator::Byte {
|
||||
|
||||
```rust
|
||||
let event = Schema::union_(Discriminator::Field { name: "type" })
|
||||
.field("type", Schema::string())
|
||||
.mapping("read", Schema::ref_def("Read"))
|
||||
.mapping("write", Schema::ref_def("Write"))
|
||||
.build();
|
||||
// -> {
|
||||
// "AlkType:Union": true,
|
||||
// "kind": "union",
|
||||
// "discriminator": { "kind": "field", "name": "type" },
|
||||
// "fields": [ { "name": "type", "kind": "string" } ],
|
||||
// "mapping": { "read": {...}, "write": {...} }
|
||||
// }
|
||||
```
|
||||
|
||||
(Field-name-discriminator unions require a `fields` array declaring the
|
||||
discriminator field — D-BAST-005. The builder emits `fields` only when
|
||||
the discriminator is `Field` and at least one field was added.)
|
||||
|
||||
### `Definitions` — named `$defs` for cross-reference
|
||||
|
||||
`Definitions` is a helper for building named `$defs` that schemas can
|
||||
`$ref` by name. This is the ergonomics win for alkcall's
|
||||
`OperationSpec`, where input/output/error schemas reference shared
|
||||
definitions (e.g., `FileNotFound`, `RateLimited`).
|
||||
`$ref` by name, and for assembling a complete BAST document. This is
|
||||
the ergonomics win for alkcall's `OperationSpec`, where
|
||||
input/output/error schemas reference shared definitions (e.g.,
|
||||
`FileNotFound`, `RateLimited`).
|
||||
|
||||
```rust
|
||||
pub struct Definitions { /* ... */ }
|
||||
@@ -410,50 +457,67 @@ impl Definitions {
|
||||
pub fn new() -> Self;
|
||||
|
||||
/// Define a named schema. Returns a `Schema` that produces
|
||||
/// `{"$ref": "#/$defs/<name>"}` — the JSON Pointer form that
|
||||
/// `jsonschema` and `AlkTypeEngine::compile` expect (after
|
||||
/// `normalize_refs`, which the engine runs at compile time).
|
||||
/// `{"$ref": "#/$defs/<name>"}` — the JSON Pointer form BAST
|
||||
/// requires (no `normalize_refs` step; refs are always full
|
||||
/// pointers).
|
||||
pub fn define(&mut self, name: &str, schema: Schema) -> Schema;
|
||||
|
||||
/// Like `define`, but the schema is an existing `Value` (adopted
|
||||
/// via `Schema::from_value`).
|
||||
pub fn define_value(&mut self, name: &str, value: Value) -> Schema;
|
||||
|
||||
/// Produce the `{"$defs": { ... }}` object to merge into a
|
||||
/// top-level schema. Call once at the end.
|
||||
/// Produce the `{"$defs": { ... }}` object.
|
||||
pub fn build(self) -> Value;
|
||||
|
||||
/// Build a complete BAST document with `root_name` as the root
|
||||
/// type. The root schema is inserted into `$defs` alongside any
|
||||
/// previously defined entries. The resulting `Value` is ready for
|
||||
/// `AlkTypeEngine::compile(&doc, root_name, mode, ...)`.
|
||||
pub fn build_doc(self, root_name: &str, root: Schema) -> Value;
|
||||
|
||||
/// Merge the `$defs` into a top-level schema `Value`. If `top`
|
||||
/// already has a `$defs` object, the definitions are merged into
|
||||
/// it; otherwise a `$defs` key is inserted. For BAST documents,
|
||||
/// prefer `build_doc` — it places the root type inside `$defs`
|
||||
/// (where BAST requires it).
|
||||
pub fn merge_into(self, top: &mut Value);
|
||||
}
|
||||
```
|
||||
|
||||
**Usage:**
|
||||
**Usage (complete BAST document):**
|
||||
|
||||
```rust
|
||||
let mut defs = Definitions::new();
|
||||
|
||||
let file_not_found = defs.define("FileNotFound",
|
||||
Schema::object()
|
||||
.field("path", Schema::string_())
|
||||
.field("errno", Schema::integer())
|
||||
.required(["path", "errno"])
|
||||
);
|
||||
defs.define("Init", Schema::struct_().field("version", Schema::uint32()));
|
||||
defs.define("Read", Schema::struct_()
|
||||
.field("handle", Schema::bytes())
|
||||
.field("offset", Schema::uint64())
|
||||
.field("len", Schema::uint32()));
|
||||
|
||||
let rate_limited = defs.define("RateLimited",
|
||||
Schema::object()
|
||||
.field("retry_after_ms", Schema::integer().minimum(0))
|
||||
.required(["retry_after_ms"])
|
||||
);
|
||||
|
||||
let read_file_error = Schema::object()
|
||||
.field("code", Schema::string_())
|
||||
.field("details", Schema::any()) // one of the defined errors
|
||||
.required(["code"])
|
||||
.build();
|
||||
|
||||
// Merge $defs into the top-level schema that references them
|
||||
let mut top = Schema::object()
|
||||
.field("error", read_file_error)
|
||||
.build();
|
||||
top.as_object_mut().unwrap().insert("$defs".to_string(), defs.build());
|
||||
let doc = defs.build_doc("Packet", Schema::struct_()
|
||||
.field("payload", Schema::union_(Discriminator::Byte {
|
||||
offset: 0,
|
||||
disc_type: AlkTypeKind::Uint8,
|
||||
})
|
||||
.mapping("1", Schema::ref_def("Init"))
|
||||
.mapping("5", Schema::ref_def("Read"))));
|
||||
// -> {
|
||||
// "$defs": {
|
||||
// "Init": { "kind": "struct", "fields": [ { "name": "version", "kind": "uint32" } ] },
|
||||
// "Read": { "kind": "struct", "fields": [ ... ] },
|
||||
// "Packet": { "kind": "struct", "fields": [
|
||||
// { "name": "payload", "kind": {
|
||||
// "kind": "union",
|
||||
// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
// "mapping": { "1": { "$ref": "#/$defs/Init" }, "5": { "$ref": "#/$defs/Read" } }
|
||||
// } }
|
||||
// ] }
|
||||
// }
|
||||
// }
|
||||
//
|
||||
// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None)
|
||||
// then validate incoming frames via engine.validate_bytes(&frame).
|
||||
```
|
||||
|
||||
`define` returns a `Schema` (the `$ref` to the definition), so it can
|
||||
@@ -483,31 +547,36 @@ For cases where the `Definitions::define` return value isn't handy
|
||||
|
||||
## Usage Examples
|
||||
|
||||
### Example 1: channels' 8-byte chunk header (binary layout)
|
||||
### Example 1: channels' 8-byte chunk header (binary layout, BAST)
|
||||
|
||||
```rust
|
||||
use alktype::{Schema, Endian};
|
||||
use alktype::{Schema, Endian, Definitions};
|
||||
|
||||
let chunk_header = Schema::struct_()
|
||||
.endian(Endian::Big)
|
||||
.field("channel_id", Schema::uint32())
|
||||
.field("length", Schema::uint32())
|
||||
.build();
|
||||
.field("length", Schema::uint32());
|
||||
|
||||
// Build a complete BAST document (single-type — one $defs entry).
|
||||
let doc = Definitions::new().build_doc("ChunkHeader", chunk_header);
|
||||
// -> {
|
||||
// "AlkType:Struct": true,
|
||||
// "endian": "big",
|
||||
// "properties": {
|
||||
// "channel_id": { "AlkType:Uint32": true },
|
||||
// "length": { "AlkType:Uint32": true }
|
||||
// "$defs": {
|
||||
// "ChunkHeader": {
|
||||
// "kind": "struct",
|
||||
// "endian": "big",
|
||||
// "fields": [
|
||||
// { "name": "channel_id", "kind": "uint32" },
|
||||
// { "name": "length", "kind": "uint32" }
|
||||
// ]
|
||||
// }
|
||||
// }
|
||||
// }
|
||||
//
|
||||
// Feed to AlkTypeEngine::compile(&mut chunk_header, LayoutMode::Packed)
|
||||
// Feed to AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)
|
||||
// then validate incoming frames via engine.validate_bytes(&frame).
|
||||
```
|
||||
|
||||
### Example 2: call's `OperationSpec` input schema (JSON payload)
|
||||
### Example 2: call's `OperationSpec` input schema (JSON payload, standard JSON Schema)
|
||||
|
||||
```rust
|
||||
use alktype::Schema;
|
||||
@@ -530,18 +599,19 @@ let read_file_input = Schema::object()
|
||||
// }
|
||||
//
|
||||
// Stored in OperationSpec.input_schema; validated via the standard
|
||||
// jsonschema validator (validate_json for parsed payloads, or via
|
||||
// serde_json::from_slice then validate_json for wire frames).
|
||||
// jsonschema validator (AlkTypeEngine::compile with Some(&read_file_input)
|
||||
// for the validate_json path, or serde_json::from_slice then
|
||||
// validate_json for wire frames).
|
||||
```
|
||||
|
||||
### Example 3: SFTP `Packet` union (binary layout, byte discriminator)
|
||||
|
||||
The SFTP wire shape is `[type:u8][payload-struct]` — a struct with a
|
||||
union payload field. The engine requires `AlkType:Struct` at the top
|
||||
level (`OffsetMap::compute` / `SequentialReader::new` both enforce
|
||||
this; a `Union` is a field type within a struct, not a top-level
|
||||
schema). The builder constructs the union wrapped in a struct, and
|
||||
`$defs` are merged into the top-level schema so `$ref`s resolve:
|
||||
union payload field. The engine requires a struct at the root
|
||||
(`OffsetMap::compute` / `SequentialReader::new` both enforce this; a
|
||||
`Union` is a field type within a struct, not a top-level schema). The
|
||||
builder constructs the union wrapped in a struct, and `$defs` are
|
||||
placed inside the document via `build_doc` so `$ref`s resolve:
|
||||
|
||||
```rust
|
||||
use alktype::{Definitions, Discriminator, AlkTypeKind, Schema};
|
||||
@@ -555,27 +625,21 @@ defs.define("Status", Schema::struct_().field("code", Schema::uint32()).field("m
|
||||
|
||||
// A "Packet" is a struct with one field — the union. This mirrors
|
||||
// SFTP's wire shape: [type:u8][payload-struct].
|
||||
let mut packet = Schema::struct_()
|
||||
.field(
|
||||
"payload",
|
||||
Schema::union_(Discriminator::Byte {
|
||||
offset: 0,
|
||||
disc_type: AlkTypeKind::Uint8,
|
||||
})
|
||||
.mapping("1", Schema::ref_def("Init"))
|
||||
.mapping("3", Schema::ref_def("Open"))
|
||||
.mapping("5", Schema::ref_def("Read"))
|
||||
.mapping("6", Schema::ref_def("Write"))
|
||||
.mapping("101", Schema::ref_def("Status")),
|
||||
)
|
||||
.build();
|
||||
// Merge $defs into the top-level schema so $refs resolve at compile time.
|
||||
defs.merge_into(&mut packet);
|
||||
// Feed to AlkTypeEngine::compile(&mut packet, LayoutMode::Packed)
|
||||
let doc = defs.build_doc("Packet", Schema::struct_()
|
||||
.field("payload", Schema::union_(Discriminator::Byte {
|
||||
offset: 0,
|
||||
disc_type: AlkTypeKind::Uint8,
|
||||
})
|
||||
.mapping("1", Schema::ref_def("Init"))
|
||||
.mapping("3", Schema::ref_def("Open"))
|
||||
.mapping("5", Schema::ref_def("Read"))
|
||||
.mapping("6", Schema::ref_def("Write"))
|
||||
.mapping("101", Schema::ref_def("Status"))));
|
||||
// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None)
|
||||
// then validate incoming frames via engine.validate_bytes(&frame).
|
||||
```
|
||||
|
||||
### Example 4: OperationSpec error schemas (named `$defs`)
|
||||
### Example 4: OperationSpec error schemas (named `$defs`, standard JSON Schema)
|
||||
|
||||
```rust
|
||||
use alktype::{Definitions, Schema};
|
||||
@@ -610,15 +674,19 @@ let op_errors = vec![
|
||||
http_status: Some(429),
|
||||
},
|
||||
];
|
||||
// $defs is built once and stored alongside the OperationSpec
|
||||
// `$defs` is built once and stored alongside the OperationSpec.
|
||||
// (For the validate_json path, compile with Some(&defs.build()) as the
|
||||
// json_schema argument — but typically OperationSpec schemas are
|
||||
// validated directly via jsonschema, not via AlkTypeEngine.)
|
||||
```
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
|
||||
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation shapes the builder's setters produce |
|
||||
| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 |
|
||||
| BAST format + two output formats | [ADR-BAST](decisions/bast-bast-format.md) | `struct_()` → BAST, `object()` → standard JSON Schema (D-BAST-008) |
|
||||
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation semantics the builder's setters produce (location moved to BAST type-level properties) |
|
||||
| Load-time validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | The builder does not pre-validate; compile-time is the validation point |
|
||||
|
||||
## Open Questions
|
||||
@@ -634,12 +702,15 @@ and transitively on `Schema::union_`). See
|
||||
## References
|
||||
|
||||
- [ADR-009](decisions/009-builder-api.md) — the decision this spec implements
|
||||
- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format and the
|
||||
two-output-format decision (D-BAST-008)
|
||||
- [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) —
|
||||
scope boundaries this module extends; "schemas are JSON" principle
|
||||
- [ADR-003](decisions/003-schema-annotations.md) — the annotation
|
||||
shapes the builder's setters produce
|
||||
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds the
|
||||
builder's constructors produce
|
||||
semantics the builder's setters produce
|
||||
- [`bast-format.md`](bast-format.md) — the normative BAST format
|
||||
specification (the output format for `struct_()`)
|
||||
- [schema-layer.md](schema-layer.md) — the BAST kinds and parser
|
||||
- [validation.md](validation.md) — the validation layer that consumes
|
||||
builder output (via `AlkTypeEngine::compile`)
|
||||
- `@alkdev/alknet: docs/architecture/crates/call/operation-registry.md`
|
||||
|
||||
@@ -100,7 +100,7 @@ impl SequentialReader {
|
||||
```
|
||||
|
||||
`read_field`/`write_field` on `AlkTypeEngine` work for the fixed-size
|
||||
primitive kinds and the length-prefixed `String`/`Bytes`/`Timestamp`
|
||||
primitive kinds and the length-prefixed `String`/`Bytes`
|
||||
fields. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a
|
||||
`FieldValue` carrying a layout descriptor (byte range, variant start,
|
||||
or array stride) for the consumer to recurse on — see §"FieldValue" above.
|
||||
@@ -238,21 +238,23 @@ pub struct UnionDispatch {
|
||||
}
|
||||
```
|
||||
|
||||
After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)`
|
||||
to get the variant schema, then reads the variant's fields at
|
||||
After dispatch, the consumer calls `tunion::resolve_variant(union_node, &dispatch.key)`
|
||||
to get the variant `BastType`, then reads the variant's fields at
|
||||
`dispatch.variant_offset` using the normal `data_access` functions (or a
|
||||
fresh `SequentialReader` scoped to the variant).
|
||||
fresh `SequentialReader` scoped to the variant). `$ref` variant types are
|
||||
returned as `BastType::Ref`; the caller resolves them via
|
||||
`BastDoc::resolve_typeref` when a concrete definition is needed.
|
||||
|
||||
### Byte-offset discriminator
|
||||
|
||||
```rust
|
||||
/// Read the discriminator value from a byte-offset TUnion. The discriminator
|
||||
/// is a fixed-size integer (AlkType:Uint8/Uint16/Uint32) at a known byte
|
||||
/// offset. Returns the mapping key (stringified integer) and the variant
|
||||
/// struct offset.
|
||||
/// is a fixed-size integer (uint8/uint16/uint32) at a known byte offset.
|
||||
/// Returns the mapping key (stringified integer) and the variant struct
|
||||
/// offset.
|
||||
pub fn read_byte_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &Value,
|
||||
union_node: &BastUnion<'_>,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, AlkTypeError>;
|
||||
```
|
||||
@@ -268,11 +270,11 @@ starts at `offset + discriminator_size`.
|
||||
/// Read the discriminator value from a field-name TUnion. The
|
||||
/// discriminator is a named field within the struct — the consumer
|
||||
/// provides the field's computed offset (from the OffsetMap or
|
||||
/// LayoutBuilder). Supports AlkType:String, Uint8, and Enum discriminator
|
||||
/// LayoutBuilder). Supports string, uint8, and enum discriminator
|
||||
/// fields.
|
||||
pub fn read_field_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &Value,
|
||||
union_node: &BastUnion<'_>,
|
||||
disc_field_offset: usize,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, AlkTypeError>;
|
||||
@@ -287,16 +289,19 @@ the variant's fields starting at the end of the discriminator field.
|
||||
### Variant resolution
|
||||
|
||||
```rust
|
||||
/// Look up a variant schema from the union's mapping. Inline schemas
|
||||
/// are returned directly. $ref pointers of the form "#/$defs/<name>"
|
||||
/// are resolved against the union schema's own $defs block.
|
||||
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str)
|
||||
-> Result<&'a Value, AlkTypeError>;
|
||||
/// Look up a variant type from the union's mapping. Inline struct/
|
||||
/// union/enum types are returned directly; `$ref` pointers are
|
||||
/// returned as `BastType::Ref` — the caller resolves them via
|
||||
/// `BastDoc::resolve_typeref` when a concrete definition is needed.
|
||||
pub fn resolve_variant<'a>(
|
||||
union_node: &'a BastUnion<'a>,
|
||||
key: &str,
|
||||
) -> Result<&'a BastType<'a>, AlkTypeError>;
|
||||
|
||||
/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a
|
||||
/// Get the discriminator's byte size (1/2/4 for uint8/16/32) for a
|
||||
/// byte-offset TUnion. Field-name discriminators have no fixed size
|
||||
/// and produce a AlkTypeError::Schema.
|
||||
pub fn discriminator_size(union_schema: &Value) -> Result<usize, AlkTypeError>;
|
||||
/// and produce an `AlkTypeError::Schema`.
|
||||
pub fn discriminator_size(union_node: &BastUnion<'_>) -> Result<usize, AlkTypeError>;
|
||||
```
|
||||
|
||||
### TUnion in the layout engines
|
||||
@@ -324,7 +329,7 @@ dispatch to the primitive `data_access` function for the field's kind.
|
||||
|
||||
For aligned-mode access, `AlkTypeEngine::read_field(&buffer, "header.version")`
|
||||
returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds
|
||||
the field's `AlkType:*` kind in the schema, and calls the matching
|
||||
the field's `AlkTypeKind` in the BAST typed tree, and calls the matching
|
||||
`data_access::read_*` function. `write_field` is the mirror. Composite
|
||||
kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue`
|
||||
carrying a layout descriptor; the consumer recurses with a fresh reader
|
||||
|
||||
@@ -1,7 +1,19 @@
|
||||
# ADR-001: alktype — Purpose, Scope, and the jsonschema Engine
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
**Superseded (format-specific content) by
|
||||
[ADR-BAST](bast-bast-format.md).** The crate's purpose, scope
|
||||
boundaries, and the "schema is the format" principle are **retained and
|
||||
strengthened** — BAST *is* the format. Only the *concrete format*
|
||||
(custom-keyword JSON Schema → BAST) and the *validation strategy*
|
||||
(single `jsonschema` custom-keyword validator → two-validator model)
|
||||
are superseded: the format-specific content by ADR-BAST, the
|
||||
validation-strategy content by
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md). This ADR is kept as
|
||||
the historical record of the v0.1.0 design and the purpose/scope
|
||||
decision; read it alongside ADR-BAST and ADR-VAL-SPLIT for the current
|
||||
state.
|
||||
|
||||
## Context
|
||||
|
||||
|
||||
@@ -1,7 +1,14 @@
|
||||
# ADR-002: Two Layout Modes — Packed Sequential vs Aligned Static
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
Accepted — unchanged under the BAST pivot
|
||||
([ADR-BAST](bast-bast-format.md)). Layout modes are format-agnostic:
|
||||
the input format changed from custom-keyword JSON Schema to BAST, but
|
||||
the two modes, their alignment/packing rules, and the
|
||||
`LayoutBuilder`/`SequentialReader`/`OffsetMap` API did not. The layout
|
||||
engines now walk the BAST typed tree ([`BastDoc`](../schema-layer.md))
|
||||
instead of raw JSON with `get_alktype_kind*`, but the offset
|
||||
computation algorithm is identical.
|
||||
|
||||
## Context
|
||||
|
||||
|
||||
@@ -1,7 +1,18 @@
|
||||
# ADR-003: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
**Accepted (semantics); amended (location) by
|
||||
[ADR-BAST](bast-bast-format.md).** The annotation *semantics* decided
|
||||
here — endianness default, struct/field-level alignment, the three
|
||||
variable-length encoding strategies, and the two TUnion discriminator
|
||||
kinds — **carry forward unchanged** under the BAST pivot. Only the
|
||||
annotation *location* moves: from v0.1.0's custom-keyword objects
|
||||
(`{"AlkType:String": { "encoding": "..." }}`) to BAST type-level
|
||||
properties (`{ "name": "handle", "kind": "string", "encoding": "..." }`).
|
||||
The BAST shapes are normative in
|
||||
[`bast-format.md`](../bast-format.md#variable-length-encoding); this
|
||||
ADR is kept as the semantic reference. Read it alongside ADR-BAST.
|
||||
|
||||
## Context
|
||||
|
||||
@@ -125,9 +136,9 @@ reserving worst-case space.
|
||||
|
||||
- `true` is a shorthand for the default (length-prefixed). This keeps
|
||||
the common case concise and the override explicit.
|
||||
- The `encoding` annotation and `maxLength` apply to all variable-length
|
||||
types: `AlkType:String`, `AlkType:Bytes`, `AlkType:Array`,
|
||||
`AlkType:Record`, `AlkType:Timestamp`.
|
||||
- The `encoding` annotation and `maxLength` apply to the variable-length
|
||||
primitive types `AlkType:String` and `AlkType:Bytes`. (`maxLength` on
|
||||
records was amended out by review #006 N3/M5 — see §3a.)
|
||||
|
||||
### 3a. TRecord value type
|
||||
|
||||
@@ -153,8 +164,13 @@ the `"values"` property in the schema:
|
||||
the value's size is determined by its kind (fixed-size kinds have a
|
||||
known size; variable-length kinds carry their own length prefix).
|
||||
- The count and key-length prefixes respect the schema's endianness.
|
||||
- In aligned static mode with `maxLength`, the entire record is reserved
|
||||
at `maxLength` bytes (zero-padded).
|
||||
- ~~In aligned static mode with `maxLength`, the entire record is
|
||||
reserved at `maxLength` bytes (zero-padded).~~ **Amended (review #006
|
||||
N3/M5, 2026-09-02):** `maxLength` is rejected at parse on record
|
||||
fields. The aligned materializer walks the record's inline
|
||||
count-prefixed form and never honors the reservation (M5: silent
|
||||
cross-field corruption), and no packed consumer enforced it either
|
||||
(N3: silently unenforced). `maxLength` is `string`/`bytes`-only.
|
||||
|
||||
### 4. TUnion discriminators
|
||||
|
||||
|
||||
@@ -1,7 +1,21 @@
|
||||
# ADR-004: Error Handling and Validation Strategy
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
**Accepted (error type); amended (validation strategy) by
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The `AlkTypeError`
|
||||
enum, its four variants, the load-time-build / access-time-check split,
|
||||
and the field-path-carrying errors decided here are **retained
|
||||
unchanged** under the BAST pivot (D-BAST-009 keeps
|
||||
`Validation(jsonschema::ValidationError<'static>)`). The "validation
|
||||
strategy" section — which described v0.1.0's single
|
||||
`jsonschema`-custom-keyword validator for both paths — is **refined**:
|
||||
the bytes path now uses the BAST-native validator
|
||||
(`bast_validation`), the JSON path now uses a standard
|
||||
`jsonschema::Validator` from a consumer-provided JSON Schema. See
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the
|
||||
two-validator model. This ADR is kept as the error-handling reference;
|
||||
read it alongside ADR-VAL-SPLIT for the current validation strategy.
|
||||
|
||||
## Context
|
||||
|
||||
|
||||
@@ -63,11 +63,26 @@ write-side.
|
||||
|
||||
### Cost
|
||||
|
||||
`SequentialReader::new` clones the top-level struct's field schemas (a
|
||||
`Vec<(String, Value)>` of the `properties` entries) and clones the
|
||||
schema itself. This is cheap — a struct has a small number of fields
|
||||
(SFTP's largest packet has 5). The construction cost is negligible
|
||||
compared to the cost of reading a buffer.
|
||||
`SequentialReader::new(Arc<ReadPlan>)` is a refcount bump — 15.7 ns
|
||||
(measured, alktty `wire_vs_bast` bench, 0.3.0). The reader shares the
|
||||
engine's compiled [`ReadPlan`](011-compiled-read-plan-for-packed-mode.md)
|
||||
(the packed read-side compiled form) via `Arc` instead of cloning
|
||||
schema data; construction cost is negligible compared to reading a
|
||||
buffer. The engine holds the owned `BastDoc` (ADR-012 §2a) for the
|
||||
aligned materialize path and the one-shot `*::compile` paths.
|
||||
|
||||
> **Historical note**: the original 0.2.0 framing here ("re-parse on
|
||||
> demand" — the read loop re-parsing `BastDoc::new` per field) was the
|
||||
> root cause of the 400x read-path gap measured in
|
||||
> [review #004](../../reviews/004-performance-review.md).
|
||||
> [ADR-011](011-compiled-read-plan-for-packed-mode.md) (implemented,
|
||||
> 0.3.0) retired it: the packed read loop walks `Arc<ReadPlan>` (2.27
|
||||
> µs/chunk → 98 ns/chunk), `sequential_reader()` is an `Arc::clone`,
|
||||
> and the owned `BastDoc` (ADR-012 §2a) removed the remaining
|
||||
> per-access re-parse sites in `read_field`/`write_field`/`validate_bytes`.
|
||||
> The factory decision itself (`sequential_reader() ->
|
||||
> Option<SequentialReader>`, owned fresh reader, consumer-driven
|
||||
> cursor) was retained unchanged.
|
||||
|
||||
## Consequences
|
||||
|
||||
|
||||
@@ -2,7 +2,19 @@
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
**Accepted (API surface); amended (output format) by
|
||||
[ADR-BAST](bast-bast-format.md).** The fluent builder API, the
|
||||
`Schema`/`Definitions`/`Discriminator` types, the constructor and
|
||||
setter catalog, and the "produces `serde_json::Value`, not a typed
|
||||
`Schema` enum" decision decided here are **retained unchanged** under
|
||||
the BAST pivot. Only the `build()` *output format* changes:
|
||||
`struct_()` now produces BAST JSON (`{ "kind": "struct", "fields": [...] }`)
|
||||
instead of v0.1.0's custom-keyword JSON (`{ "AlkType:Struct": true,
|
||||
"properties": {...} }`); `object()` continues to produce standard JSON
|
||||
Schema. This is D-BAST-008, recorded in ADR-BAST. The builder examples
|
||||
in [`builder.md`](../builder.md) reflect the current BAST output. This
|
||||
ADR is kept as the API-surface decision; read it alongside ADR-BAST
|
||||
for the output format.
|
||||
|
||||
## Context
|
||||
|
||||
|
||||
@@ -2,7 +2,23 @@
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
**Accepted (two-step concept); amended (validation step) by
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The
|
||||
`validate_bytes(&[u8])` entry point, the "materialize `Value` from
|
||||
bytes, then validate" two-step concept, the mode dispatch, the
|
||||
field-path-carrying errors, and the "not a `Validator` trait / not
|
||||
framing-aware / not a binary-payload validator for JSON-only schemas"
|
||||
scope boundaries decided here are **retained unchanged** under the
|
||||
BAST pivot. Only the validation *step's implementation* changes: the
|
||||
materialized `Value` is validated by the **BAST-native validator**
|
||||
(`bast_validation`) instead of v0.1.0's `jsonschema` custom-keyword
|
||||
validator. The `jsonschema` crate is no longer touched on the bytes
|
||||
path (it remains for the `validate_json` path and for BAST meta-schema
|
||||
validation). The error payload type stays
|
||||
`Validation(jsonschema::ValidationError<'static>)` (D-BAST-009). See
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the
|
||||
two-validator model. This ADR is kept as the `validate_bytes` decision;
|
||||
read it alongside ADR-VAL-SPLIT for the current validation step.
|
||||
|
||||
## Context
|
||||
|
||||
|
||||
@@ -0,0 +1,584 @@
|
||||
# ADR-011: Compiled Read Plan for Packed Mode
|
||||
|
||||
## Status
|
||||
|
||||
Accepted. Implemented in 0.3.0 (phases 1–2, 2026-09-02). Closes review
|
||||
#004 H1 + M1 (packed side) + L1 + L2;
|
||||
retires the "re-parse on demand" framing from ADR-007. A derisking
|
||||
POC on branch `readplan-poc` confirmed the `ReadPlan` shape covers
|
||||
every `BastType` arm in the current read loop before implementation
|
||||
began (see "POC coverage" at the end).
|
||||
|
||||
**Refinements on ADR-012 acceptance (2026-08-20, review #005):** the
|
||||
`CompositePlan::Union` shape was refined to carry `shared:
|
||||
Option<Box<ReadPlan>>` (field-disc union shared fields, resolving POC
|
||||
Finding 1 / review #005 H1) and to drop `VariantPlan`/`VariantKind`
|
||||
in favor of `variants: Vec<(String, CompositePlan)>` (resolving review
|
||||
#005 M1 — nested unions now work by `CompositePlan` recursion,
|
||||
restoring the 0.2.0 capability the POC rejected). The "BastDoc
|
||||
unchanged" scope statement stands as the ADR-011-only view; ADR-012
|
||||
§2a subsequently makes `BastDoc` owned. The `ValidationPlan` this
|
||||
ADR's "Out of scope" originally deferred indefinitely is now in
|
||||
0.3.0 via ADR-012 §3 (review #005 M3 reversed the deferral). These
|
||||
are pre-implementation refinements to types that do not yet exist on
|
||||
`main`; the ADR-011 decision (a compiled `ReadPlan` for packed reads)
|
||||
is unchanged.
|
||||
|
||||
**Addendum — field-disc union wire convention (2026-09-02, review #006
|
||||
H3):** the packed-mode wire layout for a field-name-discriminator
|
||||
TUnion is **shared-then-variant**: the union's declared `fields` (the
|
||||
discriminator field + any shared fields) occupy the union's start
|
||||
offset in declaration order, and the selected variant's fields follow
|
||||
immediately after all shared fields. All three packed-mode consumers
|
||||
now implement this one convention: the reader and materializer already
|
||||
walked `shared` then the variant (the `shared` sub-plan shape above);
|
||||
`LayoutBuilder` was corrected in the same pass — it previously laid out
|
||||
only the selected variant, disagreeing with the read side on span and
|
||||
field positions (review #006 H3 item 1). The convention requires that
|
||||
a variant **must not re-declare** the discriminator field or any
|
||||
shared field — `BastUnion::parse` enforces this at parse time (also:
|
||||
the discriminator field must be declared in `fields`, and `fields`
|
||||
must not contain duplicate names), so the shared walk and the variant
|
||||
walk cover disjoint fields and the wire has exactly one copy of each
|
||||
shared byte. Schemas whose variants redeclared shared fields were
|
||||
ambiguous under the old split-convention behavior and are rejected
|
||||
rather than given a silent meaning; this is a **breaking wire-format
|
||||
constraint** for any 0.2.0-era schema that relied on re-declaration,
|
||||
announced with the 0.3.x series. `DiscriminatorPlan::Field`'s disc
|
||||
read is at the disc field's position within the shared walk (the
|
||||
materializer's position-correct behavior, review #006 H3 item 2); the
|
||||
reader's plan-walk reads it there too.
|
||||
|
||||
## Context
|
||||
|
||||
Review #004 (`docs/reviews/004-performance-review.md`) measured the
|
||||
packed read path at **~400x slower per chunk** than a hand-rolled codec
|
||||
(2.27 µs/chunk vs 5.6 ns/chunk), with the cost fixed across payload
|
||||
sizes — the signature of per-field interpretive overhead, not
|
||||
payload-copy overhead. The write path is competitive (~1.1x at 4 KiB)
|
||||
because it uses a compiled form; the read path is not because it
|
||||
doesn't.
|
||||
|
||||
### The three layout-side compiled forms and the one gap
|
||||
|
||||
ADR-002 defines two layout modes. Each mode has a write-side and a
|
||||
read-side. Three of the four slots already have a **compiled form** —
|
||||
a data structure built once from the schema, held by the engine, and
|
||||
walked at access time without re-touching the schema:
|
||||
|
||||
| mode | write-side | read-side |
|
||||
|------|-----------|-----------|
|
||||
| aligned static | `OffsetMap` (used for both) | `OffsetMap` |
|
||||
| packed sequential | `PackedLayout` (`LayoutBuilder::build`) | *(none)* |
|
||||
|
||||
- **`OffsetMap`** (aligned, both sides) — a flat table of
|
||||
`(field_path, ByteRange)` pairs computed once via
|
||||
`OffsetMap::compute(&BastDoc)`. Read and write both look up a field's
|
||||
byte range and call `data_access::read_*`/`write_*` at the known
|
||||
offset. No schema walk at access time.
|
||||
- **`PackedLayout`** (packed, write-side) — a flat table of
|
||||
`(field_path, FieldPosition)` pairs computed once via
|
||||
`LayoutBuilder::build(&var_sizes)`. The write loop calls
|
||||
`data_access::write_*` at the precomputed offsets. No schema walk at
|
||||
write time.
|
||||
- **packed read-side** — `SequentialReader` walks `BastDoc`
|
||||
interpretively on every field read. There is no compiled form.
|
||||
|
||||
This is the structural reason the read path is 400x slow: it is the
|
||||
only access path in the engine with no compiled form. Every other
|
||||
mode/side pair compiles the schema once and reuses the result.
|
||||
|
||||
### Root cause: the `BastDoc<'a>` borrow constraint
|
||||
|
||||
`BastDoc<'a>` borrows `&'a Value` and `&'a str` throughout
|
||||
(`src/bast.rs:51-55`). The owning structs that need a parsed tree at
|
||||
read time — `SequentialReader` (owns a cloned `Value`),
|
||||
`AlkTypeEngine` (owns `bast_doc: Value`), `LayoutBuilder` (owns
|
||||
`doc_value: Value`) — cannot store a `BastDoc` that borrows from their
|
||||
own `Value` field. That would be a self-referential struct, which safe
|
||||
Rust cannot express.
|
||||
|
||||
The workaround chosen in ADR-007 was "re-parse on demand": the engine
|
||||
and reader retain a clone of the raw `Value` and reconstruct the
|
||||
`BastDoc` from it whenever the typed tree is needed. ADR-007's "Cost"
|
||||
section argued this was cheap because construction is a small `Vec` of
|
||||
field schemas. That is true for *construction* (once), but the decision
|
||||
did not account for `read_field_at` re-parsing `BastDoc::new` **per
|
||||
field read** — the cost that actually dominates. For an N-field struct,
|
||||
reading all fields is O(N²) in parse work (each of N reads re-parses
|
||||
all N fields).
|
||||
|
||||
### Why `OffsetMap` is a flat table but the packed read plan cannot be
|
||||
|
||||
`OffsetMap` works as a flat `(path, byte_range)` lookup table because
|
||||
aligned positions are **data-independent** — field N's offset depends
|
||||
only on the schema, not on the bytes of fields 0..N-1. Random access
|
||||
by path is free.
|
||||
|
||||
Packed positions are **data-dependent** — a variable-length field's
|
||||
extent is read from its length prefix at access time, and every
|
||||
subsequent field's position shifts accordingly. You cannot look up
|
||||
field N's offset without reading fields 0..N-1 first. So the compiled
|
||||
form for packed reads cannot be a flat lookup table; it must be a
|
||||
**read program** — a pre-resolved tree of read instructions that the
|
||||
read loop walks in order, advancing a cursor. The schema is compiled
|
||||
into the program once; the bytes are walked against it at read time.
|
||||
|
||||
This asymmetry is inherent to packed sequential layout (ADR-002) and is
|
||||
not a flaw in `OffsetMap`. The two modes need different compiled-form
|
||||
shapes because they have different position-computation semantics.
|
||||
|
||||
## Decision
|
||||
|
||||
**Introduce `ReadPlan` — the compiled read-side form for packed mode,
|
||||
symmetric to `OffsetMap` (aligned read-side) and `PackedLayout` (packed
|
||||
write-side).**
|
||||
|
||||
`AlkTypeEngine::compile` builds the `ReadPlan` once from the `BastDoc`
|
||||
(in packed mode) and holds it for the life of the engine.
|
||||
`sequential_reader()` hands out fresh `SequentialReader`s that share
|
||||
the engine's `Arc<ReadPlan>` — the plan is immutable; only the cursor
|
||||
state (`field_index`, `position`) is per-reader. The read loop walks
|
||||
the plan, never touching `BastDoc` or the raw `Value`.
|
||||
|
||||
The same `ReadPlan` is consumed by `materialize_packed` (the other
|
||||
byte-walking path), unifying the two packed read-side consumers on one
|
||||
compiled form — mirroring how `OffsetMap` unifies the aligned read and
|
||||
write sides.
|
||||
|
||||
### The `ReadPlan` shape
|
||||
|
||||
A pre-resolved tree of read instructions. Every `$ref` is resolved, every
|
||||
endianness is computed (field override or container default), every
|
||||
union variant is inlined. The read loop indexes into a `Vec`, matches
|
||||
on a `ReadKind`, and calls `data_access::read_*` with a precomputed
|
||||
`Endian` — no `resolve_typeref`, no `BastDef::parse`, no JSON node
|
||||
access on the happy path. (`format!` for error-path attribution may
|
||||
still occur on the error path; it does not run on the happy path and
|
||||
is not the cost being removed here.)
|
||||
|
||||
```rust
|
||||
pub struct ReadPlan {
|
||||
endian: Endian,
|
||||
fields: Vec<FieldPlan>,
|
||||
by_name: HashMap<String, usize>,
|
||||
}
|
||||
|
||||
pub struct FieldPlan {
|
||||
name: String,
|
||||
kind: ReadKind,
|
||||
endian: Endian,
|
||||
encoding: VariableEncoding,
|
||||
max_length: Option<usize>,
|
||||
body: Option<CompositePlan>,
|
||||
}
|
||||
|
||||
pub enum ReadKind {
|
||||
Primitive(AlkTypeKind),
|
||||
Enum,
|
||||
Struct,
|
||||
Union,
|
||||
Array,
|
||||
Record,
|
||||
}
|
||||
|
||||
pub enum CompositePlan {
|
||||
Struct(ReadPlan),
|
||||
Union {
|
||||
disc: DiscriminatorPlan,
|
||||
shared: Option<Box<ReadPlan>>,
|
||||
variants: Vec<(String, CompositePlan)>,
|
||||
},
|
||||
Array {
|
||||
element: Box<CompositePlan>,
|
||||
count: usize,
|
||||
element_stride: usize,
|
||||
},
|
||||
Record {
|
||||
value: Box<CompositePlan>,
|
||||
},
|
||||
}
|
||||
|
||||
pub enum DiscriminatorPlan {
|
||||
Byte { offset: usize, disc_type: AlkTypeKind },
|
||||
Field { name: String, field_index: usize },
|
||||
}
|
||||
```
|
||||
|
||||
Nested structs share the `ReadPlan` shape (a struct field's `body` is
|
||||
`CompositePlan::Struct(ReadPlan)`). Union variants are pre-resolved:
|
||||
each `(key, CompositePlan)` entry carries the variant's compiled body,
|
||||
so dispatch is a flat lookup + recurse — no `resolve_typeref_as_def` at
|
||||
read time. A variant may itself be `CompositePlan::Union { ... }`, so
|
||||
**nested unions** (a union variant that is itself a union, which the
|
||||
0.2.0 reader supports via `resolve_and_walk_variant`'s `Union` arm) are
|
||||
covered by ordinary recursion; no separate `VariantKind` enum is
|
||||
needed. Array element strides are precomputed (`element_stride = 0`
|
||||
signals variable-length elements, same convention as today's
|
||||
`FieldValue::Array`).
|
||||
|
||||
`Union.shared` carries the union's declared `fields` (the discriminator
|
||||
field + any shared fields) for the field-name-discriminator case —
|
||||
`DiscriminatorPlan::Field.field_index` indexes into `shared`, and the
|
||||
read loop walks `shared` first, then looks up and walks the selected
|
||||
variant's `CompositePlan` starting after the shared fields. The
|
||||
byte-offset-discriminator case has no shared fields (`shared: None`):
|
||||
the discriminator byte is read at `disc.offset` and the variant starts
|
||||
immediately after the discriminator size. The POC's `plan_read_union`
|
||||
`Field` arm stub (Finding 1) is replaced by this `shared` sub-plan;
|
||||
there is no separate `VariantPlan`/`VariantKind` type in the production
|
||||
shape — the POC's `VariantPlan { kind, plan }` wrapper is dropped in
|
||||
favor of recursing on `CompositePlan` directly, which is what makes
|
||||
nested-union support fall out for free.
|
||||
|
||||
### Construction
|
||||
|
||||
```rust
|
||||
impl ReadPlan {
|
||||
pub fn compile(bast_doc: &Value, root_name: &str) -> Result<Self, AlkTypeError>;
|
||||
}
|
||||
```
|
||||
|
||||
`compile` walks `BastDoc` once, resolves all `$ref`s eagerly, computes
|
||||
effective endianness at every node, inlines union variants, and
|
||||
builds the `FieldPlan`/`CompositePlan` tree. Malformed schemas surface
|
||||
as `AlkTypeError::Schema` — the same untrusted-input discipline
|
||||
(AGENTS.md §3) and overflow-safe arithmetic (AGENTS.md §4) as
|
||||
`BastDoc::new`.
|
||||
|
||||
### Engine integration
|
||||
|
||||
`AlkTypeEngine::compile` builds the `ReadPlan` in packed mode and
|
||||
stores `Arc<ReadPlan>`. `sequential_reader()` returns
|
||||
`SequentialReader { plan: Arc::clone(&self.plan), .. }` — an owned
|
||||
reader (ADR-007's factory decision is retained; the reader owns its
|
||||
cursor, shares the immutable plan).
|
||||
|
||||
`bast_doc: Value` (`src/engine.rs:85`) is retained on the engine
|
||||
unconditionally. In packed mode it becomes unused by the read path
|
||||
(both `sequential_reader` and `validate_bytes` consume the plan);
|
||||
in aligned mode it is still needed for `read_field`/`write_field`/
|
||||
`validate_bytes`. Keeping it always avoids a mode-conditional field
|
||||
and costs only a `Value` clone paid once at `compile`. `ReadPlan`
|
||||
must be `Send + Sync` so `Arc<ReadPlan>` can be shared from the
|
||||
`Send + Sync` engine (ADR-007); this falls out naturally from the
|
||||
plan being immutable owned data, but the implementation should add a
|
||||
`static` bound assertion test to lock it in.
|
||||
|
||||
### Public API change (breaking — version bump to 0.3.0)
|
||||
|
||||
- `SequentialReader::new(&Value, &str)` → `SequentialReader::new(Arc<ReadPlan>)`.
|
||||
The old constructor is replaced by `ReadPlan::compile(&Value, &str)`
|
||||
followed by `SequentialReader::new(Arc::from(plan))`.
|
||||
- `materialize_packed(&BastDoc<'_>, &[u8])` →
|
||||
`materialize_packed(&ReadPlan, &[u8])`.
|
||||
- `materialize_aligned` is unchanged (already takes `&OffsetMap`, a
|
||||
compiled form).
|
||||
- `ReadPlan` is a new public type, re-exported from `lib.rs`.
|
||||
- `BastDoc` and the `Bast*` types are **unchanged by this ADR** — they
|
||||
remain the validation-side typed tree, borrowed, as today. This is a
|
||||
smaller breakage than review #004's Option A (which changed
|
||||
`BastDoc<'a>` → `BastDoc` and every `Bast*` signature).
|
||||
**Note (added on ADR-012 acceptance):** ADR-012 §2a subsequently
|
||||
makes `BastDoc` owned, riding the same 0.3.0 bump. That is an
|
||||
ADR-012 change, not an ADR-011 change; ADR-011's scope statement
|
||||
stands as the ADR-011-only view. With ADR-012 §3 (ValidationPlan,
|
||||
now in 0.3.0), `bast_validation` will also stop being the permanent
|
||||
home of the `BastDoc` walk — see ADR-012.
|
||||
|
||||
The crate is pre-1.0 with two in-house downstream consumers
|
||||
(`alktty`, `alkcall`), both of which will be updated with the bump.
|
||||
|
||||
## Scope
|
||||
|
||||
### In scope (consumes the `ReadPlan`)
|
||||
|
||||
- **`SequentialReader`** — the read loop walks `FieldPlan`/`CompositePlan`
|
||||
instead of `&BastField`/`&BastType` + `&BastDoc`. The functions
|
||||
`read_field_value`, `read_typeref_value`, `walk_struct_size`,
|
||||
`read_union_value`, `read_array_value`, `read_record_value` are
|
||||
rewritten to take plan nodes. One walker, not two — the review's
|
||||
Option B concern ("duplicates the `BastType` matching logic") does
|
||||
not apply because the plan *replaces* the `BastType` matching, not
|
||||
parallels it.
|
||||
- **`materialize` (packed mode)** — `materialize_packed` takes
|
||||
`&ReadPlan` and walks it to produce `serde_json::Value`. Same read
|
||||
logic, same `data_access` calls, different input type. Unifies the
|
||||
two packed read-side consumers on one compiled form.
|
||||
- **`AlkTypeEngine::validate_bytes` (packed mode)** — calls
|
||||
`materialize_packed(&self.plan, buffer)` instead of reconstructing a
|
||||
`BastDoc`. Closes M1's `engine.rs:284` re-parse.
|
||||
|
||||
### Out of scope (stays on `BastDoc`)
|
||||
|
||||
- **`bast_validation`** — the BAST-native value-domain validator walks
|
||||
`BastDoc` to check constraints (`maxLength`, enum string values,
|
||||
union variant keys, integer ranges). These are value-domain checks,
|
||||
not byte-position walks; they don't benefit from a *read* plan and
|
||||
would require a separate "validation plan" with a different shape.
|
||||
**Not in scope for ADR-011** — but no longer deferred indefinitely:
|
||||
ADR-012 §3 brings a `ValidationPlan` into 0.3.0. The "not a hot
|
||||
loop" framing this paragraph originally relied on was re-evaluated
|
||||
and rejected (see ADR-012 §3): read+validate on untrusted streams
|
||||
makes validation hot in the same sense review #004 measured for
|
||||
the read path. For 0.3.0 as accepted by ADR-011 alone,
|
||||
`bast_validation` keeps walking `BastDoc`; ADR-012 §3 closes that.
|
||||
- **`LayoutBuilder` / `PackedLayout`** — the packed write-side already
|
||||
has a compiled form (`PackedLayout`). `LayoutBuilder::build`
|
||||
(`src/layout_builder.rs:190`) re-parses `BastDoc::new` per `build()`
|
||||
call (M1), but the typical pattern is build-once-reuse, so this is
|
||||
not a hot loop. A future `WritePlan` that lets `LayoutBuilder` cache
|
||||
the typed tree (review #004 Option A's territory) is additive and can
|
||||
follow; it is not blocking the read-path fix.
|
||||
- **`OffsetMap` / aligned mode** — already a compiled form; unchanged.
|
||||
`materialize_aligned` already takes `&OffsetMap`.
|
||||
- **Aligned-mode `read_field` / `write_field`** (`engine.rs:334,467`)
|
||||
re-parse `BastDoc::new` per call (M1). These are one-shot paths, not
|
||||
hot loops; they can adopt a compiled form later without affecting
|
||||
the packed read-path decision. Left as-is for now.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Closes the 400x read-path gap (review #004 H1).** The read loop no
|
||||
longer touches `BastDoc` or the raw `Value`. Per-field work drops
|
||||
from "re-parse the typed tree + resolve_typeref + match" to "index
|
||||
into a `Vec` + match `ReadKind` + `data_access::read_*` with a
|
||||
precomputed `Endian`." The expected per-chunk cost is in the
|
||||
hand-rolled codec's ballpark (the `data_access` calls are the same
|
||||
ones the hand-rolled codec makes).
|
||||
- **Closes M1's packed-side re-parse (`engine.rs:284`, `validate_bytes`).**
|
||||
The aligned-side M1 sites (`engine.rs:334,467`,
|
||||
`layout_builder.rs:190`) are deliberately left as-is — they are not
|
||||
hot for the packed-codec use case, and Option A remains available as
|
||||
an additive later fix if an aligned-mode hot loop ever emerges. This
|
||||
is a reversible bet, not a claim that the aligned-side M1 is a
|
||||
non-issue.
|
||||
- **Unifies the two packed read-side consumers on one compiled form.**
|
||||
`SequentialReader` and `materialize_packed` walk the same `ReadPlan`,
|
||||
mirroring how `OffsetMap` unifies the aligned read and write sides.
|
||||
The "two parallel walkers" concern from review #004 Option B does
|
||||
not apply — the plan replaces the `BastType` matching, not
|
||||
duplicates it.
|
||||
- **Symmetric with the other compiled forms.** The engine now has a
|
||||
compiled form for every mode/side pair: `OffsetMap` (aligned R/W),
|
||||
`PackedLayout` (packed W), `ReadPlan` (packed R). The "compiled form
|
||||
of a BAST document" framing in ADR-004/validation.md becomes true for
|
||||
the read path, not just the write path.
|
||||
- **ADR-007's factory gets cheaper.** Today
|
||||
`sequential_reader()` clones `doc_value: Value` (the whole BAST
|
||||
document) per reader. With the plan, it clones an `Arc<ReadPlan>`
|
||||
(refcount bump). The plan is immutable and shared across all readers
|
||||
from one engine. ADR-007's "owned fresh reader" decision is retained;
|
||||
the reader owns its cursor, shares the plan.
|
||||
- **Deterministic compile.** `ReadPlan::compile` is a pure function of
|
||||
the BAST document + root name — same input, same plan. This makes the
|
||||
"compiled form" visibly deterministic, which is a prerequisite for
|
||||
future capabilities (fingerprinting the plan for cross-run caching,
|
||||
disk-cached compiled plans, or schema-version handshakes for
|
||||
`alkcall`'s hub/spoke topology). Not implemented in this ADR and not
|
||||
needed to justify the decision; listed here only so a future ADR
|
||||
doesn't re-derive the prerequisite. See "Future capabilities" below.
|
||||
- **Retires the "re-parse on demand" framing (L2).** ADR-007's "Cost"
|
||||
section and `src/engine.rs:112-115`'s doc comment framed re-parse as
|
||||
the intended design. With the plan, the read path never re-parses;
|
||||
the framing is retired. ADR-007's "Cost" section and the doc comment
|
||||
are updated in the same commit.
|
||||
- **L1 falls out.** The dead `_field_schema: &Value` parameter and the
|
||||
`Vec<(String, Value)>` field storage (where the `Value` half is
|
||||
unused) are replaced by `Vec<FieldPlan>`. No dead `Value` clones.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Breaking public-API change (0.2.0 → 0.3.0).** `SequentialReader::new`
|
||||
and `materialize_packed` change signatures (take `ReadPlan` instead
|
||||
of `&Value`/`&BastDoc`). `ReadPlan` is a new public type. Per
|
||||
AGENTS.md, this is semver-relevant. The crate is pre-1.0 with two
|
||||
in-house downstream consumers, both updated with the bump. The
|
||||
breakage is smaller than review #004's Option A (no `Bast*` type
|
||||
changes — `BastDoc` stays borrowed, stays the validation-side tree).
|
||||
- **A parallel typed tree, not a flat lookup table.** This is the
|
||||
honest cost. `OffsetMap` and `PackedLayout` are flat `(path, range)`/
|
||||
`(path, position)` projections of `BastType`; `ReadPlan` is a full
|
||||
parallel hierarchy (`CompositePlan` mirrors `BastType`'s
|
||||
Struct/Union/Array/Record). The maintenance tax is real and higher
|
||||
than those: when schema semantics change, `BastDoc`/`BastType` and
|
||||
`ReadPlan`/`CompositePlan` move together. It is worth it because the
|
||||
perf win on composite-heavy schemas (the SFTP-shaped union-with-`$ref`
|
||||
-variants packet) justifies it — see the next bullet. This is a
|
||||
permanent tax accepted in exchange for a ~20–50x composite-dispatch
|
||||
win on top of the 400x re-parse fix, not a structural symmetry with
|
||||
the flat compiled forms.
|
||||
- **Eager `$ref` resolution at compile time.** `ReadPlan::compile`
|
||||
resolves all `$ref`s eagerly, including union variant refs. This is
|
||||
correct (the schema is fixed at compile time) and matches the
|
||||
review's Option B design, but it means a schema with a `$ref` cycle
|
||||
(which BAST forbids — refs are always `#/$defs/<name>`, no
|
||||
recursion) would loop forever. The meta-schema (`bast_meta`)
|
||||
already rejects recursive schemas; `ReadPlan::compile` inherits
|
||||
that guard. No new failure mode.
|
||||
- **Nested `Box<CompositePlan>` vs flat `Vec<Op>`.** The plan as
|
||||
specified uses nested `Box`es — idiomatic, debuggable, easy to
|
||||
build. A flat `Vec<Op>` with jump indices (true "bytecode") would be
|
||||
more cache-friendly but harder to build and read. Protocol headers
|
||||
are small N (SFTP's largest packet has 5 fields); the perf win is
|
||||
eliminating the re-parse and `resolve_typeref`, not SoA cache
|
||||
effects. Start nested; go flat only if a bench says otherwise (a
|
||||
two-way door — the plan is a private internal type; its shape can
|
||||
change without a semver bump as long as the public `ReadPlan` name
|
||||
and `compile`/`SequentialReader::new` signatures are stable).
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
- **Not a `BastDoc` replacement.** `BastDoc` stays as the
|
||||
validation-side typed tree within ADR-011's scope (borrowed from
|
||||
`&Value`, unchanged). The `Bast*` types and their signatures are not
|
||||
touched by ADR-011. Validation (`bast_validation`), aligned one-shot
|
||||
reads/writes (`engine.rs:334,467`), and `LayoutBuilder::build`
|
||||
continue to walk `BastDoc` within ADR-011's scope. **ADR-012
|
||||
subsequently revises two of these:** `BastDoc` becomes owned (§2a)
|
||||
and `bast_validation` adopts a `ValidationPlan` (§3), both riding
|
||||
the same 0.3.0 bump. ADR-011's scope statement is the ADR-011-only
|
||||
view and is not re-litigated here.
|
||||
- **Not a flat lookup table.** Packed positions are data-dependent;
|
||||
the plan is a read program (instructions to walk), not a `(path,
|
||||
offset)` table. This is inherent to packed sequential layout
|
||||
(ADR-002), not a limitation of this design.
|
||||
- **Not a validation plan.** `bast_validation`'s value-domain checks
|
||||
(maxLength, enum values, union variant keys, integer ranges) are a
|
||||
different concern and a different shape from `ReadPlan`. They are
|
||||
out of ADR-011's scope; ADR-012 §3 adds a `ValidationPlan` in 0.3.0
|
||||
rather than leaving validation on an interpretive `BastDoc` walk
|
||||
indefinitely.
|
||||
- **Not the review's Option A or Option B.** It is the "compiled form"
|
||||
path the review pointed at but did not name: Option A (make `BastDoc`
|
||||
own its data) kills the re-parse but leaves the read loop as a
|
||||
`BastDoc` tree walk; Option B (precompute an owned read plan in
|
||||
`SequentialReader::new`) is the surgical subset that closes H1 only.
|
||||
This ADR is the principled version of B — a public `ReadPlan` built
|
||||
at `compile` time, shared across readers and `materialize`, symmetric
|
||||
with `OffsetMap`/`PackedLayout` — and it closes H1 + M1 (packed side)
|
||||
+ L1 + L2.
|
||||
|
||||
## Recommended Order
|
||||
|
||||
1. **`ReadPlan` type + `compile`** — the `ReadPlan`/`FieldPlan`/
|
||||
`CompositePlan`/`ReadKind`/`DiscriminatorPlan` types and the
|
||||
`ReadPlan::compile(&Value, &str)` constructor. Pure addition; no
|
||||
existing code touched. Unit-tested against the same BAST fixtures
|
||||
the `BastDoc` tests use.
|
||||
2. **`SequentialReader` rewrite** — the read loop walks `&ReadPlan`
|
||||
instead of reconstructing `BastDoc`. `read_field_value`,
|
||||
`read_typeref_value`, `walk_struct_size`, `read_union_value`,
|
||||
`read_array_value`, `read_record_value` take plan nodes. The
|
||||
existing `sequential_reader.rs` tests (which drive `read_next`/
|
||||
`read_field`/`reset` over real buffers) pass unchanged — they
|
||||
exercise the read path through the public API, so they validate
|
||||
the rewrite without modification.
|
||||
3. **`materialize_packed` rewrite** — takes `&ReadPlan`, walks the
|
||||
plan to produce `Value`. The existing `validate_bytes` (packed)
|
||||
tests cover it end-to-end.
|
||||
4. **`AlkTypeEngine::compile` integration** — builds `Arc<ReadPlan>`
|
||||
in packed mode, stores it, `sequential_reader()` hands out
|
||||
`Arc::clone(&self.plan)`. `validate_bytes` (packed) calls
|
||||
`materialize_packed(&self.plan, buffer)`.
|
||||
5. **L2 — retire the "re-parse on demand" framing.** Update
|
||||
ADR-007's "Cost" section (replace the "re-parse on demand"
|
||||
paragraph with the `Arc<ReadPlan>` cost) and the
|
||||
`src/engine.rs:112-115` doc comment. ADR-007's status block stays
|
||||
"Accepted" for the factory decision; only the cost framing changes.
|
||||
6. **Public API bump (0.2.0 → 0.3.0).** `lib.rs` re-exports `ReadPlan`;
|
||||
`SequentialReader::new` and `materialize_packed` signatures change.
|
||||
Update `alktty`/`alkcall` in the same commit.
|
||||
7. **Verification block.** `cargo test --release`,
|
||||
`cargo clippy --all-targets -- -D warnings`, `cargo doc --no-deps`
|
||||
(new public type), `cargo build --target wasm32-unknown-unknown
|
||||
--release` (the plan touches `sequential_reader.rs` and
|
||||
`materialize.rs`, both wasm-relevant). Re-run the `alktty`
|
||||
`wire_vs_bast` bench to confirm the 400x gap closes.
|
||||
|
||||
Steps 1–2 close H1. Step 3 closes the `materialize` half of M1.
|
||||
Step 4 closes the `validate_bytes` half of M1 (packed side). Step 5
|
||||
closes L2. L1 falls out at step 2. The aligned-side M1 paths
|
||||
(`engine.rs:334,467`, `layout_builder.rs:190`) are left as-is per
|
||||
"Out of scope."
|
||||
|
||||
## References
|
||||
|
||||
- [Review #004](../../reviews/004-performance-review.md) — the
|
||||
performance finding (H1, M1, L1, L2) and the three fix options this
|
||||
ADR supersedes
|
||||
- [ADR-002](002-two-layout-modes-packed-vs-aligned.md) — the two
|
||||
layout modes; `ReadPlan` is the packed read-side compiled form that
|
||||
this ADR adds to the table
|
||||
- [ADR-007](007-packed-mode-read-factory.md) — the engine as
|
||||
`SequentialReader` factory; retained (owned fresh reader), with the
|
||||
"re-parse on demand" framing retired (L2)
|
||||
- [ADR-004](004-error-handling-validation-strategy.md) — `AlkTypeError`,
|
||||
load-time build / access-time check, field-path-carrying errors;
|
||||
`ReadPlan::compile` is a load-time build, the read loop is an
|
||||
access-time check
|
||||
- [ADR-010](010-generalized-validation-validate-bytes.md) —
|
||||
`validate_bytes` (packed) consumes the `ReadPlan` via
|
||||
`materialize_packed`
|
||||
|
||||
## Future capabilities (in 0.3.0 via ADR-012)
|
||||
|
||||
The deterministic-compile property of `ReadPlan` is a prerequisite for
|
||||
several capabilities. ADR-012 ("Plan Fingerprinting, ValidationPlan,
|
||||
and Closing the Deferred M1 Sites in 0.3.0") picks up all three items
|
||||
below into the 0.3.0 release so they ship with this ADR's breaking
|
||||
changes in one round of downstream churn, not two or three:
|
||||
|
||||
- **Fingerprinting the plan** (`#[derive(Hash)]` + a `fingerprint()`
|
||||
method) for cross-run caching of compiled plans, disk-cached plans,
|
||||
and `alkcall` schema-version handshakes. → **In 0.3.0 (ADR-012 §1).**
|
||||
- **Closing the deferred M1 sites** via an owned `BastDoc` (lifetime
|
||||
removal) for `LayoutBuilder` + extending `OffsetMap` with leaf
|
||||
metadata for the aligned `read_field`/`write_field` paths. → **In
|
||||
0.3.0 (ADR-012 §2).** Note: ADR-012 reframes the earlier "WritePlan"
|
||||
candidate listed here as "not a new type — extend the existing
|
||||
compiled forms (`PackedLayout`/`OffsetMap`) and cache the parse."
|
||||
- A `ValidationPlan` that follows the same compile-once-walk-many
|
||||
pattern for `bast_validation`. → **In 0.3.0 (ADR-012 §3).** Different
|
||||
shape (value-domain, not byte-position) but the same class of
|
||||
per-buffer re-walk cost on the read+validate-on-untrusted-input
|
||||
common case. ADR-012 owns the shape decision and the implementation
|
||||
plan scopes it. (Originally deferred by ADR-012 as "not a hot loop";
|
||||
review #005 M3 reversed the deferral — see ADR-012 §3.)
|
||||
|
||||
None of the in-0.3.0 items justify this ADR; the 400x read-path gap
|
||||
does. They are listed here as forward references and to record that
|
||||
the "WritePlan" candidate has been reframed out by ADR-012.
|
||||
|
||||
## POC coverage
|
||||
|
||||
Before this ADR was accepted, a derisking POC on branch `readplan-poc`
|
||||
walked the read loop against every `BastType` arm in
|
||||
`src/sequential_reader.rs:303-381` (and the parallel arms in
|
||||
`materialize.rs`) and confirmed the `ReadPlan`/`CompositePlan`/
|
||||
`ReadKind`/`DiscriminatorPlan` shape covers all cases, including the
|
||||
two spots where a plan arm could subtly miss a case:
|
||||
|
||||
- **Union discriminator split (`Byte` vs `Field`).** Both are covered:
|
||||
`DiscriminatorPlan::Byte { offset, disc_type }` and
|
||||
`DiscriminatorPlan::Field { name, field_index }`. The field-name case
|
||||
pre-resolves the discriminator field's `ReadKind` so dispatch reads
|
||||
it from the plan, not from a re-parsed `BastField`. The
|
||||
field-disc union's declared `fields` (discriminator + any shared
|
||||
fields) are carried as a sub-`ReadPlan` on `CompositePlan::Union`'s
|
||||
`shared` field (a refinement of the POC shape, which stubbed the
|
||||
`Field` arm — Finding 1); the production read loop walks `shared`
|
||||
first, then the selected variant's `CompositePlan`. Nested-union
|
||||
variants (a variant that is itself a union) are covered by ordinary
|
||||
`CompositePlan` recursion; the POC rejected them, the 0.2.0 reader
|
||||
accepts them, and the production shape restores parity.
|
||||
- **Array variable-element-stride (`element_stride = 0`).** Covered:
|
||||
`CompositePlan::Array { element, count, element_stride }` preserves
|
||||
the `0`-signals-variable convention, and the read loop walks
|
||||
sequentially when `element_stride == 0` (matching today's
|
||||
`walk_variable_array_size`).
|
||||
|
||||
The POC also confirmed `ReadPlan: Send + Sync` holds for the planned
|
||||
shape (immutable owned data, no interior mutability, no lifetimes).
|
||||
@@ -0,0 +1,611 @@
|
||||
# ADR-012: Plan Fingerprinting, ValidationPlan, and Closing the Deferred M1 Sites in 0.3.0
|
||||
|
||||
## Status
|
||||
|
||||
Accepted. Implemented in 0.3.0 — §3's `ValidationPlan` (phase 7,
|
||||
2026-08-31), §1's fingerprinting (phase 6), §2a's owned `BastDoc`
|
||||
(phases 3–4), §2b's `LeafMeta` (phase 5); all shipped 2026-09-02.
|
||||
Bundles three pieces of work into the 0.3.0 release so the
|
||||
crate ships one round of breaking changes, not two (or three). The
|
||||
three pieces: (a) fingerprinting `ReadPlan`/`OffsetMap`, (b) closing
|
||||
the deferred M1 sites via an owned `BastDoc` + `OffsetMap` `LeafMeta`,
|
||||
and (c) a `ValidationPlan` that retires the interpretive
|
||||
`bast_validation` walk (added by reversing the original "defer
|
||||
`ValidationPlan`" decision — see "ValidationPlan — in scope for
|
||||
0.3.0" below). Companion to [ADR-011](011-compiled-read-plan-for-packed-mode.md)
|
||||
(the `ReadPlan`) and the [0.3.0 implementation plan](../../plans/030-compiled-forms.md).
|
||||
§3's concrete shape was scoped by the follow-on design session and is
|
||||
implemented in `src/validation_plan.rs` — see §3a below.
|
||||
|
||||
## Context
|
||||
|
||||
ADR-011 accepted the `ReadPlan` as the packed read-side compiled form
|
||||
and deferred two things to "future capabilities":
|
||||
|
||||
1. **Fingerprinting the plan** for cross-run caching, disk-cached
|
||||
compiled plans, and `alkcall` hub/spoke schema-version handshakes.
|
||||
2. **A `WritePlan` and/or `ValidationPlan`** following the same
|
||||
compile-once-walk-many pattern for the deferred M1 sites and the
|
||||
validation walk.
|
||||
|
||||
ADR-011 also explicitly deferred the aligned-side M1 sites
|
||||
(`engine.rs:334,467` `read_field`/`write_field`;
|
||||
`layout_builder.rs:190` `LayoutBuilder::build`) as "a deliberate
|
||||
reversible bet that an aligned-mode hot loop won't emerge."
|
||||
|
||||
This ADR retires the deferrals in one release. The reasoning is
|
||||
timing: 0.3.0 is already a breaking bump (ADR-011 changes
|
||||
`SequentialReader::new` and `materialize_packed` signatures), and the
|
||||
crate has no real downstream consumers yet (only `alktty`/`alkcall`,
|
||||
both in-house). Doing all three pieces now costs one round of
|
||||
downstream churn instead of two or three, and the fingerprinting work
|
||||
cuts across both `ReadPlan` and `OffsetMap` — splitting would create
|
||||
a cross-release dependency that's cleaner in one release. The
|
||||
`ValidationPlan` inclusion follows the same logic applied to the
|
||||
validation walk: deferring it would create a *second* breaking change
|
||||
to `validate_bytes`/`bast_validation` after 0.3.0, which is exactly
|
||||
the round of downstream churn this release exists to retire.
|
||||
|
||||
### Reframing "WritePlan"
|
||||
|
||||
ADR-011's "Future capabilities" section listed a `WritePlan` as a
|
||||
candidate. On inspection, a new public `WritePlan` type is the wrong
|
||||
shape for the deferred M1 sites, for two reasons:
|
||||
|
||||
1. **The packed write-side already has a compiled form: `PackedLayout`.**
|
||||
`LayoutBuilder::build`'s M1 re-parse is the *builder* re-parsing
|
||||
`BastDoc::new` on each `build()` call to get the typed tree it
|
||||
walks. The fix is to cache the parsed tree on the builder at `new()`
|
||||
time — internal, non-breaking, no new public type. The compiled
|
||||
form (`PackedLayout`) is unchanged; only its construction stops
|
||||
re-parsing.
|
||||
|
||||
2. **The aligned R/W side already has a compiled form: `OffsetMap`.**
|
||||
`read_field`/`write_field`'s M1 re-parse is `lookup_leaf_field`
|
||||
walking `BastDoc` to get leaf metadata (`kind`, `encoding`,
|
||||
`endian`) that `OffsetMap` doesn't carry. The fix is to extend
|
||||
`OffsetMap`'s entries with that metadata at `compute` time —
|
||||
additive fields on an existing public type (breaking, but we're
|
||||
bumping anyway). No new public type.
|
||||
|
||||
A new `WritePlan` type would overlap with `PackedLayout` (packed
|
||||
write) and `OffsetMap` (aligned R/W) without a clean distinguishing
|
||||
shape. The honest picture: the packed write-side compiled form is
|
||||
`PackedLayout`; the aligned R/W compiled form is `OffsetMap`; the M1
|
||||
fixes are "cache the parse" and "extend the compiled form with leaf
|
||||
metadata," not "add a third compiled form." This serves the
|
||||
minimal-public-API-changes goal better than a literal `WritePlan`.
|
||||
|
||||
### `ValidationPlan` — in scope for 0.3.0 (no longer deferred)
|
||||
|
||||
The BAST-native validator (`bast_validation`) walks `BastDoc` to
|
||||
check value-domain constraints (enum value sets, integer ranges,
|
||||
`maxLength` caps, union variant keys). This is a different shape
|
||||
from `ReadPlan`/`OffsetMap` (value-domain, not byte-position), and
|
||||
the earlier framing deferred it as "not a hot loop — validation is
|
||||
opt-in per operation per AGENTS.md."
|
||||
|
||||
**That deferral is reversed.** The "not a hot loop" dismissal
|
||||
under-counted the common case: **read + validate together on
|
||||
untrusted input.** The downstream `alkcall` consumer accepts schemas
|
||||
from arbitrary internet peers in a hub/spoke topology (AGENTS.md §3);
|
||||
the common operation on an incoming frame is "read it, then validate
|
||||
it before acting." `validate_bytes` (ADR-010) is therefore called
|
||||
once per incoming buffer, and each call re-walks `BastDoc` for
|
||||
validation even after ADR-011 makes the *read* half plan-fast. That
|
||||
is the same class of per-buffer interpretive cost review #004 measured
|
||||
for the read path (400x per chunk), on a different code path, on the
|
||||
operation the untrusted-input discipline actually requires.
|
||||
|
||||
The cost-of-inaction framing that the original deferral relied on was
|
||||
also wrong: a `ValidationPlan` introduced *after* 0.3.0 would be a
|
||||
breaking change to `validate_bytes`'s contract and to the
|
||||
`bast_validation` public surface, forcing rework of `alktty`/`alkcall`
|
||||
— the exact downstream-churn this release is supposed to retire, not
|
||||
create a second round of. Shipping it in 0.3.0 pays the cost once,
|
||||
alongside the other breaking changes, while there are zero real
|
||||
consumers. The cost of action now is a static, known quantity; the
|
||||
cost of action later is the same work plus a second round of
|
||||
downstream churn plus the risk of the interpretive path being the
|
||||
one that gets used in the meantime on untrusted bytes.
|
||||
|
||||
**Decision: a `ValidationPlan` ships in 0.3.0 as §3 below.** The
|
||||
shape is a compile-once-walk-many compiled form over the BAST
|
||||
document's value-domain constraints, symmetric to `ReadPlan` (packed
|
||||
read-side) and `OffsetMap` (aligned R/W). The concrete shape,
|
||||
construction, and `validate_bytes` integration are scoped in the
|
||||
0.3.0 implementation plan (a dedicated phase) and detailed in a
|
||||
follow-on design session before implementation; this ADR commits the
|
||||
*decision* (in 0.3.0, not deferred) and the *scope* (a compiled
|
||||
validation form that retires the interpretive `BastDoc` walk in
|
||||
`bast_validation`), so the deferral black hole is closed.
|
||||
|
||||
## Decision
|
||||
|
||||
### 1. Fingerprinting — `ReadPlan: Hash + Eq`, `OffsetMap: Hash + Eq`
|
||||
|
||||
Add `#[derive(Hash, Eq)]` (alongside the existing `Debug, Clone, PartialEq`)
|
||||
to `ReadPlan` and `OffsetMap`, plus their public sub-types
|
||||
(`FieldPlan`, `CompositePlan`, `ReadKind`, `DiscriminatorPlan`,
|
||||
`ByteRange`, and the new `LeafMeta` — see §2). `VariantPlan`/
|
||||
`VariantKind` are not in the production `ReadPlan` shape (ADR-011
|
||||
was refined on acceptance to drop them — see ADR-011 status), so
|
||||
they are not derived. `ValidationPlan` (§3) gets `Hash + Eq` + its
|
||||
own `fingerprint()` as part of its public surface.
|
||||
|
||||
**`by_name` representation change.** `ReadPlan.by_name` is currently
|
||||
`HashMap<String, usize>`. `HashMap` iteration order is non-deterministic
|
||||
and `HashMap` does not implement `Hash`, which blocks `#[derive(Hash)]`
|
||||
on `ReadPlan`. Switch `by_name` to `BTreeMap<String, usize>`. Lookup
|
||||
cost at protocol-header N (~5 fields) is negligible (the `BTreeMap` is
|
||||
only used by `read_field`'s name→index lookup, not by the sequential
|
||||
`read_next` hot path). This makes the derived `Hash` cover the full
|
||||
structural state of the plan.
|
||||
|
||||
**Fingerprint contract.** Two plans with equal `Hash` (or equal under
|
||||
`PartialEq`) produce identical reads over identical bytes. Formally:
|
||||
`plan1 == plan2 ⟹ ∀ buffer. read(plan1, buffer) == read(plan2, buffer)`.
|
||||
This is the contract the downstream uses rely on:
|
||||
|
||||
- **Cross-run disk cache.** A consumer can hash a `ReadPlan`/
|
||||
`OffsetMap` and cache the compiled plan keyed by the hash, skipping
|
||||
`compile` on warm starts. Safe because the contract guarantees a
|
||||
cache hit produces identical read behavior.
|
||||
- **`alkcall` hub/spoke schema handshake.** Peers exchange plan
|
||||
fingerprints instead of full BAST documents. A peer that receives a
|
||||
fingerprint it has already compiled can skip re-transmitting the
|
||||
schema. The contract guarantees fingerprint equality implies
|
||||
behavioral equivalence, so the handshake is sound.
|
||||
- **Schema-version diagnostics.** A consumer can log a plan
|
||||
fingerprint alongside read results for reproducibility — two runs
|
||||
over "the same schema" that produce different fingerprints reveal a
|
||||
silent schema drift.
|
||||
|
||||
The contract is a *behavioral* equivalence, not a structural identity:
|
||||
two plans with different `by_name` insertion order but the same
|
||||
`fields` Vec produce the same reads, and after the `BTreeMap` change
|
||||
they also produce the same `Hash`. The contract is documented on the
|
||||
`Hash` impl and tested by a property-style test (compile the same
|
||||
schema twice, assert `plan1 == plan2` and `plan1.hash() ==
|
||||
plan2.hash()`).
|
||||
|
||||
**Fingerprint API.** No new public method is strictly needed —
|
||||
consumers call `std::hash::Hash` directly. For ergonomics and to make
|
||||
the contract visible, add a convenience method:
|
||||
|
||||
```rust
|
||||
impl ReadPlan {
|
||||
/// A stable 64-bit fingerprint of this plan's read behavior.
|
||||
///
|
||||
/// Two plans with the same fingerprint produce identical reads
|
||||
/// over identical bytes (the fingerprint contract).
|
||||
pub fn fingerprint(&self) -> u64;
|
||||
}
|
||||
impl OffsetMap {
|
||||
/// A stable 64-bit fingerprint of this offset map's read/write
|
||||
/// behavior. Same contract as `ReadPlan::fingerprint`.
|
||||
pub fn fingerprint(&self) -> u64;
|
||||
}
|
||||
```
|
||||
|
||||
Implemented via `std::hash::DefaultHasher` (or a stable hasher like
|
||||
`FxHasher` if we want cross-version stability — decision belongs to
|
||||
the implementation step, called out in the plan). The fingerprint is
|
||||
additive API, not breaking.
|
||||
|
||||
### 2. Closing the deferred M1 sites
|
||||
|
||||
#### 2a. `LayoutBuilder` — cache the parsed `BastDoc` at `new()`
|
||||
|
||||
`LayoutBuilder` currently stores `doc_value: Value` + `root_name: String`
|
||||
and re-parses `BastDoc::new(&self.doc_value, &self.root_name)` on every
|
||||
`build()` call (`layout_builder.rs:190`). The fix: store the parsed
|
||||
typed tree at `new()` time and reuse it in `build()`.
|
||||
|
||||
This requires `BastDoc` to be owned (no lifetime borrowing from
|
||||
`doc_value`). Two options:
|
||||
|
||||
- **Option α (smaller):** keep `BastDoc<'a>` borrowing, store
|
||||
`doc_value: Value` + a *pre-resolved, owned* representation of just
|
||||
what `build` needs (the field tree with `$ref`s resolved). This is
|
||||
essentially a `WritePlan` by another name — rejected per the
|
||||
reframing above.
|
||||
- **Option β (cleaner):** make `BastDoc` own its data. This is
|
||||
review #004's Option A, scoped to `LayoutBuilder` only. It's a
|
||||
larger refactor but eliminates the lifetime entanglement for the
|
||||
builder and is the prerequisite for any future owning consumer that
|
||||
wants to cache the parsed tree.
|
||||
|
||||
**Decision: Option β, scoped to `LayoutBuilder`.** The `BastDoc<'a>` →
|
||||
`BastDoc` (owned) refactor is the principled fix and is already
|
||||
breaking (the `Bast*` types are re-exported from `lib.rs`), so it
|
||||
rides the 0.3.0 bump. This does *not* change `SequentialReader` or
|
||||
`materialize_packed` (those consume `ReadPlan` per ADR-011, not
|
||||
`BastDoc`). It changes `LayoutBuilder::new` to parse once and `build`
|
||||
to reuse. The `doc_value: Value` field is removed; the builder holds
|
||||
the owned `BastDoc` directly.
|
||||
|
||||
**Note on `BastDoc` ownership scope:** ADR-011 left `BastDoc` borrowed
|
||||
and unchanged ("the validation-side typed tree"). This ADR changes
|
||||
that: `BastDoc` becomes owned. The validation-side (`bast_validation`)
|
||||
and aligned-side (`OffsetMap::compute`, `materialize_aligned`)
|
||||
consumers adapt to the owned `BastDoc` — they no longer need a
|
||||
borrowed `&Value` kept alive alongside. This is a net simplification:
|
||||
one typed-tree type, owned, used by all non-`ReadPlan` consumers. The
|
||||
POC on `readplan-poc` confirmed `ReadPlan` doesn't need `BastDoc` to
|
||||
be borrowed (it compiles from `&Value` once and discards the
|
||||
`BastDoc`), so making `BastDoc` owned doesn't regress the read path.
|
||||
|
||||
#### 2b. `OffsetMap` — carry leaf metadata
|
||||
|
||||
`OffsetMap` currently stores `Vec<(String, ByteRange)>`. The
|
||||
`read_field`/`write_field` M1 re-parse is `lookup_leaf_field` walking
|
||||
`BastDoc` to get `LeafFieldInfo { kind, encoding, endian }`
|
||||
(`engine.rs:556-600`). The fix: extend `OffsetMap`'s entries to carry
|
||||
that metadata at `compute` time.
|
||||
|
||||
```rust
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
|
||||
pub struct LeafMeta {
|
||||
pub kind: AlkTypeKind,
|
||||
pub encoding: VariableEncoding,
|
||||
pub endian: Endian,
|
||||
}
|
||||
|
||||
pub struct OffsetMap {
|
||||
fields: Vec<(String, ByteRange, LeafMeta)>, // was Vec<(String, ByteRange)>
|
||||
total_size: usize,
|
||||
}
|
||||
```
|
||||
|
||||
`OffsetMap::compute` resolves each leaf field's `LeafMeta` during the
|
||||
walk (it already walks the tree; it just doesn't currently record the
|
||||
metadata). `read_field`/`write_field` drop the `BastDoc::new` +
|
||||
`lookup_leaf_field` calls and read `LeafMeta` from the map. The
|
||||
`LeafFieldInfo` struct in `engine.rs` is removed (replaced by
|
||||
`OffsetMap`'s `LeafMeta`).
|
||||
|
||||
**Breaking changes:**
|
||||
- `OffsetMap::get` return type: `Option<&ByteRange>` →
|
||||
`Option<(&ByteRange, &LeafMeta)>` (or a small accessor struct).
|
||||
Call sites in `alktty`/`alkcall` update with the bump.
|
||||
- `ByteRange` is unchanged (still `Copy + Hash`).
|
||||
- `LeafMeta` is a new public type, re-exported from `lib.rs`.
|
||||
|
||||
This is additive on the *capability* (the map now answers questions it
|
||||
previously couldn't) but breaking on the *signature* (`get`'s return
|
||||
type changes). Rides the 0.3.0 bump.
|
||||
|
||||
### 3. `ValidationPlan` — compile-once validation form
|
||||
|
||||
`bast_validation` currently walks `BastDoc` interpretively on every
|
||||
`validate_bytes` call to check value-domain constraints (enum value
|
||||
sets, integer ranges, `maxLength` caps, union variant keys). After
|
||||
ADR-011, the *read* half of `validate_bytes` (packed) is plan-fast;
|
||||
the *validation* half is still an interpretive `BastDoc` walk per
|
||||
buffer. On the `alkcall` hub/spoke topology, `validate_bytes` is the
|
||||
gate between "bytes arrived from an untrusted peer" and "act on the
|
||||
decoded frame," so it runs once per incoming buffer and validation is
|
||||
hot in the same sense review #004 measured for the read path.
|
||||
|
||||
**Decision: a `ValidationPlan` is a compiled form over the BAST
|
||||
document's value-domain constraints, built once at `compile` time
|
||||
(symmetric to `ReadPlan`/`OffsetMap`) and walked by
|
||||
`bast_validation`/`validate_bytes` without re-touching `BastDoc`.**
|
||||
|
||||
The shape, construction, and `validate_bytes` integration are scoped
|
||||
in the 0.3.0 implementation plan as a dedicated phase and detailed in
|
||||
a follow-on design session before implementation begins. The
|
||||
properties this ADR commits to (so the plan and any implementing agent
|
||||
have a fixed contract):
|
||||
|
||||
- **Compile-once-walk-many.** `ValidationPlan::compile` walks `BastDoc`
|
||||
once; `validate_bytes` (both modes) walks the `ValidationPlan` per
|
||||
buffer, never `BastDoc`. This is the same pattern as `ReadPlan` and
|
||||
`OffsetMap`; it is the structural reason the per-buffer
|
||||
interpretive cost goes away.
|
||||
- **Value-domain, not byte-position.** The plan carries constraint
|
||||
descriptors (enum allowed-sets, integer range bounds, `maxLength`
|
||||
caps, union variant keys, and any other value-domain checks
|
||||
`bast_validation` performs today), keyed for dispatch against the
|
||||
materialized `Value` tree, not byte offsets. The shape is therefore
|
||||
different from `ReadPlan`/`OffsetMap`; the *pattern* (compiled form,
|
||||
immutable, shared via `Arc`) is the same.
|
||||
- **No new `BastDoc` walk in the hot path.** After this ADR, the only
|
||||
consumers that walk `BastDoc` interpretively are the one-shot
|
||||
`compile` paths (`ReadPlan::compile`, `OffsetMap::compute`,
|
||||
`ValidationPlan::compile`, `LayoutBuilder::new`). The per-buffer
|
||||
paths (`sequential_reader`, `materialize_packed`,
|
||||
`materialize_aligned`, `validate_bytes`) all walk compiled forms.
|
||||
This is the end state ADR-011 pointed at; this ADR closes it.
|
||||
- **Semver.** `ValidationPlan` is a new public type, re-exported from
|
||||
`lib.rs`. `validate_bytes`'s *signature* is unchanged (still
|
||||
`(buffer) -> Result<(), AlkTypeError>`); the change is internal
|
||||
(walks the plan instead of `BastDoc`). If the `ValidationPlan`
|
||||
design surfaces a need to change `validate_bytes`'s signature, that
|
||||
rides the 0.3.0 bump and is recorded in the plan's Semver Contract
|
||||
table when the shape is scoped. `bast_validation`'s public surface
|
||||
(`build_validator`, `validate_value`) is reviewed at shape-scope
|
||||
time; additive changes ride the bump, removals/renames are avoided
|
||||
unless the shape work shows they're necessary.
|
||||
- **Fingerprinting.** `ValidationPlan` is `Hash + Eq` with a
|
||||
`fingerprint()` method, same as `ReadPlan`/`OffsetMap` (§1/§4), so
|
||||
the downstream uses (cross-run cache, `alkcall` handshake,
|
||||
schema-version diagnostics) extend to the validation form without
|
||||
new API. The fingerprint contract generalizes: two validation plans
|
||||
with equal hashes accept/reject identical `(bytes)` identically.
|
||||
|
||||
**What this ADR does *not* decide** (left to the follow-on shape
|
||||
session + plan phase): the concrete `ValidationPlan` struct/enum
|
||||
shape, how `maxLength`/range/enum/union-key constraints are
|
||||
represented, whether `bast_validation`'s `validate_value` is retired
|
||||
or kept as a convenience wrapper over the plan, and whether the
|
||||
`AlkTypeKind`-driven dispatch in `bast_validation` collapses into the
|
||||
plan or stays a thin match over plan-carried descriptors. These are
|
||||
shape questions, not decision questions; the decision (in 0.3.0,
|
||||
compiled form, no per-buffer `BastDoc` walk) is fixed here.
|
||||
|
||||
### 3a. `ValidationPlan` shape — resolved by the design session
|
||||
|
||||
The follow-on design session (0.3.0 phase 7 predecessor) resolved the
|
||||
open shape questions; implemented in `src/validation_plan.rs`:
|
||||
|
||||
- **Shape.** `ValidationPlan { root: ValidNode }`, a compiled
|
||||
constraint tree — one `ValidNode` arm per value-domain check,
|
||||
mirroring the interpretive walker's arms one-to-one:
|
||||
`Int { min, max }` / `I64` / `Uint { max }` / `U64` / `Float` / `Bool`
|
||||
/ `Str { max_len }` / `Bytes { max_len }` / `Enum { count }` /
|
||||
`Struct { fields: Vec<ValidField> }` / `Union { variants:
|
||||
Vec<ValidVariant> }` / `Array { count, element }` / `Record
|
||||
{ values }`. `ValidField` carries `name + node`; `ValidVariant`
|
||||
carries `key + node`. The nodes are public (diagnostics access via
|
||||
`ValidationPlan::root()`); construction is only possible through
|
||||
`compile`. Note the union node carries *only the variant nodes* — the
|
||||
declared union `fields` (shared fields) are validated as part of the
|
||||
variant walk, because the walker dispatches on the materialized
|
||||
`__discriminator` and validates the whole object against the selected
|
||||
variant (the `ValidNode::Union` doc comment records this; the
|
||||
interpretive union arm recursed into the variant the same way).
|
||||
- **Constraint representation.** Inline scalar fields on the node arms
|
||||
(ranges as `i64`/`u64` pairs, `maxLength` as `Option<usize>`, enum
|
||||
bound as `count: u64`, union keys as owned `String`s). `maxLength` is
|
||||
resolved from the *owning field* at compile time and baked into the
|
||||
`Str`/`Bytes` leaf — the walk never consults field annotations. It
|
||||
never crosses a `$ref` (a `$ref` always targets a struct/union/enum
|
||||
`$defs` entry, so the interpretive walk could never consult it
|
||||
through one either).
|
||||
- **`compile` signature.** `ValidationPlan::compile(&BastDoc) ->
|
||||
Result<Self, AlkTypeError>` — the plan-table's `&str` root-name
|
||||
parameter was vestigial (the doc already holds its root).
|
||||
- **`validate_value` disposition.** Retained as a one-shot wrapper:
|
||||
`compile(doc)` + `validate(value)`. The interpretive walker behind it
|
||||
is *retired* (deleted) — the wrapper delegates to the plan, so there
|
||||
is one constraint implementation, not two. `bast_validation.rs` keeps
|
||||
the shared error helper (`validation_err`) and the
|
||||
`__discriminator` key constant.
|
||||
- **Error contract preserved.** The plan walk reproduces the
|
||||
interpretive error messages byte-identically: a segment stack
|
||||
(`field` / `[index]` / `[key]`) renders paths only on failure — zero
|
||||
per-node allocation on the happy path. Numeric `__discriminator`
|
||||
dispatch matches mapping keys without allocation for the u64/i64
|
||||
forms (mapping keys are stringified integers; non-integer numbers
|
||||
fall back to `Number::to_string`).
|
||||
- **Compile-time rejection of adversarial graphs.** Eager `$ref`
|
||||
resolution with a definition-level cycle set and a depth cap (128):
|
||||
a cyclic or self-referential schema is `AlkTypeError::Schema` at
|
||||
compile, not a stack overflow — the interpretive walker resolved
|
||||
`$ref`s lazily with no guard and could overflow on recursion.
|
||||
Diamond (shared, non-cyclic) refs compile fine; the cycle set is
|
||||
path-scoped.
|
||||
- **Engine integration.** `AlkTypeEngine` holds `Arc<ValidationPlan>`
|
||||
built at `compile` time in *both* modes; the accessor
|
||||
`validation_plan() -> &Arc<ValidationPlan>` is new public API.
|
||||
`validate_bytes`'s signature is unchanged. The plan compile runs
|
||||
*before* the layout build: it is the engine's reference-graph gate
|
||||
(see Consequences).
|
||||
- **Scope note.** phase-7's `Send + Sync` / `Hash + Eq` /
|
||||
`fingerprint()` requirements are structural on the types above
|
||||
(`#[derive(...)]` on plain owned data; the same `DefaultHasher`
|
||||
fingerprint as phase 6).
|
||||
|
||||
### 4. Fingerprinting `OffsetMap` (bundled with §2b)
|
||||
|
||||
Since `OffsetMap` is getting new fields (`LeafMeta`) in §2b, its
|
||||
`#[derive(Hash, Eq)]` (from §1) covers the new fields automatically.
|
||||
The fingerprint contract for `OffsetMap` is the aligned-side analog
|
||||
of `ReadPlan`'s: two offset maps with equal hashes produce identical
|
||||
aligned reads/writes over identical bytes.
|
||||
|
||||
## Scope
|
||||
|
||||
### In scope
|
||||
|
||||
- `ReadPlan: Hash + Eq` + `fingerprint()` method (§1).
|
||||
- `OffsetMap: Hash + Eq` + `fingerprint()` method (§1, §4).
|
||||
- `BastDoc<'a>` → `BastDoc` (owned) refactor, scoped to the consumers
|
||||
that currently hold `doc_value: Value` and re-parse: `LayoutBuilder`,
|
||||
`bast_validation`, `materialize_aligned`, `OffsetMap::compute`
|
||||
(§2a). `ReadPlan::compile` and the packed read path are unaffected
|
||||
(they consume `&Value` once and discard `BastDoc`).
|
||||
- `LayoutBuilder` caches the owned `BastDoc` at `new()`, `build()`
|
||||
reuses it — no re-parse (§2a).
|
||||
- `OffsetMap` carries `LeafMeta`; `read_field`/`write_field` drop
|
||||
`BastDoc::new` + `lookup_leaf_field` (§2b).
|
||||
- `LeafMeta` new public type (§2b).
|
||||
- `BTreeMap` for `ReadPlan.by_name` (§1).
|
||||
- `ValidationPlan` new public type + `compile` + `Hash + Eq` +
|
||||
`fingerprint()` (§3). `validate_bytes` (both modes) walks the
|
||||
`ValidationPlan` instead of re-walking `BastDoc`. `bast_validation`
|
||||
adopts the plan; the public `validate_value`/`build_validator`
|
||||
surface is reviewed at shape-scope time and rides the bump only if
|
||||
the shape work shows a signature change is necessary.
|
||||
- `ValidationPlan: Hash + Eq` + `fingerprint()` method (§3, §1) — the
|
||||
fingerprint contract extends to the validation form.
|
||||
|
||||
### Out of scope
|
||||
|
||||
- Disk-cache or handshake *implementations* — the fingerprint
|
||||
*contract* and method are in scope (§1, §3); the downstream uses
|
||||
(cache format, wire protocol) are the consumers' problem, not this
|
||||
ADR's.
|
||||
- The `ValidationPlan` shape — **resolved** (§3a). Decided by the
|
||||
design session and implemented in `src/validation_plan.rs`; the
|
||||
decision (in 0.3.0, compiled form, no per-buffer `BastDoc` walk) was
|
||||
fixed here.
|
||||
- Cycle-guard hardening for the *layout* walkers' own recursion
|
||||
(`LayoutBuilder`/`OffsetMap` struct recursion) beyond the engine-path
|
||||
gate described in Consequences — if a non-engine entry point walking
|
||||
those types on untrusted docs becomes a consumer pattern, the
|
||||
guards get their own change (the `AlkTypeEngine::compile` gate
|
||||
covers the supported path today).
|
||||
- Cross-version fingerprint stability — the fingerprint is stable
|
||||
within a crate version but may change across versions (a new
|
||||
`AlkTypeKind` variant, for example, changes the hash). Cross-version
|
||||
stability is a non-goal; consumers cache within a version. The
|
||||
implementation step chooses a hasher and documents the stability
|
||||
contract.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **One breaking release, not two (or three).** ADR-011's `ReadPlan`
|
||||
+ this ADR's `BastDoc`-owned + `OffsetMap` extension +
|
||||
`ValidationPlan` all ship together. The two in-house downstream
|
||||
consumers (`alktty`, `alkcall`) update once.
|
||||
- **Closes all deferred M1 sites.** `LayoutBuilder::build`
|
||||
(`layout_builder.rs:190`), `read_field` (`engine.rs:334`),
|
||||
`write_field` (`engine.rs:467`) all stop re-parsing. The packed-side
|
||||
`validate_bytes` (`engine.rs:284`) was already closed by ADR-011;
|
||||
this ADR closes the aligned-side equivalent.
|
||||
- **Retires the interpretive validation walk.** `validate_bytes` on
|
||||
untrusted streams (the `alkcall` common case) stops re-walking
|
||||
`BastDoc` per buffer. This is the latent perf cliff review #005 M3
|
||||
flagged: the read half was plan-fast after ADR-011, the validation
|
||||
half was not. Closing it here — while there are zero real consumers
|
||||
and one breaking bump already paying the downstream-churn cost —
|
||||
avoids a second breaking change to `validate_bytes`/`bast_validation`
|
||||
after 0.3.0. (Implemented: a spot benchmark of plan-validate on a
|
||||
4-field mixed frame puts the validation half at ~0.2 µs/validate;
|
||||
the compile-per-call one-shot it replaces runs ~2.7x slower before
|
||||
the walk is even counted — and the full 0.2.0 per-buffer cost
|
||||
included lazy `$ref` deep-clones that the one-shot no longer pays.
|
||||
The materialize half, not validation, remains the dominant
|
||||
`validate_bytes` cost.)
|
||||
- **Compile-time rejection of cyclic `$ref` graphs.** A side effect of
|
||||
eager plan compilation: a self-referential document is now a clean
|
||||
`Schema` error instead of a stack overflow. The plan compile runs
|
||||
*before* the layout build in `AlkTypeEngine::compile`, making it the
|
||||
engine's reference-graph gate — `LayoutBuilder`/`OffsetMap`
|
||||
struct-recursion had no cycle guard and previously could recurse
|
||||
unboundedly on such a document (a pre-existing untrusted-schema
|
||||
hazard, surfaced by the phase-7 `compile_rejects_cyclic_ref_graph`
|
||||
test). (Resolved since: review #006 H2 added the shared
|
||||
`walk_guard::check_ref_graph` guard at every standalone walker entry,
|
||||
so the trust boundary no longer depends on the engine path.)
|
||||
- **Fingerprinting enables downstream uses.** Cross-run plan caching,
|
||||
`alkcall` schema handshake, and schema-version diagnostics all
|
||||
become possible without further API work — across `ReadPlan`,
|
||||
`OffsetMap`, and `ValidationPlan`.
|
||||
- **`BastDoc` owned is a net simplification.** One typed-tree type,
|
||||
owned, used by all `*::compile` paths. No more
|
||||
lifetime-entanglement workarounds. The "re-parse on demand" framing
|
||||
from ADR-007 is fully retired across read, write, and validation
|
||||
paths.
|
||||
- **`OffsetMap` extension is additive capability.** The map now
|
||||
answers `kind`/`encoding`/`endian` questions it previously couldn't,
|
||||
enabling future aligned-side tools without re-walking `BastDoc`.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Breaking public-API changes (0.2.0 → 0.3.0).** `BastDoc<'a>` →
|
||||
`BastDoc` (owned) changes every `Bast*` signature that took `&'a`.
|
||||
`OffsetMap::get` return type changes. `LeafMeta` is new public.
|
||||
`ReadPlan` is new public (from ADR-011). `ValidationPlan` (+ the
|
||||
`ValidNode`/`ValidField`/`ValidVariant` node types) is new public
|
||||
(§3a). All ride the bump.
|
||||
- **`BastDoc` ownership refactor is broad.** Touches `bast.rs` (every
|
||||
typed node: `&'a str` → `String`/`Arc<str>`, `&'a Value` →
|
||||
`Value`/`Arc<Value>`) and every consumer (`layout_builder`,
|
||||
`offset_map`, `materialize`, `bast_validation`, `engine`). This is
|
||||
review #004's Option A, which ADR-011 deferred — this ADR picks it
|
||||
up because the `LayoutBuilder` M1 fix requires it and we're bumping
|
||||
anyway. The refactor is mechanical (lifetime removal, not logic
|
||||
rewrites); the POC on `readplan-poc` confirmed the read path is
|
||||
unaffected.
|
||||
- **Interpretive `validate_value` is compile-per-call.** The retained
|
||||
one-shot wrapper (`bast_validation::validate_value`) compiles a plan
|
||||
then validates — fine for one-off/diagnostic use, wrong for per-
|
||||
buffer use. Per-buffer callers must hold the engine (or a plan) —
|
||||
the doc comments say so. The walker it replaced had the inverse
|
||||
trade (no compile, but interpretive per call); the engine path
|
||||
(compile once) is the one that matters.
|
||||
- **Aligned-mode `maxLength`-reserved strings/bytes.** Materialization
|
||||
emits the *full reserved* (zero-padded) data for these fields. A
|
||||
`ValidationPlan` compiled from a document used in packed mode would
|
||||
apply `maxLength` to trimmed length, matching packed semantics; the
|
||||
aligned materializer's zero-padding means the value passed to
|
||||
validation can carry trailing NULs. This is pre-existing
|
||||
materialize behavior (not a plan artifact); consumers relying on
|
||||
trimmed values already see it.
|
||||
- **Fingerprint cross-version stability is not guaranteed.** A future
|
||||
`AlkTypeKind` variant changes the hash. Documented as a within-
|
||||
version contract. Consumers that need cross-version stability
|
||||
serialize the BAST document and re-compile.
|
||||
- **`BTreeMap` for `by_name` is a tiny lookup cost.** Negligible at
|
||||
protocol-header N; irrelevant to the 400x fix.
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
- **Not a `WritePlan` type.** The packed write-side compiled form is
|
||||
`PackedLayout`; the aligned R/W compiled form is `OffsetMap`. The
|
||||
M1 fixes are "cache the parse" (§2a) and "extend the compiled form
|
||||
with leaf metadata" (§2b), not "add a third compiled form."
|
||||
- **Not a `ValidationPlan` deferral.** `ValidationPlan` is in scope
|
||||
(§3) and implemented (§3a); the interpretive walker is retired.
|
||||
- **Not cross-version fingerprint stability.** Within-version only.
|
||||
- **Not a disk-cache or wire-protocol spec.** The fingerprint contract
|
||||
and method are in scope; the downstream uses are the consumers'
|
||||
concern.
|
||||
|
||||
## Recommended Order
|
||||
|
||||
See [the 0.3.0 implementation plan](../../plans/030-compiled-forms.md)
|
||||
for the step-by-step execution order. The high-level grouping:
|
||||
|
||||
1. **`ReadPlan` (ADR-011 steps 1–5)** — the packed read-path fix. Closes
|
||||
H1 + packed-side M1 + L1 + L2.
|
||||
2. **`BastDoc` owned (§2a)** — the typed-tree ownership refactor. Prerequisite
|
||||
for the `LayoutBuilder` M1 fix and for `ValidationPlan::compile`.
|
||||
3. **`LayoutBuilder` M1 fix (§2a)** — cache the owned `BastDoc` at `new()`.
|
||||
4. **`OffsetMap` extension (§2b)** — carry `LeafMeta`; close the
|
||||
aligned-side `read_field`/`write_field` M1.
|
||||
5. **`ValidationPlan` (§3)** — compiled validation form; `validate_bytes`
|
||||
walks the plan instead of `BastDoc`. Requires the owned `BastDoc`
|
||||
from step 2 for `ValidationPlan::compile`.
|
||||
6. **Fingerprinting (§1, §4)** — `Hash + Eq` + `fingerprint()` on
|
||||
`ReadPlan`, `OffsetMap`, and `ValidationPlan`. Rides on top of the
|
||||
above.
|
||||
7. **Public API bump (0.2.0 → 0.3.0)** — `lib.rs` re-exports, version
|
||||
bump, update `alktty`/`alkcall`.
|
||||
8. **Verification block** — full suite + wasm + bench.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-011](011-compiled-read-plan-for-packed-mode.md) — the
|
||||
`ReadPlan` (packed read-side compiled form). This ADR extends the
|
||||
0.3.0 release with fingerprinting, the deferred M1 fixes, and the
|
||||
`ValidationPlan`.
|
||||
- [Review #004](../../reviews/004-performance-review.md) — the
|
||||
performance finding (H1, M1, L1, L2). ADR-011 closed H1 + packed
|
||||
M1 + L1 + L2; this ADR closes the aligned-side M1.
|
||||
- [Review #005](../../reviews/005-plan-review-030.md) — the 0.3.0 plan
|
||||
review whose M3 finding reversed the `ValidationPlan` deferral.
|
||||
- [ADR-007](007-packed-mode-read-factory.md) — the "re-parse on
|
||||
demand" framing, retired across read, write, and validation paths
|
||||
by ADR-011 + this ADR.
|
||||
- [ADR-002](002-two-layout-modes-packed-vs-aligned.md) — the two
|
||||
layout modes; `OffsetMap` is the aligned R/W compiled form extended
|
||||
here with `LeafMeta`.
|
||||
- [0.3.0 implementation plan](../../plans/030-compiled-forms.md) —
|
||||
the step-by-step execution plan.
|
||||
@@ -0,0 +1,316 @@
|
||||
# ADR-BAST: BAST (Binary Abstract Syntax Tree) as the Schema Format
|
||||
|
||||
## Status
|
||||
|
||||
Accepted — supersedes the format-specific content of
|
||||
[ADR-001](001-alktype-purpose-scope-jsonschema-engine.md). ADR-001's
|
||||
purpose, scope, and "schema is the format" principle are retained and
|
||||
strengthened; only the concrete format (custom-keyword JSON Schema →
|
||||
BAST) is superseded by this ADR.
|
||||
|
||||
## Context
|
||||
|
||||
alktype v0.1.0 embedded binary layout information inside standard JSON
|
||||
Schema documents via custom keywords:
|
||||
|
||||
```json
|
||||
{
|
||||
"AlkType:Struct": true,
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"channel_id": { "AlkType:Uint32": true, "type": "integer" },
|
||||
"length": { "AlkType:Uint32": true, "type": "integer" }
|
||||
},
|
||||
"endian": "big"
|
||||
}
|
||||
```
|
||||
|
||||
This worked for the Rust engine — it walked the tree, detected
|
||||
keywords, computed offsets. But it created friction for everything
|
||||
outside Rust:
|
||||
|
||||
1. **Cross-language consumption.** A Python, Go, or TypeScript consumer
|
||||
that wanted to parse an alktype schema had to re-implement custom
|
||||
keyword detection. The format was not self-describing — you needed
|
||||
to know that `AlkType:Uint32` meant "4-byte little/big-endian
|
||||
unsigned integer" out-of-band.
|
||||
2. **No meta-schema.** The custom keywords were not part of any JSON
|
||||
Schema dialect, so `jsonschema` itself could not validate an alktype
|
||||
schema's *structure*. Editors had no autocomplete; a typo in
|
||||
`AlkType:Uint32` (e.g., `AlkType:UINT32`) was a runtime engine
|
||||
error, not a schema-validation error.
|
||||
3. **Awkward composition.** The keyword-value shape (`true` vs an
|
||||
annotation object) and the `normalize_refs` step needed to bridge
|
||||
TypeBox's bare-name `$ref` output and `jsonschema`'s JSON Pointer
|
||||
requirement were engine internals leaking into the format.
|
||||
4. **Validator coupling.** The v0.1.0 bytes path validated by
|
||||
registering 19 `jsonschema::Keyword` factories (~200 lines). The
|
||||
custom keyword integration was the only way to enforce value-domain
|
||||
constraints (integer ranges, `maxLength`, enum index bounds) on the
|
||||
materialized `Value`. The built-in `enum` keyword checked string
|
||||
membership, but the materializer emitted `Value::Number(index)` —
|
||||
so out-of-bounds enum indices *silently passed* (a dead constraint).
|
||||
|
||||
The engine's core logic (layout computation, data access, union
|
||||
dispatch, two layout modes) was format-agnostic beneath the accessor
|
||||
layer. A POC on branch `bast-validator-poc` (commit `f371fe4`,
|
||||
`src/bast_poc.rs`) proved that a `kind`-based vocabulary with
|
||||
`$defs`/`$ref`, a BAST-native validator, and lazy variant ref
|
||||
resolution could replace the custom-keyword machinery end-to-end with
|
||||
no loss of capability and a net reduction in code. The research record
|
||||
is [`docs/research/bast-pivot.md`](../../research/bast-pivot.md);
|
||||
decisions D-BAST-001 through D-BAST-009 are recorded there.
|
||||
|
||||
## Decision
|
||||
|
||||
**alktype's schema format is BAST (Binary Abstract Syntax Tree): a JSON
|
||||
document that describes binary data layouts using a `kind`-based
|
||||
vocabulary with `$defs`/`$ref` for composition.** BAST is itself a
|
||||
valid JSON Schema instance (it has a meta-schema), making it
|
||||
self-validating, editor-friendly, and trivially consumable from any
|
||||
language with a JSON parser.
|
||||
|
||||
The normative format specification is
|
||||
[`docs/architecture/bast-format.md`](../bast-format.md) (meta-schema,
|
||||
TypeRef, examples, validation model). This ADR records the decision and
|
||||
its consequences; the spec records the shape.
|
||||
|
||||
### Design principles
|
||||
|
||||
1. **BAST is a JSON Schema instance.** A BAST document is valid JSON
|
||||
that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON
|
||||
Schema). Any JSON Schema validator can check whether a BAST document
|
||||
is well-formed; editors with JSON Schema support provide autocomplete
|
||||
and inline validation for free.
|
||||
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
|
||||
top-level `$defs` block. `$ref` handles cross-references and union
|
||||
variant references — the same pattern as JSON Schema's own `$defs`
|
||||
and TypeBox's `Type.Module`. No custom reference resolution
|
||||
mechanism.
|
||||
3. **`kind`-based vocabulary.** Every type has a `kind` field whose
|
||||
value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.).
|
||||
This replaces the `AlkType:*` custom-keyword pattern with a flat,
|
||||
easily-matched string. The 18 `AlkTypeKind` enum variants are
|
||||
the BAST kinds; `AlkTypeKind::from_bast_str`/`to_bast_str` map between
|
||||
the enum and the lowercase BAST strings (D-BAST-002).
|
||||
4. **Order is explicit.** Struct fields are an ordered array, not an
|
||||
object with `properties`. Field order is unambiguous — no reliance
|
||||
on `serde_json`'s `preserve_order` for correctness — and matches the
|
||||
mental model of binary layouts.
|
||||
5. **Annotations are type-level properties.** Endianness, alignment,
|
||||
encoding, and discriminators are properties of the type definition
|
||||
or field, not custom keywords on a separate schema object. Their
|
||||
*semantics* carry forward unchanged from
|
||||
[ADR-003](003-schema-annotations.md); only their *location* moves.
|
||||
|
||||
### Document shape
|
||||
|
||||
Every BAST document has the same top-level shape:
|
||||
|
||||
```json
|
||||
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
|
||||
```
|
||||
|
||||
- The `$defs` block is **required** (D-BAST-003). Single-type documents
|
||||
are a special case with one entry.
|
||||
- The **root type name** is a required parameter to
|
||||
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001).
|
||||
Convention (first entry) is fragile and depends on JSON key order; an
|
||||
explicit parameter is used instead.
|
||||
|
||||
### TypeRef
|
||||
|
||||
`TypeRef` is the central mechanism for referencing types. Four forms:
|
||||
primitive string (`"uint32"`), `$ref` object
|
||||
(`{ "$ref": "#/$defs/Read" }`), array object
|
||||
(`{ "kind": "array", "element": "uint32", "count": 3 }`), and record
|
||||
object (`{ "kind": "record", "values": "string" }`).
|
||||
|
||||
The `$ref` form uses standard JSON Pointer syntax **restricted to
|
||||
`#/$defs/<name>`** — no external references, no fragment-only pointers,
|
||||
no bare names. The restriction keeps resolution a single hash lookup
|
||||
and eliminates the `normalize_refs` step the v0.1.0 engine needed for
|
||||
TypeBox's bare-name refs.
|
||||
|
||||
### Meta-schema
|
||||
|
||||
The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that
|
||||
validates the *structure* of BAST documents (is it well-formed?). It
|
||||
lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded
|
||||
in the crate as `BAST_META_SCHEMA` (re-exported from the crate root) for
|
||||
offline use. A *different* validator — the BAST-native validator (see
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md)) — validates *binary
|
||||
data* against a BAST document (are the bytes a valid instance?). These
|
||||
are different validators for different inputs.
|
||||
|
||||
### The typed parser
|
||||
|
||||
`src/bast.rs` parses a BAST document into a borrowed typed tree
|
||||
(`BastDoc`/`BastDef`/`BastStruct`/`BastField`/`BastType`/`BastUnion`/
|
||||
`BastEnum`/`BastArray`/`BastRecord`/`BastRef`). Three consumers (layout
|
||||
engines, materializer, BAST-native validator) walk the same tree, so a
|
||||
typed view pays for itself. See
|
||||
[`schema-layer.md`](../schema-layer.md) for the parser's surface and
|
||||
[`bast-format.md`](../bast-format.md) for the format.
|
||||
|
||||
Variant `$ref`s (union `mapping` entries) are resolved **lazily** by
|
||||
the materializer and validator via `BastDoc::resolve_typeref` — no
|
||||
compile-time inlining.
|
||||
|
||||
### Untrusted input
|
||||
|
||||
Every path that walks a BAST document returns
|
||||
`Err(AlkTypeError::Schema)` on a malformed document, never
|
||||
`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream
|
||||
`alkcall` consumer accepts schemas from arbitrary internet peers in its
|
||||
hub/spoke topology).
|
||||
|
||||
### Bug fix: enum index bounds
|
||||
|
||||
The v0.1.0 engine had a dead constraint on the bytes path — the
|
||||
built-in `enum` keyword checked string membership, but the materializer
|
||||
emitted `Value::Number(index)`, which never matched. The BAST-native
|
||||
validator checks the materialized index against `values.len()` bounds,
|
||||
fixing this. Net improvement, recorded as intended behavior in the
|
||||
test suite.
|
||||
|
||||
## What is removed
|
||||
|
||||
Under the BAST pivot, the v0.1.0 custom-keyword machinery is removed:
|
||||
|
||||
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
|
||||
factories) — replaced by the BAST-native validator (~250 lines, a
|
||||
flat match with no factories, no trait objects, no sub-validator
|
||||
pre-computation). See [ADR-VAL-SPLIT](val-split-two-validator-model.md).
|
||||
- `normalize_refs()` / `inline_union_variant_refs()` — BAST refs are
|
||||
always `#/$defs/<name>`; one hash lookup. Variant refs resolve lazily.
|
||||
- `get_alktype_kind*` family — superseded by the parser's `kind`-string
|
||||
dispatch.
|
||||
- `parse_encoding`/`parse_align`/`parse_max_length`/`parse_endian`/
|
||||
`parse_discriminator` + `DiscriminatorKind` — replaced by the typed
|
||||
`BastField`/`BastDiscriminator` views and the parser's internal
|
||||
BAST-property-form copies.
|
||||
- `resolve_ref`/`resolve_ref_or_inline` — replaced by
|
||||
`BastDoc::lookup_def`/`resolve_typeref`.
|
||||
- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`,
|
||||
`BYTE_DISCRIMINATOR_TYPES` — replaced by `from_bast_str`/`to_bast_str`
|
||||
and the parser's typed views.
|
||||
- The custom-keyword `build_validator` path — `build_validator` is
|
||||
repurposed to build a *standard* `jsonschema::Validator` from a
|
||||
consumer-provided JSON Schema (no custom keywords). See
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md).
|
||||
|
||||
The `jsonschema` crate **remains a direct dependency** for
|
||||
`validate_json` and for validating BAST documents against the BAST
|
||||
meta-schema. The only thing removed is the custom keyword integration
|
||||
path.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Self-describing, cross-language format.** A BAST document carries
|
||||
its type vocabulary in a meta-schema'd JSON Schema instance. Any
|
||||
language with a JSON parser and a JSON Schema validator can validate
|
||||
BAST document structure without knowing alktype's Rust internals.
|
||||
Editors with `$schema` support provide autocomplete and inline
|
||||
validation for free.
|
||||
- **Simpler `$ref` story.** One restricted form (`#/$defs/<name>`), one
|
||||
hash lookup, no normalization pass. TypeBox interop is a serialization
|
||||
concern (TypeBox → BAST JSON), not an engine concern.
|
||||
- **Explicit field order.** The `fields` array makes byte order
|
||||
unambiguous — no reliance on `serde_json`'s `preserve_order` for
|
||||
correctness (it remains a dependency for builder output and for
|
||||
`mapping` iteration order, but layout correctness no longer depends
|
||||
on it).
|
||||
- **Architecture simplification.** The BAST-native validator is a flat
|
||||
recursive match — no factories, no trait objects, no sub-validator
|
||||
pre-computation, no `with_keyword` registration. ~200 lines of
|
||||
custom-keyword validators become ~250 lines of straightforward
|
||||
pattern matching.
|
||||
- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed.
|
||||
- **Wasm binary-size win.** The `validate_bytes` path no longer touches
|
||||
`jsonschema` for validation (it still uses `jsonschema`'s
|
||||
`ValidationError::custom` type for the error payload, per
|
||||
D-BAST-009 — but no validator compilation, no keyword registration,
|
||||
no sub-validators).
|
||||
|
||||
### Negative
|
||||
|
||||
- **Breaking change to the v0.1.0 public surface.** `compile`'s
|
||||
signature changes (new `root_name` param, drops `&mut`, takes a BAST
|
||||
document not a custom-keyword JSON Schema). `validate_json`/
|
||||
`is_valid_json` change contract (validate against a consumer-provided
|
||||
JSON Schema, not the alktype schema). `Schema::build`/
|
||||
`Definitions::build` output format changes. The ~13 `schema::*`
|
||||
helper re-exports are removed. `build_validator` is repurposed. The
|
||||
crate is on crates.io at 0.1.0 with zero real consumers, so the bump
|
||||
is free — but the contract is explicit (see the implementation plan's
|
||||
Semver Contract table).
|
||||
- **Two output formats from the builder.** `struct_()` → BAST,
|
||||
`object()` → standard JSON Schema. The construction API is the same;
|
||||
only the serialization differs. This is deliberate (D-BAST-008) but
|
||||
is a thing consumers must learn.
|
||||
- **`AlkTypeKind::Display` is backed by `to_bast_str`.** The v0.1.0
|
||||
`as_str`/`Display` rendered the `"AlkType:Uint8"` keyword; the new
|
||||
`Display` renders the BAST canonical string (`"uint8"`). Error
|
||||
messages across six modules surface the new name. This is the right
|
||||
name to surface now, but it is a visible change in error output.
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
- **Not a replacement for JSON Schema for JSON validation.** A BAST
|
||||
document cannot validate a JSON payload — it describes binary data
|
||||
layouts and value-domain constraints for bytes. For JSON validation,
|
||||
consumers use standard JSON Schema documents (which may be derived
|
||||
from BAST via future codegen, or authored separately). See
|
||||
[ADR-VAL-SPLIT](val-split-two-validator-model.md).
|
||||
- **Not a code generator.** BAST is a data format, not a Rust source
|
||||
generator. ADR-001's scope boundary stands.
|
||||
- **Not a schema-evolution / Value system.** TypeBox's `Value.Diff`,
|
||||
`Value.Migrate`, `Value.Convert` remain out of scope (ADR-001).
|
||||
- **Not a framing format.** BAST describes one struct/union/enum
|
||||
instance; it does not strip length prefixes or handle multi-frame
|
||||
buffers. Framing stays in the consumer (ADR-010).
|
||||
|
||||
## Decisions (D-BAST-001..009)
|
||||
|
||||
The BAST format is grounded in decisions D-BAST-001 through D-BAST-009,
|
||||
recorded in
|
||||
[the pivot research record](../../research/bast-pivot.md#decisions).
|
||||
Summary:
|
||||
|
||||
| Decision | Summary |
|
||||
|----------|---------|
|
||||
| [D-BAST-001](../../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention |
|
||||
| [D-BAST-002](../../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase |
|
||||
| [D-BAST-003](../../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape |
|
||||
| [D-BAST-004](../../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) |
|
||||
| [D-BAST-005](../../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` |
|
||||
| [D-BAST-006](../../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed |
|
||||
| [D-BAST-007](../../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema |
|
||||
| [D-BAST-008](../../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema |
|
||||
| [D-BAST-009](../../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths |
|
||||
|
||||
## References
|
||||
|
||||
- [`bast-format.md`](../bast-format.md) — the normative BAST format
|
||||
specification
|
||||
- [`schema-layer.md`](../schema-layer.md) — the BAST parser
|
||||
implementation
|
||||
- [ADR-VAL-SPLIT](val-split-two-validator-model.md) — the two-validator
|
||||
model (BAST-native for bytes, standard `jsonschema` for JSON)
|
||||
- [ADR-001](001-alktype-purpose-scope-jsonschema-engine.md) — purpose,
|
||||
scope, and the "schema is the format" principle (format-specific
|
||||
content superseded by this ADR; purpose/scope retained)
|
||||
- [ADR-003](003-schema-annotations.md) — annotation semantics (carry
|
||||
forward unchanged; only location moves)
|
||||
- [ADR-009](009-builder-api.md) — builder API (output format amended
|
||||
to BAST / standard JSON Schema)
|
||||
- [ADR-010](010-generalized-validation-validate-bytes.md) —
|
||||
`validate_bytes` (validation step amended to the BAST-native
|
||||
validator)
|
||||
- [BAST pivot research record](../../research/bast-pivot.md) —
|
||||
motivation, POC scope and result, decisions D-BAST-001..009, risks
|
||||
- [BAST pivot implementation plan](../../plans/bast-implementation.md)
|
||||
— ordered steps, semver contract, ADR-sync checklist
|
||||
@@ -0,0 +1,228 @@
|
||||
# ADR-VAL-SPLIT: Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON
|
||||
|
||||
## Status
|
||||
|
||||
Accepted — refines the "validation strategy" section of
|
||||
[ADR-004](004-error-handling-validation-strategy.md) and the "validation
|
||||
step" of [ADR-010](010-generalized-validation-validate-bytes.md) for
|
||||
the BAST pivot. Records decisions D-BAST-006, D-BAST-007, and
|
||||
D-BAST-009.
|
||||
|
||||
## Context
|
||||
|
||||
alktype v0.1.0 used a single validation mechanism — the `jsonschema`
|
||||
crate with 19 custom keyword validators — for both the JSON path
|
||||
(`validate_json(&Value)`) and the bytes path (`validate_bytes(&[u8])`).
|
||||
The bytes path materialized a `serde_json::Value` tree from the buffer,
|
||||
then ran the same `jsonschema::Validator` against it.
|
||||
|
||||
Under the BAST pivot ([ADR-BAST](bast-bast-format.md)), the format
|
||||
changed from custom-keyword JSON Schema to BAST, and the custom-keyword
|
||||
integration was removed. This forced a re-evaluation of both validation
|
||||
paths:
|
||||
|
||||
1. **The bytes path.** BAST is the complete specification of the binary
|
||||
format — it describes both the layout (how to read) and the
|
||||
constraints (what values are valid). An external JSON Schema is not
|
||||
needed for `validate_bytes`; the BAST document *is* the validation
|
||||
spec for bytes. The natural validator is a recursive walker over the
|
||||
BAST type tree that checks the value-domain constraints the
|
||||
materializer does not (integer ranges, `maxLength`,
|
||||
enum index bounds, union variant constraints). The POC
|
||||
(`bast-validator-poc` branch, `src/bast_poc.rs`) proved this out
|
||||
end-to-end with 20 reference tests.
|
||||
|
||||
2. **The JSON path.** BAST describes bytes, not JSON shape. A JSON
|
||||
`Value` (e.g., an incoming JSON-RPC request) is the wrong input for
|
||||
a BAST document; the right validator is a standard
|
||||
`jsonschema::Validator` built from a standard JSON Schema document
|
||||
the consumer provides. BAST is not involved on this path. This is
|
||||
the path alkcall uses for its `OperationSpec` JSON validation.
|
||||
|
||||
The two paths have different inputs (bytes vs JSON `Value`), different
|
||||
schema sources (the BAST document vs a consumer-provided JSON Schema),
|
||||
and different validators (a flat recursive match vs a compiled
|
||||
`jsonschema::Validator`). But they share the same error variant —
|
||||
`AlkTypeError::Validation` — so consumers handling both (alkcall uses
|
||||
`validate_json` for channel 0 JSON-RPC and `validate_bytes` for binary
|
||||
channels) match one arm.
|
||||
|
||||
## Decision
|
||||
|
||||
**alktype has two validators for two input types:**
|
||||
|
||||
| Path | Input | Validator | Schema source |
|
||||
|------|-------|-----------|---------------|
|
||||
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator (`bast_validation`) | The BAST document (binary layout + value constraints) |
|
||||
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
|
||||
|
||||
### `validate_bytes` — BAST-native validator (D-BAST-006)
|
||||
|
||||
`src/bast_validation.rs` is a recursive walker
|
||||
(`validate_value(doc, &value)`) over the BAST typed tree
|
||||
([`crate::bast::BastDoc`]/[`BastType`]). The materializer
|
||||
(`src/materialize.rs`) produces a structurally-correct `Value` tree
|
||||
from bytes (all declared fields present, types correct, bounds checked,
|
||||
UTF-8 valid, discriminator in mapping, boolean byte 0 or 1). The
|
||||
validator enforces only the **value-domain constraints expressed in the
|
||||
BAST document** — the ones the materializer can't see from the bytes
|
||||
alone:
|
||||
|
||||
| Constraint | Validator arm |
|
||||
|------------|---------------|
|
||||
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` |
|
||||
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` |
|
||||
| Float finiteness (Float32/64) | `validate_float` |
|
||||
| String `maxLength` (byte length) | `check_string` |
|
||||
| Bytes `maxLength` (array length) | `check_bytes` (accepts `Value::String` and `Value::Array`) |
|
||||
| Enum index bounds | `validate_enum` — **fixes the v0.1.0 dead constraint** |
|
||||
| Union variant dispatch | `validate_union` reads `__discriminator`, resolves the variant, recurses |
|
||||
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
|
||||
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
|
||||
| Record values | `validate_record` recurses into each value's `values` type |
|
||||
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
|
||||
|
||||
The validator is a flat `match` — no factories, no trait objects, no
|
||||
sub-validator pre-computation, no `jsonschema` involvement. ~250 lines
|
||||
replace ~200 lines of v0.1.0 custom-keyword factories.
|
||||
|
||||
No external JSON Schema is required. The BAST document is the complete
|
||||
specification of the binary format. An optional external JSON Schema
|
||||
can be layered on top for constraints BAST doesn't express (cross-field
|
||||
consistency, regex patterns on string content) — additive, not
|
||||
load-bearing.
|
||||
|
||||
### `validate_json` — standard jsonschema (D-BAST-007)
|
||||
|
||||
`validate_json(&Value)` / `is_valid_json(&Value)` validate a JSON
|
||||
`Value` against a standard `jsonschema::Validator` compiled at
|
||||
`AlkTypeEngine::compile` time from a consumer-provided JSON Schema
|
||||
(`compile`'s `json_schema: Option<&Value>` parameter). No custom
|
||||
keywords, no BAST involvement. The JSON Schema is independent of the
|
||||
BAST document — BAST describes bytes, not JSON shape. It may be
|
||||
authored separately or derived from BAST via future codegen.
|
||||
|
||||
If no JSON Schema was supplied to `compile`, `validate_json` returns
|
||||
`AlkTypeError::Schema` and `is_valid_json` returns `false`.
|
||||
|
||||
`build_validator` (in `src/validation.rs`) is **repurposed**: it builds
|
||||
a *standard* `jsonschema::Validator` from a plain JSON Schema (no
|
||||
custom keywords). The engine calls it internally during `compile` when
|
||||
`json_schema` is `Some`. Consumers that only need a one-off validator
|
||||
may call `jsonschema::options().build(schema)` directly; `build_validator`
|
||||
exists so the engine's error mapping (`jsonschema` build error →
|
||||
`AlkTypeError::Schema`) is reused. The v0.1.0 custom-keyword
|
||||
`build_validator` is removed.
|
||||
|
||||
The `jsonschema` crate remains a direct dependency for this path and
|
||||
for validating BAST documents against the BAST meta-schema.
|
||||
|
||||
### Error payload (D-BAST-009)
|
||||
|
||||
`AlkTypeError::Validation(jsonschema::ValidationError<'static>)` is
|
||||
**retained** as the error variant for both paths. The bytes path no
|
||||
longer uses `jsonschema` for validation, so its error payload is
|
||||
constructed via `jsonschema::ValidationError::custom` purely to keep
|
||||
the variant's type unchanged. The rationale is consumer ergonomics on
|
||||
the *combined* path: consumers like alkcall use both `validate_json`
|
||||
and `validate_bytes` and handle `AlkTypeError::Validation` in one
|
||||
place. A single uniform payload type means one match arm covers both
|
||||
sources.
|
||||
|
||||
The alternative (`Validation(String)`) was rejected — it would force
|
||||
`validate_json` to flatten its structured errors (instance path, schema
|
||||
path, keyword) to a `String` via `Display`. The more information-rich
|
||||
path would lose data to accommodate the less rich one. That is the
|
||||
wrong direction.
|
||||
|
||||
The `no_std`/minimal-build angle (OQ-002) that the alternative was
|
||||
meant to enable is moot: `validate_json` requires `jsonschema`
|
||||
regardless, so a bytes-only `no_std` build already has to give up
|
||||
`validate_json` as a separate, larger decision. The right place to
|
||||
revisit is when/if OQ-002 is actually pursued.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Right validator for each input.** Bytes are validated by the BAST
|
||||
document that describes them; JSON values are validated by a JSON
|
||||
Schema that describes them. No forced isomorphism between two
|
||||
different input types.
|
||||
- **No external JSON Schema needed for `validate_bytes`.** The BAST
|
||||
document is both the layout spec and the validation spec for bytes.
|
||||
This is the "schema is the format" principle from ADR-001, now fully
|
||||
realized.
|
||||
- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed
|
||||
— the BAST-native validator checks the materialized index against
|
||||
`values.len()` directly.
|
||||
- **Per-variant constraint enforcement (OQ-008) without custom
|
||||
keywords.** The validator recurses into the selected variant's BAST
|
||||
definition on `__discriminator` lookup, enforcing every field
|
||||
constraint the variant declares (e.g., `maxLength` on a `bytes` field
|
||||
inside a variant struct).
|
||||
- **Wasm binary-size win.** The `validate_bytes` path no longer touches
|
||||
`jsonschema` for validation (it still uses `ValidationError::custom`
|
||||
for the error payload type, per D-BAST-009 — but no validator
|
||||
compilation, no keyword registration, no sub-validators).
|
||||
- **Architecture simplification.** ~200 lines of custom-keyword
|
||||
factories become ~250 lines of straightforward pattern matching. No
|
||||
`with_keyword` registration; no `inline_union_variant_refs` compile
|
||||
step.
|
||||
- **Uniform error payload.** Consumers handle one
|
||||
`AlkTypeError::Validation` match arm for both paths (D-BAST-009).
|
||||
|
||||
### Negative
|
||||
|
||||
- **Two validators, not one.** The engine struct carries an
|
||||
`Option<jsonschema::Validator>` (for `validate_json`) and re-parses
|
||||
the BAST typed tree on each `validate_bytes` call (the BAST-native
|
||||
validator is not pre-built — it's a recursive walker over the
|
||||
on-demand `BastDoc`). This is a small cost; the validators serve
|
||||
different inputs and don't share structure.
|
||||
- **`validate_json` requires a consumer-provided JSON Schema.** The
|
||||
engine no longer builds a validator from the alktype schema; the
|
||||
consumer must supply a JSON Schema at `compile` time (or accept that
|
||||
`validate_json` returns `AlkTypeError::Schema`). This is a behavioral
|
||||
break from v0.1.0, intentional under the pivot.
|
||||
- **`AlkTypeError::Validation` payload is `jsonschema`'s type even on
|
||||
the bytes path.** The bytes path constructs it via
|
||||
`ValidationError::custom`, which is slightly awkward but keeps the
|
||||
variant uniform. The `no_std` revisit (OQ-002) is the place to
|
||||
reconsider if a bytes-only minimal build ever materializes.
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
- **Not a `Validator` trait abstraction.** Two methods on one struct,
|
||||
not a trait with impls for JSON-only and BAST-binary schemas. The two
|
||||
impls share little internally (`validate_json` is a single
|
||||
`jsonschema` call; `validate_bytes` is materialize + BAST-native
|
||||
walk), so a trait would add a layer without unifying behavior. See
|
||||
[ADR-010](010-generalized-validation-validate-bytes.md) §"Not a
|
||||
`Validator` trait abstraction".
|
||||
- **Not a binary-aware validator that skips the `Value` tree.** The
|
||||
`Value`-materialization path is the validation path. A future
|
||||
"validate bytes without materializing" path is a two-way door but
|
||||
explicitly out of scope for v1 (would re-introduce a hand-rolled
|
||||
validator, ADR-001).
|
||||
- **Not framing-aware.** `validate_bytes` validates the bytes of *one*
|
||||
schema instance. Framing stays in the consumer (ADR-010).
|
||||
|
||||
## References
|
||||
|
||||
- [`bast-format.md` §Validation Model](../bast-format.md#validation-model)
|
||||
— the normative validation model
|
||||
- [ADR-BAST](bast-bast-format.md) — the BAST format decision
|
||||
- [ADR-004](004-error-handling-validation-strategy.md) — error handling
|
||||
and validation strategy (load-time build, access-time check,
|
||||
`AlkTypeError` enum — retained; validation-strategy section refined
|
||||
by this ADR)
|
||||
- [ADR-010](010-generalized-validation-validate-bytes.md) —
|
||||
`validate_bytes` (the two-step concept retained; the validation step
|
||||
amended to the BAST-native validator by this ADR)
|
||||
- [`validation.md`](../validation.md) — the validation layer
|
||||
documentation
|
||||
- `src/bast_validation.rs` — the BAST-native validator implementation
|
||||
- `src/validation.rs` — the `build_validator` helper
|
||||
- [BAST pivot research record](../../research/bast-pivot.md) —
|
||||
D-BAST-006, D-BAST-007, D-BAST-009
|
||||
@@ -7,8 +7,8 @@ last_updated: 2026-07-22
|
||||
|
||||
The layout engine: offset computation, the two layout modes (packed
|
||||
sequential vs aligned static), alignment, endianness, and variable-length
|
||||
field handling. This is the novel code — the recursive walk of the schema
|
||||
JSON that computes byte positions for each field.
|
||||
field handling. This is the novel code — the recursive walk of the BAST
|
||||
typed tree that computes byte positions for each field.
|
||||
|
||||
## The Two Layout Modes
|
||||
|
||||
@@ -24,8 +24,8 @@ protocols.
|
||||
|
||||
**Components:**
|
||||
|
||||
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `AlkType:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
|
||||
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
|
||||
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(bast_doc, root_name)` (requires a `struct` at the root), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
|
||||
- **`SequentialReader`** — constructed via `engine.sequential_reader()` (shares the engine's compiled `ReadPlan` via `Arc` — see [ADR-011](decisions/011-compiled-read-plan-for-packed-mode.md)), then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
|
||||
|
||||
**How it works:**
|
||||
|
||||
@@ -69,7 +69,7 @@ and safetensors.
|
||||
|
||||
**Component:**
|
||||
|
||||
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, AlkTypeError>` (requires `AlkType:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
|
||||
- **`OffsetMap`** — constructed via `OffsetMap::compute(&doc) -> Result<Self, AlkTypeError>` (requires a `struct` at the root). Walks the BAST typed tree once, computes fixed byte positions for each field based on type sizes and alignment, and resolves each leaf's `LeafMeta` (kind, encoding, effective endianness — ADR-012 §2b). The output is a flat table of `(field_path, OffsetEntry)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
|
||||
|
||||
**How it works:**
|
||||
|
||||
@@ -113,19 +113,24 @@ with a `AlkTypeError::Offset` — the `OffsetMap` reserves only 4 bytes
|
||||
(the length prefix), but `data_access::write_string` writes prefix +
|
||||
data inline, which would clobber subsequent fields. Non-final variable
|
||||
fields must use `maxLength` (fixed-size reservation) or
|
||||
`"encoding": "offset-indirect"`. See
|
||||
`"encoding": "offset-indirect"` — except `record` fields, for which
|
||||
neither remedy is available (`maxLength` is rejected at parse — review
|
||||
#006 N3 — and `offset-indirect` is rejected for records in aligned
|
||||
mode — review #006 M5), so a non-final record field cannot be repaired
|
||||
and must move to the last position. See
|
||||
[ADR-006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md).
|
||||
|
||||
## Offset Computation Algorithm
|
||||
|
||||
The offset computation is a recursive walk of the schema JSON. The
|
||||
The offset computation is a recursive walk of the BAST typed tree
|
||||
([`BastDoc`](schema-layer.md#the-bast-parser-bast-module)). The
|
||||
algorithm is the same for both modes; the difference is whether alignment
|
||||
padding is inserted between fields.
|
||||
|
||||
### Fixed-size types
|
||||
|
||||
For each fixed-size type, the algorithm:
|
||||
1. Determines the type's byte size from the `AlkType:*` kind.
|
||||
1. Determines the type's byte size from the `AlkTypeKind`.
|
||||
2. In aligned mode: inserts padding to satisfy the type's alignment
|
||||
(or the field's `align` annotation, or the struct's `align` default).
|
||||
3. Records the field's `(start, end)` range.
|
||||
@@ -133,14 +138,14 @@ For each fixed-size type, the algorithm:
|
||||
|
||||
### Composite types
|
||||
|
||||
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
|
||||
**`struct`:** Recurse into the struct's `fields` array. The inner fields
|
||||
are computed relative to the struct's start offset. The struct's total
|
||||
size is the sum of its fields' sizes (plus alignment padding in aligned
|
||||
mode). The struct itself may have an `align` annotation that rounds up
|
||||
its total size.
|
||||
|
||||
**`TUnion`:** TUnion is supported in packed sequential mode only. In
|
||||
aligned static mode, `OffsetMap::compute` rejects `TUnion` fields with
|
||||
**`union`:** TUnion is supported in packed sequential mode only. In
|
||||
aligned static mode, `OffsetMap::compute` rejects `union` fields with
|
||||
`AlkTypeError::Offset` — see
|
||||
[ADR-008](decisions/008-reject-tunion-in-aligned-mode.md). Unions
|
||||
are the protocol dispatch pattern (SFTP type bytes, call protocol event
|
||||
@@ -161,12 +166,12 @@ total size. The `SequentialReader` reads the discriminator first, looks
|
||||
up the variant schema, then reads the variant struct sequentially — it
|
||||
doesn't need to know the union's total size upfront.
|
||||
|
||||
**`TArray` of fixed-size elements:** Element stride = element size (plus
|
||||
**`array` of fixed-size elements:** Element stride = element size (plus
|
||||
alignment padding in aligned mode). Element `i` starts at
|
||||
`array_offset + i × stride`. The array's total size is `count × stride`.
|
||||
|
||||
**`TArray` of variable-length-element structs:** Deferred for v1
|
||||
(OQ-001).
|
||||
**`array` of variable-length-element structs:** Deferred for v1
|
||||
(OQ-001, D-BAST-004 — BAST arrays require `count` in v1).
|
||||
|
||||
### Variable-length types
|
||||
|
||||
@@ -193,6 +198,12 @@ annotation shapes).
|
||||
only. The engine uses strategy 1 (inline length-prefixing) because
|
||||
protocols don't benefit from fixed-size reservation.
|
||||
|
||||
`maxLength` applies to `string` and `bytes` fields only. The parser
|
||||
rejects it on any other kind (review #006 N3): the validation plan
|
||||
bakes it into string/bytes leaves only, so on a record (or any other
|
||||
kind) the annotation did nothing — and in aligned mode a record
|
||||
reservation was silently corrupt (review #006 M5).
|
||||
|
||||
**Strategy 3: Offset indirection (`"encoding": "offset-indirect"`).**
|
||||
1. The field is a struct `{offset: u32, length: u32}`.
|
||||
2. The `OffsetMap` records the position of this struct.
|
||||
@@ -289,7 +300,7 @@ pub struct FieldPosition {
|
||||
A field's computed position in a packed layout, produced by
|
||||
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
|
||||
length prefix); for fixed-size fields, `size` is the type's byte size.
|
||||
`kind` records the field's `AlkType:*` kind so the consumer can dispatch
|
||||
`kind` records the field's `AlkTypeKind` so the consumer can dispatch
|
||||
to the correct `data_access` read/write function.
|
||||
|
||||
### `PackedLayout` (packed mode)
|
||||
@@ -311,22 +322,46 @@ discriminators, the discriminator is recorded under the synthetic path
|
||||
(schema `properties` order, with nested struct fields appearing inline
|
||||
under their parent's path prefix).
|
||||
|
||||
### `OffsetMap` (aligned mode)
|
||||
|
||||
A flat table of `(field_path, byte_range)` pairs computed from a schema.
|
||||
### `LeafMeta` / `OffsetEntry` (aligned mode, 0.3.0)
|
||||
|
||||
```rust
|
||||
impl OffsetMap {
|
||||
pub fn compute(schema: &Value) -> Result<Self, AlkTypeError>;
|
||||
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
|
||||
pub fn total_size(&self) -> usize;
|
||||
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
|
||||
pub struct LeafMeta {
|
||||
pub kind: AlkTypeKind,
|
||||
pub encoding: VariableEncoding,
|
||||
pub endian: Endian,
|
||||
}
|
||||
|
||||
pub struct OffsetEntry {
|
||||
pub range: ByteRange,
|
||||
pub meta: LeafMeta,
|
||||
}
|
||||
```
|
||||
|
||||
`compute` requires a `AlkType:Struct` at the top level. `total_size`
|
||||
includes trailing alignment padding. `iter` yields fields in insertion
|
||||
order (schema `properties` order, nested struct fields appearing inline).
|
||||
`OffsetMap::compute` resolves each leaf's read/write metadata (kind,
|
||||
variable-length encoding, effective endianness — field override else
|
||||
container default, propagated the aligned-materializer way) alongside
|
||||
its byte range, so `read_field`/`write_field` dispatch on the entry
|
||||
without re-walking the BAST tree per access (ADR-012 §2b).
|
||||
|
||||
### `OffsetMap` (aligned mode)
|
||||
|
||||
A flat table of `(field_path, OffsetEntry)` pairs computed from a schema.
|
||||
|
||||
```rust
|
||||
impl OffsetMap {
|
||||
pub fn compute(doc: &BastDoc) -> Result<Self, AlkTypeError>;
|
||||
pub fn get(&self, field_path: &str) -> Option<&OffsetEntry>;
|
||||
pub fn total_size(&self) -> usize;
|
||||
pub fn iter(&self) -> impl Iterator<Item = (&str, &OffsetEntry)>;
|
||||
pub fn fingerprint(&self) -> u64;
|
||||
}
|
||||
```
|
||||
|
||||
`compute` requires a `struct` at the root. `total_size`
|
||||
includes trailing alignment padding. `iter` yields fields in the BAST
|
||||
`fields` array order (nested struct fields appearing inline). The map
|
||||
carries `Hash + Eq` (ADR-012 §1); `fingerprint()` is the
|
||||
stable-within-version hash for caching and schema handshakes.
|
||||
|
||||
## Design Decisions
|
||||
|
||||
@@ -355,7 +390,7 @@ See [open-questions.md](open-questions.md) for full details.
|
||||
the two layout modes decision
|
||||
- [ADR-003](decisions/003-schema-annotations.md) — schema
|
||||
annotations
|
||||
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds and their
|
||||
- [schema-layer.md](schema-layer.md) — the 18 AlkType kinds and their
|
||||
byte sizes
|
||||
- [data-access.md](data-access.md) — read/write functions that use the
|
||||
computed offsets
|
||||
@@ -106,31 +106,35 @@ architect's desk" is answerable at a glance.
|
||||
### OQ-006: Builder spec Example 3 — wrap `Union` in a `Struct` — RESOLVED
|
||||
|
||||
- **Status**: resolved. [builder.md](builder.md) Example 3 now wraps
|
||||
the `Union` in a `Schema::struct_().field("payload", ...)` and merges
|
||||
`$defs` via `Definitions::merge_into`. Matches the realistic SFTP
|
||||
wire shape and the engine's `AlkType:Struct`-at-root constraint.
|
||||
the `Union` in a `Schema::struct_().field("payload", ...)` and builds
|
||||
a complete BAST document via `Definitions::build_doc`. Matches the
|
||||
realistic SFTP wire shape and the engine's struct-at-root constraint
|
||||
(the root `$defs` entry must be a `struct`).
|
||||
- **Full file**: [OQ-006](questions/006-builder-spec-example-3-wrap-union.md)
|
||||
|
||||
### OQ-007: `Bytes` materialization — lossy UTF-8 conversion — RESOLVED
|
||||
|
||||
- **Status**: resolved. Array of u8: the materializer produces
|
||||
`Value::Array` of `Value::Number` (one entry per byte, 0..=255) for
|
||||
`AlkType:Bytes` fields. The `BytesValidator` accepts both
|
||||
`Value::String` (for `validate_json`) and `Value::Array` (for
|
||||
`validate_bytes`). `maxLength` = max byte count. Implemented in
|
||||
`src/materialize.rs` and `src/validation.rs`.
|
||||
`bytes` fields. The BAST-native validator (`bast_validation::check_bytes`)
|
||||
accepts both `Value::String` (for `validate_json`-style inputs) and
|
||||
`Value::Array` (for `validate_bytes`). `maxLength` = max byte count.
|
||||
Implemented in `src/materialize.rs` and `src/bast_validation.rs`.
|
||||
- **Full file**: [OQ-007](questions/007-bytes-materialization-lossy-utf8.md)
|
||||
|
||||
### OQ-008: `UnionValidator` variant dispatch — RESOLVED
|
||||
|
||||
- **Status**: resolved. `UnionValidator` now builds a sub-validator for
|
||||
each variant at factory time and dispatches on `__discriminator` at
|
||||
validation time. `AlkTypeEngine::compile` calls
|
||||
`schema::inline_union_variant_refs` before `build_validator` to inline
|
||||
`$ref`s in union `mapping` entries (necessary because the
|
||||
`union_factory` receives the union node, but `$defs` live at the
|
||||
schema root). Implemented in `src/validation.rs`, `src/schema.rs`,
|
||||
and `src/engine.rs`.
|
||||
- **Status**: resolved. Under the BAST pivot, the BAST-native validator
|
||||
(`bast_validation::validate_union`) reads `__discriminator`, looks up
|
||||
the variant `BastType` in the union's `mapping`, and recurses into the
|
||||
variant's BAST definition via `validate_typeref` — enforcing every
|
||||
field constraint the variant declares (e.g. `maxLength` on a `bytes`
|
||||
field inside a variant struct). Variant `$ref`s resolve lazily via
|
||||
`BastDoc::resolve_typeref` — no `inline_union_variant_refs` compile
|
||||
step (removed under BAST). No custom keywords, no `jsonschema`
|
||||
involvement on the bytes path. Implemented in `src/bast_validation.rs`
|
||||
and `src/bast.rs`. See
|
||||
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
|
||||
- **Full file**: [OQ-008](questions/008-unionvalidator-variant-dispatch.md)
|
||||
|
||||
## Deferred / Blocked
|
||||
|
||||
+114
-77
@@ -1,12 +1,12 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-08-11
|
||||
status: accepted
|
||||
last_updated: 2026-08-15
|
||||
---
|
||||
|
||||
# alktype — Overview
|
||||
|
||||
The binary struct engine: a small Rust crate that takes a JSON Schema
|
||||
with `AlkType:*` custom keywords and produces an offset map, read/write
|
||||
The binary struct engine: a small Rust crate that takes a BAST (Binary
|
||||
Abstract Syntax Tree) document and produces an offset map, read/write
|
||||
functions, and validation — all driven by the schema. The schema is the
|
||||
format definition; the engine is generic.
|
||||
|
||||
@@ -16,29 +16,43 @@ Component details are in the sibling documents.
|
||||
|
||||
## What
|
||||
|
||||
`alktype` is a library crate that consumes JSON Schemas annotated
|
||||
with `AlkType:*` custom keywords (the same kinds defined in TypeBox's
|
||||
`typedef.ts`, plus `AlkType:Bytes`, `AlkType:Int64`, and `AlkType:Uint64`
|
||||
as alktype additions) and produces three capabilities:
|
||||
`alktype` is a library crate that consumes BAST documents and produces
|
||||
three capabilities:
|
||||
|
||||
1. **An offset map** — walks the schema, computes byte offsets for each
|
||||
field based on type sizes, field order, and alignment.
|
||||
1. **An offset map** — walks the BAST typed tree, computes byte offsets
|
||||
for each field based on type sizes, field order, and alignment.
|
||||
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
|
||||
read the field's bytes at its offset (zero-copy for fixed-size types).
|
||||
Given a `&mut [u8]` buffer, write a value at its offset.
|
||||
3. **Validation** — via `jsonschema` custom keywords, validates that a
|
||||
buffer's bytes match the schema's type constraints.
|
||||
3. **Validation** — two validators for two input types:
|
||||
- `validate_bytes(&[u8])` uses the BAST-native validator (a recursive
|
||||
walker over the BAST type tree) to check the value-domain
|
||||
constraints the materializer doesn't (integer ranges, `maxLength`,
|
||||
enum index bounds, union variant constraints).
|
||||
- `validate_json(&Value)` uses a standard `jsonschema::Validator`
|
||||
compiled from a consumer-provided JSON Schema (BAST is not involved
|
||||
— BAST describes bytes, not JSON shape).
|
||||
|
||||
The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||
`serde_json` (schema parsing). The novel code is the offset computation
|
||||
— a recursive walk of the schema JSON that computes byte positions for
|
||||
each field. The custom keyword implementations are small (a few lines
|
||||
each, generated from shared macros — see [validation.md](validation.md)).
|
||||
BAST is a JSON document that describes binary data layouts using a
|
||||
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
|
||||
itself a valid JSON Schema instance (it has a meta-schema), making it
|
||||
self-validating, editor-friendly, and trivially consumable from any
|
||||
language with a JSON parser. See [ADR-BAST](decisions/bast-bast-format.md)
|
||||
and [`bast-format.md`](bast-format.md).
|
||||
|
||||
The heavy lifting is done by the `jsonschema` crate (the
|
||||
`validate_json` path and BAST document meta-schema validation) and
|
||||
`serde_json` (BAST document parsing). The novel code is the offset
|
||||
computation — a recursive walk of the BAST typed tree that computes
|
||||
byte positions for each field — and the BAST-native validator — a flat
|
||||
recursive match over the same tree. See
|
||||
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md)
|
||||
(purpose/scope) and [ADR-BAST](decisions/bast-bast-format.md) (format).
|
||||
|
||||
The crate replaces two prior attempts that built their own jsonschema
|
||||
engines — typebox-rs (~8,400 lines) and the @alkdev/alktype prototype
|
||||
(~5,600 lines) — with `jsonschema` + an offset map + small custom keyword
|
||||
implementations. See
|
||||
(~5,600 lines) — with a BAST parser + an offset map + a BAST-native
|
||||
validator + the `jsonschema` crate for the JSON-validation path. See
|
||||
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
## Why
|
||||
@@ -48,21 +62,20 @@ read or write binary data at computed offsets. Instead of per-protocol
|
||||
serde structs (russh-sftp's 29 packet types), per-handler wire format
|
||||
code (TTY's 5-byte format parser), or per-format offset computation
|
||||
(metatensor's tensor access), all of these become instances of the same
|
||||
engine with different schemas.
|
||||
engine with different BAST documents.
|
||||
|
||||
The guiding insight:
|
||||
|
||||
> **The schema is the format.** A JSON Schema with `AlkType:Float32`,
|
||||
> `AlkType:Struct`, `AlkType:Union` etc. is both the validation spec and
|
||||
> the layout spec. No separate format definition, no separate parser, no
|
||||
> separate validator. One schema, three uses: validate, compute offsets,
|
||||
> access data.
|
||||
> **The schema is the format.** A BAST document is both the layout spec
|
||||
> and the validation spec for bytes. No separate format definition, no
|
||||
> separate parser, no separate validator. One schema, three uses:
|
||||
> validate, compute offsets, access data.
|
||||
|
||||
This is the convergence of three threads identified in the
|
||||
call-channels-unification research: the `typedef.ts` schema kinds from
|
||||
TypeBox, the russh-sftp protocol packets, and the metatensor format. The
|
||||
common pattern: a JSON Schema describes the shape of binary data, and
|
||||
the binary data is the struct's bytes at computed offsets.
|
||||
common pattern: a schema describes the shape of binary data, and the
|
||||
binary data is the struct's bytes at computed offsets.
|
||||
|
||||
The crate was bumped up in the timeline when the call-channels-unification
|
||||
research surfaced that channels, TTY, and the binary call protocol are
|
||||
@@ -75,41 +88,45 @@ read/write the binary payload."
|
||||
|
||||
## The "Schema Is the Format" Principle
|
||||
|
||||
A JSON Schema with `AlkType:*` custom keywords serves three roles
|
||||
simultaneously:
|
||||
A BAST document serves three roles simultaneously:
|
||||
|
||||
| Role | Mechanism | When |
|
||||
|------|-----------|------|
|
||||
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
|
||||
| **Validation spec (bytes)** | BAST-native validator (recursive walker over the BAST type tree) | Load time (parse typed tree), access time (`validate_bytes`) |
|
||||
| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) |
|
||||
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) |
|
||||
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
|
||||
|
||||
No separate format definition, no separate parser, no separate validator.
|
||||
The schema is the single source of truth for the binary format. Adding a
|
||||
new field to a protocol is adding a property to the schema JSON — the
|
||||
engine computes the new offsets automatically.
|
||||
No separate format definition, no separate parser, no separate
|
||||
validator. The BAST document is the single source of truth for the
|
||||
binary format. Adding a new field to a protocol is adding an entry to
|
||||
the BAST `fields` array — the engine computes the new offsets
|
||||
automatically.
|
||||
|
||||
This is the same principle as `#[repr(C)]` struct field access, but at
|
||||
runtime from a portable JSON Schema instead of at compile-time from
|
||||
language-specific annotations. The schema is the ABI contract.
|
||||
runtime from a portable JSON document instead of at compile-time from
|
||||
language-specific annotations. The BAST document is the ABI contract.
|
||||
|
||||
## Dependencies
|
||||
|
||||
```
|
||||
alktype
|
||||
├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support
|
||||
├── serde_json (with preserve_order) — schema parsing; field order is load-bearing
|
||||
└── (no tokio, no platform deps) — WASM-clean by construction
|
||||
├── jsonschema (v0.46, Draft 2020-12, default-features=false) — validate_json path + BAST meta-schema validation
|
||||
├── serde_json (with preserve_order) — BAST document parsing; mapping iteration order is load-bearing
|
||||
└── (no tokio, no platform deps) — WASM-clean by construction
|
||||
```
|
||||
|
||||
`alktype` is dependency-light: `jsonschema` + `serde_json` only.
|
||||
No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for
|
||||
browser use. The `jsonschema` crate is already in the workspace at
|
||||
`@alkdev/alknet: jsonschema/` — alktype is its first consumer.
|
||||
browser use. The `validate_bytes` path does not touch `jsonschema` for
|
||||
validation (it uses `jsonschema::ValidationError::custom` only for the
|
||||
error payload type, D-BAST-009) — a small wasm binary-size win.
|
||||
|
||||
`serde_json` requires the `preserve_order` feature because field order
|
||||
is load-bearing for binary layouts. The order of properties in the
|
||||
schema JSON determines the order of fields in the binary struct.
|
||||
`serde_json`'s `preserve_order` feature remains a dependency. Under
|
||||
BAST, struct field order is explicit (the `fields` array), so layout
|
||||
correctness no longer depends on it; but `mapping` iteration order and
|
||||
the `Definitions` `$defs` block order are still load-bearing for the
|
||||
parser's lazy resolution and the builder's output.
|
||||
|
||||
## Consumers
|
||||
|
||||
@@ -133,14 +150,17 @@ schema roles (binary layout + JSON payloads) from one library. See
|
||||
The russh-sftp case is the most instructive and the highest-value POC
|
||||
target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written
|
||||
dispatch on a type byte followed by serde deserialization. Under alktype,
|
||||
the dispatch is `TUnion` with a byte-offset discriminator — the schema
|
||||
says "byte 0 is the discriminator, bytes 1..N are the variant struct."
|
||||
The engine reads the discriminator, looks up the variant schema, computes
|
||||
offsets, reads fields. Same result, no per-packet-type code.
|
||||
the dispatch is a BAST `union` with a byte-offset discriminator — the
|
||||
schema says "byte 0 is the discriminator, bytes 1..N are the variant
|
||||
struct." The engine reads the discriminator, looks up the variant
|
||||
schema, computes offsets, reads fields. Same result, no per-packet-type
|
||||
code.
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
|
||||
These boundaries are decided in
|
||||
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md)
|
||||
and [ADR-BAST](decisions/bast-bast-format.md).
|
||||
|
||||
- **Not metatensor.** alktype is the binary struct *engine*. Metatensor
|
||||
is a *format* (8-byte header + JSON header + binary data) that uses the
|
||||
@@ -150,60 +170,73 @@ These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-js
|
||||
should not do anything that explicitly blocks adding a Value system
|
||||
later.
|
||||
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
|
||||
concern. The alktype engine consumes schemas; it does not generate them.
|
||||
concern. The alktype engine consumes BAST documents; it does not
|
||||
generate them.
|
||||
- **Schema builder is in scope as of v0.1.0.** A fluent Rust API for
|
||||
constructing schemas at runtime, producing `serde_json::Value`, is
|
||||
shipped in v0.1.0 ([ADR-009](decisions/009-builder-api.md), resolves
|
||||
OQ-003). The builder covers AlkType kinds and standard JSON Schema;
|
||||
see [builder.md](builder.md). Schemas may still be authored in
|
||||
TypeBox, generated by ujsx components, or hand-written — the builder
|
||||
is an additional construction path, not a replacement.
|
||||
constructing BAST documents and standard JSON Schemas at runtime,
|
||||
producing `serde_json::Value`, is shipped in v0.1.0
|
||||
([ADR-009](decisions/009-builder-api.md), resolves OQ-003). The
|
||||
builder covers BAST kinds (`struct_()`) and standard JSON Schema
|
||||
(`object()`); see [builder.md](builder.md). BAST documents may still
|
||||
be authored in TypeBox, generated by ujsx components, or hand-written
|
||||
— the builder is an additional construction path, not a replacement.
|
||||
- **Not a serialization framework.** The alktype engine is not a
|
||||
general-purpose serde replacement. It operates on raw byte buffers at
|
||||
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||
with a known schema, use alktype.
|
||||
computed offsets — no intermediate `Value` tree (except for the
|
||||
`validate_bytes` materialization step), no reflection, no dynamic
|
||||
dispatch per field. For JSON data, use serde. For binary data with a
|
||||
known schema, use alktype.
|
||||
- **Not a JSON-payload validator.** BAST describes bytes, not JSON
|
||||
shape. `validate_json` validates a JSON `Value` against a
|
||||
consumer-provided standard JSON Schema, not against the BAST document.
|
||||
See [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
|
||||
|
||||
## Architecture (component pointers)
|
||||
|
||||
- **[schema-layer.md](schema-layer.md)** — the 19 `AlkType:*` kinds,
|
||||
jsonschema custom keyword integration, TypeBox interop, schema
|
||||
annotations (endianness, alignment, encoding, TUnion discriminators).
|
||||
- **[schema-layer.md](schema-layer.md)** — the BAST parser (the typed
|
||||
surface every engine module walks), the 18 BAST kinds, the
|
||||
`AlkTypeKind` enum, and the foundational annotation types.
|
||||
- **[`bast-format.md`](bast-format.md)** — the normative BAST format
|
||||
specification (meta-schema, TypeRef, examples, validation model).
|
||||
- **[layout-engine.md](layout-engine.md)** — offset computation, the two
|
||||
layout modes (packed sequential vs aligned static), alignment,
|
||||
endianness, variable-length field handling.
|
||||
- **[data-access.md](data-access.md)** — read/write functions, TUnion
|
||||
dispatch, field paths, zero-copy access for fixed-size types,
|
||||
length-prefix reading for variable-length types.
|
||||
- **[validation.md](validation.md)** — custom keyword validators for all
|
||||
19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time
|
||||
validation, `AlkTypeEngine` as the compiled form of a schema.
|
||||
`validate_json` for JSON values; `validate_bytes` for binary buffers
|
||||
(ADR-010).
|
||||
- **[builder.md](builder.md)** — fluent Rust API for constructing
|
||||
alktype JSON Schemas at runtime, producing `serde_json::Value`.
|
||||
Covers AlkType kinds and standard JSON Schema (ADR-009).
|
||||
- **[validation.md](validation.md)** — the two-validator model
|
||||
(BAST-native for `validate_bytes`, standard `jsonschema` for
|
||||
`validate_json`), `AlkTypeError`, load-time vs access-time validation,
|
||||
`AlkTypeEngine` as the compiled form of a BAST document.
|
||||
- **[builder.md](builder.md)** — fluent Rust API for constructing BAST
|
||||
documents and standard JSON Schemas at runtime, producing
|
||||
`serde_json::Value`. Covers BAST kinds and standard JSON Schema
|
||||
(ADR-009, D-BAST-008).
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
|
||||
| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries (format-specific content superseded by ADR-BAST) |
|
||||
| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | BAST as the schema format; meta-schema, `$defs`/`$ref`/`kind` vocabulary; supersedes ADR-001's format-specific content |
|
||||
| Two layout modes | [ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats |
|
||||
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) |
|
||||
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
|
||||
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) — semantics carry forward; location moved to BAST type-level properties |
|
||||
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping (validation-strategy section refined by ADR-VAL-SPLIT) |
|
||||
| Two-validator model | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | BAST-native validator for `validate_bytes`; standard `jsonschema::Validator` for `validate_json`; D-BAST-006/007/009 |
|
||||
| Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (SFTP offsets, metatensor data_offsets) |
|
||||
| Non-final inline variable fields | [ADR-006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields) |
|
||||
| Packed-mode read factory | [ADR-007](decisions/007-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader |
|
||||
| TUnion in aligned mode | [ADR-008](decisions/008-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) |
|
||||
| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
|
||||
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate |
|
||||
| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 (output format amended to BAST / standard JSON Schema by ADR-BAST) |
|
||||
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate (validation step amended to the BAST-native validator by ADR-VAL-SPLIT) |
|
||||
|
||||
## Open Questions
|
||||
|
||||
See [open-questions.md](open-questions.md) for full details.
|
||||
|
||||
- **OQ-001** (deferred(scope)): Arrays of variable-length-element structs.
|
||||
- **OQ-001** (deferred(scope)): Arrays of variable-length-element
|
||||
structs — BAST arrays require `count` in v1 (D-BAST-004), aligning
|
||||
with this deferral.
|
||||
- **OQ-002** (deferred(scope)): `no_std` + `alloc` support.
|
||||
- **OQ-003** (resolved by [ADR-009](decisions/009-builder-api.md)):
|
||||
Builder API for schema construction. Shipped in v0.1.0; see
|
||||
@@ -223,8 +256,12 @@ See [open-questions.md](open-questions.md) for full details.
|
||||
- `@alkdev/alknet: alknet-typedef-poc/` — the POC code (disposable)
|
||||
- `@alkdev/alknet: typebox-rs/` — prior attempt, replaced by alktype
|
||||
- `@alkdev/alknet: alktype-prototype/` — prior attempt (the @alkdev/alktype prototype; not to be confused with this crate, which reuses the name but is backed by the `jsonschema` crate)
|
||||
- [BAST pivot research record](../research/bast-pivot.md) — motivation,
|
||||
POC scope and result, decisions D-BAST-001..009, risks
|
||||
- [BAST pivot implementation plan](../plans/bast-implementation.md) —
|
||||
ordered implementation steps, semver contract, ADR-sync checklist
|
||||
|
||||
> **Note**: The research findings, POC code, and prior-attempt paths above
|
||||
> refer to the parent `@alkdev/alknet` workspace where this crate originated.
|
||||
> They are preserved here as historical context for the architectural
|
||||
> decisions; the artifacts themselves are not part of this standalone repo.
|
||||
> decisions; the artifacts themselves are not part of this standalone repo.
|
||||
@@ -1,5 +1,19 @@
|
||||
# OQ-008: `UnionValidator` variant dispatch — validate variant fields against variant schema
|
||||
|
||||
> **Note (post-BAST-pivot):** The v0.1.0 resolution below —
|
||||
> `UnionValidator` + `inline_union_variant_refs` + custom-keyword
|
||||
> `jsonschema` sub-validators — was superseded by the BAST pivot. The
|
||||
> current implementation is the BAST-native validator
|
||||
> (`bast_validation::validate_union`), which reads `__discriminator`,
|
||||
> looks up the variant `BastType` in the union's `mapping`, and
|
||||
> recurses via `validate_typeref` into the variant's BAST definition.
|
||||
> Variant `$ref`s resolve lazily via `BastDoc::resolve_typeref` — no
|
||||
> `inline_union_variant_refs` compile step (removed under BAST). No
|
||||
> custom keywords, no `jsonschema` involvement on the bytes path. See
|
||||
> [ADR-VAL-SPLIT](../decisions/val-split-two-validator-model.md). The
|
||||
> v0.1.0 resolution text is preserved below as the historical record
|
||||
> of how the question was originally resolved.
|
||||
|
||||
- **Origin**: Raised during the v0.1.0 POC round 2 (SFTP Packet
|
||||
`validate_bytes` POC). Surfaced when the over-`maxLength` `Bytes`
|
||||
test failed: the materializer read the bytes correctly, but the
|
||||
|
||||
+214
-425
@@ -1,58 +1,60 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
status: accepted
|
||||
last_updated: 2026-08-15
|
||||
---
|
||||
|
||||
# alktype — Schema Layer
|
||||
|
||||
The schema layer: the 19 `AlkType:*` custom type kinds, their mapping to
|
||||
Rust types and byte sizes, the `jsonschema` custom keyword integration,
|
||||
TypeBox interop, and the concrete JSON shapes for schema-level
|
||||
annotations.
|
||||
The schema layer: the BAST (Binary Abstract Syntax Tree) format and the
|
||||
typed parser that the layout engines, materializer, and BAST-native
|
||||
validator walk. BAST replaces the v0.1.0 `AlkType:*` custom-keyword JSON
|
||||
Schema format decided in ADR-001; the pivot is recorded in
|
||||
[ADR-BAST](decisions/bast-bast-format.md) and grounded in
|
||||
[D-BAST-001..009](../research/bast-pivot.md#decisions).
|
||||
|
||||
## The 19 AlkType Kinds
|
||||
The **normative format specification** is
|
||||
[`bast-format.md`](bast-format.md) (meta-schema, TypeRef, examples,
|
||||
validation model). This document describes the *implementation* — the
|
||||
typed parser in `src/bast.rs` and the foundational `AlkTypeKind` enum
|
||||
in `src/schema.rs` — and points at the format spec for shape details.
|
||||
|
||||
These are the custom schema kinds defined in TypeBox's `typedef.ts`
|
||||
(`@alkdev/alknet: typebox/example/typedef/typedef.ts`, 619 lines) and
|
||||
ported to Rust via `jsonschema` custom keywords. Each kind carries binary
|
||||
layout semantics — a known byte size (for fixed-size types) or a known
|
||||
encoding strategy (for variable-length types).
|
||||
## The 18 BAST Kinds
|
||||
|
||||
| Kind | TypeBox key | Rust type | Size | Category |
|
||||
|------|-------------|-----------|------|----------|
|
||||
| `TFloat32` | `AlkType:Float32` | `f32` | 4 | fixed |
|
||||
| `TFloat64` | `AlkType:Float64` | `f64` | 8 | fixed |
|
||||
| `TInt8` | `AlkType:Int8` | `i8` | 1 | fixed |
|
||||
| `TInt16` | `AlkType:Int16` | `i16` | 2 | fixed |
|
||||
| `TInt32` | `AlkType:Int32` | `i32` | 4 | fixed |
|
||||
| `TInt64` | `AlkType:Int64` | `i64` | 8 | fixed |
|
||||
| `TUint8` | `AlkType:Uint8` | `u8` | 1 | fixed |
|
||||
| `TUint16` | `AlkType:Uint16` | `u16` | 2 | fixed |
|
||||
| `TUint32` | `AlkType:Uint32` | `u32` | 4 | fixed |
|
||||
| `TUint64` | `AlkType:Uint64` | `u64` | 8 | fixed |
|
||||
| `TBoolean` | `AlkType:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed |
|
||||
| `TString` | `AlkType:String` | length-prefixed UTF-8 | variable | variable |
|
||||
| `TBytes` | `AlkType:Bytes` | length-prefixed raw bytes | variable | variable |
|
||||
| `TStruct` | `AlkType:Struct` | record of fields | sum of field sizes | composite |
|
||||
| `TUnion` | `AlkType:Union` | tagged union | discriminator + variant | composite |
|
||||
| `TArray` | `AlkType:Array` | repeated element | count × element size | composite |
|
||||
| `TEnum` | `AlkType:Enum` | u32 index into enum values | 4 (fixed) | fixed |
|
||||
| `TRecord` | `AlkType:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
|
||||
| `TTimestamp` | `AlkType:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
|
||||
BAST uses lowercase `kind` strings (`"uint32"`, `"struct"`, `"union"`,
|
||||
etc.). The engine represents them as the `AlkTypeKind` Rust enum — one
|
||||
variant per kind — providing compile-time exhaustiveness checking and
|
||||
integer-discriminant dispatch (a jump table) instead of string
|
||||
comparison at every field access.
|
||||
|
||||
`AlkType:Int64` and `AlkType:Uint64` are alktype additions —
|
||||
TypeBox's `typedef.ts` tops out at 32-bit integers. They are required by
|
||||
the primary POC targets: SFTP `Read`/`Write` packets have `offset: u64`,
|
||||
and metatensor `data_offsets` are `u64`. See
|
||||
| BAST kind | `AlkTypeKind` | Rust type | Size | Category |
|
||||
|-----------|---------------|-----------|------|----------|
|
||||
| `int8` | `Int8` | `i8` | 1 | fixed |
|
||||
| `int16` | `Int16` | `i16` | 2 | fixed |
|
||||
| `int32` | `Int32` | `i32` | 4 | fixed |
|
||||
| `int64` | `Int64` | `i64` | 8 | fixed |
|
||||
| `uint8` | `Uint8` | `u8` | 1 | fixed |
|
||||
| `uint16` | `Uint16` | `u16` | 2 | fixed |
|
||||
| `uint32` | `Uint32` | `u32` | 4 | fixed |
|
||||
| `uint64` | `Uint64` | `u64` | 8 | fixed |
|
||||
| `float32` | `Float32` | `f32` | 4 | fixed |
|
||||
| `float64` | `Float64` | `f64` | 8 | fixed |
|
||||
| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed |
|
||||
| `string` | `String` | length-prefixed UTF-8 | variable | variable |
|
||||
| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable |
|
||||
| `struct` | `Struct` | record of fields | sum of field sizes | composite |
|
||||
| `union` | `Union` | tagged union | discriminator + variant | composite |
|
||||
| `array` | `Array` | repeated element | count × element size | composite |
|
||||
| `enum` | `Enum` | u32 index into enum values | 4 (fixed) | fixed |
|
||||
| `record` | `Record` | count-prefixed (key, value) pairs | variable | variable |
|
||||
|
||||
`int64`/`uint64` are alktype additions — TypeBox's `typedef.ts` tops
|
||||
out at 32-bit integers. Required by SFTP `Read`/`Write` `offset: u64`
|
||||
and metatensor `data_offsets`. See
|
||||
[ADR-005](decisions/005-int64-uint64-first-class-kinds.md).
|
||||
|
||||
### The `AlkTypeKind` enum
|
||||
|
||||
The engine represents the 19 kinds as a Rust enum — `AlkTypeKind` — with
|
||||
one variant per kind (`AlkTypeKind::Float32`, `AlkTypeKind::Struct`, etc.).
|
||||
The enum provides compile-time exhaustiveness checking and integer
|
||||
discriminant dispatch (a jump table) instead of string comparison at
|
||||
every field access. It is `pub` and re-exported from the crate root.
|
||||
`src/schema.rs` defines the enum:
|
||||
|
||||
```rust
|
||||
pub enum AlkTypeKind {
|
||||
@@ -60,7 +62,7 @@ pub enum AlkTypeKind {
|
||||
Uint8, Uint16, Uint32, Uint64,
|
||||
Float32, Float64,
|
||||
Boolean, Enum,
|
||||
String, Bytes, Timestamp,
|
||||
String, Bytes,
|
||||
Struct, Union, Array, Record,
|
||||
}
|
||||
```
|
||||
@@ -69,439 +71,226 @@ The enum carries the kind's binary-layout metadata as inherent methods:
|
||||
|
||||
| Method | Returns | Notes |
|
||||
|--------|---------|-------|
|
||||
| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"AlkType:Uint8"` |
|
||||
| `to_bast_str(self)` | `&'static str` | The lowercase BAST kind string (`"uint32"`) |
|
||||
| `from_bast_str(s)` | `Result<AlkTypeKind, AlkTypeError>` | Parses a lowercase BAST kind string; `AlkTypeError::Schema` for unknowns |
|
||||
| `type_size(self)` | `Option<usize>` | `Some(N)` for fixed-size kinds; `None` for variable/composite |
|
||||
| `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array |
|
||||
| `is_fixed_size(self)` | `bool` | True for the 12 fixed-size primitive kinds |
|
||||
| `is_composite(self)` | `bool` | True for Struct, Union, Array, Record |
|
||||
| `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record |
|
||||
| `is_variable_length(self)` | `bool` | True for String, Bytes, Record |
|
||||
| `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter |
|
||||
|
||||
`AlkTypeKind` implements `Display` (renders the keyword string) and
|
||||
`FromStr` (parses the keyword string back into the variant, returning
|
||||
`AlkTypeError::Schema` for unknown kinds). The layout engines and the
|
||||
validator dispatch on the enum, not on strings.
|
||||
`AlkTypeKind` implements `Display`, backed by `to_bast_str` so the
|
||||
layout engines, materializer, validator, and parser surface the
|
||||
canonical BAST name in error messages. `from_bast_str` is the inverse
|
||||
and is the dispatch point the BAST parser uses to map a `kind` string
|
||||
to the enum variant (D-BAST-002).
|
||||
|
||||
### Fixed-size types
|
||||
### Foundational annotation types
|
||||
|
||||
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
|
||||
`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset
|
||||
computation uses these sizes directly. Read/write is zero-copy pointer
|
||||
cast for these types.
|
||||
|
||||
**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other
|
||||
values are invalid and produce a `AlkTypeError::Access` on read.
|
||||
|
||||
**`TEnum` binary representation:** A `u32` index into the enum's declared
|
||||
values, in declaration order. The first declared value is index 0, the
|
||||
second is index 1, etc. The enum's values are declared via the standard
|
||||
JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`).
|
||||
The `AlkType:Enum` custom keyword signals that the type is an enum for
|
||||
layout purposes; the built-in `enum` keyword provides the value list.
|
||||
|
||||
**Design note:** TypeBox's `TEnum` is a string enum (variable-length). The
|
||||
alktype engine uses a `u32` index instead — a deliberate deviation from
|
||||
TypeBox fidelity in favor of binary efficiency. Most enums have a small
|
||||
number of variants (e.g., the call protocol's 5 event types); a `u32`
|
||||
index is compact, fixed-size, and sufficient for any realistic enum. The
|
||||
JSON representation (for validation) remains a string; the binary
|
||||
representation is the `u32` index.
|
||||
The `u32` index follows the schema's endianness annotation (ADR-003), like
|
||||
all other fixed-size types. In little-endian mode the index is
|
||||
`u32::from_le_bytes`; in big-endian mode it is `u32::from_be_bytes`.
|
||||
|
||||
### Variable-length types
|
||||
|
||||
`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes.
|
||||
The alktype engine supports three strategies for handling variable-length
|
||||
types in binary layouts, selected by the `encoding` annotation and the
|
||||
standard JSON Schema `maxLength` keyword:
|
||||
|
||||
| Strategy | Encoding annotation | Layout behavior | Use case |
|
||||
|----------|-------------------|-----------------|----------|
|
||||
| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) |
|
||||
| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) |
|
||||
| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) |
|
||||
|
||||
**Strategy 1: Inline length-prefixing (default).** The field's fixed
|
||||
portion is a 4-byte length prefix at a computed offset. The variable data
|
||||
follows immediately after. In packed sequential mode, the length prefix
|
||||
determines the position of subsequent fields. In aligned static mode, the
|
||||
length prefix is at a known offset; the variable data is not included in
|
||||
the static layout. This is the universal pattern used by channels, SFTP,
|
||||
TTY, and most binary protocols.
|
||||
|
||||
**Strategy 2: Fixed-size reservation.** When a variable-length field
|
||||
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
|
||||
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
|
||||
than `maxLength` is zero-padded; data longer than `maxLength` is a
|
||||
validation error. This makes the field fixed-size from the layout
|
||||
perspective — subsequent fields have known, unchanging offsets. This is
|
||||
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
|
||||
pattern for fields with known maximum sizes.
|
||||
|
||||
In packed sequential mode, `maxLength` is a validation constraint only —
|
||||
the engine still uses inline length-prefixing (strategy 1) because
|
||||
protocols don't benefit from fixed-size reservation.
|
||||
|
||||
**Strategy 3: Offset indirection.** The field is a struct
|
||||
`{offset: u32, length: u32}` at a known position. The consumer provides
|
||||
the data region separately; the engine reads the offset and length, then
|
||||
slices the data region. This is the metatensor blob tensor pattern — the
|
||||
index struct lives in one region, the blob data lives in another. Enables
|
||||
mmap-friendly random access to variable-length data without parsing
|
||||
length prefixes and without reserving worst-case space.
|
||||
|
||||
**Default strategy selection:**
|
||||
- In packed sequential mode: always strategy 1 (inline length-prefixing).
|
||||
`maxLength` is a validation constraint only.
|
||||
- In aligned static mode: strategy 2 (fixed-size reservation) if
|
||||
`maxLength` is declared; strategy 3 (offset indirection) if
|
||||
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
|
||||
length-prefixing) otherwise.
|
||||
|
||||
**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3)
|
||||
respects the schema's `"endian"` annotation (ADR-003). In little-endian
|
||||
mode, the length is `u32::from_le_bytes`. In big-endian mode, the length
|
||||
is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have
|
||||
consistent byte order for both field values and length prefixes.
|
||||
|
||||
**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`.
|
||||
Otherwise identical to `TString` in layout (same three strategies).
|
||||
|
||||
**Design note:** `AlkType:Bytes` is an alktype addition — it does
|
||||
not exist in TypeBox's `typedef.ts` (which defines 16 kinds). It is
|
||||
included because raw byte arrays are a common binary protocol primitive
|
||||
(SFTP data payloads, channels payloads, tensor data) and are semantically
|
||||
distinct from UTF-8 strings. In the binary representation, TBytes is raw
|
||||
bytes with no encoding (not base64, not hex). In the JSON representation
|
||||
(for validation), TBytes is a string (JSON has no native byte type).
|
||||
|
||||
**`TRecord`:** A string-keyed map. The value type is declared via the
|
||||
schema's `"values"` property (e.g., `"values": { "AlkType:Float32": true }`).
|
||||
Binary layout is a count-prefixed sequence of `(key, value)` pairs:
|
||||
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
|
||||
The count is the number of entries. Each key is a length-prefixed UTF-8
|
||||
string. Each value is encoded according to its declared `AlkType:*` kind
|
||||
— a `Record<Uint32>` value is 4 raw bytes; a `Record<String>` value is
|
||||
itself a length-prefixed string; a `Record<Struct>` value is the struct's
|
||||
fields laid out inline. There is **no separate `value_len` prefix** —
|
||||
the value's size is determined by its kind (fixed-size kinds have a
|
||||
known size; variable-length kinds carry their own length prefix). The
|
||||
count and key-length prefixes respect the schema's endianness. In
|
||||
aligned static mode with `maxLength`, the entire record is reserved at
|
||||
`maxLength` bytes (zero-padded).
|
||||
|
||||
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
|
||||
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
|
||||
fixed-size reservation (strategy 2 with `maxLength`). The data-access
|
||||
layer treats timestamps as opaque length-prefixed strings — it does not
|
||||
parse or validate the timestamp format. The jsonschema custom keyword
|
||||
validator checks RFC 3339 conformance at the JSON level (see
|
||||
[validation.md](validation.md)).
|
||||
|
||||
`TArray` is variable-length when the element type is variable-length or
|
||||
when the count is not known at schema time. For fixed-size element arrays
|
||||
with a known count, the size is `element_size × count`.
|
||||
|
||||
**`TArray` count declaration:** The array count is declared via the
|
||||
standard JSON Schema `"minItems"` and `"maxItems"` keywords. When
|
||||
`minItems == maxItems`, the array has a fixed count known at schema time.
|
||||
When they differ or are absent, the count is variable and the array uses
|
||||
a length-prefixed encoding: `[count: u32][element_0]...[element_N]`.
|
||||
The count prefix respects the schema's endianness.
|
||||
|
||||
### Composite types
|
||||
|
||||
`TStruct` and `TUnion` are composite — their size is the sum of their
|
||||
fields' sizes (plus alignment padding in aligned static mode). The offset
|
||||
computation recurses into their properties.
|
||||
|
||||
## Schema-Layer Public API
|
||||
|
||||
The `schema` module exposes the foundational types and functions every
|
||||
other module depends on. These are re-exported from the crate root.
|
||||
|
||||
### `get_alktype_kind` vs `get_alktype_kind_loose`
|
||||
|
||||
The engine recognizes a `AlkType:*` kind on a schema node two ways,
|
||||
because the keyword value may be either a boolean (`true`) or an
|
||||
annotation object (`{ "encoding": "..." }`):
|
||||
|
||||
| Function | Recognizes | Returns |
|
||||
|----------|------------|---------|
|
||||
| `get_alktype_kind(node) -> Option<&str>` | Boolean form only (`{ "AlkType:String": true }`) | The keyword string, e.g. `"AlkType:String"` |
|
||||
| `get_alktype_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string |
|
||||
| `get_alktype_kind_enum(node) -> Option<AlkTypeKind>` | Boolean form only | The parsed enum variant |
|
||||
| `get_alktype_kind_loose_enum(node) -> Option<AlkTypeKind>` | Boolean form **and** object form | The parsed enum variant |
|
||||
|
||||
The boolean-form-only functions are used by the validator factories
|
||||
(which reject the object form as a schema error) and the top-level
|
||||
kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new`
|
||||
(which require `AlkType:Struct` at the root). The "loose" variants are
|
||||
used by the layout engines during field traversal, so that a variable-
|
||||
length field with an `encoding` annotation (`{ "AlkType:String":
|
||||
{ "encoding": "offset-indirect" } }`) is still recognized as a `String`.
|
||||
|
||||
### Annotation parsers
|
||||
|
||||
Each schema-level annotation has a dedicated parser that reads it from a
|
||||
`serde_json::Value` node and returns a sensible default when absent:
|
||||
|
||||
| Function | Annotation | Default |
|
||||
|----------|------------|---------|
|
||||
| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` |
|
||||
| `parse_align(node) -> Option<usize>` | `"align"` | `None` |
|
||||
| `parse_max_length(node) -> Option<usize>` | `"maxLength"` | `None` |
|
||||
| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` |
|
||||
| `parse_discriminator(node) -> Result<DiscriminatorKind, AlkTypeError>` | `"discriminator"` | (required — returns `AlkTypeError::Schema` if absent) |
|
||||
|
||||
### Public enums
|
||||
`src/schema.rs` also defines the two annotation enums (semantics
|
||||
unchanged from ADR-003; only their *location* in the document moved —
|
||||
see [ADR-BAST](decisions/bast-bast-format.md) and
|
||||
[bast-format.md §Variable-Length Encoding](bast-format.md#variable-length-encoding)):
|
||||
|
||||
```rust
|
||||
pub enum Endian { Little, Big }
|
||||
pub enum VariableEncoding { LengthPrefixed, OffsetIndirect }
|
||||
pub enum DiscriminatorKind {
|
||||
Byte { offset: usize, disc_type: AlkTypeKind },
|
||||
Field { name: String },
|
||||
}
|
||||
```
|
||||
|
||||
`DiscriminatorKind::Byte` carries the byte position (`offset`) and the
|
||||
discriminator's `AlkType:*` kind (`disc_type`, restricted to `Uint8`/
|
||||
`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator
|
||||
field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for
|
||||
how these drive dispatch.
|
||||
The `Discriminator` builder enum lives in
|
||||
[`src/builder.rs`](../../src/builder.rs) (the builder's domain); the BAST
|
||||
parser's typed discriminator view is
|
||||
[`BastDiscriminator`](#the-bast-parser-bast-module).
|
||||
|
||||
### `$ref` resolution and normalization
|
||||
## The BAST Parser (`bast` module)
|
||||
|
||||
| Function | Purpose |
|
||||
|----------|---------|
|
||||
| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/<name>"`. Idempotent. Runs once at `AlkTypeEngine::compile` time. |
|
||||
| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. |
|
||||
| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). |
|
||||
`src/bast.rs` is the typed surface over a BAST document. Three
|
||||
consumers walk the same tree — the layout engines
|
||||
([`offset_map`](layout-engine.md), [`layout_builder`](layout-engine.md),
|
||||
[`sequential_reader`](layout-engine.md)), the
|
||||
[`materialize`](data-access.md) layer, and the
|
||||
[`bast_validation`](validation.md) validator — so a typed view pays for
|
||||
itself: each walks matched arms over `BastType` instead of re-parsing
|
||||
raw JSON at every node. Borrowing (not cloning) the source
|
||||
`serde_json::Value` keeps the parse allocation-free beyond the small
|
||||
typed nodes themselves.
|
||||
|
||||
`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s
|
||||
JSON Pointer requirement. The layout engines call `resolve_ref_or_inline`
|
||||
on every `$ref`-bearing node they encounter during traversal.
|
||||
### Document shape
|
||||
|
||||
## jsonschema Custom Keyword Integration
|
||||
Every BAST document has the same top-level shape:
|
||||
|
||||
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
|
||||
via the `with_keyword` API. Each `AlkType:*` kind is registered as a
|
||||
custom keyword:
|
||||
|
||||
```rust
|
||||
let validator = jsonschema::options()
|
||||
.with_keyword("AlkType:Float32", factory)
|
||||
.with_keyword("AlkType:Int32", factory)
|
||||
.with_keyword("AlkType:Struct", factory)
|
||||
// ... all 19 kinds
|
||||
.build(&schema)?;
|
||||
```json
|
||||
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
|
||||
```
|
||||
|
||||
The factory closure receives the parent schema object, the keyword's
|
||||
value, and the schema path — enabling cross-keyword awareness. The
|
||||
`AlkType:Struct` validator, for example, inspects the parent's
|
||||
`properties` to validate each field against its declared `AlkType:*` kind.
|
||||
- The `$defs` block is **required** (D-BAST-003). Single-type documents
|
||||
are a special case with one entry.
|
||||
- The **root type name** is a required parameter to
|
||||
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001).
|
||||
It selects which `$defs` entry is the top-level type; convention
|
||||
(first entry) is fragile and depends on JSON key order, so an
|
||||
explicit parameter is used instead.
|
||||
|
||||
Each custom keyword implementation is ~10 lines. The `jsonschema` crate
|
||||
handles all structural validation (object properties, required fields,
|
||||
array items, enum values) — the custom keywords only need to validate
|
||||
the leaf type constraints. See [validation.md](validation.md) for the
|
||||
validator implementations.
|
||||
See [`bast-format.md`](bast-format.md) for the normative TypeDef shapes
|
||||
(Struct, Union, Enum, FieldDef, TypeRef) and the meta-schema.
|
||||
|
||||
This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side.
|
||||
Same semantics, different language, same JSON Schema wire format. A
|
||||
TypeBox schema serialized to JSON feeds into the alktype engine after a
|
||||
single pre-processing step: normalizing `$ref` values (see below).
|
||||
### Typed tree
|
||||
|
||||
## TypeBox Interop
|
||||
The parser produces a borrowed typed tree:
|
||||
|
||||
TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox
|
||||
schema like:
|
||||
| Type | Role |
|
||||
|------|------|
|
||||
| `BastDoc<'a>` | The parsed document: the root `Value`, the chosen root name, and the parsed root `BastDef`. Entry point via `BastDoc::new(root, root_name)`. |
|
||||
| `BastDef<'a>` | A named `$defs` entry — `{ name, kind: BastDefKind, source }`. Only `struct`/`union`/`enum` can live at the top level. |
|
||||
| `BastDefKind<'a>` | `Struct(BastStruct)` / `Union(BastUnion)` / `Enum(BastEnum)`. |
|
||||
| `BastStruct<'a>` | `{ endian, align, fields: Vec<BastField>, source }`. Field order is the `fields` array order (BAST design principle #4 — no reliance on `serde_json`'s `preserve_order`). |
|
||||
| `BastField<'a>` | `{ name, ty: BastType, endian, align, encoding, max_length, source }`. Annotations are field-level properties (ADR-003 semantics, BAST location). |
|
||||
| `BastUnion<'a>` | `{ endian, discriminator, fields, mapping: Vec<(key, BastType)>, source }`. Variant `$ref`s resolve **lazily** — no compile-time inlining. |
|
||||
| `BastDiscriminator<'a>` | `Byte { offset, disc_type }` / `Field { name }`. The typed view of the `discriminator` object. |
|
||||
| `BastEnum<'a>` | `{ values: Vec<&'a str>, source }`. Non-empty (enforced). |
|
||||
| `BastType<'a>` | A TypeRef — `Primitive(AlkTypeKind)` / `Ref(BastRef)` / `Array(BastArray)` / `Record(BastRecord)` / `Struct(...)` / `Union(...)` / `Enum(...)`. The central mechanism for typing fields, array elements, record values, and union variants. |
|
||||
| `BastRef<'a>` | A `$ref` restricted to `#/$defs/<name>`. Carries just the name. |
|
||||
| `BastArray<'a>` | `{ element: Box<BastType>, count, source }`. `count` is required in v1 (D-BAST-004). |
|
||||
| `BastRecord<'a>` | `{ values: Box<BastType>, source }`. |
|
||||
|
||||
```typescript
|
||||
const TensorRef = Type.Object({
|
||||
dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]),
|
||||
shape: Type.Array(Type.Number()),
|
||||
data_offsets: Type.Tuple([Type.Number(), Type.Number()])
|
||||
});
|
||||
```
|
||||
All of these are re-exported from the crate root (`pub use bast::{...}`
|
||||
in `src/lib.rs`).
|
||||
|
||||
serialized to JSON is a standard JSON Schema with `type: "object"`,
|
||||
`properties`, and `required`. That JSON feeds into the alktype engine
|
||||
after `$ref` normalization. The `AlkType:*` custom keywords are added by
|
||||
TypeBox's `TypeRegistry.Set` — they appear in the serialized JSON as
|
||||
additional properties on the schema object.
|
||||
### `$ref` resolution
|
||||
|
||||
### `$ref` normalization
|
||||
BAST `$ref`s are always full JSON Pointers restricted to
|
||||
`#/$defs/<name>` — no external references, no fragment-only pointers,
|
||||
no bare names (rejected by the parser). The restriction keeps
|
||||
resolution a single hash lookup and eliminates the v0.1.0
|
||||
`normalize_refs` pass that rewrote TypeBox's bare-name refs.
|
||||
|
||||
TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`),
|
||||
referencing sibling definitions within the same `$defs` block. The
|
||||
`jsonschema` crate requires full JSON Pointer paths (e.g.,
|
||||
`"$ref": "#/$defs/Read"`). The alktype engine normalizes TypeBox-style
|
||||
refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization)
|
||||
— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to
|
||||
`"#/$defs/<name>"`. The normalization is idempotent — full JSON Pointer
|
||||
refs pass through unchanged. It runs once at `AlkTypeEngine::compile`
|
||||
time, before the schema is passed to `jsonschema` or the offset
|
||||
computation.
|
||||
`BastDoc` exposes three resolution helpers:
|
||||
|
||||
**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs
|
||||
with `Resource 'Read' is not present in a registry`. Full JSON Pointer
|
||||
refs (`#/$defs/Read`) resolve correctly. The normalization step bridges
|
||||
the gap between TypeBox's output and jsonschema's input.
|
||||
| Method | Purpose |
|
||||
|--------|---------|
|
||||
| `lookup_def(name) -> Result<&'a Value, AlkTypeError>` | The single hash lookup into `$defs`. |
|
||||
| `resolve_ref(r: &BastRef) -> Result<BastDef, AlkTypeError>` | Resolve a `BastRef` to its `BastDef`. |
|
||||
| `resolve_typeref(ty: &BastType) -> Result<BastType, AlkTypeError>` | Deref one `$ref` level, or return the inline type unchanged. The composite-walkers call this. |
|
||||
| `resolve_typeref_as_def(ty, path) -> Result<BastDef, AlkTypeError>` | Resolve a `BastType` to a `BastDef`, wrapping inline composites in a synthetic def. Convenient for the validator/materializer. |
|
||||
|
||||
The alktype engine does not depend on TypeBox or any JS toolchain. It
|
||||
consumes JSON — whether that JSON was authored in TypeBox, generated by
|
||||
a ujsx component, or hand-written. The schema is the interface.
|
||||
Variant `$ref`s (union `mapping` entries) are resolved **lazily** by
|
||||
the materializer and validator via these helpers — no
|
||||
`inline_union_variant_refs` compile step (removed under BAST). The
|
||||
parser only records the `BastRef` target name.
|
||||
|
||||
### Untrusted input
|
||||
|
||||
Every path that walks a BAST document returns
|
||||
`Err(AlkTypeError::Schema)` on a malformed document, never
|
||||
`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream
|
||||
`alkcall` consumer accepts schemas from arbitrary internet peers in its
|
||||
hub/spoke topology). Overflow-safe arithmetic (`checked_add`,
|
||||
`usize::try_from`) is used for any offset/count cast (AGENTS.md §4).
|
||||
|
||||
A malformed document (missing `$defs`, missing `kind`, unknown kind
|
||||
string, dangling `$ref`, empty `mapping`, non-struct/union/enum at the
|
||||
top level, a field-name union without a `fields` array, etc.) surfaces
|
||||
as `AlkTypeError::Schema` with a path-annotated message.
|
||||
|
||||
### What the parser does *not* do
|
||||
|
||||
- **No meta-schema validation.** `BastDoc::new` parses structurally
|
||||
(every reachable def parses to a typed `BastDef`) but does not run
|
||||
the BAST meta-schema. Consumers that want full structural validation
|
||||
can run the meta-schema via `jsonschema` directly
|
||||
([`BAST_META_SCHEMA`](bast-format.md#the-meta-schema) is re-exported
|
||||
from the crate root). The parser's structural checks catch the cases
|
||||
that matter for layout/materialize/validate; the meta-schema is the
|
||||
authoritative well-formedness check.
|
||||
- **No eager full-document parse.** Only the root definition and the
|
||||
definitions it (transitively) references are parsed eagerly; orphan
|
||||
`$defs` entries are not checked. Lazy `$ref` resolution reaches the
|
||||
rest at access time.
|
||||
- **No annotation interpretation.** The parser *records* `endian`/
|
||||
`align`/`encoding`/`maxLength` on `BastField`/`BastStruct`; the
|
||||
layout engines and validator *interpret* them (ADR-003 semantics).
|
||||
|
||||
## Schema Annotations
|
||||
|
||||
Schema-level annotations control binary layout behavior. These are
|
||||
decided in [ADR-003](decisions/003-schema-annotations.md).
|
||||
Annotation *semantics* carry forward unchanged from ADR-003; only their
|
||||
*location* moved from v0.1.0's custom-keyword objects to BAST
|
||||
type-level properties. The concrete BAST shapes are in
|
||||
[`bast-format.md`](bast-format.md):
|
||||
|
||||
### Endianness
|
||||
- [Endianness](bast-format.md#endianness) — struct/union-level `endian`
|
||||
with field-level override.
|
||||
- [Alignment](bast-format.md#alignment) — struct/field-level `align`
|
||||
(aligned mode only).
|
||||
- [Variable-length encoding](bast-format.md#variable-length-encoding) —
|
||||
field-level `encoding` and `maxLength`.
|
||||
- [Union discriminators](bast-format.md#union) — `discriminator` object
|
||||
on the union def (`byte` or `field`).
|
||||
|
||||
Schema-level annotation with a default of little-endian:
|
||||
The `maxLength` keyword is *not* a BAST invention — it is the standard
|
||||
JSON Schema `maxLength`, repurposed as a byte-length cap. In aligned
|
||||
mode it reserves a fixed-size slot; in packed mode it is a validation
|
||||
constraint only. It applies to `string`/`bytes` fields only: the parser
|
||||
rejects it on any other kind (review #006 N3 — elsewhere it was
|
||||
silently unenforced), and in aligned mode a record reservation was
|
||||
silently corrupt (review #006 M5). See [bast-format.md §Variable-Length
|
||||
Encoding](bast-format.md#variable-length-encoding) and
|
||||
[ADR-003](decisions/003-schema-annotations.md).
|
||||
|
||||
```json
|
||||
{ "AlkType:Struct": true, "endian": "big", "properties": { ... } }
|
||||
```
|
||||
## What Was Removed
|
||||
|
||||
- `"endian": "little"` (default) — read/write in little-endian byte order.
|
||||
- `"endian": "big"` — read/write in big-endian byte order.
|
||||
- Applies to the entire schema and all nested types.
|
||||
The v0.1.0 custom-keyword accessor layer was removed in step 8 of the
|
||||
BAST pivot. The `schema` module retains only the foundational types
|
||||
(`AlkTypeKind`, `Endian`, `VariableEncoding`, shared constants); the
|
||||
BAST parser is the typed surface every engine module walks. Removed:
|
||||
|
||||
### Alignment
|
||||
- `get_alktype_kind` / `get_alktype_kind_enum` / `get_alktype_kind_loose`
|
||||
/ `get_alktype_kind_loose_enum` — superseded by the parser's
|
||||
`kind`-string dispatch.
|
||||
- `normalize_refs` / `inline_union_variant_refs` (+ recursive helpers)
|
||||
— BAST refs are always `#/$defs/<name>`; one hash lookup, variant
|
||||
refs resolve lazily.
|
||||
- `parse_encoding` / `parse_align` / `parse_max_length` / `parse_endian`
|
||||
— `bast.rs` has its own BAST-property-form copies (internal to the
|
||||
parser).
|
||||
- `parse_discriminator` + `DiscriminatorKind` — replaced by
|
||||
`bast::BastDiscriminator`; the builder has its own `Discriminator`
|
||||
enum.
|
||||
- `resolve_ref` / `resolve_ref_or_inline` — replaced by
|
||||
`BastDoc::lookup_def` / `resolve_typeref`.
|
||||
- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`,
|
||||
`BYTE_DISCRIMINATOR_TYPES`, and their unit tests.
|
||||
|
||||
Both struct-level and field-level, with field-level overriding:
|
||||
|
||||
```json
|
||||
{
|
||||
"AlkType:Struct": true,
|
||||
"align": 256,
|
||||
"properties": {
|
||||
"weight": { "AlkType:Float32": true, "align": 16 }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- Struct-level `"align"` sets the default for all fields.
|
||||
- Field-level `"align"` overrides the struct default.
|
||||
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/
|
||||
enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix),
|
||||
1 for struct/union/array.
|
||||
- Only meaningful in aligned static mode (ADR-002). Ignored in packed
|
||||
sequential mode.
|
||||
|
||||
### Variable-length encoding
|
||||
|
||||
The alktype engine supports three strategies for variable-length types
|
||||
(see §Variable-length types above for full details). The strategy is
|
||||
selected by the `encoding` annotation and the standard JSON Schema
|
||||
`maxLength` keyword:
|
||||
|
||||
```json
|
||||
// Strategy 1: Inline length-prefixing (default, shorthand)
|
||||
{ "AlkType:String": true }
|
||||
|
||||
// Strategy 1: Explicit inline length-prefixing
|
||||
{ "AlkType:String": { "encoding": "length-prefixed" } }
|
||||
|
||||
// Strategy 2: Fixed-size reservation (uses standard maxLength)
|
||||
{ "AlkType:String": true, "maxLength": 256 }
|
||||
|
||||
// Strategy 3: Offset indirection (opt-in)
|
||||
{ "AlkType:String": { "encoding": "offset-indirect" } }
|
||||
```
|
||||
|
||||
- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at
|
||||
computed offset, variable data follows immediately. Used by protocol
|
||||
wire formats.
|
||||
- `maxLength` (standard JSON Schema keyword) — in aligned static mode,
|
||||
reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the
|
||||
field fixed-size from the layout perspective. In packed sequential
|
||||
mode, `maxLength` is a validation constraint only.
|
||||
- `"encoding": "offset-indirect"` — field is a struct
|
||||
`{offset: u32, length: u32}` pointing into a separate data region.
|
||||
The consumer provides the data region separately. Used by metatensor
|
||||
blob tensors.
|
||||
- Applies to all variable-length types: `AlkType:String`, `AlkType:Bytes`,
|
||||
`AlkType:Array`, `AlkType:Record`, `AlkType:Timestamp`.
|
||||
|
||||
### TUnion discriminators
|
||||
|
||||
Two discriminator kinds: byte-offset (protocol dispatch) and field-name
|
||||
(typedef.ts pattern).
|
||||
|
||||
**Byte-offset discriminator** (SFTP type bytes, call protocol event types):
|
||||
|
||||
```json
|
||||
{
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "AlkType:Uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" },
|
||||
"101": { "$ref": "#/$defs/Status" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `"offset"` — byte position of the discriminator.
|
||||
- `"type"` — the `AlkType:*` kind of the discriminator (typically
|
||||
`AlkType:Uint8`).
|
||||
- Mapping keys are stringified integers. The variant struct starts at
|
||||
`offset + discriminator_size`.
|
||||
|
||||
**Field-name discriminator** (typedef.ts pattern):
|
||||
|
||||
```json
|
||||
{
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "field",
|
||||
"name": "type"
|
||||
},
|
||||
"mapping": {
|
||||
"read": { "$ref": "#/$defs/Read" },
|
||||
"write": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `"name"` — the field name holding the discriminator value.
|
||||
- Mapping keys are string values matching the discriminator field's value.
|
||||
- The discriminator field is just another field in the struct.
|
||||
|
||||
Mapping values may be either inline schemas or `$ref` pointers. Both work.
|
||||
See [ADR-BAST](decisions/bast-bast-format.md) §"What is removed" and
|
||||
[`bast-format.md` §What is removed](bast-format.md#what-is-removed).
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators |
|
||||
| BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary | [ADR-BAST](decisions/bast-bast-format.md) | Supersedes ADR-001's format-specific content; records D-BAST-001..009 |
|
||||
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Annotation semantics (carry forward unchanged; only location moves) |
|
||||
| Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (required by SFTP offsets and metatensor data_offsets) |
|
||||
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle |
|
||||
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle (format-specific content superseded by ADR-BAST) |
|
||||
|
||||
## Open Questions
|
||||
|
||||
See [open-questions.md](open-questions.md) for full details.
|
||||
|
||||
- **OQ-003** (deferred(scope)): Builder API for schema construction.
|
||||
- **OQ-001** (deferred(scope)): Arrays of variable-length-element
|
||||
structs — BAST arrays require `count` in v1 (D-BAST-004), aligning
|
||||
with this deferral.
|
||||
|
||||
## References
|
||||
|
||||
- `@alkdev/alknet: typebox/example/typedef/typedef.ts` — the TypeBox
|
||||
schema kinds (619 lines)
|
||||
- `@alkdev/alknet: jsonschema/` — the jsonschema crate (v0.46.5, Draft
|
||||
2020-12)
|
||||
- [ADR-003](decisions/003-schema-annotations.md) — schema
|
||||
annotation shapes
|
||||
- [validation.md](validation.md) — custom keyword validator implementations
|
||||
- [`bast-format.md`](bast-format.md) — the normative BAST format
|
||||
specification (meta-schema, TypeRef, examples, validation model)
|
||||
- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format decision
|
||||
- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics
|
||||
- [BAST pivot research record](../research/bast-pivot.md) — motivation,
|
||||
POC scope and result, decisions D-BAST-001..009
|
||||
- [validation.md](validation.md) — the BAST-native validator and the
|
||||
`validate_json` JSON-Schema path
|
||||
- [`src/bast.rs`](../../src/bast.rs) — the parser implementation
|
||||
- [`src/schema.rs`](../../src/schema.rs) — the `AlkTypeKind` enum and
|
||||
foundational annotation types
|
||||
+304
-288
@@ -1,63 +1,163 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-08-11
|
||||
status: accepted
|
||||
last_updated: 2026-08-31
|
||||
---
|
||||
|
||||
# alktype — Validation
|
||||
|
||||
The validation layer: custom keyword validators for all 19 `AlkType:*`
|
||||
kinds, the `AlkTypeError` enum, load-time vs access-time validation
|
||||
strategy, and the `AlkTypeEngine` as the compiled form of a schema.
|
||||
The validation layer: two validators for two input types, the
|
||||
`AlkTypeError` enum, the load-time-build / access-time-check strategy,
|
||||
and the `AlkTypeEngine` as the compiled form of a BAST document.
|
||||
|
||||
## The Validator Split
|
||||
|
||||
BAST separates two concerns that the v0.1.0 format conflated, and in
|
||||
doing so reveals that the engine has **two distinct validation paths**
|
||||
with different inputs and guarantees. This is the validator split,
|
||||
decided in [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model),
|
||||
[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model),
|
||||
and
|
||||
[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape),
|
||||
and recorded in [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
|
||||
|
||||
| Path | Input | Validator | Schema source |
|
||||
|------|-------|-----------|---------------|
|
||||
| `validate_bytes(&[u8])` | Raw bytes | Compiled `ValidationPlan` walk | The BAST document (binary layout + value constraints) |
|
||||
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
|
||||
|
||||
### `validate_bytes` — bytes in, BAST is the validator
|
||||
|
||||
The materializer produces a `serde_json::Value` tree from bytes
|
||||
(walking the layout engine). By construction, this `Value` is
|
||||
*structurally correct*: all declared fields are present (the
|
||||
materializer iterates the field list), types are correct (`read_u32`
|
||||
produces `Value::Number`), bounds are checked (via
|
||||
`data_access::check_bounds`), UTF-8 is valid (via `from_utf8`), the
|
||||
discriminator is in the mapping, and the boolean byte is 0 or 1.
|
||||
|
||||
What the materializer does NOT check — and what the validation half
|
||||
checks afterward — are **value-domain constraints expressed in the BAST
|
||||
document**. Since ADR-012 §3 (0.3.0), those constraints are not walked
|
||||
interpretively per buffer: they are **compiled once** into a
|
||||
`ValidationPlan` ([`src/validation_plan.rs`](../../src/validation_plan.rs))
|
||||
at `AlkTypeEngine::compile` time, and each `validate_bytes` call walks
|
||||
the compiled constraint tree against the materialized `Value` — no
|
||||
`$ref` re-resolution, no schema re-parse, no per-node path formatting
|
||||
(error paths render only on failure). The plan's constraint nodes
|
||||
implement exactly the table below (the constraint set is unchanged from
|
||||
the retired interpretive walker):
|
||||
|
||||
| Constraint | Plan node (`ValidNode`) |
|
||||
|------------|-------------------------|
|
||||
| Integer range (Int8..Uint32) | `Int { min, max }` / `Uint { max }` |
|
||||
| Int64/Uint64 (full range) | `I64` / `U64` (JSON precision caveat per ADR-005) |
|
||||
| Float finiteness (Float32/64) | `Float` with `as_f64().is_finite()` |
|
||||
| String `maxLength` (byte length) | `Str { max_len }` — `maxLength` baked in from the owning field at compile time |
|
||||
| Bytes `maxLength` (array length) | `Bytes { max_len }` — accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
|
||||
| Enum index bounds | `Enum { count }` checks `idx < count` — **fixes the v0.1.0 dead constraint** |
|
||||
| Union variant dispatch | `Union { variants }` reads `__discriminator`, dispatches on the compiled variant nodes |
|
||||
| Struct fields | `Struct { fields }` requires each declared field present, recurses |
|
||||
| Array count | `Array { count, element }` checks `arr.len() == count` and recurses per element |
|
||||
| Record values | `Record { values }` recurses into each value |
|
||||
| Boolean | `Bool` (materializer already rejects non-0/1 bytes) |
|
||||
|
||||
The plan is a public type (`ValidationPlan`, `Debug + Clone +
|
||||
PartialEq + Eq + Hash + Send + Sync`): `engine.validation_plan()`
|
||||
exposes it for consumers that validate their own materialized `Value`
|
||||
trees or want its `fingerprint()` for caching / schema handshakes
|
||||
(ADR-012 §1). The one-shot `bast_validation::validate_value(&doc,
|
||||
&value)` remains as a convenience wrapper (compile + validate) for
|
||||
callers holding a BAST document without an engine.
|
||||
|
||||
No external JSON Schema is required for `validate_bytes`. The BAST
|
||||
document is the complete specification of the binary format — it
|
||||
describes both the layout (how to read) and the constraints (what
|
||||
values are valid). This is the "schema is the format" principle from
|
||||
ADR-001, now fully realized.
|
||||
|
||||
An optional external JSON Schema can be layered on top for constraints
|
||||
BAST doesn't express (cross-field consistency, regex patterns on string
|
||||
content). This is additive, not load-bearing.
|
||||
|
||||
### `validate_json` — JSON in, JSON Schema is the validator
|
||||
|
||||
The consumer provides a JSON `Value` (e.g., an incoming JSON-RPC
|
||||
request). The BAST document is irrelevant — BAST describes bytes, not
|
||||
JSON shape. The right validator for a JSON value is a standard
|
||||
`jsonschema::Validator` built from a standard JSON Schema document the
|
||||
consumer provides at `AlkTypeEngine::compile` time. This is the path
|
||||
alkcall uses for its `OperationSpec` JSON validation. No custom
|
||||
keywords; BAST is not involved.
|
||||
|
||||
If no JSON Schema was supplied to `compile`, the JSON-validation
|
||||
methods return `AlkTypeError::Schema` (`validate_json`) or `false`
|
||||
(`is_valid_json`).
|
||||
|
||||
### `AlkTypeError::Validation` payload shape (D-BAST-009)
|
||||
|
||||
The `validate_bytes` path no longer uses `jsonschema`, so its error
|
||||
payload is constructed via `jsonschema::ValidationError::custom` purely
|
||||
to keep the `Validation` variant's type unchanged. The rationale is
|
||||
consumer ergonomics on the *combined* path: consumers like alkcall use
|
||||
both `validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary
|
||||
channels) and handle `AlkTypeError::Validation` in one place. A single
|
||||
uniform payload type means one match arm covers both sources.
|
||||
|
||||
The alternative (`Validation(String)`) would force `validate_json` to
|
||||
flatten its structured errors (instance path, schema path, keyword) to
|
||||
a `String` via `Display` — the more information-rich path loses data to
|
||||
accommodate the less rich one. That is the wrong direction.
|
||||
|
||||
### What is removed
|
||||
|
||||
Under the BAST pivot, the v0.1.0 validation machinery is removed from
|
||||
the `validate_bytes` path:
|
||||
|
||||
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
|
||||
factories) — replaced by the BAST-native validator (~250 lines, a
|
||||
flat match with no factories, no trait objects, no sub-validator
|
||||
pre-computation).
|
||||
- `inline_union_variant_refs()` — union variant refs are resolved lazily
|
||||
by the validator and materializer.
|
||||
- The custom-keyword `build_validator` path — `build_validator` is
|
||||
repurposed to build a *standard* `jsonschema::Validator` from a
|
||||
consumer-provided JSON Schema (no custom keywords). See
|
||||
[`build_validator`](#build_validator).
|
||||
|
||||
The `jsonschema` crate **remains a direct dependency** for
|
||||
`validate_json` and for validating BAST documents against the BAST
|
||||
meta-schema. The only thing removed is the custom keyword integration
|
||||
path. The `validate_bytes` path no longer touches `jsonschema` — a
|
||||
small wasm binary-size win in addition to the architecture
|
||||
simplification.
|
||||
|
||||
## Validation Strategy
|
||||
|
||||
Validation is delegated to the `jsonschema` crate (v0.46.5, Draft
|
||||
2020-12). The alktype engine does not implement its own validation —
|
||||
it registers custom keyword validators for each `AlkType:*` kind and
|
||||
lets `jsonschema` handle the structural validation (object properties,
|
||||
required fields, array items, enum values).
|
||||
The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md)
|
||||
and refined by [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md):
|
||||
|
||||
The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md):
|
||||
|
||||
1. **Load time:** Parse the schema JSON, build the layout engine, build the
|
||||
jsonschema validator. This is the `AlkTypeEngine::compile(schema)` constructor.
|
||||
1. **Load time:** Parse the BAST document into the typed tree, compile
|
||||
the `ValidationPlan` (the value-domain constraint tree), compute
|
||||
the layout engine, and (optionally) build the standard
|
||||
`jsonschema::Validator` for the JSON-validation path. This is the
|
||||
`AlkTypeEngine::compile` constructor.
|
||||
2. **Access time:** Use the compiled engine for repeated read/write
|
||||
operations. Validation is opt-in per operation.
|
||||
|
||||
### What validation validates
|
||||
|
||||
The jsonschema validator operates on `serde_json::Value` instances — it
|
||||
validates JSON representations of data, not raw byte buffers. This is
|
||||
the correct separation of concerns:
|
||||
|
||||
- **JSON validation** (jsonschema): validates that a JSON document
|
||||
conforms to the schema. Used for validating hand-written schemas,
|
||||
TypeBox output, JSON payloads, or the JSON representation of a binary
|
||||
struct after deserialization.
|
||||
- **Binary access validation** (data access layer): the read/write
|
||||
functions perform type-level validation at access time — range checks
|
||||
for integers, UTF-8 validity for strings, buffer bounds checking.
|
||||
These return `AlkTypeError::Access` with field paths.
|
||||
|
||||
The "schema is the format" principle means the same schema describes
|
||||
both the JSON shape and the binary layout. The jsonschema validator
|
||||
checks the JSON shape; the data access layer checks the binary layout.
|
||||
A consumer that wants to validate a binary buffer end-to-end reads the
|
||||
buffer into a `Value` tree via the data access layer, then validates
|
||||
that `Value` against the jsonschema validator. This is a two-step
|
||||
process, not a single `validate(buffer)` call.
|
||||
operations. Validation is opt-in per operation: the validation half
|
||||
walks the compiled `ValidationPlan`, never the BAST document.
|
||||
|
||||
### The `AlkTypeEngine` struct
|
||||
|
||||
The `AlkTypeEngine` is the compiled form of a schema. It supports both
|
||||
layout modes (ADR-002) via an internal `Layout` enum:
|
||||
The `AlkTypeEngine` is the compiled form of a BAST document. It
|
||||
supports both layout modes (ADR-002) via an internal `Layout` enum:
|
||||
|
||||
```rust
|
||||
pub struct AlkTypeEngine {
|
||||
layout: Layout, // packed or aligned (private enum)
|
||||
validator: jsonschema::Validator, // compiled once at load time
|
||||
endian: Endian, // parsed from the schema's "endian" annotation
|
||||
schema: Value, // the normalized schema (refs resolved)
|
||||
layout: Layout, // packed or aligned (private enum)
|
||||
json_validator: Option<jsonschema::Validator>, // None when no JSON Schema supplied
|
||||
validation_plan: Arc<ValidationPlan>, // compiled value-domain constraints (ADR-012 §3)
|
||||
endian: Endian, // parsed from the root struct's "endian"
|
||||
bast_doc: Value, // retained for sequential_reader/read_field
|
||||
root_name: String, // the selected $defs entry
|
||||
}
|
||||
|
||||
// Private — the consumer selects via LayoutMode at compile time.
|
||||
@@ -68,205 +168,106 @@ enum Layout {
|
||||
```
|
||||
|
||||
The consumer selects the mode at construction time via `LayoutMode`
|
||||
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout`
|
||||
enum is private — the engine exposes mode-appropriate accessors instead:
|
||||
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The
|
||||
`Layout` enum is private — the engine exposes mode-appropriate
|
||||
accessors instead:
|
||||
|
||||
```rust
|
||||
impl AlkTypeEngine {
|
||||
pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result<Self, AlkTypeError>;
|
||||
pub fn compile(
|
||||
bast_doc: &Value,
|
||||
root_name: &str,
|
||||
mode: LayoutMode,
|
||||
json_schema: Option<&Value>,
|
||||
) -> Result<Self, AlkTypeError>;
|
||||
pub fn mode(&self) -> LayoutMode;
|
||||
pub fn endian(&self) -> Endian;
|
||||
pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode
|
||||
pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode
|
||||
pub fn sequential_reader(&self) -> Option<SequentialReader>; // owned fresh reader (ADR-007)
|
||||
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // ADR-004
|
||||
pub fn is_valid_json(&self, instance: &Value) -> bool; // ADR-004
|
||||
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // ADR-010
|
||||
pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str)
|
||||
-> Result<FieldValue<'a>, AlkTypeError>; // aligned mode
|
||||
pub fn write_field(&self, buffer: &mut [u8], field_path: &str,
|
||||
value: &FieldValue<'_>) -> Result<(), AlkTypeError>; // aligned mode
|
||||
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // D-BAST-007
|
||||
pub fn is_valid_json(&self, instance: &Value) -> bool; // D-BAST-007
|
||||
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // D-BAST-006
|
||||
pub fn validation_plan(&self) -> &Arc<ValidationPlan>; // compiled constraints (ADR-012 §3)
|
||||
}
|
||||
```
|
||||
|
||||
`compile` takes `&mut Value` because it normalizes `$ref` values in place
|
||||
(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization))
|
||||
before computing the layout and building the validator. The `schema`
|
||||
field retains the normalized schema for `read_field`'s kind lookup and
|
||||
for `sequential_reader()`'s factory construction. The validator is
|
||||
mode-agnostic (it operates on `Value`, not raw bytes).
|
||||
`compile` takes `&Value` (not `&mut Value`) — BAST needs no in-place
|
||||
`normalize_refs`. `root_name` selects which `$defs` entry is the
|
||||
top-level type (D-BAST-001). `json_schema` is the optional
|
||||
consumer-provided standard JSON Schema for the `validate_json` path
|
||||
(D-BAST-007); pass `None` when JSON validation is not needed. The
|
||||
engine retains a clone of the BAST `Value` so `sequential_reader` and
|
||||
`read_field` can re-parse the typed tree on demand without lifetime
|
||||
entanglement with the caller's `Value`.
|
||||
|
||||
The `Layout::Packed` variant stores only the `LayoutBuilder` (write-side).
|
||||
The `SequentialReader` (read-side) is not stored — it has mutable cursor
|
||||
state that the consumer owns, so `sequential_reader()` constructs a fresh
|
||||
reader on each call (ADR-007).
|
||||
The `Layout::Packed` variant stores only the `LayoutBuilder`
|
||||
(write-side). The `SequentialReader` (read-side) is not stored — it has
|
||||
mutable cursor state that the consumer owns, so `sequential_reader()`
|
||||
constructs a fresh reader on each call (ADR-007).
|
||||
|
||||
The `read_field`/`write_field` methods on `AlkTypeEngine` are the
|
||||
aligned-mode data-access API — see [data-access.md](data-access.md)
|
||||
§"Higher-level read/write".
|
||||
|
||||
## Custom Keyword Validators
|
||||
## `build_validator`
|
||||
|
||||
Each `AlkType:*` kind gets a `Keyword` implementation registered via
|
||||
`jsonschema::options().with_keyword(...)`. The validators check leaf
|
||||
type constraints; `jsonschema` handles all structural validation.
|
||||
|
||||
### Numeric type validators
|
||||
|
||||
**`AlkType:Float32` / `AlkType:Float64`:**
|
||||
- Value must be a finite number.
|
||||
- For `Float32`: value must be representable as `f32` (no precision loss
|
||||
beyond `f32`'s mantissa).
|
||||
|
||||
**`AlkType:Int8` / `AlkType:Int16` / `AlkType:Int32`:**
|
||||
- Value must be an integer within the type's range.
|
||||
- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
|
||||
|
||||
**`AlkType:Uint8` / `AlkType:Uint16` / `AlkType:Uint32`:**
|
||||
- Value must be a non-negative integer within the type's range.
|
||||
- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295.
|
||||
|
||||
### String and binary validators
|
||||
|
||||
**`AlkType:String`:**
|
||||
- Value must be a valid UTF-8 string.
|
||||
- If `maxLength` is specified in the schema, the string's byte length
|
||||
must not exceed it.
|
||||
|
||||
**`AlkType:Bytes`:**
|
||||
- Value must be a string (the JSON form for `validate_json` consumers)
|
||||
or an array of integers 0..=255 (the materialized form for
|
||||
`validate_bytes`). JSON has no native byte type; the string form is
|
||||
the JSON convention, the array form is the round-trippable form for
|
||||
non-UTF-8 bytes (see [OQ-007](questions/007-bytes-materialization-lossy-utf8.md)).
|
||||
- If `maxLength` is specified, the byte length must not exceed it. For
|
||||
the string form, this is the string's byte length; for the array
|
||||
form, this is the array length (one entry per byte).
|
||||
- **Binary representation:** In the binary layout, `TBytes` is raw bytes
|
||||
with no encoding (not base64, not hex). The JSON representation (for
|
||||
validation) uses a string or array; the binary representation (for
|
||||
data access) uses `&[u8]` directly.
|
||||
|
||||
**`AlkType:Enum`:**
|
||||
- The `AlkType:Enum` custom keyword signals that the type is an enum for
|
||||
*layout* purposes (the engine needs to know it's a fixed-size u32 index,
|
||||
not a variable-length string). The built-in `enum` keyword provides the
|
||||
value list and handles value-membership validation. The custom keyword
|
||||
validator is a no-op beyond the built-in check — it exists solely for
|
||||
the layout engine to recognize the type.
|
||||
|
||||
**`AlkType:Timestamp`:**
|
||||
- Value must be a valid RFC 3339 timestamp string (the internet profile
|
||||
of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
|
||||
|
||||
### Composite type validators
|
||||
|
||||
**`AlkType:Struct`:**
|
||||
- Value must be an object.
|
||||
- Each property must match its declared `AlkType:*` kind.
|
||||
- Required fields must be present.
|
||||
- The `jsonschema` crate's built-in `properties` and `required` keywords
|
||||
handle the structural checks — the custom keyword only needs to
|
||||
validate that each field's value matches its `AlkType:*` kind.
|
||||
|
||||
**`AlkType:Union`:**
|
||||
- The instance must be an object with a `__discriminator` field
|
||||
carrying the mapping key (stringified discriminator value for
|
||||
byte-offset discriminators, string value for field-name
|
||||
discriminators). This is the shape the materializer produces for
|
||||
`validate_bytes`; `validate_json` consumers produce the same shape
|
||||
when validating a union instance.
|
||||
- The `UnionValidator` builds a sub-validator for each variant at
|
||||
factory time (when the parent validator tree is constructed) and
|
||||
dispatches on `__discriminator` at validation time, validating the
|
||||
full instance (including the variant fields) against the selected
|
||||
variant's schema. This closes the OQ-008 gap: variant field
|
||||
constraints (e.g. `maxLength` on a `Bytes` field inside a variant)
|
||||
are checked.
|
||||
- `$ref`s in the union's `mapping` are inlined by
|
||||
`schema::inline_union_variant_refs` during `AlkTypeEngine::compile`
|
||||
(before `build_validator`), so the `union_factory` sees full inline
|
||||
variant schemas. See [OQ-008](questions/008-unionvalidator-variant-dispatch.md).
|
||||
|
||||
**`AlkType:Array`:**
|
||||
- Value must be an array.
|
||||
- Each element must match the array's declared element type.
|
||||
- If `minItems`/`maxItems` is specified, the array length must be within
|
||||
bounds.
|
||||
|
||||
### Other validators
|
||||
|
||||
**`AlkType:Boolean`:**
|
||||
- Value must be `true` or `false`.
|
||||
|
||||
**`AlkType:Record`:**
|
||||
- Value must be an object.
|
||||
- All values must match the record's declared value type (specified via
|
||||
the `"values"` property in the schema, e.g.,
|
||||
`"values": { "AlkType:Float32": true }`).
|
||||
|
||||
### Validator implementation pattern
|
||||
|
||||
Each custom keyword implementation is ~10 lines. Example for
|
||||
`AlkType:Float32`:
|
||||
`src/validation.rs` exposes one function:
|
||||
|
||||
```rust
|
||||
struct Float32Validator;
|
||||
|
||||
impl Keyword for Float32Validator {
|
||||
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
|
||||
match instance {
|
||||
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
|
||||
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
|
||||
}
|
||||
}
|
||||
fn is_valid(&self, instance: &Value) -> bool {
|
||||
instance.as_f64().map_or(false, |f| f.is_finite())
|
||||
}
|
||||
}
|
||||
pub fn build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError>;
|
||||
```
|
||||
|
||||
Registration:
|
||||
Under the pivot this is **repurposed** (D-BAST-007): it builds a
|
||||
*standard* `jsonschema::Validator` from a consumer-provided plain JSON
|
||||
Schema — no custom keywords, no BAST involvement. The engine calls it
|
||||
internally during `compile` when `json_schema` is `Some`. Consumers
|
||||
that only need a one-off validator may call `jsonschema::options().build(schema)`
|
||||
directly; `build_validator` exists so the engine's error mapping
|
||||
(`jsonschema` build error → `AlkTypeError::Schema`) is reused.
|
||||
|
||||
```rust
|
||||
let validator = jsonschema::options()
|
||||
.with_keyword("AlkType:Float32", |parent, value, path| {
|
||||
Ok(Box::new(Float32Validator))
|
||||
})
|
||||
.build(&schema)?;
|
||||
```
|
||||
|
||||
The factory closure receives the parent schema object, the keyword's
|
||||
value, and the schema path. This enables cross-keyword awareness — for
|
||||
example, a `AlkType:Struct` validator can inspect the parent's
|
||||
`properties` to validate each field against its declared `AlkType:*` kind.
|
||||
The v0.1.0 custom-keyword `build_validator` (registered 19
|
||||
`with_keyword(...)` factories) is removed.
|
||||
|
||||
## AlkTypeError
|
||||
|
||||
A single `AlkTypeError` enum covers all error conditions across the
|
||||
engine's three phases (schema parsing, offset computation, read/write)
|
||||
plus validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md).
|
||||
engine's phases (schema parsing, offset computation, read/write) plus
|
||||
validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md);
|
||||
the variant shapes are unchanged under the pivot (D-BAST-009).
|
||||
|
||||
```rust
|
||||
pub enum AlkTypeError {
|
||||
/// Schema parsing errors (invalid JSON, missing keywords, unknown AlkType kinds).
|
||||
/// Schema parsing errors (malformed BAST, dangling $ref, unknown kind).
|
||||
Schema(String),
|
||||
/// Offset computation errors (field not found, unsupported type).
|
||||
Offset { field_path: String, reason: String },
|
||||
/// Read/write errors (buffer too short, invalid UTF-8, value out of range).
|
||||
Access { field_path: String, reason: String },
|
||||
/// Validation errors (delegated to jsonschema).
|
||||
Validation(ValidationError<'static>),
|
||||
/// Validation errors (both paths — D-BAST-009 uniform payload).
|
||||
Validation(jsonschema::ValidationError<'static>),
|
||||
}
|
||||
```
|
||||
|
||||
- **`Schema`** — for errors during `AlkTypeEngine::compile()`. Invalid
|
||||
JSON, missing required keywords, unknown `AlkType:*` kinds.
|
||||
- **`Schema`** — for errors during `AlkTypeEngine::compile()` or any
|
||||
BAST-walking path. Malformed BAST, missing `$defs`, unknown `kind`
|
||||
string, dangling `$ref`, empty `mapping`, etc.
|
||||
- **`Offset`** — for errors during offset computation. Field not found
|
||||
in the schema, type not supported for offset computation, recursive
|
||||
depth exceeded. Carries the field path.
|
||||
- **`Access`** — for errors during read/write. Buffer too short, invalid
|
||||
UTF-8 in a string field, value out of range for the target type.
|
||||
Carries the field path.
|
||||
- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The
|
||||
`'static` lifetime is correct — the validator owns its schema reference
|
||||
and lives for the lifetime of the `AlkTypeEngine`.
|
||||
in the BAST tree, type not supported for offset computation. Carries
|
||||
the field path.
|
||||
- **`Access`** — for errors during read/write. Buffer too short,
|
||||
invalid UTF-8 in a string field, value out of range for the target
|
||||
type. Carries the field path.
|
||||
- **`Validation`** — wraps a `jsonschema::ValidationError<'static>`.
|
||||
On the `validate_json` path, this is the `jsonschema` crate's own
|
||||
structured error. On the `validate_bytes` path, it is constructed via
|
||||
`jsonschema::ValidationError::custom` from the BAST-native
|
||||
validator's path + reason string. The `'static` lifetime is correct —
|
||||
the payload owns its data.
|
||||
|
||||
### Field-path-carrying errors
|
||||
|
||||
@@ -287,20 +288,64 @@ you exactly which field failed and why.
|
||||
### Load time: `AlkTypeEngine::compile()`
|
||||
|
||||
The expensive work happens once at schema load time:
|
||||
1. Normalize `$ref` values in the schema (`normalize_refs`).
|
||||
2. Parse the schema's `"endian"` annotation.
|
||||
3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned).
|
||||
4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
|
||||
|
||||
The result is a `AlkTypeEngine` that can be used for repeated operations.
|
||||
1. Parse the BAST document into the typed tree (`BastDoc::new`).
|
||||
2. Parse the root struct's `"endian"` annotation.
|
||||
3. Compile the `ValidationPlan` — the value-domain constraint tree,
|
||||
with eager `$ref` resolution. Its compile walk rejects cyclic `$ref`
|
||||
graphs with a clean `Schema` error *before* the layout computation.
|
||||
(The layout walkers now also guard themselves — each standalone
|
||||
entry point runs the shared reference-graph check
|
||||
(`walk_guard::check_ref_graph`, review #006 H2) — so the plan-first
|
||||
ordering is belt-and-suspenders at engine compile, and the trust
|
||||
boundary no longer depends on the call path.)
|
||||
4. Compute the layout (`LayoutBuilder` for packed, `OffsetMap` for
|
||||
aligned).
|
||||
5. If `json_schema` is `Some`, build the standard
|
||||
`jsonschema::Validator` via `validation::build_validator`.
|
||||
|
||||
The result is an `AlkTypeEngine` that can be used for repeated
|
||||
operations. The validation half is pre-built: the engine holds an
|
||||
`Arc<ValidationPlan>` and walks it per buffer without re-touching the
|
||||
BAST document (ADR-012 §3).
|
||||
|
||||
### Access time: `engine.validate_bytes(&[u8])`
|
||||
|
||||
For binary-layout schemas, `validate_bytes` runs the two phases in
|
||||
sequence (D-BAST-006):
|
||||
|
||||
1. **Materialize `Value` from bytes.** `materialize::materialize_packed`
|
||||
or `materialize::materialize_aligned` walks the buffer against the
|
||||
BAST typed tree and the engine's `Endian`, producing a
|
||||
`serde_json::Value` tree. Composites are recursed into (`Struct` →
|
||||
object of field values; `Array` → array of element values; `Union`
|
||||
→ dispatch then recurse; `Record` → object of key/value entries).
|
||||
The read phase reuses the existing data-access functions and returns
|
||||
`AlkTypeError::Access` (with field paths) on read failures.
|
||||
2. **Validate the `Value` against the `ValidationPlan`.** The
|
||||
materialized `Value` is walked against the compiled constraint tree
|
||||
(`engine.validation_plan().validate(&value)`), producing
|
||||
`AlkTypeError::Validation` on the first violated value-domain
|
||||
constraint. This is the ADR-012 §3 end state: the only per-buffer
|
||||
schema-touching step is the materialize half (the bytes must be
|
||||
decoded against the tree); the validation half is plan-fast.
|
||||
|
||||
Mode dispatch:
|
||||
|
||||
- **Packed mode** — materializes fields in declaration order.
|
||||
- **Aligned mode** — uses the `OffsetMap` to read fields at their
|
||||
computed offsets.
|
||||
|
||||
Both modes produce the same `Value` form; the validation plan is
|
||||
mode-agnostic.
|
||||
|
||||
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
|
||||
|
||||
Validation is opt-in per operation. The consumer calls
|
||||
`engine.validate_json(instance)` when validation is desired, or
|
||||
`engine.is_valid_json(instance)` for a boolean check. The jsonschema
|
||||
validator is already compiled — these are fast checks against the
|
||||
compiled validator.
|
||||
`engine.is_valid_json(instance)` for a boolean check. The
|
||||
`jsonschema::Validator` is already compiled — these are fast checks
|
||||
against the compiled validator.
|
||||
|
||||
```rust
|
||||
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>;
|
||||
@@ -308,79 +353,37 @@ pub fn is_valid_json(&self, instance: &Value) -> bool;
|
||||
```
|
||||
|
||||
The argument is a `serde_json::Value` (the JSON representation of the
|
||||
data), not a raw byte buffer — see §"What validation validates" above.
|
||||
To validate a binary buffer end-to-end, the consumer reads it into a
|
||||
`Value` tree via the data access layer, then validates that `Value`.
|
||||
data), not a raw byte buffer. `validate_json` validates against the
|
||||
consumer-provided JSON Schema supplied at `compile` time (D-BAST-007);
|
||||
the BAST document is not involved. If no JSON Schema was supplied,
|
||||
`validate_json` returns `AlkTypeError::Schema` and `is_valid_json`
|
||||
returns `false`.
|
||||
|
||||
High-throughput paths can skip validation. Security-sensitive paths
|
||||
(parsing incoming frames from untrusted peers) can validate every frame.
|
||||
The choice is the consumer's.
|
||||
(parsing incoming frames from untrusted peers) can validate every
|
||||
frame. The choice is the consumer's.
|
||||
|
||||
### Access time: `engine.validate_bytes(&[u8])` — binary buffer validation
|
||||
|
||||
For binary-layout schemas (schemas declaring `AlkType:*` kinds), the
|
||||
engine offers a single-call form of the two-step dance: walk the bytes
|
||||
against the layout to materialize a `Value` tree, then validate that
|
||||
`Value` against the compiled jsonschema validator. Decided in
|
||||
[ADR-010](decisions/010-generalized-validation-validate-bytes.md).
|
||||
|
||||
```rust
|
||||
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>;
|
||||
```
|
||||
|
||||
`validate_bytes` runs the existing machinery in sequence:
|
||||
|
||||
1. **Materialize `Value` from bytes.** A new internal helper
|
||||
(`materialize_value`, alongside `SequentialReader::read_field_value`
|
||||
in `src/sequential_reader.rs`) walks the buffer against the schema
|
||||
and the engine's `Endian`, producing a `serde_json::Value` tree.
|
||||
Composites are recursed into (`Struct` → object of field values;
|
||||
`Array` → array of element values; `Union` → dispatch then recurse;
|
||||
`Record` → object of key/value entries). The read phase reuses the
|
||||
existing data-access functions and returns `AlkTypeError::Access`
|
||||
(with field paths) on read failures.
|
||||
2. **Validate the `Value`.** The materialized `Value` is passed to the
|
||||
existing `self.validator.validate(&value)`, producing
|
||||
`AlkTypeError::Validation` on failure.
|
||||
|
||||
Mode dispatch:
|
||||
|
||||
- **Packed mode** — walks with a fresh `SequentialReader` (the engine
|
||||
is already a reader factory per ADR-007), materializing fields in
|
||||
declaration order.
|
||||
- **Aligned mode** — uses the `OffsetMap` to read fields at their
|
||||
computed offsets, then materializes composites by recursing into the
|
||||
offset map's nested entries.
|
||||
|
||||
Both modes produce the same `Value` form; the validator is
|
||||
mode-agnostic (it operates on `Value`, not bytes — ADR-004).
|
||||
|
||||
#### When to use which entry point
|
||||
### When to use which entry point
|
||||
|
||||
| Entry point | Schema form | Input form | When |
|
||||
|-------------|--------------|------------|------|
|
||||
| `validate_json(&Value)` | Any (AlkType or plain JSON Schema) | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); TypeBox output; anything off `serde_json::from_slice` / `from_str` |
|
||||
| `validate_bytes(&[u8])` | AlkType binary-layout schema | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs |
|
||||
| `validate_json(&Value)` | Consumer-provided standard JSON Schema | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); anything off `serde_json::from_slice` / `from_str` |
|
||||
| `validate_bytes(&[u8])` | BAST document (binary layout) | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs |
|
||||
|
||||
`validate_bytes` requires the engine's schema to declare `AlkType:*`
|
||||
kinds — it materializes `Value` via the layout engine, which needs
|
||||
binary-layout semantics. A pure JSON Schema (call's `input_schema`,
|
||||
no AlkType kinds) compiled via `AlkTypeEngine::compile` would fail at
|
||||
the materialize step (no `AlkType:Struct` at the root). For pure JSON
|
||||
payloads, the consumer uses `serde_json::from_slice` then
|
||||
`validate_json`. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md)
|
||||
§"Not a binary-payload validator for JSON-only schemas".
|
||||
`validate_bytes` requires the engine's root type to be a struct (the
|
||||
layout engine enforces this) — it materializes `Value` via the layout
|
||||
engine, which needs binary-layout semantics. For pure JSON payloads,
|
||||
the consumer uses `serde_json::from_slice` then `validate_json`. See
|
||||
[ADR-010](decisions/010-generalized-validation-validate-bytes.md) and
|
||||
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
|
||||
|
||||
#### What `validate_bytes` is not
|
||||
|
||||
- **Not a new validation engine.** It runs the existing `jsonschema`
|
||||
validator against the existing materialized `Value`. No new
|
||||
validator code, no parallel validation path (ADR-001).
|
||||
- **Not framing-aware.** It validates the bytes of *one* schema
|
||||
instance. It does not strip length prefixes, parse
|
||||
`[length: u32][payload]` framing, or handle multiple frames in a
|
||||
buffer. That's the consumer's job. alktype validates what one
|
||||
schema describes; it does not parse the wire envelope around it.
|
||||
buffer. That's the consumer's job. alktype validates what one schema
|
||||
describes; it does not parse the wire envelope around it.
|
||||
- **Not a `Validator` trait.** Two methods on one struct, not a trait
|
||||
abstraction. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md)
|
||||
§"Not a `Validator` trait abstraction".
|
||||
@@ -390,24 +393,25 @@ payloads, the consumer uses `serde_json::from_slice` then
|
||||
Validation and data access are independent operations on the same data.
|
||||
The consumer can:
|
||||
|
||||
1. Validate the JSON representation of a buffer to ensure it conforms to
|
||||
the schema.
|
||||
1. Validate the bytes of a buffer to ensure it conforms to the BAST
|
||||
document's value constraints.
|
||||
2. Read fields from the binary buffer at computed offsets.
|
||||
3. Both — validate the JSON representation first, then read the binary
|
||||
buffer (defense in depth).
|
||||
3. Both — validate first, then read (defense in depth).
|
||||
|
||||
The engine does not couple validation and access. A consumer that trusts
|
||||
its data source can skip validation and go straight to read/write. A
|
||||
consumer that parses untrusted input can validate the JSON
|
||||
representation first, then access the binary buffer.
|
||||
The engine does not couple validation and access. A consumer that
|
||||
trusts its data source can skip validation and go straight to
|
||||
read/write. A consumer that parses untrusted input can validate first,
|
||||
then access the binary buffer.
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
|
||||
| Two-validator model (BAST-native + standard jsonschema) | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | `validate_bytes` uses the compiled `ValidationPlan`; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema; D-BAST-006/007/009 |
|
||||
| Error handling and validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
|
||||
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation (materialize `Value` from bytes, then validate); two methods on one struct, not a trait |
|
||||
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine |
|
||||
| Compiled `ValidationPlan` | [ADR-012](decisions/012-plan-fingerprinting-and-m1-closure.md) | The value-domain constraint tree is compiled once at `compile` (eager `$ref` resolution, cycle rejection) and walked per buffer; `Hash + Eq` + `fingerprint()`; retires the interpretive `BastDoc` walk |
|
||||
| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | The BAST document is the complete binary-format spec (layout + value constraints) |
|
||||
|
||||
## Open Questions
|
||||
|
||||
@@ -418,11 +422,23 @@ see [builder.md](builder.md).
|
||||
|
||||
## References
|
||||
|
||||
- `@alkdev/alknet: docs/research/alknet-typedef/findings.md`
|
||||
§"Validation" — the POC's custom keyword validators for all 17 kinds
|
||||
- [`bast-format.md` §Validation Model](bast-format.md#validation-model)
|
||||
— the normative validation model
|
||||
- [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) — the
|
||||
two-validator decision
|
||||
- [ADR-012](decisions/012-plan-fingerprinting-and-m1-closure.md) — the
|
||||
`ValidationPlan` decision (§3)
|
||||
- [ADR-004](decisions/004-error-handling-validation-strategy.md) —
|
||||
error handling and validation strategy
|
||||
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds that the
|
||||
validators check
|
||||
- [data-access.md](data-access.md) — read/write functions that operate
|
||||
on the same buffers
|
||||
- [ADR-010](decisions/010-generalized-validation-validate-bytes.md) —
|
||||
`validate_bytes` (the collapsed two-step dance)
|
||||
- [schema-layer.md](schema-layer.md) — the BAST parser that the
|
||||
plan compiler consumes
|
||||
- [data-access.md](data-access.md) — read/write functions and the
|
||||
materializer that produce the `Value` the plan checks
|
||||
- [`src/validation_plan.rs`](../../src/validation_plan.rs) — the
|
||||
compiled `ValidationPlan` implementation
|
||||
- [`src/bast_validation.rs`](../../src/bast_validation.rs) — the
|
||||
one-shot wrapper (`validate_value`) and shared error helpers
|
||||
- [`src/validation.rs`](../../src/validation.rs) — the `build_validator`
|
||||
helper
|
||||
@@ -0,0 +1,983 @@
|
||||
---
|
||||
status: done
|
||||
created: 2026-08-19
|
||||
last_updated: 2026-09-02
|
||||
adr: ADR-011, ADR-012
|
||||
---
|
||||
|
||||
# 0.3.0 — Compiled Forms: ReadPlan, Owned BastDoc, OffsetMap LeafMeta, ValidationPlan, Fingerprinting
|
||||
|
||||
This is the execution plan for the 0.3.0 release: the compiled-form
|
||||
rollup that closes review #004's 400x read-path gap (ADR-011) *and*
|
||||
the deferred M1 sites (ADR-012) *and* retires the interpretive
|
||||
validation walk via a `ValidationPlan` (ADR-012 §3, reversing the
|
||||
original deferral per review #005 M3) *and* adds plan fingerprinting
|
||||
(ADR-012 §1/§4) in one breaking bump. It is the **entry point** an
|
||||
implementing agent reads first.
|
||||
|
||||
Companion documents:
|
||||
|
||||
- [ADR-011](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md)
|
||||
— the `ReadPlan` decision (packed read-side compiled form).
|
||||
- [ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md)
|
||||
— fingerprinting + owned `BastDoc` + `OffsetMap` `LeafMeta` +
|
||||
`ValidationPlan` (this release's other three pieces).
|
||||
- [Review #004](../reviews/004-performance-review.md) — the
|
||||
performance finding being closed.
|
||||
- [POC findings](../../poc/readplan/FINDINGS.md) (branch `readplan-poc`)
|
||||
— the derisking POC that confirmed the `ReadPlan` shape and surfaced
|
||||
two findings (field-disc union read shape; struct-array stride).
|
||||
|
||||
**Working order:** read this plan top-to-bottom. The Semver Contract
|
||||
section is the scope-creep guardrail — consult it before each step.
|
||||
Each step links to its ADR and lists its verification gate. Implement
|
||||
phases in order; within a phase, steps are ordered by dependency.
|
||||
|
||||
## Phases vs sessions
|
||||
|
||||
This plan is deliberately larger than one session's work. The eight
|
||||
phases are the session boundaries — each phase is a coherent unit
|
||||
that leaves the tree building and tests green, so any one session
|
||||
can pick up a phase without needing context from the previous one.
|
||||
Phase boundaries are also commit boundaries (and push boundaries per
|
||||
AGENTS.md). If a phase is large enough to span sessions, the steps
|
||||
within it are the sub-session boundaries.
|
||||
|
||||
## Semver Contract
|
||||
|
||||
The crate is on crates.io at 0.2.0 with zero real consumers (only
|
||||
`alktty`/`alkcall`, both in-house path dev-deps). A breaking bump to
|
||||
0.3.0 is free but the contract is explicit so the implementation
|
||||
doesn't drift. Per AGENTS.md, the public surface is the `lib.rs`
|
||||
re-exports.
|
||||
|
||||
| Public item (from `lib.rs` re-exports) | Class | Change |
|
||||
|---|---|---|
|
||||
| `AlkTypeEngine::compile` | **Breaking (internal)** | Signature unchanged `(bast_doc: &Value, root_name: &str, mode, json_schema) -> Result<Self, AlkTypeError>`. Internally builds a `ReadPlan` (packed) or extended `OffsetMap` (aligned) and stores it. The `bast_doc: Value` clone is retained (ADR-011 §Engine integration). |
|
||||
| `AlkTypeEngine::sequential_reader` | **Breaking (return type)** | Returns `Option<SequentialReader>` (unchanged type), but the reader is now constructed from `Arc<ReadPlan>`, not from `&bast_doc`. The reader's public methods (`read_next`/`read_field`/`reset`/`position`/`endian`/`schema`) keep their signatures. `schema()` returns the `&Value` the plan was compiled from (retained on the engine). |
|
||||
| `AlkTypeEngine::read_field` / `write_field` | **Unchanged (signature)** | Still `(buffer, field_path) -> Result<FieldValue, AlkTypeError>`. Internally reads `LeafMeta` from the extended `OffsetMap` instead of re-parsing `BastDoc`. |
|
||||
| `AlkTypeEngine::validate_bytes` | **Unchanged (signature)** | Packed mode calls `materialize_packed(&self.plan, buffer)` (ADR-011); aligned mode calls `materialize_aligned(&doc, buffer, &self.offset_map)` with the owned `BastDoc`. |
|
||||
| `LayoutMode`, `AlkTypeEngine` | **Unchanged** | — |
|
||||
| `BastDoc`, `BastDef`, `BastDefKind`, `BastStruct`, `BastField`, `BastType`, `BastUnion`, `BastDiscriminator`, `BastEnum`, `BastArray`, `BastRecord`, `BastRef` | **Breaking (lifetime removal)** | `BastDoc<'a>` → `BastDoc` (owned). Every `&'a str` → `String` (or `Arc<str>` — decision in phase 3). Every `&'a Value` → `Value` (or `Arc<Value>`). Every method signature that took/returned `&'a` changes. The `Bast*` types are re-exported from `lib.rs` so this is a public break. |
|
||||
| `OffsetMap` | **Breaking (`get` return type)** | `get(field_path) -> Option<&ByteRange>` → `get(field_path) -> Option<&OffsetEntry>` where `OffsetEntry { range: ByteRange, meta: LeafMeta }` (or two accessors). Additive capability. |
|
||||
| `ByteRange` | **Unchanged** | Still `Copy + PartialEq + Eq + Hash`. |
|
||||
| `LeafMeta` | **New public type** | `{ kind: AlkTypeKind, encoding: VariableEncoding, endian: Endian }`, re-exported from `lib.rs`. `Copy + PartialEq + Eq + Hash`. |
|
||||
| `ReadPlan` | **New public type** | From ADR-011. Re-exported from `lib.rs`. `Debug + Clone + PartialEq + Eq + Hash`. |
|
||||
| `SequentialReader` | **Breaking (constructor + return type)** | `SequentialReader::new(&Value, &str) -> Result<Self, AlkTypeError>` → `SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible — just stores the `Arc`; the `BastDoc` parse moved to `ReadPlan::compile`). Public methods (`read_next`/`read_field`/`reset`/`position`/`endian`/`schema`) unchanged. `schema()` returns `&Value` retained on the plan (see phase 2 — the plan stores `Arc<Value>`, not `&Value`, to avoid the self-referential struct ADR-011 rejects). `engine.rs`'s `.ok()` on the old `Result` correspondingly goes away. |
|
||||
| `FieldValue` | **Unchanged** | — |
|
||||
| `materialize_packed` | **Breaking (signature)** | `materialize_packed(&BastDoc<'_>, &[u8])` → `materialize_packed(&ReadPlan, &[u8])`. |
|
||||
| `materialize_aligned` | **Breaking (signature)** | `materialize_aligned(&BastDoc<'_>, &[u8], &OffsetMap)` → `materialize_aligned(&BastDoc, &[u8], &OffsetMap)` (owned `BastDoc`, no lifetime). |
|
||||
| `ValidationPlan` | **New public type** | From ADR-012 §3. Re-exported from `lib.rs`. `Debug + Clone + PartialEq + Eq + Hash`. `compile(&BastDoc, &str) -> Result<Self, AlkTypeError>`, `fingerprint() -> u64`. Shape scoped in phase 7 (and a preceding design session); the contract is fixed in ADR-012 §3. |
|
||||
| `AlkTypeEngine::validate_bytes` | **Unchanged (signature)** | Still `(buffer) -> Result<(), AlkTypeError>`. Internally walks `&ValidationPlan` (both modes) instead of re-walking `BastDoc` for value-domain checks. The `materialize` half is unchanged from the ADR-011/§2b work (packed: `materialize_packed(&self.plan, ...)`; aligned: `materialize_aligned(&self.doc, ..., &self.offset_map)`). |
|
||||
| `bast_validation` (`build_validator`, `validate_value`) | **Additive (reviewed at phase 7)** | `validate_value` is expected to become a thin wrapper over `ValidationPlan` (or be retired if the shape work shows it's redundant). Additive changes ride the bump; renames/removals are avoided unless phase 7's shape work shows they're necessary. The BAST meta-schema validator (`build_validator`/`BAST_META_SCHEMA`) used at `compile` time is unaffected. |
|
||||
| `LayoutBuilder`, `PackedLayout`, `FieldPosition` | **Unchanged (signature)** | `LayoutBuilder::new`/`build` signatures unchanged. Internally caches the owned `BastDoc` instead of re-parsing. |
|
||||
| `AlkTypeKind`, `Endian`, `VariableEncoding` | **Unchanged** | — |
|
||||
| `AlkTypeError` | **Unchanged** | — |
|
||||
| `Schema`, `Definitions`, `Discriminator` builders | **Unchanged** | — |
|
||||
| `UnionDispatch`, `build_validator`, `BAST_META_SCHEMA` | **Unchanged** | — |
|
||||
| `data_access::*` functions | **Unchanged** | — |
|
||||
|
||||
**Net breaking surface:** `BastDoc` and all `Bast*` types (lifetime
|
||||
removal), `OffsetMap::get` (return type), `SequentialReader::new`
|
||||
(constructor + `Result` drop), `materialize_packed`/`materialize_aligned`
|
||||
(signatures). **Net additive:** `ReadPlan`, `LeafMeta`, `ValidationPlan`,
|
||||
`fingerprint()` methods, `Hash + Eq` derives on `ReadPlan`/`OffsetMap`/
|
||||
`ValidationPlan`. **Net unchanged:** the builder, `AlkTypeEngine`
|
||||
accessors (signatures), `FieldValue`, `AlkTypeKind`, `AlkTypeError`,
|
||||
`data_access`, the `Schema`/`Definitions` builders, `validate_bytes`
|
||||
(signature).
|
||||
|
||||
### Decisions deferred to their implementation phases
|
||||
|
||||
1. **`Arc<str>` vs `String` for owned `BastDoc` names** (phase 3): `Arc<str>`
|
||||
shares allocation for repeated names (e.g. union variant keys appearing
|
||||
in multiple places); `String` is simpler. The POC used `String`. Lean:
|
||||
`String` unless a bench shows `Arc<str>` matters — the typed tree is
|
||||
built once, not hot. Decided in phase 3.
|
||||
2. **`OffsetMap::get` return shape** (phase 5): `Option<&OffsetEntry>` (a
|
||||
new accessor struct) vs two methods `range(path) -> Option<&ByteRange>`
|
||||
+ `meta(path) -> Option<&LeafMeta>`. The struct is fewer calls; the two
|
||||
methods preserve back-compat shape for callers that only want the range.
|
||||
Lean: struct — it's a breaking bump anyway and the struct is cleaner.
|
||||
Decided in phase 5.
|
||||
3. **Fingerprint hasher** (phase 6): `DefaultHasher` (std, stable within a
|
||||
version) vs `FxHasher` (faster, also stable). Cross-version stability
|
||||
is a non-goal (ADR-012). Lean: `DefaultHasher` — no new dep, the
|
||||
fingerprint isn't hot. Decided in phase 6.
|
||||
4. **Struct-array stride** (phase 2): the POC found the existing reader
|
||||
returns `element_stride: 0` for fixed-size struct arrays
|
||||
(`sequential_reader.rs:567`). The `ReadPlan` correctly computes the
|
||||
stride. Decision: preserve the existing `0` behavior in `ReadPlan` for
|
||||
back-compat with `SequentialReader`'s consumer contract, *or* fix it
|
||||
and document the behavioral change. Lean: fix it — the `0` is a latent
|
||||
bug, the stride is behaviorally observable, and we're bumping. The
|
||||
plan step calls this out explicitly. Decided in phase 2.
|
||||
|
||||
## Phase 1 — `ReadPlan` type + `compile` (ADR-011 step 1) — **DONE (2026-09-02)**
|
||||
|
||||
> **Status: implemented.** `src/read_plan.rs` builds the refined
|
||||
> `CompositePlan::Union` shape (`shared: Option<Box<ReadPlan>>` +
|
||||
> `variants: Vec<(String, CompositePlan)>`, no `VariantPlan`/`VariantKind`)
|
||||
> with eager `$ref` resolution, `BTreeMap` `by_name`, field-disc `shared`
|
||||
> sub-plans, nested-union variant support, and true array strides
|
||||
> (deferred decision 4 resolved: fixed struct arrays compute their real
|
||||
> stride, not the 0.2.0 reader's `0`). Two parity notes recorded as
|
||||
> compile-behavior locks in tests: (a) the plan propagates the
|
||||
> *referring field's* effective endianness into nested structs/unions —
|
||||
> exactly what the 0.2.0 packed reader/materializer do — rather than
|
||||
> consulting nested containers' own `endian` annotations (the POC baked
|
||||
> `s.endian()` there; its equivalence tests never covered a nested
|
||||
> annotation, so the divergence was latent); (b) `compile` carries its
|
||||
> own depth cap (128) + definition-level cycle set, so standalone
|
||||
> `ReadPlan::compile` is untrusted-input-safe independent of the
|
||||
> meta-schema and the engine's `ValidationPlan` gate.
|
||||
|
||||
**Goal:** Add the `ReadPlan`/`FieldPlan`/`CompositePlan`/`ReadKind`/
|
||||
|
||||
**ADR reference:** [ADR-011 §The `ReadPlan` shape](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#the-readplan-shape),
|
||||
[ADR-011 §Construction](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#construction).
|
||||
The `CompositePlan::Union` shape in ADR-011 was refined (vs the
|
||||
accepted-at-POC shape) to carry the field-disc union's shared fields
|
||||
and to drop `VariantPlan`/`VariantKind`; this phase implements the
|
||||
refined shape.
|
||||
|
||||
**POC reference:** `poc/readplan/src/lib.rs` (branch `readplan-poc`)
|
||||
is the reference scaffold. The production version lives in `src/` and
|
||||
adds doc comments, clippy cleanliness, the field-name-discriminator
|
||||
union read shape the POC stubbed (POC Finding 1), and nested-union
|
||||
variant support the POC rejected but 0.2.0 accepts (POC "What this
|
||||
POC does not cover" → nested unions). Both are resolved by the
|
||||
refined `CompositePlan::Union` shape — see below.
|
||||
|
||||
**Files:** New `src/read_plan.rs`. Update `src/lib.rs` to add
|
||||
`pub mod read_plan;` and re-export `ReadPlan` (and the plan sub-types
|
||||
that are part of the public surface — `ReadKind`, `CompositePlan`,
|
||||
etc. if the ADR's public-API section calls for them; the ADR lists
|
||||
`ReadPlan` as the public type, sub-types can stay `pub` in-module if
|
||||
consumers don't need to name them).
|
||||
|
||||
**Implementation notes:**
|
||||
- `compile` walks `BastDoc` once (via the existing borrowed `BastDoc`,
|
||||
which still exists at this phase — the owned-`BastDoc` refactor is
|
||||
phase 3). Resolves all `$ref`s eagerly, computes effective endianness
|
||||
at every node, inlines union variants. Malformed schemas surface as
|
||||
`AlkTypeError::Schema` (AGENTS.md §3); overflow-safe arithmetic
|
||||
(AGENTS.md §4).
|
||||
- `by_name: BTreeMap<String, usize>` (not `HashMap` — ADR-012 §1
|
||||
requires `Hash` on `ReadPlan`, and `HashMap` blocks derive). The
|
||||
POC used `HashMap`; swap to `BTreeMap`.
|
||||
- **Union shape (ADR-011 refined — resolves POC Finding 1 and the
|
||||
nested-union gap):** `CompositePlan::Union { disc, shared,
|
||||
variants: Vec<(String, CompositePlan)> }`.
|
||||
- `shared: Option<Box<ReadPlan>>` — the union's declared `fields`
|
||||
(the discriminator field + any shared fields) for the
|
||||
field-name-discriminator case. `DiscriminatorPlan::Field.field_index`
|
||||
indexes into `shared`. The read loop walks `shared` first, then
|
||||
looks up the selected variant and walks its `CompositePlan`
|
||||
starting after the shared fields. The byte-offset-discriminator
|
||||
case sets `shared: None` (no shared fields; the variant starts
|
||||
immediately after the discriminator size). This replaces the
|
||||
POC's stubbed `plan_read_union` `Field` arm.
|
||||
- `variants: Vec<(String, CompositePlan)>` — **not** the POC's
|
||||
`Vec<(String, VariantPlan)>`. Dropping `VariantPlan`/`VariantKind`
|
||||
means a variant's body is just a `CompositePlan`, so **nested
|
||||
unions** (a variant that is itself a union, which the 0.2.0
|
||||
reader supports via `resolve_and_walk_variant`'s `Union` arm at
|
||||
`sequential_reader.rs:800`) work by ordinary `CompositePlan`
|
||||
recursion — a variant can be `CompositePlan::Union { ... }`. No
|
||||
separate `VariantKind::Union` arm, no behavioral drop vs 0.2.0,
|
||||
no Semver Contract entry for a capability regression. The POC's
|
||||
`VariantPlan { kind, plan }` wrapper is not carried forward.
|
||||
- `compile_union` must reject a variant that is neither a struct
|
||||
nor a union with `AlkTypeError::Schema` (mirroring
|
||||
`resolve_and_walk_variant`'s `other => Err(...)` arm), so the
|
||||
eager-resolution path keeps the untrusted-input discipline.
|
||||
- Do *not* wire `ReadPlan` into `SequentialReader` or `materialize` yet
|
||||
— that's phase 2. Phase 1 is the type + `compile` only, unit-tested
|
||||
against the same BAST fixtures `bast.rs` uses (the existing `bast.rs`
|
||||
tests are a ready source of fixtures).
|
||||
|
||||
**Verification:** `cargo test --release` (new unit tests for `compile`
|
||||
covering every `BastType` arm — port the `cov_*` tests from
|
||||
`poc/readplan/tests/coverage.rs`, **plus** a nested-union-variant
|
||||
test asserting `compile_union` produces `CompositePlan::Union` whose
|
||||
variant body is itself `CompositePlan::Union`, restoring the 0.2.0
|
||||
capability the POC rejected). `cargo clippy --all-targets -- -D
|
||||
warnings`. `cargo doc --no-deps` (new public type). `cargo build
|
||||
--target wasm32-unknown-unknown --release` (`read_plan.rs` is
|
||||
wasm-relevant). **Add a `fn read_plan_is_send_sync()` assertion
|
||||
test** (a `const _: fn() = || { fn assert_send_sync<T: Send + Sync>()
|
||||
{}; assert_send_sync::<ReadPlan>(); };` static-bound assertion, as
|
||||
ADR-011 §"Engine integration" requires) to lock in `ReadPlan: Send +
|
||||
Sync` so a future change can't break it silently — mirror the POC's
|
||||
`readplan_is_send_sync` test.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — `SequentialReader` + `materialize_packed` consume `ReadPlan` (ADR-011 steps 2–4) — **DONE (2026-09-02)**
|
||||
|
||||
> **Status: implemented.** The packed read loop walks `Arc<ReadPlan>`:
|
||||
> `SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible; the old
|
||||
> fallible constructor's work moved to `ReadPlan::compile`), the reader
|
||||
> holds `plan: Arc<ReadPlan>` + cursor only, `schema()` returns the
|
||||
> `Arc<Value>` retained on the plan (review #005 H2 closed — no
|
||||
> self-referential struct), and a new `plan()` accessor exposes the
|
||||
> shared plan. `materialize_packed(&ReadPlan, &[u8])` walks the same
|
||||
> plan; the aligned materialize path keeps walking `BastDoc` with the
|
||||
> retained `dummy_field_for`/`ty_source`/`materialize_typeref_packed`
|
||||
> helpers (phase 5 Scope Boundary). Engine: `Layout::Packed` carries
|
||||
> `Arc<ReadPlan>`; `sequential_reader()` is an `Arc::clone` (15.7 ns,
|
||||
> was a whole-document `Value` clone); packed `validate_bytes` calls
|
||||
> `materialize_packed(&self.plan, ...)`. The temporary validation
|
||||
> bridge (reconstruct `BastDoc` for the validator) is still in place —
|
||||
> phase 7 already retired it on `main`'s ValidationPlan; this phase's
|
||||
> `validate_bytes` edit merged cleanly onto that state.
|
||||
>
|
||||
> **Stride (deferred decision 4):** fixed struct/nested-array elements
|
||||
> now report their true stride through `FieldValue::Array`
|
||||
> (0.2.0 returned `0`); doc comment updated; no existing test asserted
|
||||
> the `0`, so no test needed changing — the plan-compile tests lock the
|
||||
> new values.
|
||||
>
|
||||
> **Two parity subtleties found and preserved** (both invisible to the
|
||||
> existing test suite, both now locked by tests or by construction):
|
||||
> (a) the materializer unwraps the plan's anonymous single-field
|
||||
> wrapper for primitive array elements/record values — without this,
|
||||
> materialized records/arrays would nest each leaf under a synthetic
|
||||
> object and `validate_bytes` would fail its own parity suite (caught
|
||||
> by `materialize_record_packed_count_prefixed_pairs`); (b) the
|
||||
> field-disc union's materialized key order keeps `__discriminator`
|
||||
> first (matching 0.2.0's `Map` insertion order, observable under
|
||||
> `preserve_order`).
|
||||
>
|
||||
> **Bench (alktty `wire_vs_bast`, 1024 chunks/stream):** read p64
|
||||
> 2.27 µs/chunk (review #004) → **98 ns/chunk** (~23x; gap 400x →
|
||||
> ~17x vs hand-rolled's 5.6 ns); read p4k → 100 ns/chunk. `engine.
|
||||
> sequential_reader()` construction 15.7 ns (was a full `Value` clone).
|
||||
> The residual gap is dominated by the per-field `String` allocation
|
||||
> mandated by the unchanged `(String, FieldValue)` `read_next` return
|
||||
> signature (2 allocs/chunk) plus `data_access` bounds checks — both
|
||||
> outside this phase's scope (the signature is pinned by the Semver
|
||||
> Contract).
|
||||
|
||||
**Goal:** Rewrite the packed read loop to walk `&ReadPlan` instead of
|
||||
reconstructing `BastDoc`. `SequentialReader` stores `Arc<ReadPlan>` +
|
||||
cursor state; `materialize_packed` takes `&ReadPlan`. Closes H1
|
||||
(the 400x gap) + the packed-side M1 + L1.
|
||||
|
||||
**ADR reference:** [ADR-011 §Scope](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#scope),
|
||||
[ADR-011 §Recommended Order](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#recommended-order)
|
||||
steps 2–4.
|
||||
|
||||
**Files:** `src/sequential_reader.rs` (rewrite the read loop, change
|
||||
`new`'s signature), `src/materialize.rs` (`materialize_packed` takes
|
||||
`&ReadPlan`), `src/engine.rs` (`compile` builds `Arc<ReadPlan>` in
|
||||
packed mode, `sequential_reader()` hands out `Arc::clone`, packed
|
||||
`validate_bytes` calls `materialize_packed(&self.plan, ...)`).
|
||||
|
||||
**Implementation notes:**
|
||||
- `SequentialReader::new(&Value, &str) -> Result<Self, AlkTypeError>` →
|
||||
`SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible — the
|
||||
fallible `BastDoc` parse moved to `ReadPlan::compile` in phase 1;
|
||||
`new` just stores the `Arc`). The reader stores `plan: Arc<ReadPlan>`,
|
||||
`field_index: usize`, `position: usize`. `endian()` reads
|
||||
`self.plan.endian()`.
|
||||
- **`schema()` ownership (resolves review #005 H2):** `schema()`
|
||||
returns `&Value` retained on the plan, but the plan stores an
|
||||
**`Arc<Value>`**, not a `&Value`. ADR-011 §"Root cause" rejects the
|
||||
self-referential struct pattern (a `ReadPlan` storing `&Value`
|
||||
borrowing from the engine's `bast_doc: Value` would make the engine
|
||||
self-referential — exactly the construction ADR-007 worked around
|
||||
and ADR-011's `Arc<ReadPlan>` was meant to retire). The fix:
|
||||
`ReadPlan` carries `schema: Arc<Value>`; `ReadPlan::compile` clones
|
||||
the input `&Value` into `Arc<Value>` once; `schema()` returns
|
||||
`&self.schema`. The engine stores `bast_doc: Arc<Value>` internally
|
||||
(one allocation, shared via refcount between the engine and all
|
||||
plans it builds) — this is an internal change, not a public
|
||||
signature change (`compile` still takes `&Value`). `Arc<Value>`
|
||||
implements `Hash + Eq` (`serde_json::Value: Hash + Eq` as of the
|
||||
pinned `serde_json` with `preserve_order`; `Map::hash` sorts keys
|
||||
for determinism), so phase 6's `#[derive(Hash)]` on `ReadPlan` is
|
||||
not blocked by carrying the schema. **Note:** if a future
|
||||
`serde_json` version regresses `Value: Hash`, phase 6 would need
|
||||
`ReadPlan`'s hash to exclude the `schema` field; that is a phase-6
|
||||
concern, not a phase-2 blocker.
|
||||
- `read_field_at`/`read_field_value`/`read_typeref_value`/
|
||||
`walk_struct_size`/`read_union_value`/`read_array_value`/
|
||||
`read_record_value` are rewritten to take plan nodes
|
||||
(`&FieldPlan`/`&CompositePlan`/`&ReadKind`) instead of
|
||||
`&BastField`/`&BastType`/`&BastDoc`. Port `plan_read_field_at`/
|
||||
`plan_walk_struct_size`/etc. from `poc/readplan/src/lib.rs` — they're
|
||||
the reference implementations. The `read_union_value` rewrite
|
||||
handles both discriminator kinds via the refined `CompositePlan::Union`
|
||||
shape from phase 1: byte-disc reads the discriminator at
|
||||
`disc.offset` then walks the variant (no `shared`); field-disc walks
|
||||
`shared` first, reads the discriminator field at
|
||||
`disc.field_index` within `shared`, looks up the variant, and walks
|
||||
it starting after the shared fields. Nested unions (variant body is
|
||||
itself `CompositePlan::Union`) recurse naturally — no special arm.
|
||||
- **Struct-array stride (deferred decision 4):** the POC computes the
|
||||
true fixed-struct stride; the existing reader returns `0`. The
|
||||
production `ReadPlan::compile_array` should compute the true stride
|
||||
(the POC's `fixed_struct_size` helper). Document the behavioral
|
||||
change in the `FieldValue::Array` doc comment: `element_stride` is
|
||||
now the true stride for fixed-size struct elements, not `0`. This is
|
||||
a breaking behavioral change; rides the bump. Update the
|
||||
`eq_array_ref_element`-style test to assert the new stride.
|
||||
- **`materialize_packed` rewrite + packed/aligned split (resolves
|
||||
review #005 L1 and L2):** `materialize_packed(&BastDoc<'_>, &[u8])`
|
||||
→ `materialize_packed(&ReadPlan, &[u8])`. The materializer walks the
|
||||
plan instead of `BastDoc`. **Scoped removal of helpers:** only the
|
||||
*packed-side* call sites of `dummy_field_for`/`ty_source`
|
||||
(`materialize.rs:249, 316, 351, 391` — the packed
|
||||
`materialize_*_packed` paths) go away when packed-materialize moves
|
||||
to the plan. The helpers themselves **stay**, because
|
||||
`materialize_leaf_at` (`materialize.rs:631`, which calls
|
||||
`dummy_field_for`) is on the **aligned** path — it's called by
|
||||
`materialize_struct_aligned` (`:475`), `materialize_array_aligned`
|
||||
(`:544`), `materialize_variable_aligned` (`:613`). Aligned
|
||||
`materialize` keeps walking `BastDoc` through 0.3.0 (see the Scope
|
||||
Boundary note in phase 5), so `dummy_field_for`/`ty_source` must
|
||||
stay. **`materialize_typeref_packed` split:** this function is
|
||||
currently shared by both packed and aligned paths (aligned's
|
||||
`materialize_leaf_at` calls it to read leaves, and aligned's record
|
||||
path at `:498-506` calls it directly). After phase 2,
|
||||
packed-materialize gets a new plan-walking function;
|
||||
`materialize_typeref_packed` stays for aligned's
|
||||
`materialize_leaf_at` and the aligned record path (renamed or not —
|
||||
implementer's choice; the function is private). This is two
|
||||
mode-specific paths — the existing design — not a "parallel walker"
|
||||
in the maintenance-tax sense ADR-011 §"Negative" cautions against
|
||||
(that caution is about packed read-side `SequentialReader` +
|
||||
`materialize_packed` sharing one plan, which this preserves).
|
||||
- `AlkTypeEngine::compile` (packed branch): build `Arc<ReadPlan>` via
|
||||
`ReadPlan::compile(bast_doc, root_name)`, store it in `Layout::Packed`.
|
||||
`sequential_reader()` returns
|
||||
`Some(SequentialReader::new(Arc::clone(&self.plan)))`.
|
||||
`validate_bytes` (packed) calls
|
||||
`materialize_packed(&self.plan, buffer)` then runs validation on the
|
||||
materialized `Value`. **Validation path through phase 2:** until
|
||||
phase 7 (ValidationPlan), `validate_bytes` reconstructs a `BastDoc`
|
||||
for the validator only — the *read* path uses the plan, the
|
||||
*validation* path uses `BastDoc`. This is a temporary bridge: phase 7
|
||||
replaces it with a `ValidationPlan` walk (ADR-012 §3), retiring the
|
||||
per-call `BastDoc` reconstruction. The bridge is acceptable for
|
||||
phases 2–6 because the ValidationPlan work is committed in this
|
||||
release (not deferred), so the bridge has a known removal point in
|
||||
phase 7.
|
||||
|
||||
**Verification:** `cargo test --release` — the existing
|
||||
`sequential_reader.rs` and `materialize.rs` tests drive `read_next`/
|
||||
`read_field`/`reset`/`validate_bytes` through the public API, so they
|
||||
validate the rewrite without modification. If any test breaks, the
|
||||
rewrite diverged from the existing behavior — investigate before
|
||||
patching the test. `cargo clippy --all-targets -- -D warnings`.
|
||||
`cargo build --target wasm32-unknown-unknown --release`. **Re-run the
|
||||
alktty `wire_vs_bast` bench** to confirm the 400x gap closes (this is
|
||||
the headline result; record the before/after numbers in the commit
|
||||
message).
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — Owned `BastDoc` (ADR-012 §2a) — **DONE (2026-09-02)**
|
||||
|
||||
> **Status: implemented.** Every `Bast*` type dropped its `<'a>`:
|
||||
> `&'a str` → `String`, `&'a Value` → `Value` (deferred decision 1
|
||||
> resolved: plain `String`/`Value` — the tree is built once, name
|
||||
> sharing via `Arc<str>` needs a bench justification that doesn't
|
||||
> exist). `BastDoc::new(&Value, &str)` still takes references in and
|
||||
> clones into owned storage; `BastDoc` gained `Clone`. `resolve_*`
|
||||
> return owned `BastDef`/`BastType`. Consumers adapted:
|
||||
> `OffsetMap::compute(&BastDoc)`, `materialize_aligned(&BastDoc, ...)`
|
||||
> (no lifetime), `BuildCtx`/`ComputeCtx` hold `&'d BastDoc`.
|
||||
> **Engine ownership flip:** `AlkTypeEngine` now holds the owned
|
||||
> `BastDoc` (replacing `bast_doc: Value` + `root_name: String` —
|
||||
> `root_name()` delegates to the doc), killing its three per-call
|
||||
> `BastDoc::new` re-parses (`validate_bytes` aligned path,
|
||||
> `read_field`, `write_field` — review #004 M1 sites by construction;
|
||||
> phase 5 removes the `lookup_leaf_field` walk itself). New public
|
||||
> accessor `AlkTypeEngine::root_name()` (additive). **Bonus cleanup:**
|
||||
> `materialize_typeref_packed`'s dead `_field` param dropped (phase 2
|
||||
> left it dangling); under ownership it would have forced a deep
|
||||
> `Value` clone per array element/record value/union variant via
|
||||
> `dummy_field_for` — the param, `dummy_field_for`, and `ty_source`
|
||||
> are gone (no behavior change; the `_field` arg was already
|
||||
> ignored). `BastField::synthetic` retains an owned-signature
|
||||
> `#[allow(dead_code)]` definition (no remaining callers; kept for the
|
||||
> phase-4/5-aligned materializer helpers if they need it). Existing
|
||||
> `BastDoc` consumers (`LayoutBuilder`'s `doc_value` re-parse cache)
|
||||
> are unchanged pending phase 4. Engine stays `Send + Sync` with the
|
||||
> owned doc — the engine's thread-share test now asserts it directly.
|
||||
|
||||
**Goal:** Make `BastDoc` own its data (drop the `<'a>` lifetime).
|
||||
`&'a str` → `String`, `&'a Value` → `Value` (or `Arc<str>`/`Arc<Value>`
|
||||
— deferred decision 1). This is the prerequisite for the `LayoutBuilder`
|
||||
M1 fix (phase 4) and simplifies all owning consumers. Broad but
|
||||
mechanical refactor.
|
||||
|
||||
**ADR reference:** [ADR-012 §2a](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2a-layoutbuilder--cache-the-parsed-bastdoc-at-new).
|
||||
|
||||
**Files:** `src/bast.rs` (every `Bast*` type), every consumer:
|
||||
`src/layout_builder.rs`, `src/offset_map.rs`, `src/materialize.rs`,
|
||||
`src/bast_validation.rs`, `src/engine.rs`, `src/tunion.rs`,
|
||||
`src/bast_meta.rs` (if it walks `BastDoc`), `src/builder.rs` (if it
|
||||
consumes `Bast*`). `src/lib.rs` re-exports (signatures change but
|
||||
names stay).
|
||||
|
||||
**Implementation notes:**
|
||||
- `BastDoc<'a>` → `BastDoc`. Fields: `root: Value` (was `&'a Value`),
|
||||
`root_name: String` (was `&'a str`), `root_def: BastDef` (was
|
||||
`BastDef<'a>`). `new(root: &Value, root_name: &str)` takes references
|
||||
*in* (the caller still owns the input `Value`) but clones into owned
|
||||
storage. The `&Value` → `Value` clone is the cost of ownership; it
|
||||
happens once at `compile`/`new`, not per-field.
|
||||
- `BastDef<'a>` → `BastDef`: `name: String`, `kind: BastDefKind`,
|
||||
`source: Value`.
|
||||
- `BastStruct<'a>` → `BastStruct`: `endian: Endian`, `align: Option<usize>`,
|
||||
`fields: Vec<BastField>`, `source: Value`.
|
||||
- `BastField<'a>` → `BastField`: `name: String`, `ty: BastType`,
|
||||
`endian: Option<Endian>`, `align: Option<usize>`,
|
||||
`encoding: VariableEncoding`, `max_length: Option<usize>`,
|
||||
`source: Value`. `synthetic` constructor takes owned `BastType` +
|
||||
`Value`.
|
||||
- `BastType<'a>` → `BastType`: `Primitive(AlkTypeKind)`,
|
||||
`Ref(BastRef)`, `Array(BastArray)`, `Record(BastRecord)`,
|
||||
`Struct(BastStruct)`, `Union(BastUnion)`, `Enum(BastEnum)`.
|
||||
- `BastUnion<'a>` → `BastUnion`: `endian: Endian`,
|
||||
`discriminator: BastDiscriminator`, `fields: Vec<BastField>`,
|
||||
`mapping: Vec<(String, BastType)>` (was `Vec<(&'a str, BastType)>`),
|
||||
`source: Value`.
|
||||
- `BastDiscriminator::Field { name: String }` (was `name: &'a str`).
|
||||
- `BastEnum<'a>` → `BastEnum`: `values: Vec<String>` (was
|
||||
`Vec<&'a str>`), `source: Value`.
|
||||
- `BastArray<'a>` → `BastArray`: `element: Box<BastType>`, `count: usize`,
|
||||
`source: Value`.
|
||||
- `BastRecord<'a>` → `BastRecord`: `values: Box<BastType>`,
|
||||
`source: Value`.
|
||||
- `BastRef<'a>` → `BastRef`: `name: String`.
|
||||
- **`resolve_typeref` / `resolve_ref` / `lookup_def`** now return owned
|
||||
`BastType`/`BastDef`/`Value` instead of borrowed. The `clone()` in
|
||||
the current `resolve_typeref` passthrough (`other => Ok(other.clone())`)
|
||||
is no longer needed for the borrow case (everything is owned) but
|
||||
the logic is unchanged — `BastType` is `Clone` either way.
|
||||
- **Consumers adapt:** any code that held `&'a Value` alongside a
|
||||
`BastDoc<'a>` (e.g. `LayoutBuilder.doc_value`, `AlkTypeEngine.bast_doc`,
|
||||
`SequentialReader.doc_value` — though the reader is already on
|
||||
`ReadPlan` after phase 2) drops the separate `Value` and holds the
|
||||
owned `BastDoc` directly. `materialize_packed` is already on
|
||||
`ReadPlan` (phase 2) and doesn't need `BastDoc` — unaffected.
|
||||
`materialize_aligned` takes `&BastDoc` (owned, no lifetime).
|
||||
- **`Arc<str>` vs `String` (deferred decision 1):** default to
|
||||
`String`. The typed tree is built once; name sharing via `Arc<str>`
|
||||
is a micro-optimization not justified without a bench. If phase 4's
|
||||
`LayoutBuilder` work shows name allocation is measurable, revisit.
|
||||
|
||||
**Verification:** `cargo test --release` — the existing `bast.rs` tests
|
||||
are the primary validation (they exercise every parser path). All
|
||||
`Bast*`-consuming tests must pass unchanged (they go through public
|
||||
APIs that still take `&Value`/`&str` in, just return owned types out).
|
||||
`cargo clippy --all-targets -- -D warnings`. `cargo doc --no-deps`.
|
||||
`cargo build --target wasm32-unknown-unknown --release` (`bast.rs` is
|
||||
wasm-relevant).
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — `LayoutBuilder` caches the owned `BastDoc` (ADR-012 §2a) — **DONE (2026-09-02)**
|
||||
|
||||
> **Status: implemented.** `LayoutBuilder` stores `doc: BastDoc` +
|
||||
> `endian` (the `doc_value: Value` + `root_name: String` cache is
|
||||
> gone); `new` parses once, `build` walks `&self.doc` — the
|
||||
> `layout_builder.rs` re-parse (M1) is retired. The `build`-time
|
||||
> root-is-struct re-check replaced its `unreachable!()` with a clean
|
||||
> `Schema` error (AGENTS.md §3 never-panic; the invariant is
|
||||
> unchanged — `new` already rejects non-struct roots). Boxing fallout:
|
||||
> the builder now holds the full owned tree, so `Layout::Packed`
|
||||
> boxes it (`builder: Box<LayoutBuilder>`) to keep the engine's
|
||||
> `Layout` enum variant sizes balanced (clippy
|
||||
> `large_enum_variant`); `layout_builder()` still returns
|
||||
> `Option<&LayoutBuilder>` via auto-deref, public API unchanged.
|
||||
|
||||
**Goal:** `LayoutBuilder::new` parses the owned `BastDoc` once and
|
||||
stores it; `build` reuses it. Removes the `layout_builder.rs:190`
|
||||
re-parse (M1). Non-breaking from the public API perspective
|
||||
(`new`/`build` signatures unchanged); the change is internal.
|
||||
|
||||
**ADR reference:** [ADR-012 §2a](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2a-layoutbuilder--cache-the-parsed-bastdoc-at-new).
|
||||
|
||||
**Files:** `src/layout_builder.rs`.
|
||||
|
||||
**Implementation notes:**
|
||||
- `LayoutBuilder` currently stores `doc_value: Value` + `root_name:
|
||||
String` + `endian: Endian`. After phase 3, it stores `doc: BastDoc`
|
||||
(owned) + `endian: Endian`. `new` calls `BastDoc::new` once;
|
||||
`build(&self, var_sizes)` uses `&self.doc` directly — no
|
||||
`BastDoc::new` call inside `build`.
|
||||
- The `BuildCtx<'d>` struct (currently `doc: &'d BastDoc<'d>`) becomes
|
||||
`doc: &BastDoc` (no lifetime, or a single lifetime for the borrow
|
||||
from `&self`). The walk logic is unchanged.
|
||||
- The `doc_value: Value` clone is removed; the builder holds the owned
|
||||
`BastDoc` directly. `endian` is read from the doc at `new` time
|
||||
(already is).
|
||||
|
||||
**Verification:** `cargo test --release` — the existing
|
||||
`layout_builder.rs` tests pass unchanged (they go through
|
||||
`LayoutBuilder::new` + `build`). `cargo clippy --all-targets -- -D
|
||||
warnings`. `cargo build --target wasm32-unknown-unknown --release`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — `OffsetMap` carries `LeafMeta` (ADR-012 §2b) — **DONE (2026-09-02)**
|
||||
|
||||
> **Status: implemented.** Prerequisite first (review #005 M2):
|
||||
> `Hash` added to `Endian`/`VariableEncoding` derives in `schema.rs`
|
||||
> (additive; both are fieldless `Eq` enums). New public types
|
||||
> `LeafMeta { kind, encoding, endian }` (`Copy + PartialEq + Eq +
|
||||
> Hash`) and `OffsetEntry { range, meta }` (with `start()`/`end()`
|
||||
> convenience accessors), both re-exported from `lib.rs`; `ByteRange`
|
||||
> gained `Hash` (additive). Storage is `Vec<(String, OffsetEntry)>`
|
||||
> (deferred decision 2 resolved: struct — `get` returns
|
||||
> `Option<&OffsetEntry>`, `iter` yields `(&str, &OffsetEntry)`).
|
||||
> `LeafMeta` is computed at `compute` time with **effective** endian
|
||||
> threaded through the walk: container default → field override per
|
||||
> field, propagated into nested-struct probes and array elements via
|
||||
> the referring field (the same propagation the aligned materializer
|
||||
> uses). **Parity note:** this replaces `engine.rs`'s
|
||||
> `lookup_leaf_field` walk, which computed nested-struct defaults from
|
||||
> the *nested struct's own* `endian` annotation — the two paths
|
||||
> diverged whenever a nested struct declared `endian` and its
|
||||
> referring field also declared one (the map now agrees with the
|
||||
> aligned materializer and the packed `ReadPlan`; the old divergence
|
||||
> was unreachable through `read_field` only when a nested annotation
|
||||
> existed, and no test pinned it). `read_field`/`write_field` dispatch
|
||||
> on the entry's `LeafMeta` — the `BastDoc` re-parse +
|
||||
> `lookup_leaf_field`/`LeafFieldInfo` per access are gone (the last
|
||||
> two M1 sites, engine.rs `read_field`/`write_field`). Behavior
|
||||
> change: `read_field` on a path absent from the map (e.g. a
|
||||
> whole-struct field path) now errors with `Offset` ("field not found
|
||||
> in offset map") instead of `Access` ("does not support composite
|
||||
> types") — the composite-path test already accepted either variant.
|
||||
> `materialize_aligned`'s four `offset_map.get` call sites updated to
|
||||
> `.range.start`. `alktty`/`alkcall` untouched (the bench never uses
|
||||
> `OffsetMap::get`; alkcall has no dependency yet).
|
||||
|
||||
**Goal:** Extend `OffsetMap`'s entries with `LeafMeta { kind, encoding,
|
||||
endian }` computed at `compute` time. `read_field`/`write_field` drop
|
||||
the `BastDoc::new` + `lookup_leaf_field` calls (M1 aligned-side).
|
||||
Closes the last two M1 sites (`engine.rs:334,467`).
|
||||
|
||||
**ADR reference:** [ADR-012 §2b](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2b-offsetmap--carry-leaf-metadata).
|
||||
|
||||
**Files:** `src/offset_map.rs` (extend entries, compute `LeafMeta`),
|
||||
`src/engine.rs` (rewrite `read_field`/`write_field` to use the map's
|
||||
`LeafMeta`, remove `lookup_leaf_field` + `LeafFieldInfo`), `src/lib.rs`
|
||||
(re-export `LeafMeta`).
|
||||
|
||||
**Implementation notes:**
|
||||
- **`Hash` on `Endian`/`VariableEncoding` (resolves review #005 M2 —
|
||||
do this first, it's a prerequisite):** `src/schema.rs:205` (`Endian`)
|
||||
and `:212` (`VariableEncoding`) currently derive only `Debug, Clone,
|
||||
Copy, PartialEq, Eq` — no `Hash`. `LeafMeta` (below) requires all
|
||||
its fields to be `Hash` for `#[derive(Hash)]`, and phase 6's
|
||||
`#[derive(Hash)]` on `ReadPlan`/`OffsetMap` requires `FieldPlan`'s
|
||||
`endian: Endian` + `encoding: VariableEncoding` to be `Hash`. Add
|
||||
`Hash` to both derives in `src/schema.rs`. Both are fieldless enums
|
||||
already at `Eq + PartialEq`, so this is additive and semver-safe —
|
||||
no behavioral change. Trivial, but it's an unstated prerequisite
|
||||
the original plan omitted.
|
||||
- New public type `LeafMeta { kind: AlkTypeKind, encoding:
|
||||
VariableEncoding, endian: Endian }`. `Copy + PartialEq + Eq + Hash`
|
||||
(all fields are `Copy + Hash` once the sub-step above is done —
|
||||
`AlkTypeKind` already derives `Hash`; `Endian`/`VariableEncoding`
|
||||
get it from the sub-step above).
|
||||
- `OffsetMap` storage: `fields: Vec<(String, ByteRange, LeafMeta)>`
|
||||
(was `Vec<(String, ByteRange)>`). The `compute` walk already resolves
|
||||
each leaf's type; add the `LeafMeta` extraction at the point where
|
||||
the leaf `ByteRange` is recorded.
|
||||
- **`OffsetMap::get` return type (deferred decision 2):** change to
|
||||
`get(field_path) -> Option<&OffsetEntry>` where `pub struct
|
||||
OffsetEntry { range: ByteRange, meta: LeafMeta }`. Add
|
||||
`OffsetEntry` to `lib.rs` re-exports. Callers that used
|
||||
`map.get(path).unwrap().start` become
|
||||
`map.get(path).unwrap().range.start`. Update `alktty`/`alkcall` call
|
||||
sites (in-house).
|
||||
- `engine.rs::read_field`/`write_field`: drop the
|
||||
`BastDoc::new(&self.bast_doc, &self.root_name)?` +
|
||||
`lookup_leaf_field(&doc, field_path)?` calls. Read `LeafMeta`
|
||||
from `offset_map.get(field_path)?.meta`. The `kind`/`encoding`/
|
||||
`endian` match arms in `read_field`/`write_field` are unchanged
|
||||
(they already dispatch on `AlkTypeKind`/`VariableEncoding`/`Endian`).
|
||||
- Remove `LeafFieldInfo` and `lookup_leaf_field` from `engine.rs`
|
||||
(subsumed by `LeafMeta` on the map).
|
||||
- `materialize_aligned` also uses `OffsetMap` — it currently calls
|
||||
`offset_map.get(&path)?.start` for leaf reads. Update those call
|
||||
sites to `.range.start`. The materializer's `resolve_typeref` calls
|
||||
for composite walks stay (composites aren't in the offset map as
|
||||
leaves; they're walked recursively). The `BastDoc` argument to
|
||||
`materialize_aligned` is now owned (phase 3) — no signature change
|
||||
beyond the lifetime drop.
|
||||
- **Scope Boundary — aligned `materialize`'s `BastDoc` structure walk
|
||||
(resolves review #005 L3):** `materialize_struct_aligned`
|
||||
(`materialize.rs:451-521`) walks `BastDoc` to traverse
|
||||
struct/array/record *structure*, using `OffsetMap` only for leaf
|
||||
byte positions. This is the **permanent design for 0.3.0**, not a
|
||||
deferral: after phase 3 the walk is over owned data (no re-parse,
|
||||
not O(N²)), and aligned `validate_bytes` is one structure walk per
|
||||
call (not per-field), so there is no perf driver analogous to review
|
||||
#004's packed per-chunk gap. ADR-011 §"Out of scope" is half-true
|
||||
here (aligned materialize takes `&OffsetMap` *and* `&BastDoc`) —
|
||||
this note owns the decision: aligned materialize keeps walking owned
|
||||
`BastDoc` for structure through 0.3.0. An `AlignedPlan` that
|
||||
compiles the structure walk is **not** in scope; if a future bench
|
||||
shows an aligned-mode hot loop, it gets its own ADR (tracked as an
|
||||
open question, not a silent gap). Phase 7's `ValidationPlan` does
|
||||
not change this — validation is value-domain, orthogonal to the
|
||||
aligned structure walk.
|
||||
|
||||
**Verification:** `cargo test --release` — existing `offset_map.rs`
|
||||
and `engine.rs` `read_field`/`write_field` tests pass (they go through
|
||||
public APIs). `cargo clippy --all-targets -- -D warnings`. `cargo doc
|
||||
--no-deps` (new public `LeafMeta`/`OffsetEntry`). `cargo build --target
|
||||
wasm32-unknown-unknown --release`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6 — Fingerprinting `ReadPlan`/`OffsetMap` (ADR-012 §1, §4) — **DONE (2026-09-02)**
|
||||
|
||||
> **Status: implemented.** `#[derive(Hash, Eq)]` added to `ReadPlan`,
|
||||
> `FieldPlan`, `CompositePlan`, `ReadKind`, `DiscriminatorPlan`
|
||||
> (`ReadPlan`'s `schema: Arc<Value>` hashes fine — `serde_json::Value:
|
||||
> Hash + Eq` under the pinned `preserve_order` serde_json) and to
|
||||
> `OffsetMap` (`Clone` added alongside; its `LeafMeta`/`OffsetEntry`/
|
||||
> `ByteRange` payload gained `Hash` in phase 5 / this phase). The
|
||||
> POC's `VariantPlan`/`VariantKind` don't exist in the production
|
||||
> shape (phase 1 dropped them). `fingerprint() -> u64` on both via
|
||||
> `DefaultHasher` (deferred decision 3 resolved: std `DefaultHasher`,
|
||||
> no new dep; the fingerprint isn't hot; cross-version stability is a
|
||||
> non-goal per ADR-012). Fingerprint contract tests on both: same
|
||||
> schema twice → equal `PartialEq` + equal fingerprint; field-kind
|
||||
> change, field-order change, and endianness change each → different
|
||||
> fingerprints; (ReadPlan) different root names over the same document
|
||||
> → different fingerprints. `ValidationPlan` already carries its own
|
||||
> `Hash + Eq` + `fingerprint` + contract test (phase 7).
|
||||
|
||||
**Goal:** Add `Hash + Eq` derives + `fingerprint() -> u64` to `ReadPlan`
|
||||
and `OffsetMap`. Enables cross-run caching, `alkcall` schema handshake,
|
||||
schema-version diagnostics. (`ValidationPlan` gets the same treatment
|
||||
in phase 7, where it's built — it carries its own `Hash + Eq` +
|
||||
`fingerprint()` as part of its public surface.)
|
||||
|
||||
**ADR reference:** [ADR-012 §1](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#1-fingerprinting--readplan-hash--eq-offsetmap-hash--eq),
|
||||
[ADR-012 §4](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#4-fingerprinting-offsetmap-bundled-with-2b).
|
||||
|
||||
**Files:** `src/read_plan.rs` (derives + `fingerprint`), `src/offset_map.rs`
|
||||
(derives + `fingerprint`), `src/lib.rs` (no new re-exports — `Hash`/`Eq`
|
||||
are trait derives, `fingerprint` is an inherent method).
|
||||
|
||||
**Implementation notes:**
|
||||
- `ReadPlan` already uses `BTreeMap` for `by_name` (phase 1), so
|
||||
`#[derive(Hash, Eq)]` works. Add it alongside the existing
|
||||
`Debug, Clone, PartialEq`. Same for `FieldPlan`, `CompositePlan`,
|
||||
`ReadKind`, `DiscriminatorPlan`. (The POC's `VariantPlan`/`VariantKind`
|
||||
are not in the production shape — phase 1 dropped them — so they are
|
||||
not derived here.) `ReadPlan` also carries `schema: Arc<Value>` from
|
||||
phase 2; `Arc<Value>: Hash + Eq` because `serde_json::Value: Hash +
|
||||
Eq` (with `preserve_order`, `Map::hash` sorts keys deterministically),
|
||||
so the `schema` field does not block the derive. If a future
|
||||
`serde_json` version regresses `Value: Hash`, exclude `schema` from
|
||||
the derived `Hash` via a manual `impl Hash for ReadPlan` that hashes
|
||||
every field except `schema` — phase-6 concern, not a blocker.
|
||||
- `OffsetMap` already carries `LeafMeta` (phase 5), and `LeafMeta` is
|
||||
`Copy + Hash + Eq` (phase 5 added `Hash` to `Endian`/`VariableEncoding`).
|
||||
Add `#[derive(Hash, Eq)]` to `OffsetMap`,
|
||||
`OffsetEntry`, `ByteRange` (already `Eq + Hash`), `LeafMeta`.
|
||||
- **Fingerprint hasher (deferred decision 3):** `DefaultHasher` (std,
|
||||
no new dep). The fingerprint isn't hot; cross-version stability is a
|
||||
non-goal. `fingerprint()`:
|
||||
```rust
|
||||
pub fn fingerprint(&self) -> u64 {
|
||||
use std::hash::{Hash, Hasher};
|
||||
let mut h = std::hash::DefaultHasher::new();
|
||||
self.hash(&mut h);
|
||||
h.finish()
|
||||
}
|
||||
```
|
||||
- **Fingerprint contract test:** compile the same schema twice, assert
|
||||
`plan1 == plan2` and `plan1.fingerprint() == plan2.fingerprint()`.
|
||||
Compile a schema with one field changed, assert fingerprints differ.
|
||||
This is the contract test for ADR-012 §1's "two plans with equal
|
||||
hashes produce identical reads over identical bytes."
|
||||
|
||||
**Verification:** `cargo test --release` (new contract tests).
|
||||
`cargo clippy --all-targets -- -D warnings`. `cargo doc --no-deps`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7 — `ValidationPlan` (ADR-012 §3) — **DONE (2026-08-31)**
|
||||
|
||||
> **Status: implemented.** The design session ran and the shape landed
|
||||
> in `src/validation_plan.rs`. Summary of what was decided and built
|
||||
> (full detail in ADR-012 §3a):
|
||||
>
|
||||
> - **Shape:** `ValidationPlan { root: ValidNode }` — a constraint tree
|
||||
> with one arm per value-domain check (`Int`/`I64`/`Uint`/`U64`/
|
||||
> `Float`/`Bool`/`Str`/`Bytes`/`Enum`/`Struct`/`Union`/`Array`/
|
||||
> `Record`), `ValidField { name, node }`, `ValidVariant { key, node }`.
|
||||
> `maxLength` baked into leaf nodes from the owning field at compile
|
||||
> time. Unions compile to variant nodes only (the declared union
|
||||
> fields are validated via the variant walk, matching the interpretive
|
||||
> arm's dispatch-on-`__discriminator` semantics).
|
||||
> - **`compile` signature:** `compile(&BastDoc) -> Result<Self,
|
||||
> AlkTypeError>` (the plan table's `&str` param was vestigial).
|
||||
> - **`bast_validation`:** the interpretive walker is *retired*
|
||||
> (deleted, not just bypassed); `validate_value` survives as a
|
||||
> compile-once-per-call wrapper over the plan (one-shot/diagnostic
|
||||
> use); shared error helper retained.
|
||||
> - **Engine:** `Arc<ValidationPlan>` built at `compile` in both modes;
|
||||
> new accessor `validation_plan()`. `validate_bytes` walks the plan.
|
||||
> **Bonus:** `ValidationPlan::compile` runs before the layout build and
|
||||
> serves as the engine's cyclic-`$ref` gate. (As of the review #006 H2
|
||||
> fix, the layout walkers also carry their own guard —
|
||||
> `walk_guard::check_ref_graph` runs at each standalone entry — so this
|
||||
> ordering is now belt-and-suspenders rather than the only defense; the
|
||||
> plan text below predates that fix.)
|
||||
> - **Remaining phase-7 bench work** (a `validate_bytes`-stream bench
|
||||
> in alktty) moves with the bench work into phase 8; a spot check
|
||||
> during development measured plan-validate at ~0.2 µs/call vs ~0.6
|
||||
> µs for the compile-per-call one-shot it replaced.
|
||||
|
||||
**Goal:** Retire the interpretive `bast_validation` walk. Introduce a
|
||||
`ValidationPlan` — a compile-once-walk-many compiled form over the
|
||||
BAST document's value-domain constraints — built once at `compile`
|
||||
time and walked by `validate_bytes` (both modes) per buffer instead
|
||||
of re-walking `BastDoc`. Closes the latent perf cliff review #005 M3
|
||||
flagged: after ADR-011 the *read* half of `validate_bytes` is
|
||||
plan-fast, but the *validation* half still re-walks `BastDoc` per
|
||||
buffer, which is hot on the `alkcall` read+validate-on-untrusted-stream
|
||||
common case.
|
||||
|
||||
**ADR reference:** [ADR-012 §3](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#3-validationplan--compile-once-validation-form).
|
||||
|
||||
**Predecessor for this phase:** a **design session** to scope the
|
||||
concrete `ValidationPlan` shape (constraint representation,
|
||||
`compile`/walk structure, `bast_validation` public-surface review)
|
||||
**before** implementation begins. ADR-012 §3 fixes the decision (in
|
||||
0.3.0, compiled form, no per-buffer `BastDoc` walk, `Hash + Eq` +
|
||||
`fingerprint`) and lists what is *not* decided (the struct/enum
|
||||
shape, the constraint descriptors, whether `validate_value` is
|
||||
retired or kept as a wrapper). This phase implements whatever the
|
||||
design session scopes; the contract below holds regardless of shape.
|
||||
|
||||
**Files:** New `src/validation_plan.rs` (the `ValidationPlan` type,
|
||||
`compile`, walk entry points). `src/bast_validation.rs` (adopt the
|
||||
plan; `validate_value` either becomes a thin wrapper over the plan
|
||||
or is retired per the design session's call). `src/engine.rs`
|
||||
(`compile` builds `Arc<ValidationPlan>` in both modes, stores it;
|
||||
`validate_bytes` walks `&self.validation_plan` instead of
|
||||
reconstructing a `BastDoc` for the validator — this removes the
|
||||
temporary bridge from phase 2). `src/lib.rs` (re-export
|
||||
`ValidationPlan`).
|
||||
|
||||
**Implementation contract (fixed by ADR-012 §3, independent of
|
||||
shape):**
|
||||
- **Compile-once-walk-many.** `ValidationPlan::compile` walks the
|
||||
owned `BastDoc` once (phase 3 made it owned); `validate_bytes`
|
||||
walks the `ValidationPlan` per buffer, never `BastDoc`. The only
|
||||
consumers that walk `BastDoc` interpretively after this phase are
|
||||
the one-shot `*::compile` paths (`ReadPlan::compile`,
|
||||
`OffsetMap::compute`, `ValidationPlan::compile`,
|
||||
`LayoutBuilder::new`).
|
||||
- **Value-domain, not byte-position.** The plan carries constraint
|
||||
descriptors (enum allowed-sets, integer range bounds, `maxLength`
|
||||
caps, union variant keys, and any other value-domain checks
|
||||
`bast_validation` performs today), keyed for dispatch against the
|
||||
materialized `Value` tree. The shape is different from
|
||||
`ReadPlan`/`OffsetMap`; the pattern (compiled form, immutable,
|
||||
shared via `Arc`) is the same.
|
||||
- **`Send + Sync`.** `ValidationPlan: Send + Sync` (immutable owned
|
||||
data, no interior mutability) so `Arc<ValidationPlan>` shares from
|
||||
the `Send + Sync` engine. Add a `static` bound assertion test
|
||||
mirroring phase 1's `read_plan_is_send_sync`.
|
||||
- **`Hash + Eq` + `fingerprint()`.** `ValidationPlan` derives
|
||||
`Debug, Clone, PartialEq, Eq, Hash` and has
|
||||
`fingerprint() -> u64` (same `DefaultHasher` implementation as
|
||||
phase 6). The fingerprint contract generalizes: two validation
|
||||
plans with equal hashes accept/reject identical `(bytes)`
|
||||
identically. Add a fingerprint contract test (compile the same
|
||||
schema twice, assert `plan1 == plan2` and `plan1.fingerprint() ==
|
||||
plan2.fingerprint()`; change one constraint, assert fingerprints
|
||||
differ).
|
||||
- **Untrusted-input discipline.** `compile` surfaces malformed
|
||||
schemas as `AlkTypeError::Schema` (AGENTS.md §3); overflow-safe
|
||||
arithmetic (AGENTS.md §4). No `unsafe`, no `async`, no new deps,
|
||||
wasm-clean (AGENTS.md §5–§11).
|
||||
|
||||
**What this phase does *not* include (shape-dependent, scoped by the
|
||||
design session):** the concrete `ValidationPlan` struct/enum, the
|
||||
constraint-descriptor representation, the `bast_validation`
|
||||
public-surface decision (`validate_value` retire-vs-wrapper), and any
|
||||
`AlkTypeError::Validation` variant changes. These are shape questions
|
||||
the design session resolves; they are *not* a re-opening of the
|
||||
"ship in 0.3.0" decision, which is fixed in ADR-012 §3.
|
||||
|
||||
**Verification:** `cargo test --release` — the existing
|
||||
`bast_validation.rs` and `engine.rs` `validate_bytes` tests are the
|
||||
primary validation (they drive validation through the public API and
|
||||
must pass unchanged, confirming behavioral parity with the
|
||||
interpretive walk). New unit tests for `ValidationPlan::compile`
|
||||
covering every constraint kind. New `Send + Sync` assertion test.
|
||||
New fingerprint contract tests. `cargo clippy --all-targets -- -D
|
||||
warnings`. `cargo doc --no-deps` (new public type). `cargo build
|
||||
--target wasm32-unknown-unknown --release` (`validation_plan.rs` is
|
||||
wasm-relevant). **Re-run the alktty `wire_vs_bast` bench** and, if
|
||||
the design session scopes one, a `validate_bytes`-on-untrusted-stream
|
||||
bench alongside `wire_vs_bast` to confirm the validation half of
|
||||
`validate_bytes` no longer dominates per-buffer.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8 — Public API bump, docs, verification (ADR-011 step 6, ADR-012) — **DONE (2026-09-02)**
|
||||
|
||||
> **Status: implemented.** Version flipped 0.2.0 → 0.3.0; `lib.rs`
|
||||
> re-exports complete (`ReadPlan` + sub-types, `LeafMeta`,
|
||||
> `OffsetEntry`, `ValidationPlan` + sub-types from earlier phases).
|
||||
> Docs: ADR-007 "Cost" section rewritten to the `Arc<ReadPlan>` cost
|
||||
> (15.7 ns) with the old framing as a historical note (review #004 L2
|
||||
> closed); ADR-011/012 status blocks flipped to implemented; the
|
||||
> architecture README ADR table rows updated; `layout-engine.md`
|
||||
> rewritten for the 0.3.0 surface (`SequentialReader` construction via
|
||||
> the engine factory, `OffsetMap` `OffsetEntry`/`LeafMeta`/`fingerprint`
|
||||
> public-types section, `OffsetMap::compute(&BastDoc)` owned signature);
|
||||
> `SequentialReader` module doc now points at the engine factory.
|
||||
> Reviews #004 and #005 status flipped to closed. Bench (alktty
|
||||
> `wire_vs_bast`, re-run on the 0.3.0 tree): read p64 98 ns/chunk
|
||||
> (hand-rolled 5.7 µs/stream — parity held from phase 2), read p4k
|
||||
> unchanged, `alktype_layout_build` **180 ns** (was ~1.2 µs — the
|
||||
> phase-4 owned-doc cache removed the per-build re-parse, ~7x),
|
||||
> `sequential_reader_new` 15.7 ns (unchanged), write p64 −3%
|
||||
> (37.9 µs), `engine_compile` 590 µs (unchanged; dominated by
|
||||
> meta-schema validation). No `validate_bytes`-stream bench was added:
|
||||
> the phase-7 spot check (~0.2 µs/call plan-validate vs ~0.6 µs
|
||||
> compile-per-call) stands as the validation-half measurement; a
|
||||
> dedicated bench remains a follow-up if `alkcall` profiling motivates
|
||||
> it. Downstream: `alktty` compiles against the path dep unchanged
|
||||
> (the bench uses `LayoutBuilder::new`/`build` and
|
||||
> `engine.sequential_reader()` — no touched signatures);
|
||||
> `alkcall` has no dependency yet.
|
||||
|
||||
**Goal:** Flip the version to 0.3.0, update `lib.rs` re-exports, update
|
||||
the architecture docs (ADR-007 "Cost" rewrite, ADR-011/012 status flip
|
||||
if not already, README ADR table), update in-house downstream
|
||||
consumers, run the full verification block.
|
||||
|
||||
**ADR reference:** [ADR-011 §Public API change](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#public-api-change-breaking--version-bump-to-030),
|
||||
[ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md).
|
||||
|
||||
**Files:** `Cargo.toml` (version 0.2.0 → 0.3.0), `src/lib.rs`
|
||||
(re-export `ReadPlan`, `LeafMeta`, `OffsetEntry`, `ValidationPlan`),
|
||||
`docs/architecture/` (README ADR table, ADR-007 "Cost" section,
|
||||
ADR-011/012 status), `docs/architecture/validation.md` /
|
||||
`layout-engine.md` (mention `ReadPlan`/`LeafMeta`/`ValidationPlan`
|
||||
where relevant), in-house downstream repos (`alktty`, `alkcall` —
|
||||
update call sites for `OffsetMap::get`, `SequentialReader::new`,
|
||||
`materialize_packed`, `BastDoc` owned, `validate_bytes` internal
|
||||
change if any signature change surfaced in phase 7's shape work).
|
||||
|
||||
**Implementation notes:**
|
||||
- **ADR-007 "Cost" section (L2 from review #004):** rewrite the
|
||||
"re-parse on demand" paragraph to describe the `Arc<ReadPlan>` cost
|
||||
and the owned-`BastDoc` cache. The factory decision itself stays
|
||||
"Accepted." This is the last loose end from review #004.
|
||||
- **`src/engine.rs:112-115` doc comment (L2):** rewrite the "re-parse
|
||||
the typed tree on demand" comment to describe the compiled-form
|
||||
architecture (`ReadPlan` for packed reads, `OffsetMap`+`LeafMeta`
|
||||
for aligned, `ValidationPlan` for validation, owned `BastDoc` for
|
||||
the builder and the `*::compile` paths).
|
||||
- **`lib.rs` re-exports:** add `ReadPlan`, `LeafMeta`, `OffsetEntry`,
|
||||
`ValidationPlan`. `BastDoc` and `Bast*` stay re-exported (signatures
|
||||
changed in phase 3, names unchanged). `materialize_packed`/
|
||||
`materialize_aligned` stay re-exported (signatures changed).
|
||||
`SequentialReader` stays re-exported (`new` signature changed).
|
||||
- **Downstream updates:** `alktty`'s bench (`benches/wire_vs_bast.rs`)
|
||||
updates `SequentialReader::new` call + any `OffsetMap::get` usage.
|
||||
`alkcall` updates similarly. Both are in-house path dev-deps; the
|
||||
updates ride this release's commits (or follow-on commits in those
|
||||
repos — they're separate repos, but the path dev-dep means a local
|
||||
update is immediate).
|
||||
- **`Cargo.toml` version bump:** `0.2.0` → `0.3.0`. The workspace
|
||||
section added for the POC (`[workspace] members = ["poc/readplan"]`)
|
||||
stays on the `readplan-poc` branch and is *not* merged to main — the
|
||||
POC branch is derisking-only, like `bast-validator-poc`. If the POC
|
||||
files ever merge to main, drop the workspace section (the POC is
|
||||
disposable).
|
||||
|
||||
**Verification block (run all, all must pass):**
|
||||
```bash
|
||||
cargo test --release # full suite
|
||||
cargo clippy --all-targets -- -D warnings
|
||||
cargo doc --no-deps # new public types
|
||||
cargo build --target wasm32-unknown-unknown --release # wasm-clean
|
||||
cargo publish --dry-run --allow-dirty # before publish
|
||||
```
|
||||
Plus: **re-run the alktty `wire_vs_bast` bench** and record the
|
||||
before/after numbers in the release commit message. The 400x gap
|
||||
should close to within ~2–5x of hand-rolled (the `data_access` calls
|
||||
are the same; the remaining gap is the `match` dispatch + `Arc` refcount
|
||||
vs hand-rolled's direct calls). The SFTP-shaped union case (the one
|
||||
ADR-011's framing argument cared about) should close further because
|
||||
eager `$ref` resolution removes the `resolve_typeref_as_def` per-
|
||||
variant dispatch cost.
|
||||
|
||||
---
|
||||
|
||||
## Cross-phase invariants
|
||||
|
||||
- **The tree builds and tests pass at every phase boundary.** No phase
|
||||
leaves the crate in a non-compiling state. Phases 1 (add `ReadPlan`),
|
||||
6 (add `Hash`/`Eq` derives to `ReadPlan`/`OffsetMap`), and 7 (add
|
||||
`ValidationPlan`) are pure additions; phases 2–5 are rewrites that
|
||||
must leave tests green; phase 8 is the bump/docs.
|
||||
- **The POC on `readplan-poc` is the reference scaffold for phases 1–2.**
|
||||
It is *not* merged to main; it stays on the branch as the derisking
|
||||
record, like `bast-validator-poc`. If a phase 1–2 implementation
|
||||
question arises about the plan shape, consult the POC. (The POC's
|
||||
`VariantPlan`/`VariantKind` and its nested-union rejection are
|
||||
**not** carried forward — phase 1's refined `CompositePlan::Union`
|
||||
shape supersedes both; see phase 1.)
|
||||
- **Review #004 is the closure target.** H1 → phase 2; packed M1 →
|
||||
phase 2; aligned M1 → phases 4–5; L1 → phase 2 (falls out); L2 →
|
||||
phase 8 (doc rewrite). The review's status flips to "closed" in the
|
||||
phase 8 commit.
|
||||
- **Review #005 is the closure target for the plan-spec issues.** H1
|
||||
→ phase 1 (refined union shape in ADR-011 + plan); H2 → phase 2
|
||||
(`Arc<Value>` on the plan); M1 → phase 1 (nested unions via
|
||||
`CompositePlan` recursion, no behavioral drop); M2 → phase 5 (`Hash`
|
||||
on `Endian`/`VariableEncoding`); M3 → ADR-012 §3 + phase 7
|
||||
(`ValidationPlan` shipped in 0.3.0, deferral reversed); L1/L2/L3 →
|
||||
phase 2 / phase 5 Scope Boundary; N1 → typo; N2 → phase 1
|
||||
`Send + Sync` assertion test; N3 → Semver Contract table row.
|
||||
- **No `unsafe`, no `async`, no new deps, no feature flags** (AGENTS.md
|
||||
§5–§11). The owned-`BastDoc` refactor uses `String`/`Value`, not
|
||||
`unsafe` self-referential tricks. `DefaultHasher` is std. Wasm-clean
|
||||
throughout. `ValidationPlan` follows the same constraints.
|
||||
- **`preserve_order` stays load-bearing** (AGENTS.md §8). The owned-
|
||||
`BastDoc` refactor must not sort schema object keys anywhere; field
|
||||
order in the `Value` still determines byte order in packed mode and
|
||||
iteration order in both modes. `ValidationPlan::compile` inherits
|
||||
this — value-domain checks that depend on field ordering (e.g. union
|
||||
discriminator field lookup) respect `preserve_order`.
|
||||
|
||||
## What this plan is *not*
|
||||
|
||||
- **Not a disk-cache or wire-protocol spec.** The fingerprint contract
|
||||
and method are in scope (phase 6 for `ReadPlan`/`OffsetMap`, phase 7
|
||||
for `ValidationPlan`); downstream uses are the consumers' concern.
|
||||
- **Not cross-version fingerprint stability.** Within-version only
|
||||
(ADR-012). The fingerprint may change across versions if a new
|
||||
`AlkTypeKind` variant is added; consumers cache within a version.
|
||||
- **Not a perf bench.** The bench lives in alktty; this plan re-runs it
|
||||
at phase 2 and phase 8 to confirm the gap closes. The plan itself
|
||||
only asserts correctness/coverage.
|
||||
- **Not an `AlignedPlan`.** Aligned `materialize`'s `BastDoc` structure
|
||||
walk is the permanent 0.3.0 design (phase 5 Scope Boundary). An
|
||||
`AlignedPlan` is out of scope; if a future bench motivates one, it
|
||||
gets its own ADR.
|
||||
@@ -0,0 +1,549 @@
|
||||
---
|
||||
status: complete
|
||||
created: 2026-08-15
|
||||
last_updated: 2026-08-15
|
||||
---
|
||||
|
||||
# BAST Pivot — Implementation Plan
|
||||
|
||||
**Status: complete.** All 10 steps are implemented and pushed to
|
||||
`origin/main` (steps 1–8 in commits `66ab9d7` → `54fd112`; step 9 was
|
||||
a no-op — steps 4–8 converted the tests as they went, leaving only the
|
||||
intentional `from_bast_str` rejection test referencing the
|
||||
`"AlkType:Uint32"` string; step 10 synced the architecture docs and
|
||||
ADRs in this commit). The two new ADRs
|
||||
([ADR-BAST](../architecture/decisions/bast-bast-format.md),
|
||||
[ADR-VAL-SPLIT](../architecture/decisions/val-split-two-validator-model.md))
|
||||
record the decisions; the amended ADRs (001, 002, 003, 004, 009, 010)
|
||||
carry supersession/amendment notes. The research record
|
||||
([`bast-pivot.md`](../research/bast-pivot.md)) is flipped to
|
||||
`implemented`. What follows is the original plan, preserved as the
|
||||
historical execution record.
|
||||
|
||||
---
|
||||
|
||||
This is the execution plan for the BAST pivot: replacing alktype's
|
||||
v0.1.0 `AlkType:*` custom-keyword JSON Schema format with the BAST
|
||||
(Binary Abstract Syntax Tree) format. It is the **entry point** an
|
||||
implementing agent reads first.
|
||||
|
||||
Companion documents:
|
||||
|
||||
- [`docs/architecture/bast-format.md`](../architecture/bast-format.md) —
|
||||
the normative BAST format spec (meta-schema, TypeRef, examples,
|
||||
validation model). Read this for *what* the format is.
|
||||
- [`docs/research/bast-pivot.md`](../research/bast-pivot.md) — the
|
||||
research record: motivation, POC scope and result, decisions
|
||||
D-BAST-001..009, risks. Read this for *why* and *what was proved*.
|
||||
The POC lives on branch `bast-validator-poc` (commit `f371fe4`) as
|
||||
`src/bast_poc.rs` — reference scaffolding, deliberately not merged.
|
||||
|
||||
**Working order:** read this plan top-to-bottom. The Semver Contract
|
||||
section is the scope-creep guardrail — consult it before each step.
|
||||
Each step links to the specific spec section it implements and the
|
||||
relevant D-BAST-* decision anchor. Implement steps in order; each step
|
||||
lists its verification gate.
|
||||
|
||||
## Semver Contract
|
||||
|
||||
The crate is on crates.io at 0.1.0 with zero real consumers, so a
|
||||
breaking bump is free — but the contract is explicit so the
|
||||
implementation doesn't drift. Per AGENTS.md, the 0.1.0 public surface
|
||||
is the items re-exported from `src/lib.rs`. This table is the
|
||||
authoritative scope-creep guardrail for the pivot.
|
||||
|
||||
| Public item (from `lib.rs` re-exports) | Class | Change |
|
||||
|---|---|---|
|
||||
| `AlkTypeKind` (enum + variants + methods) | **Additive** | Unchanged. 19 variants, same methods. New `from_str()`/`to_str()` mapping for lowercase BAST kind strings (`"uint32"` ↔ `AlkTypeKind::Uint32`) — additive methods. |
|
||||
| `Endian`, `VariableEncoding`, `DiscriminatorKind` | **Unchanged** | — |
|
||||
| `AlkTypeEngine::compile` | **Breaking** | Signature: `compile(schema: &mut Value, mode)` → `compile(bast_doc: &Value, root_name: &str, mode)`. Adds required `root_name` param (D-BAST-001); drops `&mut` (BAST needs no in-place `normalize_refs`); input is a BAST document, not a custom-keyword JSON Schema. |
|
||||
| `AlkTypeEngine::validate_json` | **Breaking (behavioral)** | Signature unchanged `(instance: &Value) -> Result<...>`, but the validator it runs is now a standard `jsonschema::Validator` from a consumer-provided JSON Schema, not a custom-keyword validator built from the alktype schema. The *contract* of what schema validates the instance changes. |
|
||||
| `AlkTypeEngine::validate_bytes` | **Unchanged (contract)** | Same signature. Internally the validation step switches from `jsonschema::Validator` to the BAST-native validator. Error type unchanged (D-BAST-009). |
|
||||
| `AlkTypeEngine::is_valid_json` | **Breaking (behavioral)** | Same caveat as `validate_json` — validates against the consumer JSON Schema, not the alktype schema. |
|
||||
| `AlkTypeEngine` accessors (`endian`, `mode`, `offset_map`, `layout_builder`, `sequential_reader`, `read_field`, `write_field`, etc.) | **Unchanged** | Layout-layer accessors are format-agnostic. |
|
||||
| `LayoutMode`, `OffsetMap`, `ByteRange` | **Unchanged** | — |
|
||||
| `LayoutBuilder`, `PackedLayout`, `FieldPosition` | **Unchanged** | — |
|
||||
| `SequentialReader`, `FieldValue` | **Unchanged** | — |
|
||||
| `UnionDispatch` | **Unchanged** | — |
|
||||
| `data_access::*` functions | **Unchanged** | — |
|
||||
| `AlkTypeError` (all 4 variants) | **Unchanged** | D-BAST-009 keeps `Validation(jsonschema::ValidationError<'static>)`. |
|
||||
| `Schema` builder (`struct_`, `object`, `field`, `build`, all setters) | **Breaking (output format)** | Public method signatures unchanged. `build()` output changes from custom-keyword JSON to BAST JSON (for `struct_`) / standard JSON Schema (for `object`). Callers that introspect the built `Value` break; callers that pass it straight to `compile` are source-compatible once `compile` takes BAST. |
|
||||
| `Definitions` builder (`new`, `define`, `define_value`, `build`, `merge_into`) | **Breaking (output format)** | Same as `Schema` — signatures unchanged, `build()`/`merge_into()` output shape changes to BAST `$defs`. |
|
||||
| `Discriminator` builder enum | **Unchanged** | — |
|
||||
| `build_validator` (from `validation`) | **Breaking (signature or removal)** | Currently `build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError>` builds a custom-keyword validator. Under the pivot it either (a) is removed (consumers call `jsonschema` directly for standard JSON Schema) or (b) is repurposed to build a standard `jsonschema::Validator` from a consumer-provided standard JSON Schema (no custom keywords). Decision belongs to step 6. Either way the current signature's contract breaks. |
|
||||
| `get_alktype_kind`, `get_alktype_kind_enum`, `get_alktype_kind_loose`, `get_alktype_kind_loose_enum`, `normalize_refs`, `inline_union_variant_refs`, `resolve_ref`, `resolve_ref_or_inline`, `parse_align`, `parse_discriminator`, `parse_encoding`, `parse_endian`, `parse_max_length` | **Breaking (removal or rework)** | All currently re-exported from `lib.rs`. `normalize_refs` and `inline_union_variant_refs` are removed (BAST needs neither). The `get_alktype_kind*` family is removed (replaced by direct `kind` parsing). The `parse_*` and `resolve_*` functions are reworked to read BAST properties instead of keyword-value objects, or removed if subsumed by the BAST parser. **Open: which of these stay public vs become internal.** Current leaning — drop all from `lib.rs` re-exports (they're engine-internal accessors, not consumer API); the BAST parser exposes a new typed surface instead. Confirmed during step 3. |
|
||||
|
||||
**Net breaking surface:** `compile`, `validate_json`/`is_valid_json`
|
||||
(contract), `Schema::build`/`Definitions::build` (output format),
|
||||
`build_validator` (signature/removal), and the ~13 `schema::*` helper
|
||||
re-exports. **Net additive:** BAST parser, BAST-native validator,
|
||||
`AlkTypeKind::from_str`/`to_str`. **Net unchanged:** the entire layout
|
||||
+ data-access + materialize + tunion layer, `AlkTypeError`, the
|
||||
`Discriminator` builder, `AlkTypeKind` variants.
|
||||
|
||||
### Decisions deferred to their implementation steps
|
||||
|
||||
These are small enough to decide when the step is reached, but are
|
||||
flagged here so they don't become drive-by semver changes:
|
||||
|
||||
1. **`validate_json` JSON Schema source** (step 6): does the consumer
|
||||
pass the JSON Schema to `compile` (engine carries a second
|
||||
validator) or to `validate_json` at call time? The former preserves
|
||||
the current single-call ergonomics; the latter is more flexible. Not
|
||||
semver-relevant either way if `validate_json`'s signature can absorb
|
||||
a new param or stay as-is — needs the call-site analysis.
|
||||
2. **`build_validator` fate** (step 6): removed vs repurposed. If
|
||||
repurposed, its signature stays but its contract (no custom
|
||||
keywords) changes — a behavioral break, not a type break.
|
||||
3. **`schema::*` helper re-exports** (step 3): drop from `lib.rs`
|
||||
(engine-internal) vs keep public for consumers that walk schemas.
|
||||
Leaning: drop — they're accessors for the old format, and the BAST
|
||||
parser exposes a cleaner typed surface. Confirmed during step 3.
|
||||
|
||||
## Steps
|
||||
|
||||
### Step 1 — Add `AlkTypeKind::from_str`/`to_str` for BAST kind strings
|
||||
|
||||
**Goal:** Add the lowercase-string mapping (`"uint32"` ↔
|
||||
`AlkTypeKind::Uint32`) that the BAST parser and validator dispatch on.
|
||||
This is the additive-only, zero-risk foundation — no existing code
|
||||
changes.
|
||||
|
||||
**Spec reference:** [bast-format.md §Primitives](../architecture/bast-format.md#primitives),
|
||||
[D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set).
|
||||
|
||||
**Files:** `src/schema.rs` (the `AlkTypeKind` impl block). No `lib.rs`
|
||||
change needed — the methods are inherent on the already-re-exported
|
||||
enum.
|
||||
|
||||
**Implementation notes:**
|
||||
- `to_str(self) -> &'static str` returns the lowercase BAST string.
|
||||
- `from_str(s: &str) -> Result<AlkTypeKind, AlkTypeError>` returns
|
||||
`AlkTypeError::Schema` for unknown strings. This is a new inherent
|
||||
method, distinct from the existing `FromStr` impl that parses the
|
||||
v0.1.0 `"AlkType:Uint32"` keyword form. Do not remove the existing
|
||||
`FromStr` yet — step 8 removes the v0.1.0 accessors.
|
||||
- Cover all 14 primitive kinds plus `struct`, `union`, `array`,
|
||||
`record`, `enum` (19 total, matching the enum variants). The
|
||||
lowercase strings are in the [primitives table](../architecture/bast-format.md#primitives);
|
||||
composite kinds are `"struct"`, `"union"`, `"array"`, `"record"`,
|
||||
`"enum"`.
|
||||
|
||||
**Verification:** `cargo test --release` (new unit tests for the
|
||||
mapping, both directions; existing tests unaffected). `cargo clippy
|
||||
--all-targets -- -D warnings`.
|
||||
|
||||
---
|
||||
|
||||
### Step 2 — Embed the BAST meta-schema
|
||||
|
||||
**Goal:** Embed the BAST meta-schema as a `serde_json::Value` constant
|
||||
in the crate, available for validating BAST documents at compile time
|
||||
and for publishing at `https://alk.dev/bast/v1/schema`.
|
||||
|
||||
**Spec reference:** [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
|
||||
|
||||
**Files:** New `src/bast_meta.rs` (or a `const` in `src/schema.rs` —
|
||||
match existing module conventions). Re-export the meta-schema `Value`
|
||||
from `lib.rs` if consumers should be able to validate BAST documents
|
||||
themselves (likely yes — additive, not semver-relevant).
|
||||
|
||||
**Implementation notes:**
|
||||
- The meta-schema JSON is in [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
|
||||
Copy it verbatim into a `serde_json::json! {...}` macro invocation or
|
||||
parse it from an embedded string via `serde_json::from_str`.
|
||||
- No feature flags (AGENTS.md §6). The meta-schema is a compile-time
|
||||
constant, no I/O.
|
||||
- WASM-clean: no `include_str!` of an external file is needed if the
|
||||
`json!` macro is used; either way is wasm-safe.
|
||||
|
||||
**Verification:** `cargo test --release`. `cargo build --target
|
||||
wasm32-unknown-unknown --release` (meta-schema is a `Value` constant —
|
||||
wasm-relevant). `cargo clippy --all-targets -- -D warnings`.
|
||||
|
||||
---
|
||||
|
||||
### Step 3 — BAST document parser
|
||||
|
||||
**Goal:** Implement the BAST document parser that the layout engines
|
||||
and materializer use instead of the `get_alktype_kind*` custom-keyword
|
||||
accessors. This is the natural entry point for the pivot — the largest
|
||||
step, and the one the rest of the steps build on.
|
||||
|
||||
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
|
||||
(the whole document — the parser implements the format spec).
|
||||
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection),
|
||||
[D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement),
|
||||
[D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions).
|
||||
|
||||
**Files:** New `src/bast.rs` (the parser). The existing `src/schema.rs`
|
||||
stays for now — steps 4–8 migrate callers off it. Update `src/lib.rs`
|
||||
to add `pub mod bast;` and re-export the parser's public surface.
|
||||
|
||||
**Implementation notes:**
|
||||
- The parser reads `kind`/`fields`/annotation properties from BAST
|
||||
nodes. It produces a typed surface (a small `BastNode` enum or
|
||||
equivalent) that the layout engines, materializer, and validator can
|
||||
walk without re-parsing the raw JSON at every node. The POC parsed
|
||||
lazily from raw JSON in both passes to keep the model honest; a typed
|
||||
tree is a straightforward follow-on optimization (POC observation 5).
|
||||
Either is acceptable for the production version; the typed tree is
|
||||
recommended since three consumers (layout, materialize, validate)
|
||||
walk the same tree.
|
||||
- `$ref` resolution: `#/$defs/<name>` only — a single hash lookup. No
|
||||
`normalize_refs` (BAST refs are always full JSON Pointers), no
|
||||
`inline_union_variant_refs` (union variant refs resolved lazily by
|
||||
the validator and materializer). See [bast-format.md §TypeRef](../architecture/bast-format.md#typeref).
|
||||
- Untrusted input: every path that walks a BAST document must return
|
||||
`Err(AlkTypeError::Schema)` on a malformed document, never `panic!`/
|
||||
`unreachable!` (AGENTS.md §3). The POC's
|
||||
`malformed_document_produces_schema_error_not_panic` test is the
|
||||
template.
|
||||
- **Decide deferred decision #3 here:** drop the `schema::*` helper
|
||||
re-exports from `lib.rs`, or keep them public. Leaning: drop. The
|
||||
BAST parser exposes a cleaner typed surface; the v0.1.0 accessors
|
||||
are engine-internal and not consumer API.
|
||||
|
||||
**Verification:** `cargo test --release` (port the POC's parser tests
|
||||
— the malformed-document test, the type-ref resolution tests).
|
||||
`cargo clippy --all-targets -- -D warnings`. The layout engines don't
|
||||
use the parser yet (step 4 wires it in), so the existing suite still
|
||||
passes on the old path.
|
||||
|
||||
---
|
||||
|
||||
### Step 4 — Wire `compile()` to accept a BAST document + root name
|
||||
|
||||
**Goal:** Change `AlkTypeEngine::compile` to the new signature and
|
||||
have it use the BAST parser instead of the custom-keyword accessors.
|
||||
The layout engines (`offset_map`, `layout_builder`,
|
||||
`sequential_reader`) consume the BAST parser's typed output instead of
|
||||
walking raw JSON with `get_alktype_kind*`.
|
||||
|
||||
**Spec reference:** [bast-format.md §Document Shape](../architecture/bast-format.md#document-shape),
|
||||
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection).
|
||||
Semver contract: `compile` is **Breaking**.
|
||||
|
||||
**Files:** `src/engine.rs` (the `compile` signature and body). The
|
||||
layout modules (`src/offset_map.rs`, `src/layout_builder.rs`,
|
||||
`src/sequential_reader.rs`) — their schema-walking code changes from
|
||||
`get_alktype_kind*` calls to BAST parser calls. `src/lib.rs` if the
|
||||
parser's public surface needs re-exporting (step 3 may have done this).
|
||||
|
||||
**Implementation notes:**
|
||||
- New signature: `pub fn compile(bast_doc: &Value, root_name: &str,
|
||||
mode: LayoutMode) -> Result<Self, AlkTypeError>`. Note `&Value` (not
|
||||
`&mut Value`) — BAST needs no in-place `normalize_refs`.
|
||||
- The engine stores the BAST document (or the parsed typed tree) for
|
||||
`sequential_reader()`'s factory construction and `read_field`'s kind
|
||||
lookup. The `Layout` enum and mode dispatch are unchanged.
|
||||
- `parse_endian`, `parse_align`, `parse_encoding`, `parse_discriminator`
|
||||
are reworked to read BAST properties (struct/field-level) instead of
|
||||
keyword-value objects. Their *semantics* are unchanged (ADR-003);
|
||||
only their *input location* moves. Whether they stay as free
|
||||
functions or become methods on the typed `BastNode` is an
|
||||
implementation choice — the POC read properties inline.
|
||||
- The layout engines are format-agnostic beneath the accessors
|
||||
(checked offset arithmetic, the two modes, union dispatch). This
|
||||
step is an accessor swap, not a layout-engine rewrite.
|
||||
|
||||
**Verification:** `cargo test --release` (test inputs must be converted
|
||||
to BAST format — see step 9 for the full test conversion; this step
|
||||
converts the layout tests as a sanity check). `cargo clippy
|
||||
--all-targets -- -D warnings`. `cargo build --target
|
||||
wasm32-unknown-unknown --release` (layout/wasm-relevant).
|
||||
|
||||
---
|
||||
|
||||
### Step 5 — BAST-native validator (production version)
|
||||
|
||||
**Goal:** Port the POC's BAST-native validator into a production module
|
||||
and wire it into `validate_bytes` as the validation step, replacing the
|
||||
`jsonschema::Validator` call on the bytes path.
|
||||
|
||||
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
|
||||
[D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model),
|
||||
[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape).
|
||||
POC reference: `src/bast_poc.rs` on branch `bast-validator-poc`.
|
||||
|
||||
**Files:** New `src/bast_validation.rs`. `src/engine.rs`
|
||||
(`validate_bytes` body — swap the `self.validator.validate(&value)` call
|
||||
for the BAST-native validator). `src/lib.rs` — add `pub mod
|
||||
bast_validation;` (the validator is engine-internal; whether it's
|
||||
re-exported is an implementation choice, leaning no).
|
||||
|
||||
**Implementation notes:**
|
||||
- The POC is the reference. The validator is a single recursive
|
||||
function (`validate_typeref`) that dispatches on the BAST `kind`. The
|
||||
constraint table is in [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model).
|
||||
- Construct `AlkTypeError::Validation` via
|
||||
`jsonschema::ValidationError::custom` — the variant's payload type is
|
||||
unchanged (D-BAST-009). The bytes path no longer touches `jsonschema`
|
||||
for validation, but the error type retains the `jsonschema` type for
|
||||
uniformity with the `validate_json` path.
|
||||
- The validator and materializer share the BAST-walking code structure.
|
||||
If step 3 produced a typed `BastNode` tree, both consume it. If step
|
||||
3 parses lazily, the validator parses lazily too (POC approach).
|
||||
- Enum index bounds: check the materialized index against
|
||||
`values.len()` — this **fixes the v0.1.0 dead constraint** (the
|
||||
built-in `enum` keyword checked string membership, but the
|
||||
materializer emits `Value::Number(index)`, which never matched). Net
|
||||
improvement.
|
||||
- Union variant dispatch: read `__discriminator`, look up the variant's
|
||||
BAST definition, recurse. Recovers OQ-008 per-variant constraint
|
||||
enforcement without custom keywords.
|
||||
|
||||
**Verification:** `cargo test --release` — the existing `validate_bytes`
|
||||
tests are the regression target (test *inputs* change to BAST format
|
||||
in step 9; expected validation outcomes must be identical). The POC's
|
||||
20 tests are the reference. `cargo clippy --all-targets -- -D warnings`.
|
||||
`cargo build --target wasm32-unknown-unknown --release`.
|
||||
|
||||
---
|
||||
|
||||
### Step 6 — `validate_json` against a consumer-provided JSON Schema
|
||||
|
||||
**Goal:** Update `validate_json`/`is_valid_json` to validate against a
|
||||
standard `jsonschema::Validator` compiled from a consumer-provided JSON
|
||||
Schema, not a custom-keyword validator built from the alktype schema.
|
||||
|
||||
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
|
||||
[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model).
|
||||
Semver contract: `validate_json`/`is_valid_json` are **Breaking
|
||||
(behavioral)**; `build_validator` is **Breaking (signature or
|
||||
removal)**.
|
||||
|
||||
**Files:** `src/engine.rs` (`validate_json`/`is_valid_json` bodies, and
|
||||
the engine's stored validator field if the JSON Schema is supplied at
|
||||
compile time). `src/validation.rs` (`build_validator` — repurposed or
|
||||
removed). `src/lib.rs` (the `build_validator` re-export if removed).
|
||||
|
||||
**Implementation notes:**
|
||||
- **Decide deferred decision #1 here:** does the consumer pass the JSON
|
||||
Schema to `compile` (engine carries a second validator) or to
|
||||
`validate_json` at call time? The former preserves single-call
|
||||
ergonomics; the latter is more flexible. Needs the alkcall call-site
|
||||
analysis. Not semver-relevant either way if the signature can absorb
|
||||
the change.
|
||||
- **Decide deferred decision #2 here:** `build_validator` removed vs
|
||||
repurposed. If repurposed, its signature stays but its contract
|
||||
changes (no custom keywords) — a behavioral break. If removed, drop
|
||||
the `lib.rs` re-export.
|
||||
- The `jsonschema` crate remains a direct dependency (for `validate_json`
|
||||
and for validating BAST documents against the meta-schema). Only the
|
||||
custom keyword integration is removed.
|
||||
- The engine may carry two validators: the BAST-native validator (for
|
||||
`validate_bytes`, from step 5) and the standard `jsonschema::Validator`
|
||||
(for `validate_json`, from this step). Or `validate_json` takes the
|
||||
JSON Schema at call time and builds a transient validator. The
|
||||
decision shapes the engine struct's fields.
|
||||
|
||||
**Verification:** `cargo test --release` (new tests for the
|
||||
consumer-provided JSON Schema path; existing `validate_json` tests
|
||||
converted — their schemas were custom-keyword, now standard). `cargo
|
||||
clippy --all-targets -- -D warnings`.
|
||||
|
||||
---
|
||||
|
||||
### Step 7 — Builder API produces BAST JSON
|
||||
|
||||
**Goal:** Update the builder's `build()` methods to produce BAST JSON
|
||||
(for `struct_()`) and standard JSON Schema (for `object()`). Public
|
||||
method signatures are unchanged; only the output `Value` shape changes.
|
||||
|
||||
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
|
||||
(the output format), [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats).
|
||||
Semver contract: `Schema::build`/`Definitions::build` are **Breaking
|
||||
(output format)**.
|
||||
|
||||
**Files:** `src/builder.rs`. `src/lib.rs` if the builder's public
|
||||
surface changes (it shouldn't — method signatures are unchanged).
|
||||
|
||||
**Implementation notes:**
|
||||
- `Schema::struct_().field(...).build()` → BAST JSON (a `$defs` entry
|
||||
with `kind: "struct"`, ordered `fields` array, type-level
|
||||
annotations).
|
||||
- `Schema::object().field(...).build()` → standard JSON Schema (no
|
||||
`AlkType:*` keywords, no BAST `kind` — just `type`/`properties`/
|
||||
`required`).
|
||||
- `Definitions::build()`/`merge_into()` → a BAST `$defs` block.
|
||||
- The builder already distinguishes AlkType kinds from JSON Schema types
|
||||
via naming conventions (`string()` vs `string_()`). The construction
|
||||
API is the same; only the serialization differs.
|
||||
- The `Discriminator` builder is unchanged (semver contract:
|
||||
**Unchanged**).
|
||||
|
||||
**Verification:** `cargo test --release` (builder tests assert on the
|
||||
output `Value` — update the expected shapes). `cargo clippy
|
||||
--all-targets -- -D warnings`.
|
||||
|
||||
---
|
||||
|
||||
### Step 8 — Remove v0.1.0 custom-keyword machinery
|
||||
|
||||
**Goal:** Remove the dead code now that all callers use the BAST parser
|
||||
and BAST-native validator.
|
||||
|
||||
**Spec reference:** [bast-format.md §What is removed](../architecture/bast-format.md#what-is-removed).
|
||||
Semver contract: the ~13 `schema::*` helper re-exports are **Breaking
|
||||
(removal or rework)** (decision #3, confirmed in step 3).
|
||||
|
||||
**Files:** `src/schema.rs` (remove `get_alktype_kind*`,
|
||||
`normalize_refs`, `inline_union_variant_refs`; rework or remove
|
||||
`parse_*`/`resolve_*`). `src/validation.rs` (remove the 19
|
||||
`jsonschema::Keyword` implementations if not already removed in step 5/6).
|
||||
`src/lib.rs` (drop the removed items from the `pub use` block).
|
||||
|
||||
**Implementation notes:**
|
||||
- Remove: all 19 `jsonschema::Keyword` implementations (~200 lines),
|
||||
`normalize_refs()`, `inline_union_variant_refs()`, the
|
||||
`get_alktype_kind*` family.
|
||||
- Rework or remove: `parse_align`, `parse_discriminator`,
|
||||
`parse_encoding`, `parse_endian`, `parse_max_length`, `resolve_ref`,
|
||||
`resolve_ref_or_inline`. If the BAST parser subsumes them (likely),
|
||||
remove them. If any remain useful as free functions over the typed
|
||||
`BastNode`, keep them internal (not re-exported from `lib.rs`).
|
||||
- The `jsonschema` crate's `with_keyword(...)` registration calls are
|
||||
removed from `compile`/`build_validator`. The crate itself stays.
|
||||
- Drop the removed items from `lib.rs`'s `pub use schema::{ ... }`
|
||||
block. The BAST parser's public surface replaces them.
|
||||
|
||||
**Verification:** `cargo test --release`. `cargo clippy --all-targets
|
||||
-- -D warnings`. `cargo doc --no-deps` (the public API surface
|
||||
changed — doc comments must build). `cargo build --target
|
||||
wasm32-unknown-unknown --release` (removing code shouldn't add
|
||||
platform deps).
|
||||
|
||||
---
|
||||
|
||||
### Step 9 — Convert all tests to BAST format
|
||||
|
||||
**Goal:** Update the full test suite to use BAST format for inputs.
|
||||
Test assertions (expected validation outcomes, expected offsets,
|
||||
expected materialized values) must be identical — only the input
|
||||
schema shape changes.
|
||||
|
||||
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
|
||||
(input format).
|
||||
|
||||
**Files:** `tests/*.rs` (integration tests), `src/*.rs` inline `#[cfg(test)]`
|
||||
modules (unit tests).
|
||||
|
||||
**Implementation notes:**
|
||||
- This may be partially done by steps 4–8 (each step converts the tests
|
||||
it touches as a sanity check). This step is the sweep: every test
|
||||
using `AlkType:*` keywords converts to BAST `kind`/`fields`.
|
||||
- The POC's 20 tests are the reference for BAST-shaped test inputs.
|
||||
- Expected validation outcomes are the regression target. The
|
||||
enum-index-bounds test is new behavior (the v0.1.0 dead constraint
|
||||
is now enforced) — that test's expectation *changes* (was: silently
|
||||
passed; now: `AlkTypeError::Validation`). This is the intended fix,
|
||||
not a regression.
|
||||
- Coverage: 310 crate + 86 integration tests (~396 total). All must
|
||||
pass.
|
||||
|
||||
**Verification:** `cargo test --release` (the full suite — this is the
|
||||
gate). `cargo clippy --all-targets -- -D warnings`.
|
||||
|
||||
---
|
||||
|
||||
### Step 10 — Sync architecture docs and ADRs
|
||||
|
||||
**Goal:** Sync the descriptive docs and ADRs to the shipped code. This
|
||||
is the final step — per AGENTS.md, ADRs are written post-implementation,
|
||||
grounded in shipped code.
|
||||
|
||||
**Spec reference:** [Semver Contract §ADR impact](#adr-impact-checklist)
|
||||
below.
|
||||
|
||||
**Files:** `docs/architecture/README.md`, `docs/architecture/overview.md`,
|
||||
`docs/architecture/schema-layer.md` (rewrite for the BAST parser),
|
||||
`docs/architecture/validation.md` (rewrite for the validator split),
|
||||
`docs/architecture/builder.md` (update `build()` output examples),
|
||||
`src/lib.rs` (module doc comment). New ADRs: ADR-BAST, ADR-VAL-SPLIT.
|
||||
Amended ADRs: 001 (superseded), 003, 004, 009, 010.
|
||||
|
||||
**Implementation notes:**
|
||||
- Rewrite `schema-layer.md` to describe the BAST parser (replaces the
|
||||
custom-keyword accessor walk-through). The current `schema-layer.md`
|
||||
content is the v0.1.0 reference; `bast-format.md` already contains
|
||||
the target spec. Either fold `bast-format.md` into `schema-layer.md`
|
||||
or keep both with `schema-layer.md` pointing at `bast-format.md` for
|
||||
the format and describing the parser module.
|
||||
- Rewrite `validation.md` for the validator split (the [bast-format.md
|
||||
§Validation Model](../architecture/bast-format.md#validation-model)
|
||||
content moves here, expanded with the production validator's
|
||||
details).
|
||||
- Update `builder.md` output examples to BAST JSON.
|
||||
- Update `src/lib.rs` module doc comment: "Takes a JSON Schema with
|
||||
`AlkType:*` custom keywords" → "Takes a BAST document".
|
||||
- Update `docs/architecture/README.md` index — the document table, the
|
||||
ADR table (new ADRs, superseded ADR-001), the key design principles
|
||||
(#1, #2, #7, #10 change wording).
|
||||
- Remove stale TODOs referencing custom-keyword normalization,
|
||||
`inline_union_variant_refs`, or the rejected bare-name-ref design
|
||||
(AGENTS.md §"Architecture Context").
|
||||
- `docs/research/bast-pivot.md` is the research record — its status
|
||||
flips from `draft` to `accepted`/`implemented` and it gains a pointer
|
||||
to the ADRs that superseded its decisions.
|
||||
|
||||
**Verification:** `cargo doc --no-deps` (doc comments build).
|
||||
Cross-reference check: every link in this plan, `bast-format.md`, and
|
||||
the new/updated ADRs resolves. `cargo test --release` (no code change,
|
||||
but the doc sweep shouldn't break anything).
|
||||
|
||||
## ADR Impact Checklist
|
||||
|
||||
Sync these ADRs when step 10 lands. Per AGENTS.md, ADRs are written
|
||||
post-implementation, grounded in shipped code.
|
||||
|
||||
| ADR | Action | Reason |
|
||||
|---|---|---|
|
||||
| [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) (purpose, scope, "schema is the format") | **Supersede** | The "schema is the format" principle is retained and strengthened (BAST *is* the format), but the concrete format changes from custom-keyword JSON Schema to BAST. A new ADR (ADR-BAST) records the BAST format as the realization of the principle. ADR-001 Status → Superseded by ADR-BAST. |
|
||||
| [ADR-002](../architecture/decisions/002-two-layout-modes-packed-vs-aligned.md) (two layout modes) | **Unchanged** | Layout modes are format-agnostic. One-line note that the input format changed but the modes didn't. |
|
||||
| [ADR-003](../architecture/decisions/003-schema-annotations.md) (annotations) | **Amend** | Annotation *semantics* carry forward unchanged; annotation *location* moves from custom-keyword objects to BAST type-level properties. Amend the "where annotations live" sections, keep the semantics. |
|
||||
| [ADR-004](../architecture/decisions/004-error-handling-validation-strategy.md) (error handling, validation strategy) | **Amend** | Error enum shape unchanged (D-BAST-009). The "validation strategy" section updates: bytes path uses BAST-native validator, JSON path uses standard `jsonschema`. The load-time/access-time split is retained. |
|
||||
| [ADR-005](../architecture/decisions/005-int64-uint64-first-class-kinds.md) (Int64/Uint64) | **Unchanged** | Kinds carry forward; JSON precision caveat unchanged. |
|
||||
| [ADR-006](../architecture/decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) (reject non-final inline in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
|
||||
| [ADR-007](../architecture/decisions/007-packed-mode-read-factory.md) (packed-mode read factory) | **Unchanged** | Reader factory semantics are format-agnostic. |
|
||||
| [ADR-008](../architecture/decisions/008-reject-tunion-in-aligned-mode.md) (reject TUnion in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
|
||||
| [ADR-009](../architecture/decisions/009-builder-api.md) (builder API) | **Amend** | Public method surface unchanged; `build()` output format changes (BAST for `struct_`, standard JSON Schema for `object`). Amend the "output format" section; keep the method catalog. |
|
||||
| [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md) (`validate_bytes`) | **Amend** | The two-step concept (materialize → validate) is retained. The validation step's *implementation* changes from `jsonschema` custom keywords to the BAST-native validator. Amend the "validation step" section; add a pointer to D-BAST-006/D-BAST-009 and ADR-VAL-SPLIT. |
|
||||
|
||||
**New ADRs to write (post-implementation, grounded in shipped code):**
|
||||
- **ADR-BAST** — the BAST format, meta-schema, and `$defs`/`$ref`/
|
||||
`kind` vocabulary. Supersedes ADR-001's format-specific content.
|
||||
- **ADR-VAL-SPLIT** (or fold into ADR-004's amend) — the two-validator
|
||||
model: BAST-native for `validate_bytes`, standard `jsonschema` for
|
||||
`validate_json`. Records D-BAST-006, D-BAST-007, D-BAST-009.
|
||||
|
||||
**Descriptive docs to sync (post-implementation):**
|
||||
- `docs/architecture/schema-layer.md` — rewrite for the BAST parser
|
||||
(replaces the custom-keyword accessor walk-through).
|
||||
- `docs/architecture/validation.md` — rewrite for the validator split.
|
||||
- `docs/architecture/builder.md` — update the `build()` output examples
|
||||
to BAST JSON.
|
||||
- `src/lib.rs` module doc comment — update the "Takes a JSON Schema
|
||||
with `AlkType:*` custom keywords" preamble to BAST.
|
||||
- `docs/architecture/README.md` — update the document table, ADR table,
|
||||
and key design principles for the pivot.
|
||||
- `docs/architecture/overview.md` — update the "what" and "why" for
|
||||
BAST (the crate now takes a BAST document, not a custom-keyword JSON
|
||||
Schema).
|
||||
|
||||
**Stale TODOs to remove:** any TODO referencing custom-keyword
|
||||
normalization, `inline_union_variant_refs`, or the rejected
|
||||
bare-name-ref design — align with the ADRs as AGENTS.md §"Architecture
|
||||
Context" requires.
|
||||
|
||||
## Verification Commands
|
||||
|
||||
Run these before committing each step. All must pass. Per AGENTS.md:
|
||||
|
||||
```bash
|
||||
cargo test --release # full suite (~396 tests: 310 crate + 86 integration)
|
||||
cargo clippy --all-targets -- -D warnings
|
||||
cargo doc --no-deps # if docs changed (step 8, step 10)
|
||||
cargo build --target wasm32-unknown-unknown --release # if layout/wasm-relevant code changed (step 2, 4, 5, 8)
|
||||
cargo publish --dry-run --allow-dirty # before a release (post-step 10)
|
||||
```
|
||||
+279
-895
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,419 @@
|
||||
---
|
||||
status: open
|
||||
last_updated: 2026-08-15
|
||||
reviewed_artifacts:
|
||||
- src/lib.rs
|
||||
- src/bast.rs
|
||||
- src/bast_meta.rs
|
||||
- src/bast_validation.rs
|
||||
- src/builder.rs
|
||||
- src/data_access.rs
|
||||
- src/engine.rs
|
||||
- src/error.rs
|
||||
- src/layout_builder.rs
|
||||
- src/materialize.rs
|
||||
- src/offset_map.rs
|
||||
- src/schema.rs
|
||||
- src/sequential_reader.rs
|
||||
- src/tunion.rs
|
||||
- src/validation.rs
|
||||
- src/macros.rs
|
||||
- tests/{engine_integration,error_paths,poc_roundtrip,tunion_dispatch}.rs
|
||||
- Cargo.toml
|
||||
tool: manual source read + cargo test/clippy + cargo-llvm-cov
|
||||
reviewer: post-BAST-pivot code review
|
||||
---
|
||||
|
||||
# Code Review #003 — Post-BAST-Pivot Review
|
||||
|
||||
## Purpose
|
||||
|
||||
First logic/correctness review after the BAST pivot (the v0.1.0
|
||||
`AlkType:*` custom-keyword JSON Schema format was replaced with the BAST
|
||||
format; see `docs/plans/bast-implementation.md`). The pivot touched
|
||||
every schema-walking path, so this pass re-reads the whole crate for
|
||||
correctness, code smell, panic safety, and coverage — the same scope as
|
||||
review #002, but against the new BAST surface.
|
||||
|
||||
Two things motivated this review beyond the routine sweep:
|
||||
|
||||
1. The pivot was a large, multi-step change (10 steps, 8 commits). A
|
||||
couple of pre-existing bugs were fixed *during* the pivot (the enum
|
||||
index-bounds dead constraint, the `write_bytes` u32 truncation), so
|
||||
the same class of bug could be lurking in the newly-rewritten paths.
|
||||
2. The publisher asked specifically for a coverage pass
|
||||
(`cargo-llvm-cov`) with an eye toward *important* things being
|
||||
covered rather than raw numbers.
|
||||
|
||||
## Methodology
|
||||
|
||||
- Full read of all 16 `src/*.rs` files (production + test modules) and
|
||||
all 4 integration test files.
|
||||
- `cargo test --release`, `cargo clippy --all-targets -- -D warnings`.
|
||||
- `cargo llvm-cov --release` (summary + per-file + uncovered-lines) to
|
||||
attribute coverage gaps to specific code paths.
|
||||
- Targeted reproduction of the suspicious paths (field-level endian
|
||||
override, aligned-mode variable-length/array materialization) via
|
||||
throwaway integration tests.
|
||||
- Cross-reference every error path against its caller to confirm errors
|
||||
propagate (not swallowed) and carry useful attribution.
|
||||
- Read `docs/reviews/002-code-review.md` for prior context and
|
||||
resolved/unresolved items.
|
||||
|
||||
## Verification Baseline
|
||||
|
||||
All verification run on the reviewed tree (commit `562284f`):
|
||||
|
||||
- `cargo test --release`: **389 tests pass** (312 crate unit tests +
|
||||
77 integration tests across 4 files). Zero failures.
|
||||
- `cargo clippy --all-targets -- -D warnings`: **clean**.
|
||||
- `cargo llvm-cov --release`: **90.14% line coverage** (7903/8682),
|
||||
**86.68% function coverage** (743/842). Per-file breakdown below.
|
||||
- No `unsafe` anywhere in the crate.
|
||||
- No `TODO`/`FIXME`/`HACK`/`XXX` markers in source.
|
||||
- All `unwrap`/`expect`/`panic!`/`unreachable!` are confined to
|
||||
`#[cfg(test)]` modules, verified by line-context cross-reference.
|
||||
|
||||
### Coverage breakdown
|
||||
|
||||
| Module | Lines | Functions |
|
||||
|---|---:|---:|
|
||||
| bast.rs | 88.3% | 90.3% |
|
||||
| bast_validation.rs | 93.0% | 90.5% |
|
||||
| builder.rs | 91.3% | 87.6% |
|
||||
| data_access.rs | 83.5% | 76.2% |
|
||||
| engine.rs | 96.9% | 98.3% |
|
||||
| layout_builder.rs | 91.6% | 81.8% |
|
||||
| materialize.rs | **81.8%** | **76.1%** |
|
||||
| offset_map.rs | 89.5% | 79.2% |
|
||||
| sequential_reader.rs | **84.6%** | **75.0%** |
|
||||
| tunion.rs | 92.7% | 91.7% |
|
||||
| **TOTAL** | **90.1%** | **86.7%** |
|
||||
|
||||
The low-function-count modules are not test-helper noise — they are
|
||||
exactly where the correctness bugs below live. The uncovered lines in
|
||||
`materialize.rs` and `sequential_reader.rs` are the aligned-mode
|
||||
variable-length/array paths and the field-level-endian paths, which are
|
||||
**untested and broken** (see M1, M2). The `data_access.rs` 76% function
|
||||
coverage is mostly the `read_*_indirect` family, which has no production
|
||||
caller (see L1).
|
||||
|
||||
## Summary Statistics
|
||||
|
||||
| Severity | Count |
|
||||
|----------|------:|
|
||||
| Critical | 0 |
|
||||
| Medium | 3 (M1, M2, M3) |
|
||||
| Low | 2 (L1, L2) |
|
||||
| Nit | 4 (N1, N2, N3, N4) |
|
||||
|
||||
No critical findings. The crate is in good shape, but the three Medium
|
||||
findings are **silent data-corruption / silent-misinterpretation bugs**
|
||||
in the newly-rewritten paths — they must be fixed before the next
|
||||
release. The Low findings are dead code and a robustness gap; the Nits
|
||||
are hygiene.
|
||||
|
||||
---
|
||||
|
||||
## Findings
|
||||
|
||||
### M1. Field-level `endian` override is ignored by the reader and aligned materializer
|
||||
|
||||
**Files**: `src/sequential_reader.rs:290`, `src/engine.rs:338,455`,
|
||||
`src/materialize.rs:461`
|
||||
|
||||
**Problem**: The spec documents per-field endian override
|
||||
(`docs/architecture/bast-format.md` §Endianness — "Field-level `endian`
|
||||
overrides the struct/union default"), and the packed materializer honors
|
||||
it (`materialize.rs:120` uses `field.effective_endian(endian)`). But
|
||||
three paths use only the struct-level endian:
|
||||
|
||||
- `sequential_reader.rs:290` `read_field_value` — uses `self.endian`,
|
||||
never `field.effective_endian`.
|
||||
- `engine.rs:338` `read_field` and `engine.rs:455` `write_field` —
|
||||
`let endian = self.endian;`.
|
||||
- `materialize.rs:461` `materialize_struct_aligned` — passes the struct
|
||||
`endian` to `materialize_leaf_at`, never `field.effective_endian`.
|
||||
|
||||
**Reproduction** (throwaway integration test, confirmed): a big-endian
|
||||
struct with a `"crc": { "kind": "uint32", "endian": "little" }` field
|
||||
reads `0x01020304` as `67305985` (big-endian interpretation) in both
|
||||
`sequential_reader` and `read_field`. The bytes are correct; the
|
||||
interpretation is wrong — silent data corruption.
|
||||
|
||||
**Fix**: thread `field.effective_endian(endian)` through all three paths.
|
||||
`read_field_value` already receives the `BastField`; `engine::read_field`
|
||||
/ `write_field` need to look up the field's effective endian (they
|
||||
already walk the BAST tree via `lookup_field_kind`); `materialize_struct_aligned`
|
||||
needs to pass `field.effective_endian(endian)` to `materialize_leaf_at`
|
||||
instead of the struct default.
|
||||
|
||||
**Lift**: closes a silent-corruption path on a documented feature. Small
|
||||
effort (~10 lines + regression tests).
|
||||
|
||||
---
|
||||
|
||||
### M2. Aligned-mode `validate_bytes` is broken for arrays, `maxLength` fields, and `offset-indirect` fields
|
||||
|
||||
**File**: `src/materialize.rs:461-512` (`materialize_struct_aligned`)
|
||||
|
||||
**Problem**: `materialize_struct_aligned` routes fixed-size and
|
||||
variable-length leaves through `materialize_leaf_at`, which calls
|
||||
`materialize_typeref_packed` — i.e. it always reads a **length-prefixed**
|
||||
value. But the aligned `OffsetMap` stores three different shapes:
|
||||
|
||||
- `maxLength` fields are a raw reservation (no length prefix) — the
|
||||
materializer reads the first 4 bytes of the *data* as a length prefix.
|
||||
Reproduced: `"hello"` in an 8-byte reservation read a length of
|
||||
`1819043180` and failed with a bounds error.
|
||||
- `offset-indirect` fields are an 8-byte `{offset, length}` pair — read
|
||||
as a length prefix, garbage.
|
||||
- arrays are recorded as `vals[0]`/`vals[1]` entries only, so
|
||||
`offset_map.get("vals")` returns `None` → `Offset` error. Reproduced.
|
||||
|
||||
Only the default inline length-prefixed variable field (and only as the
|
||||
final field, per ADR-006) works in aligned mode. There are **no tests**
|
||||
covering aligned `validate_bytes` with arrays or non-default variable
|
||||
encodings — that is why this slipped through the pivot.
|
||||
|
||||
**Fix**: two options, decide with the publisher:
|
||||
|
||||
1. **Implement** aligned materialization for the three shapes: read
|
||||
`maxLength` fields as a fixed-size slice, `offset-indirect` fields via
|
||||
`read_*_indirect` (with a data region), and arrays by iterating the
|
||||
`vals[i]` offset-map entries.
|
||||
2. **Reject** these combinations at compile time (return
|
||||
`AlkTypeError::Schema`/`Offset` from `OffsetMap::compute` or
|
||||
`compile`) if they are out of scope for v1, so the failure is loud
|
||||
and at load time rather than a silent misread at access time.
|
||||
|
||||
Option 2 is the smaller, safer fix and matches the existing ADR-006/
|
||||
ADR-008 pattern of rejecting unsupported aligned-mode combinations. The
|
||||
`offset-indirect` encoding is already dead code on the read path (see
|
||||
L1), which argues for rejecting it in aligned mode until it is actually
|
||||
implemented.
|
||||
|
||||
**Lift**: closes a silent-misread path. Medium effort either way.
|
||||
|
||||
---
|
||||
|
||||
### M3. The BAST meta-schema is never applied at compile time; annotation parsers silently tolerate malformed values
|
||||
|
||||
**Files**: `src/engine.rs:123` (`compile`), `src/bast.rs:899-926`
|
||||
(`parse_endian_opt`, `parse_align`, `parse_encoding`, `parse_max_length`)
|
||||
|
||||
**Problem**: `BAST_META_SCHEMA` is exported and self-tested, but
|
||||
`AlkTypeEngine::compile` never validates the document against it. The
|
||||
parser (`bast.rs`) is the only gate, and it silently tolerates malformed
|
||||
annotations:
|
||||
|
||||
- `parse_endian_opt` (`bast.rs:899`) — `"endian": "middle"` → `None` →
|
||||
silently defaults to little.
|
||||
- `parse_encoding` (`bast.rs:921`) — unknown encoding → silently
|
||||
`LengthPrefixed`.
|
||||
- `parse_align` / `parse_max_length` — non-integer / negative → silently
|
||||
`None`.
|
||||
|
||||
These are exactly the cases the meta-schema's `enum` / `minimum`
|
||||
constraints exist to reject. Per AGENTS.md §3, schemas are untrusted
|
||||
input (the `alkcall` consumer accepts them from arbitrary internet
|
||||
peers). A malicious peer can send `"endian": "bogus"` and get a
|
||||
silently-misinterpreted layout instead of a `Schema` error.
|
||||
|
||||
**Fix**: validate the document against `BAST_META_SCHEMA` in `compile`
|
||||
(one-time, cheap — the meta-schema is a `LazyLock<Value>`), *or* make
|
||||
the annotation parsers return `Err(AlkTypeError::Schema)` on unknown
|
||||
values. The meta-schema route is preferred: it is the single source of
|
||||
truth and catches the whole class of malformed-annotation bugs at once.
|
||||
|
||||
**Lift**: closes a silent-misinterpretation path on untrusted input.
|
||||
Small effort (~5 lines + tests).
|
||||
|
||||
---
|
||||
|
||||
### L1. `offset-indirect` is dead code on the read path
|
||||
|
||||
**Files**: `src/data_access.rs:307-355`, `src/engine.rs:388-399`
|
||||
|
||||
**Problem**: `data_access::read_string_indirect` / `read_bytes_indirect`
|
||||
have no production caller (only their own unit tests). `engine.read_field`
|
||||
always calls `read_string` / `read_bytes` (length-prefixed) regardless of
|
||||
the field's `encoding`. So a schema declaring
|
||||
`"encoding": "offset-indirect"` compiles and lays out correctly in the
|
||||
offset map, but can never be read back.
|
||||
|
||||
This is the same root cause as M2's `offset-indirect` arm. Decide
|
||||
together with M2: either wire `read_*_indirect` into the read path (and
|
||||
the aligned materializer), or drop the `offset-indirect` encoding
|
||||
entirely until a consumer needs it. Leaving it half-wired is the worst
|
||||
state — it looks supported but silently misreads.
|
||||
|
||||
**Lift**: removes dead code or completes a feature. Small effort.
|
||||
|
||||
---
|
||||
|
||||
### L2. `materialize_packed` / `materialize_aligned` take a dead `endian` parameter
|
||||
|
||||
**File**: `src/materialize.rs:44-88`
|
||||
|
||||
**Problem**: both functions take `endian: Endian` and immediately
|
||||
`let _ = endian;`, using `struct_node.endian()` instead. The caller's
|
||||
`self.endian` (from `engine.rs:285,287`) is ignored. The signature is
|
||||
misleading — a reader assumes the passed endian is honored.
|
||||
|
||||
**Fix**: drop the parameter and read the endian from the root struct
|
||||
inside the function (it already does). ~4 lines. Purely a clarity fix;
|
||||
no behavior change.
|
||||
|
||||
---
|
||||
|
||||
### N1. `number_from_f64` maps NaN/Inf to `Value::Null`
|
||||
|
||||
**File**: `src/materialize.rs:451-455`
|
||||
|
||||
**Problem**: `serde_json::Number::from_f64` returns `None` for NaN/Inf,
|
||||
so `number_from_f64` substitutes `Value::Null`. A NaN float in the buffer
|
||||
then surfaces as "expected a number" from `validate_float`
|
||||
(`bast_validation.rs:176`) rather than "expected a finite number". The
|
||||
error is misleading, though the outcome (rejection) is correct.
|
||||
|
||||
**Fix** (optional): have the materializer propagate a non-finite float
|
||||
as an `AlkTypeError::Access` at read time, or leave as-is and accept the
|
||||
slightly-off error message. Not a correctness bug.
|
||||
|
||||
---
|
||||
|
||||
### N2. `BastType::alk_kind()` returns `Struct` for any `$ref`
|
||||
|
||||
**File**: `src/bast.rs:715-725`
|
||||
|
||||
**Problem**: `BastType::Ref(_) => AlkTypeKind::Struct` is documented but
|
||||
a footgun — a `$ref` to a union/enum misreports its kind unless the
|
||||
caller resolves first. Most callers do resolve first, but the invariant
|
||||
is fragile and easy to break in a future edit.
|
||||
|
||||
**Fix** (optional): leave as-is (documented) or make `alk_kind` return
|
||||
`Option<AlkTypeKind>` / require resolution. Defer unless it bites.
|
||||
|
||||
---
|
||||
|
||||
### N3. `check_bytes` accepts both `String` and `Array` forms
|
||||
|
||||
**File**: `src/bast_validation.rs:213-256`
|
||||
|
||||
**Problem**: `check_bytes` handles `Value::String` and `Value::Array`,
|
||||
but the materializer only ever emits `Array` for bytes
|
||||
(`materialize.rs:212`). The `String` arm is dead/legacy. Harmless, but
|
||||
it widens the accepted surface for no reason.
|
||||
|
||||
**Fix** (optional): drop the `String` arm, or keep it if a future
|
||||
materializer emits bytes as a string. Defer.
|
||||
|
||||
---
|
||||
|
||||
### N4. Stale ADR references in doc comments
|
||||
|
||||
**Files**: `src/error.rs:3` ("ADR-098"), `src/tunion.rs:1` ("ADR-097"),
|
||||
`src/engine.rs:51,196` ("ADR-101"), `src/offset_map.rs:1`,
|
||||
`src/layout_builder.rs:1`, `src/sequential_reader.rs:1` ("ADR-096")
|
||||
|
||||
**Problem**: none of these ADR numbers exist in
|
||||
`docs/architecture/decisions/` (which has 001–010 + `bast-bast-format` +
|
||||
`val-split-two-validator-model`). The pivot renumbered/renamed ADRs but
|
||||
the code comments were not synced. A reader following the reference hits
|
||||
a dead end.
|
||||
|
||||
**Fix**: map each stale reference to the correct ADR (e.g. "ADR-096" →
|
||||
ADR-002 for the two layout modes, "ADR-101" → ADR-007 for the packed
|
||||
read factory, "ADR-098" → ADR-004 for error handling, "ADR-097" →
|
||||
ADR-003 for annotations) and update the comments. ~6 lines.
|
||||
|
||||
---
|
||||
|
||||
## The `Timestamp` kind
|
||||
|
||||
`AlkTypeKind::Timestamp` is a first-class kind that is byte-identical to
|
||||
`String` everywhere (length-prefixed UTF-8), and its only distinguishing
|
||||
behavior is `is_rfc3339_timestamp` (`bast_validation.rs:406`) — a
|
||||
hand-rolled, non-strict check that the doc itself admits "Feb 31 passes;
|
||||
seconds range isn't checked; leap seconds aren't handled." The parsing
|
||||
is fragile (the `rfind('-')` timezone-offset heuristic, no
|
||||
fractional-second handling).
|
||||
|
||||
It is documented as matching v0.1.0, so it is not a regression, but it
|
||||
is the weakest part of the validator and adds a 19th kind plus a
|
||||
`needs_endian` / `is_variable_length` / `natural_alignment` arm, all to
|
||||
validate a string that a consumer could validate with a standard JSON
|
||||
Schema `format: "date-time"` on the `validate_json` path.
|
||||
|
||||
**Decision (publisher)**: remove it. It is a residual from an early
|
||||
research reference that included a timestamp; it is largely irrelevant
|
||||
at the BAST level, and JSON-level timestamp validation is `jsonschema`'s
|
||||
job, not alktype's. Tracked as a follow-up task, not part of this
|
||||
review's findings.
|
||||
|
||||
---
|
||||
|
||||
## What's Good
|
||||
|
||||
The crate is in notably good shape after the pivot. Highlights:
|
||||
|
||||
- **The BAST parser is clean and defensive.** `bast.rs` returns
|
||||
`AlkTypeError::Schema` on every malformed-document path, never panics,
|
||||
and uses `checked_add` / `usize::try_from` for all count/offset casts.
|
||||
The typed tree (`BastDoc`/`BastDef`/`BastType`) is a real improvement
|
||||
over the v0.1.0 raw-JSON accessors.
|
||||
- **The enum index-bounds fix is correct.** `validate_enum`
|
||||
(`bast_validation.rs:274`) checks the materialized index against
|
||||
`values.len()`, closing the v0.1.0 dead constraint. Well-tested.
|
||||
- **Overflow safety is thorough.** `checked_add` everywhere in the hot
|
||||
paths; the `write_bytes` u32 truncation from review #002 (M2) is
|
||||
fixed and the guard pattern is now the norm.
|
||||
- **Error attribution is excellent.** Every `Access`/`Offset` error
|
||||
carries a `field_path`; the BAST parser errors carry a dotted path
|
||||
into the document (`"bast: struct at .fields[2] ..."`).
|
||||
- **The two-validator split is clean.** `bast_validation` (bytes) and
|
||||
`validation` (JSON) are clearly separated, and the
|
||||
`AlkTypeError::Validation` payload stays uniform across both
|
||||
(D-BAST-009).
|
||||
- **Tests are strong where they exist.** 389 tests, good coverage of
|
||||
error paths, both endiannesses, short buffers, unknown discriminators,
|
||||
invalid UTF-8. The gaps are precisely the paths M1/M2 identify.
|
||||
- **No `unsafe`, no `TODO`/`FIXME`** — clean codebase hygiene.
|
||||
|
||||
---
|
||||
|
||||
## Recommended Order
|
||||
|
||||
1. **M1** (field-level endian override) — ~10 lines + tests, closes a
|
||||
silent-corruption path on a documented feature. Smallest and
|
||||
highest-value.
|
||||
2. **M2** (aligned-mode materialization) — decide implement-vs-reject
|
||||
with the publisher; the reject option is small and matches the
|
||||
ADR-006/ADR-008 pattern.
|
||||
3. **M3** (meta-schema at compile time) — ~5 lines + tests, closes a
|
||||
silent-misinterpretation path on untrusted input.
|
||||
4. **L1** (offset-indirect dead code) — decide together with M2.
|
||||
5. **L2** (dead `endian` parameter) — ~4 lines, clarity only.
|
||||
6. **N1–N4** — hygiene; N4 (stale ADR refs) is worth doing in the same
|
||||
pass as the `Timestamp` removal since both touch doc comments.
|
||||
|
||||
The `Timestamp` removal is a separate, self-contained task the publisher
|
||||
has already decided on; it can be done independently of the above.
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All line numbers refer to the tree at commit `562284f` (the last
|
||||
commit on `main` at review time).
|
||||
- The coverage numbers are from `cargo llvm-cov --release` on the same
|
||||
tree. The `--summary-only` and `--show-missing-lines` outputs were
|
||||
used to attribute gaps; the full HTML report is at
|
||||
`target/llvm-cov/html`.
|
||||
- This review does not cover documentation quality (README, inline docs,
|
||||
docs.rs rendering) beyond the stale-ADR-reference nit (N4). Per the
|
||||
publisher's workflow, that is a separate sweep.
|
||||
- Findings M1 and M2 were confirmed by throwaway integration tests that
|
||||
were removed after reproduction; the regression tests for the fixes
|
||||
should be added to the permanent suite.
|
||||
@@ -0,0 +1,412 @@
|
||||
---
|
||||
status: closed
|
||||
last_updated: 2026-09-02
|
||||
reviewed_artifacts:
|
||||
- src/sequential_reader.rs
|
||||
- src/bast.rs
|
||||
- src/engine.rs
|
||||
- src/layout_builder.rs
|
||||
- src/materialize.rs
|
||||
- src/offset_map.rs
|
||||
- src/data_access.rs
|
||||
- src/lib.rs
|
||||
- docs/architecture/decisions/007-packed-mode-read-factory.md
|
||||
- ../@alkdev/alktty/benches/wire_vs_bast.rs
|
||||
tool: manual source read + downstream criterion bench (`cargo bench --bench wire_vs_bast` in alktty)
|
||||
reviewer: read-path performance review (triggered by alktty wire_vs_bast bench)
|
||||
---
|
||||
|
||||
# Review #004 — Read-Path Performance: the `BastDoc` Re-Parse Gap
|
||||
|
||||
## Purpose
|
||||
|
||||
A downstream bench in `alktty` (`benches/wire_vs_bast.rs`) compared a
|
||||
hand-rolled `ChunkHeader` codec against an alktype-driven codec built
|
||||
from the same 2-field BAST struct (`stream_type: uint8`,
|
||||
`length: uint32`, big-endian). The bench builds the engine / layout /
|
||||
reader **once** outside the measured loop, then measures per-chunk read
|
||||
and write over 1024 contiguous chunks.
|
||||
|
||||
The write path is competitive (~2.8x at 64 B, ~1.1x at 4 KiB — the
|
||||
per-chunk overhead is just `data_access::write_u8`/`write_u32` at
|
||||
precomputed `PackedLayout` offsets, and the payload copy dominates at
|
||||
4 KiB). The read path is **~400x slower per chunk** (2.27 µs/chunk vs
|
||||
5.6 ns/chunk hand-rolled), and the cost is fixed — it dominates at 64 B
|
||||
*and* at 4 KiB.
|
||||
|
||||
This review traces the gap to its source, confirms it is an
|
||||
implementation gap (not inherent to the design), and lays out the fix
|
||||
options. The bench itself is honest — its in-file note
|
||||
(`wire_vs_bast.rs:36-41`) already points at the root cause; this review
|
||||
formalizes the finding and the remediation plan.
|
||||
|
||||
## Methodology
|
||||
|
||||
- Full read of the read-path code (`sequential_reader.rs`, `bast.rs`,
|
||||
`engine.rs`), the write path (`layout_builder.rs`, `data_access.rs`),
|
||||
and ADR-007 (the packed read factory decision).
|
||||
- Cross-reference every `BastDoc::new` call site in `src/` to map the
|
||||
full re-parse surface.
|
||||
- Trace the lifetime/ownership constraint that forces the re-parse
|
||||
(`BastDoc<'a>` borrows `&'a Value`; `SequentialReader` owns its
|
||||
`Value` — self-referential struct, cannot cache the parsed tree).
|
||||
- Read the downstream bench to confirm the measurement is honest (the
|
||||
re-parse is inside the measured routine; the engine/reader are built
|
||||
once outside it).
|
||||
- Read `docs/reviews/003-code-review.md` for prior context — review #003
|
||||
did not flag the re-parse (it was a correctness/coverage pass, not a
|
||||
performance pass).
|
||||
|
||||
## Verification Baseline
|
||||
|
||||
The bench numbers below are from the alktty downstream tree
|
||||
(`benches/wire_vs_bast.rs`, criterion 0.7), run against alktype at
|
||||
commit `cab4932` (v0.2.0). The alktype tree itself is unchanged — this
|
||||
is a review of existing code, not a fix.
|
||||
|
||||
| path | payload=64 B | payload=4 KiB |
|
||||
|---|---|---|
|
||||
| hand-rolled read | 5.7 µs (5.6 ns/chunk) | 12.1 µs (11.8 ns/chunk) |
|
||||
| alktype read | 2.32 ms (2.27 µs/chunk) | 2.38 ms (2.33 µs/chunk) |
|
||||
| hand-rolled write | 14.3 µs | 321 µs |
|
||||
| alktype write | 39.8 µs | 355 µs |
|
||||
|
||||
One-shot startup costs (paid once): `AlkTypeEngine::compile` 573 µs,
|
||||
`engine.sequential_reader()` 2.84 µs, `LayoutBuilder::build` 1.30 µs.
|
||||
|
||||
The read gap is ~400x per chunk and is **fixed** (does not shrink as
|
||||
the payload grows), which is the signature of per-field overhead, not
|
||||
payload-copy overhead.
|
||||
|
||||
## Summary Statistics
|
||||
|
||||
| Severity | Count |
|
||||
|----------|------:|
|
||||
| High | 1 (H1) |
|
||||
| Medium | 1 (M1) |
|
||||
| Low | 2 (L1, L2) |
|
||||
| Nit | 1 (N1) |
|
||||
|
||||
The High finding is the read-path re-parse — a severe performance
|
||||
regression that blocks the primary intended use case (alktype as a
|
||||
runtime codec for protocol wire formats). It is not a correctness bug;
|
||||
the re-parse produces the same result every time, it is just
|
||||
catastrophically wasteful. The Medium finding is the same root cause
|
||||
manifesting in four one-shot paths. The Low/Nit findings are dead code
|
||||
and a stale doc comment exposed while tracing the root cause.
|
||||
|
||||
---
|
||||
|
||||
## Findings
|
||||
|
||||
### H1. `SequentialReader::read_field_at` re-parses the BAST typed tree on every field read
|
||||
|
||||
**File**: `src/sequential_reader.rs:262`
|
||||
|
||||
**Problem**: `read_field_at` calls `BastDoc::new(&self.doc_value,
|
||||
&self.root_name)?` on every field read. `BastDoc::new`
|
||||
(`src/bast.rs:69`) recursively parses the root def: `lookup_def_raw`
|
||||
(hash lookup into the `serde_json` Map), `BastDef::parse` →
|
||||
`BastStruct::parse` → allocates a `Vec<BastField>`, iterates fields,
|
||||
`BastField::parse` each (several `node.get().as_str()`/`as_u64()`
|
||||
calls + a `format!` allocation for the field path), `BastType::parse`
|
||||
each. For the 2-field `ChunkHeader`, that is ~1 µs per call.
|
||||
|
||||
`read_next` calls `read_field_at` once per field, so a 2-field header
|
||||
incurs **two** `BastDoc::new` calls per chunk = ~2.27 µs/chunk, matching
|
||||
the bench. For an N-field struct, reading all fields is O(N²) in parse
|
||||
work (each of N reads re-parses all N fields) — the gap widens with
|
||||
struct width.
|
||||
|
||||
The re-parse is purely redundant: `SequentialReader::new`
|
||||
(`src/sequential_reader.rs:130`) already parsed the `BastDoc` once at
|
||||
construction and extracted the field list. `read_field_at` re-derives
|
||||
the same `field_node` (the `BastField` at `index`) and the same `doc`
|
||||
that were already in hand at construction time.
|
||||
|
||||
**Root cause**: a lifetime/ownership tension, not a logic error.
|
||||
`BastDoc<'a>` borrows `&'a Value` and `&'a str` throughout
|
||||
(`src/bast.rs:51-55`). `SequentialReader` **owns** a cloned `Value`
|
||||
(`doc_value`, `src/sequential_reader.rs:111`). To cache the parsed
|
||||
`BastDoc` in the reader, the `BastDoc` would have to borrow from
|
||||
`self.doc_value` — a self-referential struct, which Rust's borrow
|
||||
checker forbids and safe Rust cannot express without a self-referential
|
||||
crate (`self_cell`/`ouroboros`, both rejected: new dep per AGENTS.md
|
||||
§7, and `self_cell`'s soundness relies on `unsafe` the crate avoids per
|
||||
AGENTS.md §11). So the only way to get a `BastDoc` at read time is to
|
||||
re-parse from the owned `Value`. The engine's own doc comment
|
||||
(`src/engine.rs:112-115`) acknowledges this design explicitly:
|
||||
|
||||
> "The engine retains a clone of the BAST `Value` so that
|
||||
> `sequential_reader` and `read_field` can re-parse the typed tree on
|
||||
> demand without lifetime entanglement with the caller's `Value`."
|
||||
|
||||
The "re-parse on demand" framing was a lifetime-entanglement workaround
|
||||
that did not anticipate the hot-loop cost. ADR-007's "Cost" section
|
||||
(`decisions/007-packed-mode-read-factory.md:64-70`) argues construction
|
||||
is cheap ("a `Vec<(String, Value)>` of the `properties` entries ... a
|
||||
struct has a small number of fields") — true for *construction* (once),
|
||||
but the decision did not account for a per-field re-parse inside the
|
||||
read loop.
|
||||
|
||||
**Why the write path is fine**: `LayoutBuilder::build`
|
||||
(`src/layout_builder.rs:190`) also re-parses `BastDoc::new`, but it is
|
||||
called **once** per write — the resulting `PackedLayout` offsets are
|
||||
cached and reused. The write loop then does only `data_access::write_*`
|
||||
at fixed offsets. There is no per-field re-parse on the write hot path,
|
||||
which is why write is ~1.1x at 4 KiB.
|
||||
|
||||
**Fix**: cache the parsed typed tree so the read path does not re-parse.
|
||||
See "Fix Options" below for the three approaches and the
|
||||
recommendation.
|
||||
|
||||
**Lift**: closes a ~400x read-path gap and unblocks alktype as a runtime
|
||||
codec for protocol wire formats (its stated purpose per ADR-001). This
|
||||
is the difference between "plausible runtime codec" and "not viable."
|
||||
Large effort depending on the chosen option.
|
||||
|
||||
---
|
||||
|
||||
### M1. The same re-parse pattern exists in four one-shot paths
|
||||
|
||||
**Files**: `src/layout_builder.rs:190`, `src/engine.rs:284,334,467`
|
||||
|
||||
**Problem**: `BastDoc::new(&self.doc_value, &self.root_name)?` /
|
||||
`BastDoc::new(&self.bast_doc, &self.root_name)?` is re-called in:
|
||||
|
||||
- `LayoutBuilder::build` (`src/layout_builder.rs:190`) — once per
|
||||
`build()` call. Re-parsing on each build is wasteful if a builder is
|
||||
reused across writes, but the typical pattern is build-once-reuse,
|
||||
so this is mild.
|
||||
- `AlkTypeEngine::validate_bytes` (`src/engine.rs:284`) — once per
|
||||
validate call. For a stream of buffers, this is a per-buffer re-parse.
|
||||
- `AlkTypeEngine::read_field` (`src/engine.rs:334`, aligned mode) — once
|
||||
per field read. Same class as H1 but for aligned random access, and
|
||||
one parse per field read (not the O(N²) of the sequential reader).
|
||||
- `AlkTypeEngine::write_field` (`src/engine.rs:467`, aligned mode) —
|
||||
once per field write.
|
||||
|
||||
These are less acute than H1 (one parse per operation, not per-field-
|
||||
in-a-loop), but they share the same root cause: the engine/builder own
|
||||
a `Value` and cannot cache a borrowing `BastDoc`. Any fix that makes
|
||||
`BastDoc` cacheable on an owning struct (Fix Option A) closes these for
|
||||
free; a read-path-only fix (Option B) leaves them as-is, which is
|
||||
acceptable since they are not hot loops.
|
||||
|
||||
**Lift**: removes redundant parse work on the validate/aligned paths.
|
||||
Free with Option A; deferred with Option B.
|
||||
|
||||
---
|
||||
|
||||
### L1. `read_field_value` carries a dead `_field_schema` parameter; `fields` stores dead `Value` clones
|
||||
|
||||
**Files**: `src/sequential_reader.rs:114,294`
|
||||
|
||||
**Problem**: `SequentialReader` stores `fields: Vec<(String, Value)>`
|
||||
where the `Value` is `f.source().clone()` per field
|
||||
(`src/sequential_reader.rs:145`). `read_field_at` passes this as
|
||||
`_field_schema` to `read_field_value` (`src/sequential_reader.rs:294`),
|
||||
where it is unused (prefixed `_`). The raw `Value` clone per field is
|
||||
dead weight — only the `String` name is used (for `read_next`'s return
|
||||
and `read_field`'s lookup). This is a minor allocation cost on top of
|
||||
H1's re-parse, and it will be removed naturally when the read plan is
|
||||
precomputed (Option B) or the `BastDoc` is cached (Option A), since
|
||||
both replace `Vec<(String, Value)>` with typed/owned field data.
|
||||
|
||||
**Lift**: trivial; falls out of the H1 fix.
|
||||
|
||||
---
|
||||
|
||||
### L2. ADR-007 "Cost" section and the engine doc comment understate the re-parse
|
||||
|
||||
**Files**: `docs/architecture/decisions/007-packed-mode-read-factory.md:64-70`,
|
||||
`src/engine.rs:112-115`
|
||||
|
||||
**Problem**: ADR-007's "Cost" section argues `sequential_reader()` is
|
||||
cheap because it clones a small `Vec` of field schemas. That is true for
|
||||
the factory call (once). But the decision did not anticipate that
|
||||
`read_field_at` would re-parse `BastDoc::new` per field — the cost that
|
||||
actually dominates. The engine doc comment at `src/engine.rs:112-115`
|
||||
explicitly frames re-parse-on-demand as the intended design ("re-parse
|
||||
the typed tree on demand without lifetime entanglement"), which is the
|
||||
root cause H1 traces.
|
||||
|
||||
**Fix**: whichever fix option is chosen, update ADR-007's "Cost" /
|
||||
"Consequences" section and the engine doc comment to reflect that the
|
||||
parsed tree is now cached (Option A) or precomputed into a read plan
|
||||
(Option B), and that the "re-parse on demand" framing is retired.
|
||||
|
||||
**Lift**: documentation accuracy; prevents the same framing from
|
||||
misleading a future edit.
|
||||
|
||||
---
|
||||
|
||||
### N1. `BastType::alk_kind()` returns `Struct` for any `$ref` (carry-forward from review #003 N2)
|
||||
|
||||
**File**: `src/bast.rs:715-725`
|
||||
|
||||
**Problem**: flagged in review #003 N2 and left as "defer unless it
|
||||
bites." It does not bite here — `read_field_value` always calls
|
||||
`doc.resolve_typeref(ty)` before matching on `BastType`, so the
|
||||
misreporting `alk_kind` is never consulted on a `Ref`. Noting it only
|
||||
because this review re-read the same path; no new action beyond review
|
||||
#003's deferral.
|
||||
|
||||
---
|
||||
|
||||
## Fix Options
|
||||
|
||||
The core constraint: `BastDoc<'a>` borrows `&'a Value` / `&'a str`; an
|
||||
owning struct (`SequentialReader`, `AlkTypeEngine`, `LayoutBuilder`)
|
||||
cannot store a `BastDoc` that borrows from its own `Value` field
|
||||
(self-referential). Three ways to break the constraint:
|
||||
|
||||
### Option A — Make the typed tree own its data (principled fix)
|
||||
|
||||
Change `BastDoc<'a>` → `BastDoc` (no lifetime), `&'a str` → `Arc<str>`
|
||||
(or `String`), `&'a Value` → `Arc<Value>` (or `Value`). Then
|
||||
`SequentialReader`, `AlkTypeEngine`, and `LayoutBuilder` each hold a
|
||||
`BastDoc` directly (built once at construction), and `read_field_at`
|
||||
uses `&self.doc` — no re-parse, anywhere.
|
||||
|
||||
- **Closes**: H1, M1 (all four one-shot paths), and the engine/reader
|
||||
lifetime entanglement that ADR-007 worked around. The engine's
|
||||
`bast_doc: Value` clone (`src/engine.rs:85`) becomes redundant with
|
||||
the owned `BastDoc`.
|
||||
- **Tradeoff**: broad refactor. Touches `bast.rs` (every typed node)
|
||||
and every consumer (`layout_builder`, `offset_map`, `materialize`,
|
||||
`sequential_reader`, `tunion`, `bast_validation`, `engine`). The
|
||||
`Bast*` types are re-exported in `lib.rs:57-60`, so this is a
|
||||
**breaking public-API change** — `BastDoc<'a>` becomes `BastDoc`,
|
||||
and every method signature that took `&'a` changes. Per AGENTS.md,
|
||||
this is semver-relevant and would need a version bump (0.2.0 → 0.3.0).
|
||||
- **Dependency cost**: `Arc<str>`/`Arc<Value>` add `alloc` (already in
|
||||
use via `Vec`/`String`); no new external deps. Stays wasm-clean. The
|
||||
`preserve_order` serde_json feature remains load-bearing (AGENTS.md
|
||||
§8) — owning the `Value` does not change field-order semantics.
|
||||
- **Effort**: large but mechanical. The borrow-based design was chosen
|
||||
for "allocation-free beyond the small typed nodes" (`src/bast.rs:11-
|
||||
18`), but the re-parse-per-field already defeats that goal by
|
||||
allocating a fresh `Vec<BastField>` per read. Owning the data makes
|
||||
the "parse once, walk many times" invariant actually hold.
|
||||
|
||||
### Option B — Precompute an owned read plan in `SequentialReader::new` (surgical fix)
|
||||
|
||||
Keep `BastDoc<'a>` borrowing for the other consumers. In
|
||||
`SequentialReader::new`, parse the `BastDoc` once, resolve all `$ref`s
|
||||
eagerly, and build a flat, owned tree of read instructions
|
||||
(`Vec<FieldPlan>`) that the read loop walks with no `BastDoc`
|
||||
involvement. Each `FieldPlan` carries the field name, the resolved
|
||||
`AlkTypeKind`, the effective `Endian`, and for composites a nested
|
||||
plan (struct → sub-plans; union → discriminator + per-variant plans;
|
||||
array → element plan + count; record → value plan).
|
||||
|
||||
- **Closes**: H1 only. M1 (the one-shot re-parses) remains, which is
|
||||
acceptable since they are not hot loops.
|
||||
- **Tradeoff**: non-breaking (internal to `sequential_reader.rs`; the
|
||||
public `SequentialReader` type and its methods keep their
|
||||
signatures). Duplicates some of the `BastType` matching logic that
|
||||
`read_field_value`/`read_union_value`/etc. already encode, so there
|
||||
are two parallel walkers to maintain.
|
||||
- **Effort**: medium. Self-contained in one file but non-trivial
|
||||
(composites require recursively resolving and pre-flattening the
|
||||
type tree, including `$ref` chains into `$defs`).
|
||||
|
||||
### Option C — Cache `BastDoc` on the engine, reader borrows (rejected)
|
||||
|
||||
Have the engine own the parsed `BastDoc` and return a
|
||||
`SequentialReader<'_>` that borrows from `&self`. This requires
|
||||
`BastDoc` to be owned (Option A prerequisite) *and* changes
|
||||
`sequential_reader() -> Option<SequentialReader>` to
|
||||
`-> Option<SequentialReader<'_>>` — a breaking public-API change that
|
||||
also contradicts ADR-007's "owned fresh reader" decision. Strictly
|
||||
worse than Option A (same refactor cost, more API churn, contradicts
|
||||
an ADR). Rejected.
|
||||
|
||||
### Recommendation
|
||||
|
||||
**Option A**, given the publisher's stated willingness to make breaking
|
||||
changes ("no one is using this except us yet; ... we can change things
|
||||
now"). It is the only option that closes H1 *and* M1 and retires the
|
||||
lifetime-entanglement workaround that caused both. The refactor is
|
||||
broad but mechanical (lifetime removal, not logic rewrites), and the
|
||||
crate is pre-1.0 with only two in-house downstream consumers
|
||||
(`alktty`, `alkcall`), so the breakage cost is bounded and known.
|
||||
|
||||
Option B is the fallback if the Option A refactor is deferred — it
|
||||
closes the acute H1 gap non-breakingly while leaving M1 for later. It
|
||||
is not the recommended path because it leaves a second parallel type
|
||||
walker in the crate and does not address the root cause (the borrow-
|
||||
based `BastDoc` design), which will keep forcing re-parses anywhere a
|
||||
new owning consumer wants to cache the parsed tree.
|
||||
|
||||
Regardless of the chosen option, ADR-007's "Cost"/"Consequences"
|
||||
section and the `src/engine.rs:112-115` doc comment should be updated to
|
||||
retire the "re-parse on demand" framing (L2).
|
||||
|
||||
---
|
||||
|
||||
## What's Good
|
||||
|
||||
- **The bench is honest.** The alktty bench builds the engine/reader
|
||||
once outside the measured loop and correctly isolates the per-chunk
|
||||
logic. Its in-file note (`wire_vs_bast.rs:36-41`) already points at
|
||||
the `read_field_at` re-parse and labels it "the honest current cost of
|
||||
the alktype read path, not a bench bug." This review confirms that
|
||||
assessment.
|
||||
- **The write path is already competitive.** Once `PackedLayout` is
|
||||
built, the write loop is just `data_access::write_*` at fixed offsets
|
||||
— no schema walk, no re-parse. This validates the "build once, reuse"
|
||||
pattern that the read path should also adopt.
|
||||
- **`resolve_typeref` for primitives is cheap.** `BastType::Primitive`
|
||||
is `Copy` (`AlkTypeKind: Copy`, `src/schema.rs:25`), so
|
||||
`resolve_typeref` (`src/bast.rs:117-125`) returns `other.clone()`
|
||||
without allocation for the common case. The re-parse cost is entirely
|
||||
in `BastDoc::new`, not in the per-field type resolution — so caching
|
||||
the `BastDoc` alone closes the gap without restructuring
|
||||
`resolve_typeref`.
|
||||
- **Overflow safety and error attribution are unaffected.** The
|
||||
`checked_add` / `usize::try_from` discipline (AGENTS.md §4) and the
|
||||
`field_path`-carrying errors (review #003 "What's Good") are in the
|
||||
read functions, not the parser — a caching fix preserves them.
|
||||
|
||||
---
|
||||
|
||||
## Recommended Order
|
||||
|
||||
1. **H1 + M1 (Option A)** — the owned-typed-tree refactor. Decide
|
||||
first (this is a one-way door: breaking public-API change, version
|
||||
bump to 0.3.0). If approved, this is one refactor that closes both.
|
||||
2. **L2** — update ADR-007 and the engine doc comment in the same
|
||||
commit as the H1 fix, since the "re-parse on demand" framing is
|
||||
being retired.
|
||||
3. **L1** — falls out of the H1 fix (the dead `Value` clones are
|
||||
replaced by the cached/owned field data).
|
||||
4. **N1** — remains deferred per review #003.
|
||||
|
||||
If Option A is deferred, **H1 (Option B)** is the standalone
|
||||
alternative — non-breaking, closes the acute gap only.
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All line numbers refer to the tree at commit `cab4932` (v0.2.0, the
|
||||
BAST pivot release).
|
||||
- The bench is in the `alktty` downstream repo
|
||||
(`/workspace/@alkdev/alktty/benches/wire_vs_bast.rs`), not in alktype.
|
||||
alktty depends on alktype as a path dev-dep for the bench only; it
|
||||
does not use alktype at runtime. The path dep means
|
||||
`cargo publish --dry-run` for alktty would complain (the bench is
|
||||
exploratory and uncommitted in alktty; the alktype crate itself has
|
||||
no bench dependency).
|
||||
- This review does not cover the wasm build (`cargo build --target
|
||||
wasm32-unknown-unknown`) because the fix is not yet implemented; the
|
||||
verification block for the fix should include it per AGENTS.md, as
|
||||
the typed-tree ownership change touches `bast.rs` which is
|
||||
wasm-relevant.
|
||||
- Review #003 (post-BAST-pivot correctness review) did not flag the
|
||||
re-parse — its scope was correctness, coverage, and panic safety, not
|
||||
performance. The re-parse is not a correctness regression; the parsed
|
||||
tree is identical across calls. This review complements #003 by
|
||||
adding the performance axis.
|
||||
@@ -0,0 +1,659 @@
|
||||
---
|
||||
status: closed
|
||||
last_updated: 2026-09-02
|
||||
resolved_findings: 2026-08-20 (all 11 — see "Resolution" at the end)
|
||||
reviewed_artifacts:
|
||||
- docs/plans/030-compiled-forms.md
|
||||
- docs/architecture/decisions/011-compiled-read-plan-for-packed-mode.md
|
||||
- docs/architecture/decisions/012-plan-fingerprinting-and-m1-closure.md
|
||||
- docs/reviews/004-performance-review.md
|
||||
- src/lib.rs
|
||||
- src/bast.rs
|
||||
- src/engine.rs
|
||||
- src/sequential_reader.rs
|
||||
- src/materialize.rs
|
||||
- src/offset_map.rs
|
||||
- src/layout_builder.rs
|
||||
- src/schema.rs
|
||||
- poc/readplan/{src/lib.rs, FINDINGS.md} (branch readplan-poc)
|
||||
tool: manual source read + plan-vs-codebase cross-check + POC branch inspection
|
||||
reviewer: 0.3.0 implementation plan review (triggered before phase 1)
|
||||
---
|
||||
|
||||
# Review #005 — 0.3.0 Plan Review: Compiled Forms
|
||||
|
||||
## Purpose
|
||||
|
||||
The 0.3.0 implementation plan
|
||||
([`docs/plans/030-compiled-forms.md`](../plans/030-compiled-forms.md)) is
|
||||
the entry point an implementing agent reads first. It rolls up
|
||||
[ADR-011](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md)
|
||||
(the `ReadPlan` packed read-side compiled form),
|
||||
[ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md)
|
||||
(fingerprinting + owned `BastDoc` + `OffsetMap` `LeafMeta`), and the
|
||||
fingerprinting work into one breaking bump. The plan is deliberately
|
||||
structured as seven phases so each can be picked up by a fresh session
|
||||
without prior context.
|
||||
|
||||
This review's purpose is to find planning-spec mistakes — factual
|
||||
errors, contradictions, undocumented behavioral changes, hedges into an
|
||||
unplanned future — *before* a phase-by-phase implementation starts,
|
||||
because fresh-session implementations are reliable precisely when the
|
||||
spec is accurate. A spec that contradicts the code or an ADR forces the
|
||||
agent to either reverse-engineer the actual intent or guess, and the
|
||||
failure rate goes up.
|
||||
|
||||
The review explicitly scans for the "deferral black hole" pattern: a
|
||||
plan or ADR puts work off into a "future version/phase/downstream" with
|
||||
no concrete reactivation condition, the next agent inherits the gap,
|
||||
and the gap festers until something forces an untangle. This is a
|
||||
known LLM-planning quirk distinct from classic planning mistakes, and
|
||||
a default scan for it is part of this review's methodology.
|
||||
|
||||
## Methodology
|
||||
|
||||
- Full read of the plan and its two companion ADRs (011, 012), the
|
||||
performance review (#004) the plan closes, and the POC findings on
|
||||
branch `readplan-poc`.
|
||||
- Cross-check every line-number reference and `src/` claim in the plan
|
||||
against the actual codebase at `main` (commit `2310f6c`, v0.2.0).
|
||||
Verified: `engine.rs:112-115,284,334,467`; `sequential_reader.rs:567`;
|
||||
`layout_builder.rs:190`; `bast.rs:51-55`; `lib.rs` re-export list;
|
||||
`Cargo.toml` version; presence of `poc/` (absent on main, present on
|
||||
`readplan-poc` as expected); presence of `alktty`/`alkcall` downstream
|
||||
path dev-deps.
|
||||
- Cross-check the POC's `ReadPlan`/`CompositePlan` shape against both
|
||||
ADR-011's shape section and the plan's phase-1 description.
|
||||
- Cross-check the plan's phase 2 rewrite claims (`dummy_field_for`/
|
||||
`ty_source` "are removed") against `materialize.rs`'s actual call
|
||||
sites across both packed and aligned paths.
|
||||
- Verify the derives the plan relies on (`Hash` on `LeafMeta`,
|
||||
`ReadPlan`, `OffsetMap`) are reachable from the derives on their
|
||||
constituent types in `src/schema.rs`.
|
||||
- Scan for the deferral pattern by flagging every "future/deferred/
|
||||
later/downstream/if needed" occurrence and asking: (a) is there a
|
||||
concrete reactivation trigger? (b) is the decision owned or silent?
|
||||
(c) does inaction have a cost that the deferral framing hides?
|
||||
|
||||
## Verification Baseline
|
||||
|
||||
The plan and both ADRs were read at the tree state at commit `2310f6c`
|
||||
("Propose ADR-012 + 0.3.0 implementation plan"), which is `main` HEAD.
|
||||
The codebase is v0.2.0 (`Cargo.toml`); the POC lives on branch
|
||||
`readplan-poc` and is not merged, as the plan states. All line-number
|
||||
references in the plan were verified correct against this tree.
|
||||
|
||||
## Summary Statistics
|
||||
|
||||
| Severity | Count |
|
||||
|----------|------:|
|
||||
| High | 2 (H1, H2) |
|
||||
| Medium | 3 (M1, M2, M3) |
|
||||
| Low | 3 (L1, L2, L3) |
|
||||
| Nit | 3 (N1, N2, N3) |
|
||||
|
||||
The two High findings are correctness/contradiction issues that would
|
||||
block or mislead an implementing agent. The Mediums are either
|
||||
undocumented behavioral drops, missing implementation prerequisites, or
|
||||
a deferral worth re-evaluating. Lows and Nits are wording/typo-level.
|
||||
|
||||
---
|
||||
|
||||
## Findings
|
||||
|
||||
### H1. Phase 1's field-name-discriminator union shape exists in neither ADR-011 nor the POC
|
||||
|
||||
**File**: `docs/plans/030-compiled-forms.md:144-152`
|
||||
|
||||
**Problem**: The plan describes the field-name-discriminator union read
|
||||
shape as:
|
||||
|
||||
> `CompositePlan::Union` carries the union's declared `fields` as a
|
||||
> sub-`ReadPlan` (the discriminator field + any shared fields), and the
|
||||
> variant plans are laid out *after* the shared fields.
|
||||
|
||||
But ADR-011 §"The `ReadPlan` shape"
|
||||
(`011-compiled-read-plan-for-packed-mode.md:144-158`) defines:
|
||||
|
||||
```rust
|
||||
pub enum CompositePlan {
|
||||
Struct(ReadPlan),
|
||||
Union {
|
||||
disc: DiscriminatorPlan,
|
||||
variants: Vec<(String, VariantPlan)>,
|
||||
},
|
||||
...
|
||||
}
|
||||
```
|
||||
|
||||
There is no field for shared/declared fields on the `Union` variant.
|
||||
The POC (`readplan-poc:poc/readplan/src/lib.rs`) matches the ADR's
|
||||
shape — `CompositePlan::Union { disc, variants }` only — and its
|
||||
`compile_union` does not carry shared fields. The POC's `FINDINGS.md`
|
||||
Finding 1 (the same one the plan cites at lines 143-152) explicitly
|
||||
says:
|
||||
|
||||
> `plan_read_union`'s `Field` arm is a stub that returns an error.
|
||||
> ... The plan needs a sub-struct for the union's declared fields,
|
||||
> separate from the variant plans.
|
||||
|
||||
So the plan describes a shape that exists in **neither** the accepted
|
||||
ADR **nor** the reference POC, and presents it as "the production
|
||||
version must implement it" within the existing `CompositePlan::Union`
|
||||
shape. An implementing agent reading ADR-011 + plan + POC gets three
|
||||
different `CompositePlan::Union` shapes and no guidance on where the
|
||||
shared-fields sub-`ReadPlan` goes (a new `shared: Option<Box<ReadPlan>>`
|
||||
field? a wrapper enum? two-variant split?).
|
||||
|
||||
This is a shape extension to an accepted ADR's public type. The plan
|
||||
either needs to flag it as an ADR-011 refinement (with the ADR updated
|
||||
first) or specify the concrete shape the agent should build.
|
||||
|
||||
**Lift**: unblocks phase 1. Without resolution, the agent will either
|
||||
guess the shape and likely diverge from intent, or stop and ask.
|
||||
|
||||
---
|
||||
|
||||
### H2. `schema()` returning `&Value` from a `&Value` "stored on the plan" is a self-referential struct
|
||||
|
||||
**File**: `docs/plans/030-compiled-forms.md:188-191`
|
||||
|
||||
**Problem**: Phase 2 says:
|
||||
|
||||
> `schema()` returns a `&Value` retained on the plan (the plan stores
|
||||
> the `&Value` it was compiled from — see ADR-011 §Engine integration;
|
||||
> the `&Value` outlives the plan because the engine owns both).
|
||||
|
||||
The "the plan stores the `&Value` it was compiled from" is the
|
||||
self-referential struct pattern ADR-011 §"Root cause"
|
||||
(`011-...md:50-58`) explicitly identifies as impossible in safe Rust
|
||||
and rejects. The engine owns `bast_doc: Value` and `Arc<ReadPlan>`. If
|
||||
`ReadPlan` stores `&Value` borrowing from the engine's `bast_doc`, the
|
||||
engine is self-referential — exactly the construction ADR-007 worked
|
||||
around with "re-parse on demand" and ADR-011's `Arc<ReadPlan>` was
|
||||
meant to retire. ADR-011 line 204 specifies the plan is "immutable
|
||||
**owned** data"; it does not say the plan stores a `&Value`.
|
||||
|
||||
`schema()`'s current contract (`src/sequential_reader.rs:247`) is to
|
||||
return the raw BAST `Value` the reader was built from. To preserve
|
||||
that contract on `Arc<ReadPlan>` without a self-referential borrow,
|
||||
`ReadPlan` must store an `Arc<Value>` (engine builds `Arc<Value>` at
|
||||
compile time, hands a clone to the plan) or an owned `Value`. Then
|
||||
`schema()` returns `&self.plan.value`. The plan should specify which.
|
||||
|
||||
**Lift**: prevents an agent from getting stuck in phase 2 trying to
|
||||
make `&Value` in `Arc<ReadPlan>` work, which the borrow checker will
|
||||
reject.
|
||||
|
||||
---
|
||||
|
||||
### M1. Nested unions: POC rejects a schema 0.2.0 accepts — undocumented behavioral drop
|
||||
|
||||
**Files**: `poc/readplan/FINDINGS.md` ("What this POC does not cover"),
|
||||
`docs/plans/030-compiled-forms.md` (silent), `src/sequential_reader.rs:789-815`
|
||||
|
||||
**Problem**: The POC's `compile_union` rejects a union variant that is
|
||||
itself a union with `AlkTypeError::Schema`. The existing reader
|
||||
supports this: `resolve_and_walk_variant` at
|
||||
`src/sequential_reader.rs:800` has a live `BastDefKind::Union` arm
|
||||
that recurses via `read_union_value`. So 0.2.0 accepts and reads
|
||||
nested-union schemas; phase 1's `ReadPlan::compile` (per the POC the
|
||||
plan cites as the reference scaffold) would reject the same schema.
|
||||
|
||||
The plan's phase 1 calls out two POC findings explicitly (field-disc
|
||||
union shape → H1 above, struct-array stride → deferred decision 4)
|
||||
and says "the production version must implement/decide these." It does
|
||||
**not** call out the nested-union rejection. An agent following the
|
||||
plan would inherit the POC's reject-nested-unions behavior by default,
|
||||
silently dropping a 0.2.0 capability — a behavioral regression that
|
||||
rides the 0.3.0 bump without being listed in the Semver Contract
|
||||
table.
|
||||
|
||||
This is also the cleanest example of the deferral-black-hole pattern
|
||||
in the plan: the POC says "if a real schema needs it, the
|
||||
implementation step adds a `VariantKind::Union` read path. Not
|
||||
blocking — no current schema exercises it." The "if needed" framing
|
||||
has no trigger, no OQ, no tracking — it's a black hole. The next agent
|
||||
inherits the gap.
|
||||
|
||||
**Lift**: either (a) add `VariantKind::Union` read path in phase 1
|
||||
(small — mirrors the existing `resolve_and_walk_variant` Union arm,
|
||||
~20 lines), or (b) list it in the Semver Contract table as a
|
||||
behavioral drop with a one-line OQ tracking the deferral. Given the
|
||||
plan says there are zero real consumers, (b) is defensible, but it
|
||||
must be *stated*, not silent. (a) is cheap and avoids the regression.
|
||||
|
||||
---
|
||||
|
||||
### M2. `Endian` and `VariableEncoding` don't derive `Hash` — phases 5/6 will not compile
|
||||
|
||||
**Files**: `src/schema.rs:205,212`, `docs/plans/030-compiled-forms.md:62,357,411-415`
|
||||
|
||||
**Problem**: `src/schema.rs:205` (`Endian`) and `:212`
|
||||
(`VariableEncoding`) both derive only `Debug, Clone, Copy, PartialEq,
|
||||
Eq` — no `Hash`. The plan requires:
|
||||
- Phase 5 (line 62, 357): `LeafMeta { kind, encoding, endian }` as
|
||||
`Copy + PartialEq + Eq + Hash`.
|
||||
- Phase 6 (lines 411-415): `#[derive(Hash, Eq)]` on `ReadPlan`/
|
||||
`OffsetMap`, and `FieldPlan` carries `endian: Endian` + `encoding:
|
||||
VariableEncoding`.
|
||||
|
||||
Both derives will fail to compile: `#[derive(Hash)]` on a struct
|
||||
requires all fields to be `Hash`. The plan never mentions adding
|
||||
`Hash` to these two enums. The fix is trivial (both are fieldless
|
||||
enums, already `Eq + PartialEq`, so adding `Hash` is semver-safe —
|
||||
additive, no behavioral change), but it's a prerequisite the plan
|
||||
omits. An agent working phase 5 will hit a compile error and have to
|
||||
diagnose why.
|
||||
|
||||
**Lift**: trivial. Add a sub-step to phase 5 (or 6): "Add `Hash` to
|
||||
`Endian` and `VariableEncoding` derives in `src/schema.rs`." This is
|
||||
additive and safe to do earlier if convenient.
|
||||
|
||||
---
|
||||
|
||||
### M3. `ValidationPlan` deferral worth re-evaluating — the read+validate common case
|
||||
|
||||
**Files**: `docs/architecture/decisions/012-...md:64-73` ("Deferring
|
||||
`ValidationPlan`"), `docs/plans/030-compiled-forms.md:526-528`,
|
||||
`docs/reviews/004-performance-review.md` (the read-path perf review)
|
||||
|
||||
**Problem**: ADR-012 defers a `ValidationPlan` as "different shape
|
||||
(value-domain, not byte-position), not a hot loop, separate ADR if a
|
||||
bench motivates it." The plan inherits this deferral ("Not a
|
||||
`ValidationPlan`" at lines 526-528). The deferral framing is "if a
|
||||
bench motivates it" — a concrete trigger exists, so this is not a
|
||||
black-hole hedge in the M1 sense.
|
||||
|
||||
Flagged for re-evaluation, not because the shape argument is wrong
|
||||
(it's correct — value-domain checks are structurally different from
|
||||
byte-position walks), but because the *hot-loop* dismissal may under-
|
||||
account a common case: **read + validate together on untrusted input.**
|
||||
|
||||
Review #004 found the packed read path was 400x slow per chunk due to
|
||||
per-field `BastDoc` re-parse. ADR-011 closes that. But
|
||||
`validate_bytes`'s packed path (ADR-010) is `materialize_packed` →
|
||||
`bast_validation::validate_value` over the materialized `Value`.
|
||||
After ADR-011, `materialize_packed` walks the `ReadPlan` (fast).
|
||||
`bast_validation::validate_value` still walks `BastDoc` to check
|
||||
value-domain constraints — once per `validate_bytes` call, over the
|
||||
full tree, on every buffer.
|
||||
|
||||
For a stream of N untrusted buffers (the `alkcall` hub/spoke topology
|
||||
accepts schemas from arbitrary internet peers — AGENTS.md §3 — and
|
||||
the common case is "read incoming frame, validate it before acting"),
|
||||
`validate_bytes` is called N times. Each call does one `BastDoc`
|
||||
walk for validation. After ADR-011, the *read* half of `validate_bytes`
|
||||
is plan-fast; the *validation* half is still a `BastDoc` walk per call.
|
||||
If validation is the common companion to read on untrusted input,
|
||||
then skipping validation is risky (accepting untrusted bytes
|
||||
unchecked) and running it re-walks `BastDoc` per buffer — the same
|
||||
class of cost review #004 measured for the read path, just on a
|
||||
different code path.
|
||||
|
||||
The argument is not "ValidationPlan has the same shape as ReadPlan"
|
||||
(it doesn't). The argument is: ADR-012's "not a hot loop" dismissal
|
||||
may be incomplete, because read+validate on untrusted streams makes
|
||||
validation hot in the same sense read was hot. The deferral's
|
||||
trigger ("if a bench motivates it") should be sharpened: either (a)
|
||||
add a `validate_bytes`-on-untrusted-stream bench to alktty alongside
|
||||
`wire_vs_bast` and let the bench decide, or (b) reason from the
|
||||
existing review #004 numbers that the validation walk is
|
||||
non-trivial and should be planned, not deferred.
|
||||
|
||||
This is not a request to implement `ValidationPlan` in 0.3.0. It's a
|
||||
request to *own the decision*: either the trigger fires (and a
|
||||
follow-on ADR/phase is scoped, possibly 0.4.0) or it doesn't (and the
|
||||
deferral stands with a sharper justification than "not a hot loop").
|
||||
As written, the deferral leaves the cost in the superposition where
|
||||
it can neither be confirmed nor dismissed.
|
||||
|
||||
**Lift**: removes a latent perf cliff for the read+validate-on-
|
||||
untrusted-input case that 0.3.0 is supposed to make viable.
|
||||
|
||||
---
|
||||
|
||||
### L1. `dummy_field_for`/`ty_source` are used in aligned `materialize`, not just packed
|
||||
|
||||
**Files**: `docs/plans/030-compiled-forms.md:209-210`,
|
||||
`src/materialize.rs:249,316,351,391,631,650-663`
|
||||
|
||||
**Problem**: Phase 2 says:
|
||||
|
||||
> The `dummy_field_for`/`ty_source` helpers in `materialize.rs` are
|
||||
> removed (the plan carries everything).
|
||||
|
||||
This is factually wrong. `dummy_field_for` is called at
|
||||
`src/materialize.rs:631` inside `materialize_leaf_at`, which is called
|
||||
by the **aligned** path: `materialize_struct_aligned` (line 475),
|
||||
`materialize_array_aligned` (line 544), `materialize_variable_aligned`
|
||||
(line 613). Aligned `materialize` keeps walking `BastDoc` through
|
||||
0.3.0 (plan lines 379-385 confirm), so `dummy_field_for`/`ty_source`
|
||||
must stay. Only the packed-side call sites (lines 249, 316, 351, 391)
|
||||
go away when packed-materialize moves to the plan.
|
||||
|
||||
**Lift**: doc accuracy. An agent following the plan literally would
|
||||
remove the helpers and break aligned `materialize`.
|
||||
|
||||
---
|
||||
|
||||
### L2. `materialize_packed` rewrite scope underspecified — packed-vs-aligned split of `materialize_typeref_packed`
|
||||
|
||||
**Files**: `docs/plans/030-compiled-forms.md:207-210`,
|
||||
`src/materialize.rs:122-200, 498-506, 619-637`
|
||||
|
||||
**Problem**: `materialize_typeref_packed` is shared by both packed and
|
||||
aligned paths — aligned's `materialize_leaf_at` (line 619-637) calls
|
||||
`materialize_typeref_packed` to read leaves, and aligned's record path
|
||||
(line 498-506) calls it directly. Phase 2 says
|
||||
`materialize_packed(&ReadPlan, &[u8])` walks the plan instead of
|
||||
`BastDoc` but does not state what happens to
|
||||
`materialize_typeref_packed`.
|
||||
|
||||
The honest resolution: packed-materialize gets a new plan-walking
|
||||
function; aligned keeps `materialize_typeref_packed` via
|
||||
`materialize_leaf_at`; the function stays (renamed or not) for aligned.
|
||||
This is two mode-specific paths — the existing design — not a
|
||||
"parallel walker" in the maintenance-tax sense ADR-011 §"Negative"
|
||||
(cautioning against) discusses. ADR-011's "one walker" claim (lines
|
||||
234-237) is specifically about packed read-side (`SequentialReader` +
|
||||
`materialize_packed` sharing the plan), not packed-vs-aligned, so
|
||||
there's no ADR contradiction — just an underspecification in the plan.
|
||||
|
||||
**Lift**: prevents the agent from having to discover the split
|
||||
mid-rewrite. Add one line to phase 2: "packed-materialize gets a new
|
||||
plan-walking function; `materialize_typeref_packed` stays for
|
||||
aligned's `materialize_leaf_at` and the aligned record path."
|
||||
|
||||
---
|
||||
|
||||
### L3. `materialize_aligned`'s `BastDoc` structure walk is silent in the plan
|
||||
|
||||
**Files**: `docs/plans/030-compiled-forms.md` (silent on this),
|
||||
`src/materialize.rs:451-521`, `docs/architecture/decisions/011-...md:264`
|
||||
|
||||
**Problem**: `materialize_struct_aligned` walks `BastDoc` to traverse
|
||||
struct/array/record structure, using `OffsetMap` only for leaf byte
|
||||
positions. ADR-011 §"Out of scope" says "aligned mode is unchanged;
|
||||
`materialize_aligned` already takes `&OffsetMap`" — which is half
|
||||
true: it takes `&OffsetMap` for positions but also `&BastDoc` for
|
||||
structure. The plan inherits the half-truth silently: there's no
|
||||
statement anywhere that aligned materialize keeps walking `BastDoc`
|
||||
for structure.
|
||||
|
||||
After phase 3 (owned `BastDoc`) + phase 5 (`LeafMeta`), the walk is
|
||||
over owned data, no re-parse, not O(N²), and aligned `validate_bytes`
|
||||
is one walk per call (not per-field). There's no perf driver
|
||||
analogous to review #004's packed per-chunk gap. But the absence of
|
||||
a driver is not the same as a decision: leaving it silent is a
|
||||
deferral-by-omission. An implementing agent or future reader can't
|
||||
tell whether the silence is "this is the permanent design" or "we'll
|
||||
fix this later."
|
||||
|
||||
The decision should be owned. Either (a) add a "Scope Boundary" note
|
||||
that aligned materialize keeps walking owned `BastDoc` for structure
|
||||
as the permanent design (with an OQ if a future bench motivates an
|
||||
`AlignedPlan`), or (b) if a bench motivation is plausible, scope an
|
||||
OQ to track it. (a) is recommended — no perf driver, and after phase
|
||||
3 the walk is over owned data, so it's not the re-parse pattern.
|
||||
|
||||
**Lift**: removes a silent gap that future agents would otherwise
|
||||
have to reverse-engineer.
|
||||
|
||||
---
|
||||
|
||||
### N1. Typo: "back-comat" → "back-compat"
|
||||
|
||||
**File**: `docs/plans/030-compiled-forms.md:104-105`
|
||||
|
||||
**Problem**: "back-comat" in deferred decision 4.
|
||||
|
||||
**Lift**: trivial.
|
||||
|
||||
---
|
||||
|
||||
### N2. Phase 1 verification omits the `Send + Sync` assertion test ADR-011 requires
|
||||
|
||||
**Files**: `docs/plans/030-compiled-forms.md:158-163`,
|
||||
`docs/architecture/decisions/011-...md:204-206`
|
||||
|
||||
**Problem**: ADR-011 §"Engine integration" says "the implementation
|
||||
should add a `static` bound assertion test to lock it in" for
|
||||
`ReadPlan: Send + Sync`. Phase 1's verification block lists `cargo
|
||||
test`, `clippy`, `doc`, `wasm` but no mention of adding the assertion
|
||||
test. An agent following the plan literally won't add it; the
|
||||
property is currently true by construction but not asserted, so a
|
||||
future change could break it silently.
|
||||
|
||||
**Lift**: add "add a `fn read_plan_is_send_sync()` assertion test" to
|
||||
phase 1's verification, mirroring the POC's
|
||||
`readplan_is_send_sync` test.
|
||||
|
||||
---
|
||||
|
||||
### N3. `SequentialReader::new` return-type change (`Result` drop) undocumented
|
||||
|
||||
**Files**: `docs/plans/030-compiled-forms.md:64`,
|
||||
`src/sequential_reader.rs:129`, `src/engine.rs:205`
|
||||
|
||||
**Problem**: Currently `new(&Value, &str) -> Result<Self,
|
||||
AlkTypeError>` — fallible (BastDoc parse). After phase 2,
|
||||
`new(Arc<ReadPlan>)` is infallible (just stores the Arc) → returns
|
||||
`Self`, not `Result<Self>`. The Semver Contract table (line 64) lists
|
||||
only the argument-type change, not the `Result` drop.
|
||||
`engine.rs:205`'s `.ok()` call correspondingly goes away. Minor, but
|
||||
it's a signature change beyond what's listed.
|
||||
|
||||
**Lift**: add a row to the Semver Contract table noting the `Result`
|
||||
drop.
|
||||
|
||||
---
|
||||
|
||||
## Deferral-pattern scan (LLM-planning quirk)
|
||||
|
||||
As part of the methodology, every "future/deferred/later/downstream/if
|
||||
needed" occurrence in the plan and its ADRs was flagged and tested
|
||||
for: (a) concrete reactivation trigger, (b) decision owned or silent,
|
||||
(c) hidden cost of inaction.
|
||||
|
||||
| Item | Trigger? | Owned? | Cost of inaction | Finding |
|
||||
|---|---|---|---|---|
|
||||
| `ValidationPlan` (ADR-012) | "if a bench motivates it" | Yes (ADR + plan "What this is not") | Possible perf cliff on read+validate untrusted streams | M3 above — sharpen the trigger |
|
||||
| Nested-union `ReadPlan` support | "if a real schema needs it" (POC) | No (POC only, plan silent) | Silent 0.2.0 capability drop | M1 above — state it |
|
||||
| `materialize_aligned` structure walk | None — silent | No (silent) | Future agent ambiguity | L3 above — own the decision |
|
||||
| `Arc<str>` vs `String` (decision 1) | "if phase 4 shows it's measurable" | Yes (deferred decision 1) | None | OK — has trigger, decided in phase 3 |
|
||||
| `OffsetMap::get` shape (decision 2) | "decided in phase 5" | Yes (deferred decision 2) | None | OK |
|
||||
| Fingerprint hasher (decision 3) | "decided in phase 6" | Yes (deferred decision 3) | None | OK |
|
||||
| Struct-array stride (decision 4) | "decided in phase 2" | Yes (deferred decision 4) | None | OK |
|
||||
| `BastDoc` `Arc<Value>` vs `Value` | None — silent | No (plan doesn't address) | Agent gets stuck (H2) | H2 above |
|
||||
| Field-disc union shape (POC Finding 1) | "production version must implement" | Yes (plan phase 1) | None, but shape is undefined | H1 above — shape not in ADR |
|
||||
|
||||
The four explicit "deferred decisions" in the plan (items 4-7) all
|
||||
have concrete triggers and decision points — these are the *good*
|
||||
pattern. The black-hole pattern appears where deferrals lack triggers
|
||||
(items 1-3, 8-9): three of those became findings (M1, L3, H2), and M3
|
||||
is a deferral worth sharpening even though it has a trigger.
|
||||
|
||||
The general signal: a deferral is healthy when it has a concrete
|
||||
reactivation condition and is tracked (OQ, ADR, or in-plan deferred
|
||||
decision). A deferral is a black hole when it has no trigger, no
|
||||
tracking, and the next agent inherits the gap by default.
|
||||
|
||||
---
|
||||
|
||||
## What's Good
|
||||
|
||||
- **Line-number accuracy is perfect.** Every `src/` reference in the
|
||||
plan (`engine.rs:112-115,284,334,467`;
|
||||
`sequential_reader.rs:567`; `layout_builder.rs:190`;
|
||||
`bast.rs:51-55`; `lib.rs` re-exports) checks out against the v0.2.0
|
||||
tree. This is unusual for a plan of this length and worth noting.
|
||||
- **The Semver Contract table is a strong scope-creep guardrail.**
|
||||
Walking every public `lib.rs` re-export against the table, the
|
||||
classifications (Breaking / Unchanged / New) are correct for every
|
||||
item, with the exceptions noted in N3 (the `Result` drop on `new`)
|
||||
and M1 (the nested-union behavioral drop not listed).
|
||||
- **The four explicit "deferred decisions" are the right pattern.**
|
||||
Each has a trigger and a decision point in a named phase. This is
|
||||
what deferrals should look like.
|
||||
- **Phases are coherent session boundaries.** Phases 1 (pure
|
||||
addition), 6 (pure addition), 7 (docs/bump) are small and clean.
|
||||
Phases 3 (broad but mechanical), 4 (single file), 5 (single file +
|
||||
engine) are well-scoped. Phase 2 is the largest and the plan
|
||||
sanctions sub-session splits at the step level (lines 40-42), which
|
||||
is the right escape valve.
|
||||
- **Cross-phase invariants are stated and checkable.** "Tree builds
|
||||
and tests pass at every phase boundary" is the right invariant; the
|
||||
POC-on-`readplan-poc`-only convention is clearly separated from
|
||||
production code; AGENTS.md §5-§11 constraints (no `unsafe`, no
|
||||
`async`, no new deps, `preserve_order` load-bearing) are
|
||||
reaffirmed.
|
||||
- **The plan honestly scopes what it is not.** "Not a `ValidationPlan`",
|
||||
"Not cross-version fingerprint stability", "Not a perf bench" —
|
||||
these boundaries are stated rather than left implicit, which helps
|
||||
an implementing agent resist scope creep. (M3 above is about
|
||||
sharpening one of these, not removing the boundary.)
|
||||
- **The POC reference is disciplined.** The plan is explicit that the
|
||||
POC is "not production code," lives only on the branch, and is the
|
||||
reference scaffold for phases 1-2 only. This matches how
|
||||
`bast-validator-poc` was handled and avoids the POC leaking into
|
||||
`main`.
|
||||
|
||||
---
|
||||
|
||||
## Recommended Order
|
||||
|
||||
1. **H1 (field-disc union shape)** — update ADR-011's
|
||||
`CompositePlan::Union` to include the shared-fields sub-`ReadPlan`
|
||||
(or document the wrapper shape), then update the plan's phase 1 to
|
||||
reference the corrected ADR shape. Do this before phase 1 starts;
|
||||
otherwise the implementing agent has to guess.
|
||||
2. **H2 (`schema()` `&Value` on `Arc<ReadPlan>`)** — edit the plan's
|
||||
phase 2 to specify `ReadPlan` stores `Arc<Value>` (or owned
|
||||
`Value`), and `schema()` borrows from that. One-line edit to the
|
||||
plan; avoid a phase-2 stuck point.
|
||||
3. **M1 (nested unions)** — decide (a) implement `VariantKind::Union`
|
||||
in phase 1, or (b) list as behavioral drop + OQ. Edit the plan and
|
||||
(if b) the Semver Contract table accordingly. Decide before phase
|
||||
1.
|
||||
4. **M2 (`Hash` on `Endian`/`VariableEncoding`)** — add a sub-step
|
||||
to phase 5 or 6. Trivial.
|
||||
5. **M3 (`ValidationPlan` re-evaluation)** — either add a
|
||||
`validate_bytes`-on-untrusted-stream bench to alktty (alongside
|
||||
`wire_vs_bast`) and let the bench decide, or sharpen ADR-012's
|
||||
"not a hot loop" justification. Does not block 0.3.0; can be
|
||||
resolved in parallel with phase 1-7 work. **Flagged for
|
||||
re-evaluation, not for implementation in 0.3.0.**
|
||||
6. **L1, L2, L3** — edit the plan's phase 2 to fix the
|
||||
`dummy_field_for` wording (L1), state the packed-vs-aligned
|
||||
materialize split (L2), and add a Scope Boundary note for
|
||||
aligned-materialize's `BastDoc` structure walk (L3). All three are
|
||||
phase-2 doc edits.
|
||||
7. **N1, N2, N3** — typo, `Send + Sync` assertion test, `Result`-drop
|
||||
Semver row. Minor plan edits.
|
||||
|
||||
Items 1-3 must be resolved before phase 1 starts (they affect the
|
||||
`ReadPlan` shape or 0.2.0 behavioral surface). Items 4-7 can be
|
||||
resolved any time before their phase begins. Item 5 (M3) is
|
||||
non-blocking and can run in parallel.
|
||||
|
||||
---
|
||||
|
||||
## Notes
|
||||
|
||||
- All line numbers refer to the tree at commit `2310f6c` (the plan's
|
||||
commit) for `src/` files, and to the plan/ADR markdown as committed
|
||||
at the same tree.
|
||||
- The POC on `readplan-poc` was inspected via
|
||||
`git show readplan-poc:poc/readplan/{src/lib.rs,FINDINGS.md}`; it is
|
||||
not merged to `main` and the plan correctly states this.
|
||||
- `alktty` and `alkcall` downstream repos exist as path dev-deps
|
||||
(`/workspace/@alkdev/alktty`, `/workspace/@alkdev/alkcall`); the
|
||||
plan's claim that they're in-house and updated with the bump is
|
||||
verifiable, though this review did not inspect their call sites
|
||||
in detail.
|
||||
- This review does not re-litigate ADR-011 or ADR-012's accepted
|
||||
decisions. H1 and H2 are about the plan *contradicting* the ADRs or
|
||||
being unsound, not about the ADR decisions themselves; M3 is about
|
||||
sharpening a deferral, not about re-deciding it.
|
||||
- The deferral-pattern scan is a methodology experiment: a
|
||||
pre-declared scan for LLM-specific planning quirks (deferral black
|
||||
holes) alongside classic planning mistakes. It surfaced M1 and L3
|
||||
that a conventional severity-only review would have missed or
|
||||
under-weighted. Worth retaining as a default scan for future plan
|
||||
reviews.
|
||||
|
||||
---
|
||||
|
||||
## Resolution (2026-08-20)
|
||||
|
||||
All 11 findings resolved in one docs-only edit pass to ADR-011,
|
||||
ADR-012, and the 0.3.0 plan. No source changed; the crate still
|
||||
builds/tests at v0.2.0. The M3 deferral reversal is the one
|
||||
substantive decision change (per user direction: ship ValidationPlan
|
||||
in 0.3.0, no more hedging); the rest are spec corrections or
|
||||
pre-implementation refinements to types that do not yet exist on
|
||||
`main`.
|
||||
|
||||
- **H1 (union shape):** ADR-011 §"The `ReadPlan` shape" refined —
|
||||
`CompositePlan::Union` now carries `shared: Option<Box<ReadPlan>>`
|
||||
(field-disc shared fields) and `variants: Vec<(String,
|
||||
CompositePlan)>` (dropping `VariantPlan`/`VariantKind`). Plan
|
||||
phase 1 rewritten to implement the refined shape. The shape
|
||||
refinement is pre-implementation (the types don't exist on `main`).
|
||||
- **H2 (`schema()` `&Value`):** plan phase 2 rewritten — `ReadPlan`
|
||||
stores `schema: Arc<Value>` (not `&Value`); `schema()` returns
|
||||
`&self.schema`. Verified `serde_json::Value: Hash + Eq` holds with
|
||||
`preserve_order` (`Map::hash` sorts keys deterministically), so
|
||||
phase 6's `#[derive(Hash)]` on `ReadPlan` is not blocked.
|
||||
- **M1 (nested unions):** resolved as the review's option (a) —
|
||||
nested-union support falls out of the H1 shape refinement (a
|
||||
variant can be `CompositePlan::Union`), so no behavioral drop vs
|
||||
0.2.0 and no Semver Contract entry for a capability regression.
|
||||
Plan phase 1 adds a nested-union-variant test.
|
||||
- **M2 (`Hash` on `Endian`/`VariableEncoding`):** plan phase 5
|
||||
rewritten with an explicit first sub-step to add `Hash` to both
|
||||
derives in `src/schema.rs` (additive, semver-safe). The inaccurate
|
||||
"all fields are `Copy + Hash`" parenthetical on `LeafMeta` is
|
||||
corrected.
|
||||
- **M3 (`ValidationPlan`):** deferral **reversed** per user
|
||||
direction. ADR-012 §"Deferring `ValidationPlan`" rewritten as
|
||||
"ValidationPlan — in scope for 0.3.0"; new ADR-012 §3 commits the
|
||||
decision (compiled form, no per-buffer `BastDoc` walk, `Hash + Eq`
|
||||
+ `fingerprint()`) and lists the shape questions deferred to a
|
||||
follow-on design session + the plan's new phase 7. Plan gains a
|
||||
new phase 7 (ValidationPlan); old phase 7 (bump) renumbered to
|
||||
phase 8. ADR-011's "Out of scope" `bast_validation` bullet and
|
||||
"Scope Boundaries" `Not a validation plan` bullet updated to point
|
||||
at ADR-012 §3. Plan's "What this plan is *not*" first bullet
|
||||
removed. The deferral-black-hole pattern this review's methodology
|
||||
flagged is closed: the work is committed in the plan with a
|
||||
concrete reactivation trigger (the shape session before phase 7),
|
||||
not hedged into an unplanned future.
|
||||
- **L1 (`dummy_field_for`/`ty_source`):** plan phase 2 rewritten —
|
||||
only the packed-side call sites go away; the helpers stay for the
|
||||
aligned `materialize_leaf_at` path.
|
||||
- **L2 (`materialize_typeref_packed` split):** plan phase 2
|
||||
rewritten — packed-materialize gets a new plan-walking function;
|
||||
`materialize_typeref_packed` stays for aligned's
|
||||
`materialize_leaf_at` and the aligned record path.
|
||||
- **L3 (aligned-materialize `BastDoc` structure walk):** plan phase 5
|
||||
gains a Scope Boundary note — the walk is the permanent 0.3.0
|
||||
design; an `AlignedPlan` is out of scope, tracked as an open
|
||||
question if a future bench motivates it.
|
||||
- **N1 (typo):** "back-comat" → "back-compat" in deferred decision 4.
|
||||
- **N2 (`Send + Sync` assertion test):** plan phase 1 verification
|
||||
rewritten to add the `read_plan_is_send_sync` static-bound
|
||||
assertion test ADR-011 §"Engine integration" requires.
|
||||
- **N3 (`Result` drop on `SequentialReader::new`):** Semver Contract
|
||||
table row updated to note the constructor return-type change
|
||||
(`Result<Self, AlkTypeError>` → `Self`) alongside the argument-type
|
||||
change.
|
||||
|
||||
The deferral-pattern scan's general signal (healthy deferrals have a
|
||||
concrete reactivation condition + tracking; black holes have neither)
|
||||
is reaffirmed by the M3 reversal: the original "if a bench motivates
|
||||
it" trigger was a black hole because no bench was ever going to be
|
||||
run against a path that didn't exist yet, and the cost of inaction
|
||||
(a second breaking change to `validate_bytes`/`bast_validation` after
|
||||
0.3.0) was hidden by the "not a hot loop" framing.
|
||||
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,454 @@
|
||||
---
|
||||
status: resolved (F1, F2, C1, C2, C3, L1, L2 resolved 2026-09-02; N1/N2a/N3a/N4a are classified-no-action / deferred-by-design)
|
||||
last_updated: 2026-09-02
|
||||
reviewed_artifacts:
|
||||
- src/materialize.rs
|
||||
- src/sequential_reader.rs
|
||||
- src/read_plan.rs
|
||||
- src/offset_map.rs
|
||||
- src/layout_builder.rs
|
||||
- src/engine.rs
|
||||
- src/data_access.rs
|
||||
- src/bast.rs
|
||||
- src/builder.rs
|
||||
- src/tunion.rs
|
||||
- src/validation_plan.rs
|
||||
- src/walk_guard.rs
|
||||
- src/bast_meta.rs
|
||||
- tests/poc_roundtrip.rs
|
||||
- tests/tunion_dispatch.rs
|
||||
- tests/error_paths.rs
|
||||
- tests/engine_integration.rs
|
||||
- docs/reviews/006-implementation-review-030.md (post-fix coverage re-check)
|
||||
tool: cargo-llvm-cov 0.8.4 (--release, per-line text) + manual classification of every uncovered production line + disposable probe tests (run in-session, then deleted)
|
||||
reviewer: post-review-#006 coverage audit (session request — check test coverage for weak spots, meaningful tests, non-happy-path posture)
|
||||
---
|
||||
|
||||
# Review #007 — Post-#006 Coverage Audit
|
||||
|
||||
## Purpose
|
||||
|
||||
Review #006 closed every finding and its M4 coverage map, but the
|
||||
session-level posture (M4's item: "fold a coverage check into each fix
|
||||
session") had never been run as a *whole-tree* pass after all those
|
||||
fixes landed. This audit re-measures coverage after the eleven #006
|
||||
commits, reads every uncovered production line, and classifies it —
|
||||
the same "untested-but-fine / load-bearing / unreachable" discipline
|
||||
M4's map used. Two probes were run in disposable tests (deleted after
|
||||
the session, per #006's no-reproducer rule; neither was a crash
|
||||
hazard — both reproduce cleanly inside the default harness).
|
||||
|
||||
## Methodology
|
||||
|
||||
- `cargo llvm-cov --release` (0.8.4, same tool as #006): summary +
|
||||
per-line text. TOTAL **90.67% lines / 86.32% functions** — stable
|
||||
with #006's post-M4 numbers (90.60%), the expected drift after the
|
||||
N3 fix sessions added parse gates + tests.
|
||||
- Per-file (worst first): `materialize.rs` 85.72, `data_access.rs`
|
||||
80.32, `sequential_reader.rs` 86.19, `bast.rs` 87.41,
|
||||
`builder.rs` 91.28, `layout_builder.rs` 91.42, `tunion.rs` 92.02,
|
||||
`offset_map.rs` 92.93, `read_plan.rs` 90.69, `engine.rs` 96.44,
|
||||
`validation_plan.rs` 93.97, `walk_guard.rs` 98.04,
|
||||
`bast_meta.rs` 98.92, `bast_validation.rs`/`error.rs`/`schema.rs`/
|
||||
`macros.rs`/`validation.rs` 100.
|
||||
- Every uncovered line *outside* `#[cfg(test)]` modules (806 raw
|
||||
lines) was read and classified. Lines inside test modules (the
|
||||
`panic!("expected X, got {other:?}")` helpers) were excluded — they
|
||||
distort per-file numbers (e.g. `bast.rs`'s 87.41% is really ~96%
|
||||
production once its 60 helper lines are excluded).
|
||||
- Two suspicions were probe-verified with disposable tests:
|
||||
the F1 cross-consumer divergence and the F2 unbounded-`maxLength`
|
||||
compile. Probe transcripts quoted verbatim in the findings.
|
||||
- Happy-path posture audit: cross-checked which *error arms* adjacent
|
||||
to covered code are 0-execution, and which public surfaces have only
|
||||
success-path tests.
|
||||
|
||||
## Baseline
|
||||
|
||||
Audited at `main` HEAD `bb28ba3` ("Resolve N3"), 0.3.0, working tree
|
||||
clean. Full suite green (548 tests static + 2 ignored doctests, per
|
||||
#006's bookkeeping).
|
||||
|
||||
## Summary Statistics
|
||||
|
||||
| Severity | Count | Status |
|
||||
|----------|------:|--------|
|
||||
| High | 2 (F1, F2) | both resolved 2026-09-02 |
|
||||
| Medium | 3 (C1, C2, C3) | all resolved 2026-09-02 |
|
||||
| Low | 2 (L1, L2) | all resolved 2026-09-02 |
|
||||
| Info | 4 (N1, N2a, N3a, N4a) | classified: N1 artifact, N2a/N4a no-action, N3a deferred to pre-release review |
|
||||
|
||||
**Resolution log:**
|
||||
|
||||
- **F1 + F2 (2026-09-02):** resolved in one commit — see the
|
||||
resolution blocks on each finding. 477 lib tests green (511
|
||||
static + 2 ignored doctests across all targets), clippy
|
||||
`-D warnings` clean, wasm build green.
|
||||
- **C1 + C2 + C3 (2026-09-02):** resolved in one commit — see the
|
||||
resolution blocks. 481 lib tests green, clippy `-D warnings` clean,
|
||||
wasm build green; `compile_variant`'s cycle arm confirmed executed
|
||||
in the post-fix coverage run.
|
||||
- **L1 + L2 (2026-09-02):** resolved in one commit — see the
|
||||
resolution blocks. 488 lib tests green, clippy `-D warnings` clean,
|
||||
wasm build green.
|
||||
|
||||
---
|
||||
|
||||
## Findings
|
||||
|
||||
### F1. Zero-progress guard missing in the plan materializer — `validate_bytes` accepts what `SequentialReader` rejects (cross-consumer divergence)
|
||||
|
||||
**Files**: `src/materialize.rs:227-253` (`materialize_plan_array` — no
|
||||
guard), contrast `src/sequential_reader.rs:772-795`
|
||||
(`plan_walk_variable_array_size` — has the guard) and
|
||||
`src/materialize.rs:636-663` (`materialize_array_packed` — has the
|
||||
guard)
|
||||
|
||||
**Problem**: The H1 fix session added the zero-progress runtime guard
|
||||
("array element consumed 0 bytes") to two of the three array walkers:
|
||||
the compiled reader's variable-array size walk and the legacy BAST
|
||||
walker's packed array arm. The *plan-based* packed materializer —
|
||||
`materialize_plan_array`, the walker `validate_bytes` actually uses in
|
||||
packed mode (engine.rs:310-312) — got no guard.
|
||||
|
||||
A stride-0 array whose elements consume 0 bytes (empty-struct elements
|
||||
are legal: the meta-schema's `StructDef` has no `minItems` on
|
||||
`fields`) compiles with `element_stride: 0` and loops `count` times
|
||||
materializing empty objects without reading a single buffer byte:
|
||||
|
||||
```
|
||||
PROBE validate_bytes([]): OK — zero-progress guard MISSING in plan materializer
|
||||
PROBE reader.read_next([]): Err(access error at items[0]: array element 0 consumed 0 bytes;
|
||||
a zero-size element makes the declared count unbounded on the wire)
|
||||
```
|
||||
|
||||
Schema: `{ "items": { "kind": "array", "element": { "kind":
|
||||
"struct", "fields": [] }, "count": 8 } }`, packed mode, empty buffer.
|
||||
Same schema, same buffer, opposite verdicts — the exact
|
||||
cross-consumer-disagreement shape review #006 existed for (H3, M6).
|
||||
Severity High by #006's own keying (AGENTS.md §3): `validate_bytes`
|
||||
is the flagship untrusted-input path, and it silently accepts a
|
||||
buffer the same engine's reader rejects. The H1 resolution text
|
||||
("plan_walk_variable_array_size (reader) and materialize_array_packed
|
||||
(materializer) now error") lists only two of the three walkers — the
|
||||
plan materializer was missed because it is *not* the legacy walker
|
||||
that finding named.
|
||||
|
||||
**Not a #006 regression**: the H1 fix text itself specified only the
|
||||
reader and legacy-walker sites; the plan materializer predates the
|
||||
guard and was outside that fix's blast radius. But the divergence is
|
||||
new information — the guard's *invariant* ("a zero-progress element
|
||||
makes the declared count unbounded") belongs to the array-walk
|
||||
concept, not to two specific functions.
|
||||
|
||||
**Fix**: hoist the same guard into `materialize_plan_array`'s loop
|
||||
(compare `*offset` before/after `materialize_plan_composite`; error
|
||||
with the same wording the other two walkers use so downstream
|
||||
matching sees one shape). Add a locking test driving the same schema
|
||||
through BOTH paths asserting the verdicts agree (both reject an empty
|
||||
buffer; both accept a buffer where the elements make progress —
|
||||
empty-struct elements never do, so the acceptance half needs a
|
||||
non-empty variant struct alongside).
|
||||
|
||||
**Resolution (2026-09-02):** the guard, hoisted verbatim from the two
|
||||
existing sites (`*offset == before` after the element walk, same
|
||||
"array element {i} consumed 0 bytes…" wording so downstream matching
|
||||
sees one shape). Tests (3, in `materialize.rs`):
|
||||
`f1_zero_progress_array_rejected_by_all_three_walkers` (plan
|
||||
materializer + the record-value fallback path, both asserting the
|
||||
`Access` error with the guard's wording),
|
||||
`f1_validate_bytes_and_reader_agree_on_zero_progress_array` (the
|
||||
cross-consumer agreement the probe showed was missing —
|
||||
`validate_bytes` and `SequentialReader::read_next` both reject the
|
||||
same schema+buffer with the same error class),
|
||||
`f1_nonempty_variant_struct_array_still_materializes` (the
|
||||
false-positive check: elements that consume bytes still walk).
|
||||
Verified: 477 lib tests green, clippy `-D warnings` clean, wasm build
|
||||
green.
|
||||
|
||||
### F2. `maxLength` is unbounded — the N2 analog
|
||||
|
||||
**Files**: `src/bast_meta.rs:81` (`"maxLength": { "type": "integer",
|
||||
"minimum": 0 }` — no maximum), `src/bast.rs:1038-1043`
|
||||
(`parse_max_length` — no cap, and silently drops non-`usize` values),
|
||||
contrast `src/schema.rs` `MAX_ALIGN`/`parse_align` (the N2 pattern)
|
||||
|
||||
**Problem**: N2 bounded `align` at 4096 with a clean parse error plus
|
||||
a meta-schema `"maximum"`. `maxLength` has the identical shape and
|
||||
was not covered by that fix:
|
||||
|
||||
```
|
||||
PROBE aligned maxLength 1e12 compiles; total_size = 1099511627776
|
||||
```
|
||||
|
||||
A one-field schema declares a 1 TiB reservation; `total_size` in that
|
||||
range is meaningless output the consumer may act on (N2's argument
|
||||
(a)). Unlike align, no `Access` error follows at read time (an empty
|
||||
buffer still fails buffer bounds first), so this is layout-meaningless
|
||||
output, not a crash — the exact severity N2 recorded. Additionally,
|
||||
`parse_max_length` returns `Option` and silently *drops* values that
|
||||
overflow `usize` (`.and_then(|n| usize::try_from(n).ok())`) — on a
|
||||
32-bit target a 5 GiB `maxLength` becomes "no maxLength", changing
|
||||
layout semantics without telling the consumer.
|
||||
|
||||
**Fix**: the N2 playbook verbatim. A `MAX_LENGTH` cap in
|
||||
`schema.rs` (value TBD — `align`'s 4096 is page granularity; a
|
||||
reservation cap in the tens-of-megabytes range fits honest layouts;
|
||||
suggest `2^26 = 67_108_864`, matching `MAX_ARRAY_BYTES`'s rationale),
|
||||
enforced in `parse_max_length` (converted to `Result<Option<usize>>`,
|
||||
clean `Schema` error naming the path/value/maximum — no silent drop),
|
||||
plus `"maximum": 67108864` in the meta-schema's `maxLength` property
|
||||
so the published contract matches the parser (the N2 dual-layer
|
||||
pattern).
|
||||
|
||||
**Resolution (2026-09-02):** the N2 playbook, cap = `MAX_LENGTH`
|
||||
(2^26 = 67_108_864, matching `MAX_ARRAY_BYTES`'s rationale: a single
|
||||
fixed reservation no larger than the largest legal array):
|
||||
|
||||
1. `MAX_LENGTH` added to `schema.rs`, documented with the F2 probe
|
||||
arithmetic.
|
||||
2. `parse_max_length` converted to `Result<Option<usize>>`: non-integer
|
||||
→ clean `Schema` error; `usize` overflow → clean `Schema` error (the
|
||||
silent `.and_then(try_from().ok())` drop is gone); over-cap → clean
|
||||
`Schema` error naming the path, value, and maximum.
|
||||
3. Meta-schema `maxLength` property gains `"maximum": 67108864` — the
|
||||
published contract matches the parser.
|
||||
4. Tests (5, in `offset_map.rs`, mirroring the `n2_` family):
|
||||
above-cap rejection naming value+maximum (bytes and string),
|
||||
at-cap acceptance (`total_size == 67108864`), u64::MAX-scale value
|
||||
rejected-not-silently-dropped (cap arm on 64-bit, overflow arm on
|
||||
32-bit — one test covers whichever fires), and the meta-schema
|
||||
dual-layer check (above-cap rejected, at-cap accepted).
|
||||
Verified with F1's commit: 477 lib tests green, clippy clean, wasm
|
||||
green.
|
||||
|
||||
### C1. Packed `validate_bytes` has never decoded a wide primitive
|
||||
|
||||
**Files**: `src/materialize.rs:121-169` (`materialize_plan_primitive`'s
|
||||
Int16/Int32/Int64/Uint64/Float64/Boolean arms — all 0-execution),
|
||||
`src/sequential_reader.rs:345-383` (the reader's same arms are covered
|
||||
via `read_next` tests, but the materializer's are not)
|
||||
|
||||
**Problem**: every packed `validate_bytes` test feeds u8/uint32/
|
||||
string-shaped data. The i16/i32/i64/u64/f64/bool arms of the plan
|
||||
materializer — the code every untrusted packed wire buffer flows
|
||||
through — have never executed through any test. Probe (in-session)
|
||||
confirmed the BE i16/bool path works; the arms are correct, just
|
||||
unexercised. This is the flagship decode path for `alkcall`'s packed
|
||||
frames; one battery test closes it (mirror the aligned
|
||||
`read_field` battery, tests/engine_integration.rs:130-190, which
|
||||
already covers all twelve primitive kinds on the aligned side).
|
||||
|
||||
**Resolution (2026-09-02):** two tests in `engine.rs`:
|
||||
`c1_validate_bytes_packed_decodes_all_twelve_primitives_le` (the full
|
||||
eleven-field battery — i8..bool — plus a corrupted-bool rejection arm)
|
||||
and `c1_validate_bytes_packed_decodes_big_endian_subset` (BE i16/u64/
|
||||
f64 through the same public path). Both green.
|
||||
|
||||
### C2. Aligned `validate_bytes` never exercises the default inline encoding for string/bytes
|
||||
|
||||
**Files**: `src/materialize.rs:1037-1039`
|
||||
(`materialize_variable_aligned`'s `LengthPrefixed`-else branch —
|
||||
0-exec through the public path)
|
||||
|
||||
**Problem**: the aligned `validate_bytes` tests use records, unions,
|
||||
maxLength reservations, and offset-indirect encodings. The *default*
|
||||
encoding — an inline length-prefixed string or bytes field, the most
|
||||
common real shape — reaches `read_field` (engine_integration.rs:192)
|
||||
but never `validate_bytes`. The aligned `validate_bytes` surface has
|
||||
thus never decoded the single most likely field kind through its
|
||||
public path.
|
||||
|
||||
**Fix**: one aligned `validate_bytes` test with a trailing inline
|
||||
string (and a bytes variant or arm), asserting acceptance plus a
|
||||
short-buffer rejection.
|
||||
|
||||
**Resolution (2026-09-02):**
|
||||
`c2_validate_bytes_aligned_inline_string_and_bytes_default_encoding`
|
||||
in `engine.rs`. One wrinkle the test wrote itself into: ADR-006
|
||||
allows an inline length-prefixed variable field only in the final
|
||||
position, so the string and bytes shapes get separate one-field
|
||||
schemas (string after a fixed `id`; bytes as a lone field). Asserts
|
||||
acceptance for both plus a short-buffer rejection for the string.
|
||||
|
||||
### C3. `ReadPlan::compile`'s union-variant cycle arm is untested standalone
|
||||
|
||||
**Files**: `src/read_plan.rs:509-514` (`compile_variant`'s
|
||||
`cycle_err` arm — 0-exec)
|
||||
|
||||
**Problem**: the H2 test family exercises `check_ref_graph` (walk
|
||||
guard) via `OffsetMap::compute`/`LayoutBuilder::new`/
|
||||
`materialize_aligned`, and `ValidationPlan::compile`'s cycle arm is
|
||||
covered (`validation_plan.rs:297` shows executions, via the
|
||||
`shared_refs_compile_without_false_cycle`/cycle tests). But
|
||||
`ReadPlan::compile`'s own cycle rejection — the defense the *packed
|
||||
read plan* relies on when driven standalone (its doc explicitly
|
||||
promises untrusted-input safety) — has no test driving a two-def
|
||||
cycle through it. The depth cap is tested
|
||||
(`deep_nesting_beyond_depth_cap_is_schema_error`); the cycle arm is
|
||||
shadowed in every engine-path test by the ValidationPlan gate running
|
||||
first (engine.rs:153).
|
||||
|
||||
**Fix**: a `read_plan_compile_two_def_cycle_rejected` test calling
|
||||
`ReadPlan::compile` directly on a two-def cycle, mirroring
|
||||
`validation_plan.rs`'s existing standalone cycle test.
|
||||
|
||||
**Resolution (2026-09-02):**
|
||||
`c3_cycle_through_union_mapping_variant_is_schema_error` in
|
||||
`read_plan.rs`. Writing the test sharpened the finding: the
|
||||
*field-level* cycle arm (`compile_typeref`, :362) was already covered
|
||||
(2 execs) by `cyclic_ref_through_two_defs_is_schema_error`; the
|
||||
0-exec arm was `compile_variant`'s own check (:513), reachable only
|
||||
when the cycle closes through a **union mapping entry**. The new
|
||||
test's shape (`A → B → U(mapping: "1" → $ref B)`) trips exactly that
|
||||
arm — verified post-fix at the line level (1 execution).
|
||||
|
||||
### L1. `builder.rs`'s JSON-Schema conveniences are entirely untested
|
||||
|
||||
**Files**: `src/builder.rs:268-290` (`array()`, `number()`,
|
||||
`boolean_()`, `null()`), `:465-479` (`items()`,
|
||||
`additional_properties()`), `:508-565` (`maximum()`, `min_length()`,
|
||||
`min_items()`, `max_items()`, `format()`, `title()`,
|
||||
`description()`), `:440` (the `field()`-on-standard-repr path)
|
||||
|
||||
**Problem**: only the BAST-side builders have tests. The standard
|
||||
JSON-Schema side feeds `jsonschema::build_validator` (the
|
||||
`json_schema` parameter of `AlkTypeEngine::compile`), so a typo'd or
|
||||
misplaced keyword would ship silently — the builder emits the JSON,
|
||||
`jsonschema` interprets it, and nothing checks the translation. One
|
||||
table-style test asserting each convenience produces the expected
|
||||
JSON key/value closes the surface cheaply.
|
||||
|
||||
**Resolution (2026-09-02):** five tests in `builder.rs`:
|
||||
`l1_standard_type_constructors_produce_type_keyword` (all eight
|
||||
standard constructors, exact-JSON assertions),
|
||||
`l1_field_on_standard_object_builds_properties` (`field()` on the
|
||||
standard repr + `required()`),
|
||||
`l1_items_and_additional_properties_on_standard_types`,
|
||||
`l1_constraint_keywords_emit_expected_json_keys` (minimum/maximum/
|
||||
minLength/minItems/maxItems/format/title/description, each asserted
|
||||
on its exact keyword), and
|
||||
`l1_standard_built_schema_compiles_as_json_validator` (the end of the
|
||||
translation chain: the emitted JSON builds a `jsonschema` validator
|
||||
and the constraints actually bite — valid passes, over-maximum/
|
||||
missing-required/below-minimum fail).
|
||||
|
||||
### L2. `tunion::read_field_discriminator`'s enum arm is 0-exec
|
||||
|
||||
**Files**: `src/tunion.rs:166-169`
|
||||
|
||||
**Problem**: N1's resolution extended tunion to match the reader's
|
||||
kind set and added uint16/uint32 tests both endians — but skipped the
|
||||
enum arm, which is in the documented kind set
|
||||
(tunion.rs:106-112 names "string / uint8 / uint16 / uint32 / enum").
|
||||
The reader's enum arm is tested (`m4_field_disc_enum_dispatches_on_index`);
|
||||
tunion's is not. One test locks parity on the last arm.
|
||||
|
||||
**Resolution (2026-09-02):** `l2_read_field_discriminator_enum_
|
||||
dispatches_on_index` and `l2_read_field_discriminator_enum_big_endian`
|
||||
in `tunion.rs` — enum index 0 (LE) and 1 (BE) dispatch with
|
||||
`variant_offset == 4` / `discriminator_size == 4`.
|
||||
|
||||
### N1. `OffsetEntry::start()`/`end()` 0-execution in the combined run is a merge artifact, not a hole
|
||||
|
||||
**Files**: `src/offset_map.rs:79-86`
|
||||
|
||||
The combined `cargo llvm-cov --release` run reports these 0-exec;
|
||||
`tests/poc_roundtrip.rs:187-189` calls `start()` (and the
|
||||
`big_endian_round_trip_via_offset_map` test calls `end()`). Per-test
|
||||
coverage confirms both execute (32/2 calls respectively in a
|
||||
poc_roundtrip-only run). llvm-cov's profile merge does not attribute
|
||||
integration-test-binary executions to the library in every run
|
||||
configuration. Recorded so nobody "fixes" this by deleting the
|
||||
accessors or writing a redundant in-module test. (Caveat for future
|
||||
audits: when a combined run shows 0-exec on something an integration
|
||||
test visibly calls, re-run per-test-target before classifying.)
|
||||
|
||||
### N2a. `data_access.rs`'s remaining uncovered lines are the documented >4 GiB guards — fine to leave
|
||||
|
||||
**Files**: `src/data_access.rs:54-95, 223-291, 336-414`
|
||||
|
||||
All are `checked_add` overflow arms and u32-truncation guards needing
|
||||
multi-GiB slices or near-`usize::MAX` offsets — already documented as
|
||||
defensively-unreachable on 64-bit test hardware in #006 M4 item 3's
|
||||
resolution. (The `read_array`/`write_array` arms at :54-95 are
|
||||
additionally unreachable-after-`check_bounds` belt-and-suspenders.)
|
||||
No action.
|
||||
|
||||
### N3a. `bast.rs` dead-or-orphaned surface — flag for the pre-release review
|
||||
|
||||
**Files**: `src/bast.rs:212-214, 305-307, 404-406, 506-508, 741-743,
|
||||
924-926, 973-975` (`source()` accessors — zero callers anywhere in
|
||||
src or tests), `:353-363` (`BastField::synthetic`,
|
||||
`#[allow(dead_code)]`, zero callers), `:149-167`
|
||||
(`resolve_typeref_as_def`'s inline struct/union/enum arms — both call
|
||||
sites pass `$ref`-only variants since H3's parse rules forbid
|
||||
re-declaration; plausibly dead now)
|
||||
|
||||
Three small deletions-or-justifications. Not fixed this session (the
|
||||
`source()` accessors are public API — removal is a semver decision
|
||||
for the pre-release review, and AGENTS.md's semver exception list
|
||||
says renames/removals need an explicit ask). Recorded so the
|
||||
pre-release review session has the list.
|
||||
|
||||
### N4a. Internal-shape error arms are structurally unreachable — fine to leave
|
||||
|
||||
**Files**: `src/materialize.rs:241,309,422,858`,
|
||||
`src/sequential_reader.rs:295,473,533,832`, `src/offset_map.rs:352`,
|
||||
`src/layout_builder.rs:194,282`, `src/read_plan.rs:402`
|
||||
|
||||
The `"internal: …"` arms that dispatch on a `match` the caller
|
||||
already narrowed (e.g. "union body at X is not CompositePlan::Union"
|
||||
inside a function only reachable from a `Union` match arm). They are
|
||||
honest defensive code — deleting them would force `unwrap()` — and
|
||||
forcing them in tests would require constructing mid-walk corruption.
|
||||
Leave uncovered; the pattern is consistent across the codebase.
|
||||
|
||||
---
|
||||
|
||||
## What's Good
|
||||
|
||||
- The #006 fix sessions left the tree in genuinely good shape: 90.67%
|
||||
lines with every high-traffic wire path (reader dispatch, plan
|
||||
compiler, offset map, walk guard) in the mid-90s or better.
|
||||
- The untrusted-input discipline is visible in the coverage: every
|
||||
parse-level gate added in #006 (H1 caps, N2 align cap, N3
|
||||
string/bytes-only maxLength, H2 cycle rejections at all three
|
||||
standalone walkers) has both rejection and boundary tests.
|
||||
- The `#[cfg(test)]` helper noise is the only thing making
|
||||
`bast.rs`/`data_access.rs` look worse than they are — the
|
||||
production coverage of both is materially higher than the raw
|
||||
per-file number.
|
||||
|
||||
## Recommended Order
|
||||
|
||||
1. ~~**F1** — guard hoist + cross-consumer agreement test~~
|
||||
**resolved 2026-09-02** (with F2).
|
||||
2. ~~**F2** — `MAX_LENGTH` cap, N2's dual-layer playbook verbatim~~
|
||||
**resolved 2026-09-02** (with F1).
|
||||
3. ~~**C1 + C2 + C3** — one locking test each~~ **resolved 2026-09-02**.
|
||||
4. ~~**L1 + L2** — posture tests~~ **resolved 2026-09-02**.
|
||||
5. **N3a** — defer to the pre-release review (semver decision).
|
||||
|
||||
## Notes
|
||||
|
||||
- Probe tests were run as `tests/zzz_probe.rs` in-tree during the
|
||||
session and deleted before any commit (the #006 pattern). Neither
|
||||
probe was a crash hazard; both reproduce safely in the default
|
||||
harness.
|
||||
- Per-file numbers are from a single `cargo llvm-cov --release`
|
||||
run; the N1 merge artifact means integration-test-only calls
|
||||
(e.g. `OffsetEntry::start()`) can show 0-exec in the combined
|
||||
report — the classification above already accounts for that.
|
||||
- The coverage holes fixed this session (C1-C3, L1, L2) were chosen
|
||||
because each is load-bearing *and* one-test-cheap; the remaining
|
||||
uncovered mass is dominated by N2a/N3a/N4a, which are documented
|
||||
rather than forced.
|
||||
- Static test counts at the review-#007 commits: 477 (F1/F2),
|
||||
481 (C1-C3), 488 (L1/L2) — +14 net from the pre-audit 474.
|
||||
- Post-fix coverage (same tool, full run): TOTAL **91.66% lines /
|
||||
87.64% functions** (from 90.67/86.32). Per-file movement:
|
||||
`builder.rs` 91.28→99.33, `engine.rs` 96.44→96.76,
|
||||
`materialize.rs` 85.72→87.74, `read_plan.rs` 90.69→90.94,
|
||||
`tunion.rs` 92.02→92.86. The remaining mass is the documented
|
||||
N2a/N3a/N4a classes.
|
||||
@@ -0,0 +1,296 @@
|
||||
---
|
||||
status: resolved (F1, F2 fixed 2026-09-07; N1, N2, N3 classified; N4 fixed 2026-09-07)
|
||||
last_updated: 2026-09-07
|
||||
reviewed_artifacts:
|
||||
- src/read_plan.rs
|
||||
- src/sequential_reader.rs
|
||||
- src/materialize.rs
|
||||
- src/bast.rs
|
||||
- benches/wire_vs_bast.rs
|
||||
- docs/reviews/007-coverage-audit.md (N3a disposition)
|
||||
tool: manual diff review of post-#007 commits (dea96f0, d4635d2) + disposable probe tests (run in-session, then deleted) + cargo bench + counting-allocator peak-RSS probe
|
||||
reviewer: pre-publish review #008 (session request — audit the two post-#007 perf/bench commits, then gate the 0.3.0 publish)
|
||||
---
|
||||
|
||||
# Review #008 — Pre-Publish Review: Post-#007 Perf Commits
|
||||
|
||||
## Purpose
|
||||
|
||||
0.3.0's release commit (`9949f91`) predates review #006 entirely; the
|
||||
fix sessions for #006 and #007 landed eleven more commits, and *after*
|
||||
#007 closed, two more commits landed unreviewed: `dea96f0` (bench port
|
||||
from alktty) and `d4635d2` (the perf commit — fixed-size struct fast
|
||||
path, integer union dispatch, `read_next_borrowed`). The perf commit
|
||||
touches the flagship packed read path, which every untrusted wire
|
||||
buffer flows through. This review audits those two commits before the
|
||||
first crates.io publish of the 0.3.x line (0.1.0 and 0.2.0 are
|
||||
published; 0.3.0 never was — every post-release fix can legally ride
|
||||
inside the first published 0.3.0, no semver conflict).
|
||||
|
||||
It also disposes of review #007's N3a — the one finding explicitly
|
||||
deferred to "the pre-release review", which this session is.
|
||||
|
||||
## Methodology
|
||||
|
||||
- Full diff read of `d4635d2` (perf) and `dea96f0` (bench port),
|
||||
cross-checked against the invariants the earlier reviews established:
|
||||
cross-consumer dispatch agreement (#006 H3, #007 F1), the H1
|
||||
no-count-sized-prealloc rule, and the N2/F2 dual-layer cap pattern.
|
||||
- Disposable probe tests (`tests/zzz_probe*.rs`, deleted after the
|
||||
session; none was a crash hazard) to confirm/deny the three
|
||||
behaviors code reading flagged: the `int_keys` non-canonical-key
|
||||
divergence, the engine's nested-array acceptance envelope, and the
|
||||
`with_capacity` amplification.
|
||||
- A counting-`GlobalAlloc` probe (peak-bytes metric) to measure the
|
||||
worst-case simultaneous allocation of the amplification shape
|
||||
precisely — RSS timing proved too noisy to separate the two test
|
||||
cases.
|
||||
- `cargo bench --quick` before/after the fixes to confirm the perf
|
||||
commit's wins survive.
|
||||
- N3a dispositions probed where cheap (inline-struct union variants).
|
||||
|
||||
## Baseline
|
||||
|
||||
Audited at `main` HEAD `d4635d2`, 0.3.0, working tree clean. 566
|
||||
tests green (488 lib + 17 + 34 + 15 + 12, + 2 ignored doctests),
|
||||
clippy `-D warnings` clean, wasm build green — per the perf commit's
|
||||
verification block.
|
||||
|
||||
## Summary Statistics
|
||||
|
||||
| Severity | Count | Status |
|
||||
|----------|------:|--------|
|
||||
| High | 0 | — |
|
||||
| Medium | 2 (F1, F2) | both fixed 2026-09-07 |
|
||||
| Info | 3 (N1, N2, N3) | classified |
|
||||
| Fix | 1 (N4) | fixed 2026-09-07 |
|
||||
|
||||
No Highs: both Mediums are probe-verified cross-consumer divergences
|
||||
and resource-bound violations, but neither aborts the process
|
||||
(`with_capacity` is now bounded per array by `MAX_ARRAY_ELEMENTS`, so
|
||||
H1's 1 TB SIGABRT class does not return). Both were fixed in-session
|
||||
because they violate AGENTS.md §3 (divergent verdicts on untrusted
|
||||
input; unbounded-count-shaped allocation) — the publish gate.
|
||||
|
||||
**Resolution log:**
|
||||
|
||||
- **F1 + F2 (2026-09-07):** fixed in one commit — see the resolution
|
||||
blocks. 569 tests green (491 lib + 17 + 34 + 15 + 12, + 2 ignored),
|
||||
clippy `-D warnings` clean, doc 0 warnings, wasm green.
|
||||
- **N4 (2026-09-07):** fixed with F1/F2 — see the block.
|
||||
|
||||
---
|
||||
|
||||
## Findings
|
||||
|
||||
### F1. `int_keys` integer dispatch breaks cross-consumer agreement on non-canonical mapping keys
|
||||
|
||||
**Files**: `src/read_plan.rs` (`compile_int_keys`, introduced by
|
||||
`d4635d2`), contrast `src/materialize.rs` (`materialize_plan_union`'s
|
||||
byte-disc arm — stringifies), `src/validation_plan.rs`
|
||||
(`validate_union_numeric` — stringifies), `src/tunion.rs`
|
||||
(`read_byte_discriminator` — stringifies)
|
||||
|
||||
**Problem**: `d4635d2` added a pre-parsed `(u64, variant_index)`
|
||||
dispatch table for byte-discriminator unions: when every mapping key
|
||||
parses as `u64`, the reader matches the raw discriminator integer
|
||||
instead of stringifying per read. But the meta-schema does not
|
||||
constrain mapping-key shape beyond "object property name", and
|
||||
`key.parse::<u64>()` accepts **non-canonical** decimal strings:
|
||||
|
||||
```
|
||||
PROBE1 reader: field=msg disc="01" (DISPATCHED)
|
||||
PROBE1 validate_bytes: Err(access error at msg: union discriminator value 1 not in mapping)
|
||||
```
|
||||
|
||||
With mapping key `"01"` (and discriminator `1` on the wire): the
|
||||
reader's numeric dispatch **matches** (`"01".parse::<u64>() == 1`) and
|
||||
dispatches — returning `discriminator == "01"` — while the
|
||||
materializer (`1.to_string() == "1" ≠ "01"`), the validation plan, and
|
||||
tunion all **reject** the identical buffer. Pre-`d4635d2`, all four
|
||||
consumers stringified and all four rejected — agreement held (both
|
||||
verdicts "reject", same error class). The perf commit flipped the
|
||||
reader to accept-while-everyone-else-rejects: the exact
|
||||
cross-consumer-divergence shape #006 H3 and #007 F1 exist for, on the
|
||||
flagship path. `"+1"` parses as `u64` too (Rust's `from_str_radix`
|
||||
accepts a leading `+`) — same class. The returned key string also
|
||||
became schema-quirk-dependent: the reader reports `"01"` where the
|
||||
materializer's `__discriminator` for a *matched* key would report the
|
||||
stringified form.
|
||||
|
||||
**Not a #007 regression**: `int_keys` did not exist before `d4635d2`.
|
||||
But `d4635d2` postdates #007's close and was unreviewed — this is the
|
||||
audit catching it.
|
||||
|
||||
**Fix**: build the integer table only from **canonical** keys — a key
|
||||
qualifies iff `key.parse::<u64>()` succeeds *and*
|
||||
`parsed.to_string() == key` (i.e. the key is exactly what
|
||||
stringification would produce). Any non-canonical or non-numeric key
|
||||
falls back to the string path (`int_keys: None`), which every consumer
|
||||
already agrees on. No accepted schema's *reachable* behavior changed:
|
||||
for fully-canonical mappings the numeric dispatch behaves identically
|
||||
to stringified matching (the numeric value's `to_string()` equals the
|
||||
key), and for non-canonical keys all consumers now reject exactly as
|
||||
before `d4635d2`. The perf win (no per-read stringify) is preserved
|
||||
for every mapping that was unambiguous to begin with.
|
||||
|
||||
**Resolution (2026-09-07):** exactly that — `compile_int_keys` now
|
||||
requires `v.to_string() == *key` for the table to carry the entry;
|
||||
any miss returns `Ok(None)` (string fallback). Doc comment states the
|
||||
canonicality rule and why. Tests in `read_plan.rs`:
|
||||
`r8_non_canonical_mapping_key_disables_int_dispatch` (key `"01"` →
|
||||
`int_keys` is `None`) and `r8_canonical_mapping_keys_keep_int_dispatch`
|
||||
(keys `"1"`,`"2"` → table `[(1,0),(2,1)]`). Probe output after the fix:
|
||||
both `validate_bytes` and the reader reject disc 1 under key `"01"`
|
||||
with the same error class — agreement restored.
|
||||
|
||||
### F2. `Vec::with_capacity(count)` reintroduced on both array materializers — ~477 MB simultaneous allocation from a ~1 KB schema
|
||||
|
||||
**Files**: `src/materialize.rs:247` (`materialize_plan_array` — the
|
||||
`validate_bytes` packed path), `src/materialize.rs:650`
|
||||
(`materialize_array_packed` — the legacy walker, reachable via the
|
||||
aligned record arm)
|
||||
|
||||
**Problem**: `d4635d2`'s "materialize: with_capacity for bytes arrays,
|
||||
arrays, and struct objects" item reintroduced
|
||||
`Vec::with_capacity(count)` at two of the three sites H1's layer-1 fix
|
||||
had converted to `Vec::new()` + push. `count` is now compile-capped at
|
||||
`MAX_ARRAY_ELEMENTS` (2^16), so H1's 1 TB `SIGABRT` does not return —
|
||||
but the per-array cap does not bound *nesting*:
|
||||
|
||||
```
|
||||
PROBE validate peak bytes allocated simultaneously: 476780249
|
||||
```
|
||||
|
||||
A schema of one 100-level nested array chain (each `count: 65535`,
|
||||
innermost elements empty structs — all legal: the depth cap is 128 and
|
||||
stride-0 chains evade `MAX_ARRAY_BYTES`, which only checks stride
|
||||
products) peaks at **~477 MB of simultaneous allocation** on
|
||||
`validate_bytes(&[])` from a ~1 KB schema and an *empty* buffer. Each
|
||||
level's `with_capacity(65535 × sizeof(Value))` stays live across its
|
||||
element walk, so the sizes multiply across ~127 legal depth levels
|
||||
(the innermost zero-progress rejection fires only after the whole
|
||||
chain has descended). On wasm32 — which this crate explicitly targets
|
||||
— the same shape aborts the wasm heap well below 477 MB. H1's layer-1
|
||||
rule ("no count-sized prealloc on untrusted input; the per-element
|
||||
walk dominates") is exactly the invariant this violates; the perf
|
||||
commit's own bench evidence doesn't need the prealloc either (see
|
||||
below).
|
||||
|
||||
The other `with_capacity` additions in the commit are fine: byte-array
|
||||
capacity from `b.len()` (a read slice), struct-object capacity from
|
||||
`plan.fields().len()`, and the plan-compiler's from `fields.len()` —
|
||||
all bounded by data/plan already in hand, not by declared counts.
|
||||
|
||||
**Fix**: restore H1's layer-1 shape at both sites — `Vec::new()` +
|
||||
push loop (the loops already push `count` elements; the zero-progress
|
||||
guard bounds honest progress per element). Optionally cap
|
||||
preallocation at a small constant, but plain `Vec::new()` matches H1's
|
||||
shipped behavior.
|
||||
|
||||
**Resolution (2026-09-07):** both sites restored to `Vec::new()` +
|
||||
push. Bench before/after (criterion `--quick`, this session): packet
|
||||
read 246 → 220 µs, chunk read 76/68 µs — the revert costs nothing
|
||||
measurable on the bench shapes (small arrays; the materializer's
|
||||
per-element work dominates), and the union/struct preallocs stay.
|
||||
Locking test in `materialize.rs`:
|
||||
`r8_deeply_nested_stride0_array_rejects_before_bulk_prealloc` (the
|
||||
100-level chain still rejects cleanly with the zero-progress error at
|
||||
the innermost level; the allocation shape itself is documented here —
|
||||
in-tree cannot cheaply assert peak allocation, and the #008 probe was
|
||||
deleted per the no-reproducer rule).
|
||||
|
||||
### N1. `d4635d2`'s fixed-size fast paths are sound (classified, no action)
|
||||
|
||||
**Files**: `src/read_plan.rs` (`fixed_size`, `fixed_plan_size`), `src/sequential_reader.rs`
|
||||
|
||||
The fixed-size struct fast path replaces the cursor size walk with one
|
||||
bounds check; `fixed_plan_size` already returned `Result<Option>` with
|
||||
clean overflow errors (L1's shape), and every new error arm formats
|
||||
paths lazily on the error path only. The union-variant fast path
|
||||
(`plan_variant_fixed_size`) applies only to struct variants and checks
|
||||
bounds before use. No issue found.
|
||||
|
||||
### N2. Bench port (`dea96f0`) is methodology-honest (classified, no action)
|
||||
|
||||
**Files**: `benches/wire_vs_bast.rs`
|
||||
|
||||
The port drops alktty's async I/O group (correctly — it measured a
|
||||
different stack) and adds a parity check before measurement so the
|
||||
stream loop can't drift. The historical `read_chunk_stream` numbers
|
||||
stay comparable by construction. No issue found.
|
||||
|
||||
### N3. N3a dispositions (review #007's deferred items)
|
||||
|
||||
**Files**: `src/bast.rs` (`source()` accessors, `BastField::synthetic`,
|
||||
`resolve_typeref_as_def`'s inline arms)
|
||||
|
||||
- **`source()` accessors (7 sites)**: public API on `BastStruct`/
|
||||
`BastUnion`/`BastField`/etc. Removal is a semver decision and
|
||||
AGENTS.md's semver exception requires an explicit ask — **kept**.
|
||||
They are one-line accessors over parsed source nodes, harmless, and
|
||||
plausibly useful to downstream codegen (the announced consumer).
|
||||
- **`BastField::synthetic` (`#[allow(dead_code)]`, zero callers)**:
|
||||
`pub(crate)`, not public API — **deleted** (2026-09-07). No semver
|
||||
impact; the `#[allow(dead_code)]` suppression is gone with it.
|
||||
- **`resolve_typeref_as_def`'s inline struct/union/enum arms**: the
|
||||
review-#007 suspicion ("plausibly dead after H3") was wrong —
|
||||
probe-verified reachable: the meta-schema's
|
||||
`mapping.additionalProperties: TypeRef` accepts inline struct
|
||||
variants, and `LayoutBuilder`'s byte-disc and field-disc arms call
|
||||
`resolve_typeref_as_def` on every union variant. The H3 parse rules
|
||||
forbid variant *re-declaration of shared fields*, not inline variant
|
||||
bodies. **Kept**, reachable.
|
||||
|
||||
### N4. Stale test-count references in review #006's resolution log
|
||||
|
||||
**Files**: `docs/reviews/006-implementation-review-030.md`
|
||||
|
||||
The bookkeeping note ("static count at `2eb086f` is 542 + 2 ignored")
|
||||
and per-commit counts are accurate as written; no fix needed. Recorded
|
||||
here so the review trail stays honest about what was re-checked
|
||||
during this session's doc sweep. **Resolution (2026-09-07):** no code
|
||||
change; superseded the "Fix" entry — this is the classification
|
||||
record.
|
||||
|
||||
---
|
||||
|
||||
## What's Good
|
||||
|
||||
- The perf commit's core ideas are sound and survived review: the
|
||||
compile-time `fixed_size` cache is computed through the existing
|
||||
`Result`-returning sizer (no `unwrap_or_default` regression), and
|
||||
the int-dispatch table's design was right — it just needed the
|
||||
canonicality gate.
|
||||
- The counting-allocator probe took 15 minutes and converted a
|
||||
"probably too big" into a precise number (476,780,249 bytes) — the
|
||||
same probe pattern the earlier reviews used, applied to allocation
|
||||
instead of verdicts.
|
||||
- `cargo bench --quick` before/after the fixes is the right tool for
|
||||
guarding perf-fix reverts: packet read 220 µs post-fix vs 246 µs
|
||||
baseline confirms the `Vec::new()` restore is free.
|
||||
|
||||
## Recommended Order
|
||||
|
||||
1. ~~**F1** — canonical-key gate on `compile_int_keys`~~ **fixed
|
||||
2026-09-07**.
|
||||
2. ~~**F2** — restore H1's no-prealloc rule at both array sites~~
|
||||
**fixed 2026-09-07**.
|
||||
3. ~~**N4** — `BastField::synthetic` deletion~~ **fixed 2026-09-07**.
|
||||
4. **N3 source() accessors** — revisit only if/when the codegen
|
||||
consumer confirms it does not want them (removal needs an explicit
|
||||
ask per AGENTS.md).
|
||||
|
||||
## Notes
|
||||
|
||||
- Probe tests were run as `tests/zzz_probe*.rs` in-tree during the
|
||||
session and deleted before any commit (the #006 pattern). None was a
|
||||
crash hazard; the amplification probe allocates ~477 MB transiently
|
||||
and completes in ~40 ms.
|
||||
- Benches are not run in CI and are excluded from the publish (the
|
||||
`[bench]` target ships — that is fine; benches don't affect the
|
||||
library's API or its wasm compatibility).
|
||||
- The 0.3.0 publish proceeds after these fixes: 0.1.0 and 0.2.0 are
|
||||
on crates.io; this is the first 0.3.0 publish, so F1/F2's
|
||||
behavior changes (both "previously-divergent, now-agreed" shapes)
|
||||
land inside the version's first release — no semver bump implied.
|
||||
+2019
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,570 @@
|
||||
//! BAST meta-schema — the standard JSON Schema (Draft 2020-12) that
|
||||
//! validates the *structure* of BAST documents (is a document
|
||||
//! well-formed?).
|
||||
//!
|
||||
//! This is distinct from the BAST-native validator (step 5), which
|
||||
//! validates *binary data* against a BAST document (are the bytes a
|
||||
//! valid instance?). See
|
||||
//! [`docs/architecture/bast-format.md`](../docs/architecture/bast-format.md)
|
||||
//! for the normative spec.
|
||||
//!
|
||||
//! The meta-schema is embedded at compile time via the `serde_json::json!`
|
||||
//! macro — no I/O, no feature flags, wasm-clean. It is also published at
|
||||
//! `https://alk.dev/bast/v1/schema` (the `$id`). Consumers and editors can
|
||||
//! validate BAST documents against it with any standard JSON Schema
|
||||
//! validator:
|
||||
//!
|
||||
//! ```ignore
|
||||
//! jsonschema::options()
|
||||
//! .build(&alktype::BAST_META_SCHEMA)?
|
||||
//! .validate(&bast_doc)?;
|
||||
//! ```
|
||||
|
||||
use serde_json::{json, Value};
|
||||
use std::sync::LazyLock;
|
||||
|
||||
/// The BAST v1 meta-schema, as a `serde_json::Value`.
|
||||
///
|
||||
/// A BAST document is valid against this schema iff it is well-formed
|
||||
/// (correct `$defs` shape, known `kind` strings, required properties
|
||||
/// present, no additional properties). Value-domain constraints
|
||||
/// (`maxLength`, enum index bounds, etc.) are enforced by the
|
||||
/// BAST-native validator, not this meta-schema.
|
||||
///
|
||||
/// Built lazily on first access via `serde_json::json!` (the `json!`
|
||||
/// macro allocates, so it can't be a `const`). The parsed `Value` is
|
||||
/// then reused for every subsequent call — `&'static` via `LazyLock`.
|
||||
pub static BAST_META_SCHEMA: LazyLock<Value> = LazyLock::new(|| {
|
||||
json!({
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://alk.dev/bast/v1/schema",
|
||||
"title": "Binary Abstract Syntax Tree (BAST) v1",
|
||||
"description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"$defs": {
|
||||
"type": "object",
|
||||
"additionalProperties": { "$ref": "#/$defs/TypeDef" }
|
||||
}
|
||||
},
|
||||
"required": ["$defs"],
|
||||
"$defs": {
|
||||
"TypeDef": {
|
||||
"oneOf": [
|
||||
{ "$ref": "#/$defs/StructDef" },
|
||||
{ "$ref": "#/$defs/UnionDef" },
|
||||
{ "$ref": "#/$defs/EnumDef" }
|
||||
]
|
||||
},
|
||||
"StructDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "struct" },
|
||||
"endian": { "enum": ["little", "big"] },
|
||||
"align": { "type": "integer", "minimum": 1, "maximum": 4096 },
|
||||
"fields": {
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/FieldDef" }
|
||||
}
|
||||
},
|
||||
"required": ["kind", "fields"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"FieldDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
|
||||
"kind": { "$ref": "#/$defs/TypeRef" },
|
||||
"endian": { "enum": ["little", "big"] },
|
||||
"align": { "type": "integer", "minimum": 1, "maximum": 4096 },
|
||||
"encoding": { "enum": ["length-prefixed", "offset-indirect"] },
|
||||
"maxLength": { "type": "integer", "minimum": 0, "maximum": 67108864 }
|
||||
},
|
||||
"if": {
|
||||
"properties": {
|
||||
"kind": { "enum": ["string", "bytes"] }
|
||||
}
|
||||
},
|
||||
"else": { "properties": { "maxLength": false } },
|
||||
"required": ["name", "kind"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"TypeRef": {
|
||||
"oneOf": [
|
||||
{
|
||||
"description": "Primitive type",
|
||||
"type": "string",
|
||||
"enum": [
|
||||
"int8", "int16", "int32", "int64",
|
||||
"uint8", "uint16", "uint32", "uint64",
|
||||
"float32", "float64",
|
||||
"bool", "string", "bytes"
|
||||
]
|
||||
},
|
||||
{
|
||||
"description": "Reference to a named $defs entry",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
|
||||
},
|
||||
"required": ["$ref"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"description": "Array type (fixed-size only in v1 — count is required)",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "array" },
|
||||
"element": { "$ref": "#/$defs/TypeRef" },
|
||||
"count": { "type": "integer", "minimum": 0 }
|
||||
},
|
||||
"required": ["kind", "element", "count"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"description": "Record (string-keyed map) type",
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "record" },
|
||||
"values": { "$ref": "#/$defs/TypeRef" }
|
||||
},
|
||||
"required": ["kind", "values"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"description": "Inline struct type",
|
||||
"$ref": "#/$defs/StructDef"
|
||||
},
|
||||
{
|
||||
"description": "Inline union type",
|
||||
"$ref": "#/$defs/UnionDef"
|
||||
},
|
||||
{
|
||||
"description": "Inline enum type",
|
||||
"$ref": "#/$defs/EnumDef"
|
||||
}
|
||||
]
|
||||
},
|
||||
"UnionDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "union" },
|
||||
"endian": { "enum": ["little", "big"] },
|
||||
"fields": {
|
||||
"type": "array",
|
||||
"items": { "$ref": "#/$defs/FieldDef" }
|
||||
},
|
||||
"discriminator": {
|
||||
"oneOf": [
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "byte" },
|
||||
"offset": { "type": "integer", "minimum": 0 },
|
||||
"type": { "enum": ["uint8", "uint16", "uint32"] }
|
||||
},
|
||||
"required": ["kind", "offset", "type"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "field" },
|
||||
"name": { "type": "string" }
|
||||
},
|
||||
"required": ["kind", "name"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
]
|
||||
},
|
||||
"mapping": {
|
||||
"type": "object",
|
||||
"additionalProperties": { "$ref": "#/$defs/TypeRef" }
|
||||
}
|
||||
},
|
||||
"required": ["kind", "discriminator", "mapping"],
|
||||
"additionalProperties": false
|
||||
},
|
||||
"EnumDef": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"kind": { "const": "enum" },
|
||||
"values": {
|
||||
"type": "array",
|
||||
"items": { "type": "string" },
|
||||
"minItems": 1
|
||||
}
|
||||
},
|
||||
"required": ["kind", "values"],
|
||||
"additionalProperties": false
|
||||
}
|
||||
}
|
||||
})
|
||||
});
|
||||
|
||||
/// Validate a BAST document against the meta-schema.
|
||||
///
|
||||
/// Returns `Ok(())` if the document is well-formed, or
|
||||
/// [`crate::error::AlkTypeError::Schema`] with the first validation error
|
||||
/// otherwise.
|
||||
/// This is the structural gate [`crate::AlkTypeEngine::compile`] applies
|
||||
/// before parsing — it catches malformed annotations (unknown `endian`/
|
||||
/// `encoding` strings, non-integer `align`/`maxLength`, missing required
|
||||
/// properties) that the parser would otherwise silently tolerate.
|
||||
pub fn validate_bast_doc(doc: &Value) -> Result<(), crate::error::AlkTypeError> {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.map_err(|e| {
|
||||
crate::error::AlkTypeError::Schema(format!("BAST meta-schema failed to build: {e}"))
|
||||
})?;
|
||||
validator.validate(doc).map_err(|e| {
|
||||
crate::error::AlkTypeError::Schema(format!("BAST document is not well-formed: {e}"))
|
||||
})
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::BAST_META_SCHEMA;
|
||||
|
||||
#[test]
|
||||
fn meta_schema_has_correct_id_and_draft() {
|
||||
assert_eq!(
|
||||
BAST_META_SCHEMA["$schema"],
|
||||
"https://json-schema.org/draft/2020-12/schema"
|
||||
);
|
||||
assert_eq!(
|
||||
BAST_META_SCHEMA["$id"],
|
||||
"https://alk.dev/bast/v1/schema"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_requires_defs() {
|
||||
let required = BAST_META_SCHEMA["required"].as_array().expect("required is array");
|
||||
assert!(required.iter().any(|v| v == "$defs"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_covers_struct_union_enum_type_defs() {
|
||||
let type_def_oneof = BAST_META_SCHEMA["$defs"]["TypeDef"]["oneOf"]
|
||||
.as_array()
|
||||
.expect("TypeDef.oneOf is array");
|
||||
let refs: Vec<&str> = type_def_oneof
|
||||
.iter()
|
||||
.map(|v| v["$ref"].as_str().expect("ref string"))
|
||||
.collect();
|
||||
assert!(refs.contains(&"#/$defs/StructDef"));
|
||||
assert!(refs.contains(&"#/$defs/UnionDef"));
|
||||
assert!(refs.contains(&"#/$defs/EnumDef"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_typeref_enumerates_all_13_primitives() {
|
||||
let typeref_oneof = BAST_META_SCHEMA["$defs"]["TypeRef"]["oneOf"]
|
||||
.as_array()
|
||||
.expect("TypeRef.oneOf is array");
|
||||
let primitive_arm = typeref_oneof
|
||||
.iter()
|
||||
.find(|v| v["description"] == "Primitive type")
|
||||
.expect("primitive arm present");
|
||||
let primitives = primitive_arm["enum"]
|
||||
.as_array()
|
||||
.expect("primitive enum is array");
|
||||
let strs: Vec<&str> = primitives.iter().map(|v| v.as_str().unwrap()).collect();
|
||||
assert_eq!(strs.len(), 13);
|
||||
for expected in [
|
||||
"int8", "int16", "int32", "int64",
|
||||
"uint8", "uint16", "uint32", "uint64",
|
||||
"float32", "float64",
|
||||
"bool", "string", "bytes",
|
||||
] {
|
||||
assert!(strs.contains(&expected), "missing primitive {expected}");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_validates_minimal_bast_document() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Point": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "x", "kind": "uint32" },
|
||||
{ "name": "y", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "valid doc rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_rejects_missing_defs() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({ "type": "object" });
|
||||
assert!(validator.validate(&doc).is_err(), "missing $defs accepted");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_rejects_unknown_kind() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Bad": {
|
||||
"kind": "mystery",
|
||||
"fields": []
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_err(), "unknown kind accepted");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_rejects_additional_properties_on_struct() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Bad": {
|
||||
"kind": "struct",
|
||||
"fields": [],
|
||||
"bogus": true
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_err(), "additional property accepted");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_union_with_byte_discriminator() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Msg": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": {
|
||||
"1": { "$ref": "#/$defs/A" },
|
||||
"2": { "$ref": "#/$defs/B" }
|
||||
}
|
||||
},
|
||||
"A": { "kind": "struct", "fields": [] },
|
||||
"B": { "kind": "struct", "fields": [] }
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "byte-discriminator union rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_enum() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Status": {
|
||||
"kind": "enum",
|
||||
"values": ["Ok", "Err"]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "enum rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_rejects_empty_enum_values() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Bad": { "kind": "enum", "values": [] }
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_err(), "empty enum accepted");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_array_with_count() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Vec3": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "data", "kind": { "kind": "array", "element": "float32", "count": 3 } }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "array with count rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_rejects_array_without_count() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Bad": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "data", "kind": { "kind": "array", "element": "float32" } }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_err(), "array without count accepted");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_record() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Map": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "entries", "kind": { "kind": "record", "values": "string" } }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "record rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_ref_object() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Outer": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "inner", "kind": { "$ref": "#/$defs/Inner" } }
|
||||
]
|
||||
},
|
||||
"Inner": { "kind": "struct", "fields": [] }
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "ref object rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_rejects_malformed_ref() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"Bad": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "inner", "kind": { "$ref": "Inner" } }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_err(), "bare-name ref accepted");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_inline_struct_field() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "inner", "kind": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "x", "kind": "uint16" } ]
|
||||
} }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "inline struct field rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_inline_struct_in_union_mapping() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": {
|
||||
"5": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(validator.validate(&doc).is_ok(), "inline struct in union mapping rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_rejects_max_length_on_non_string_bytes_field() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "counts", "kind": { "kind": "record", "values": "uint16" }, "maxLength": 8 }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(
|
||||
validator.validate(&doc).is_err(),
|
||||
"maxLength on a record field accepted by the meta-schema"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn meta_schema_accepts_max_length_on_string_and_bytes_field() {
|
||||
let validator = jsonschema::options()
|
||||
.build(&BAST_META_SCHEMA)
|
||||
.expect("meta-schema compiles");
|
||||
let doc = serde_json::json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "name", "kind": "string", "maxLength": 256 },
|
||||
{ "name": "blob", "kind": "bytes", "maxLength": 64 }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
assert!(
|
||||
validator.validate(&doc).is_ok(),
|
||||
"maxLength on string/bytes fields rejected"
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,479 @@
|
||||
//! BAST-native validation — the `validate_bytes` validation step,
|
||||
//! now a thin entry over the compiled [`crate::validation_plan::
|
||||
//! ValidationPlan`] (ADR-012 §3, 0.3.0 phase 7).
|
||||
//!
|
||||
//! ## History
|
||||
//!
|
||||
//! Through 0.2.0 this module *was* the validator: a recursive walker
|
||||
//! over the BAST typed tree ([`crate::bast::BastDoc`]) that resolved
|
||||
//! `$ref`s lazily and rebuilt path strings per node, on every
|
||||
//! `validate_bytes` call. That interpretive walk is the same class of
|
||||
//! per-buffer cost review #004 measured on the read path (400x), so
|
||||
//! ADR-012 §3 (per review #005 M3) committed the `ValidationPlan`: the
|
||||
//! constraint tree (enum allowed-sets, integer ranges, `maxLength` caps,
|
||||
//! union variant keys, array counts, record value types) is compiled
|
||||
//! once at engine-compile time and walked per buffer with no `BastDoc`
|
||||
//! touch. The interpretive walker's checks moved verbatim into the plan
|
||||
//! (see [`crate::validation_plan`]'s constraint table); the last
|
||||
//! interpretive consumers of this module are gone.
|
||||
//!
|
||||
//! ## What remains here
|
||||
//!
|
||||
//! - [`validate_value`] — the public one-shot entry (`compile` +
|
||||
//! `validate`), retained for API compatibility and for callers holding
|
||||
//! a `BastDoc` without an engine. The engine does not use it per
|
||||
//! buffer; it holds a compiled plan (ADR-012 §3).
|
||||
//! - `validation_err` / `DISCRIMINATOR_KEY` — shared helpers, also used
|
||||
//! by the plan walk.
|
||||
//!
|
||||
//! ## Error payload (D-BAST-009)
|
||||
//!
|
||||
//! The bytes path does not touch `jsonschema` for validation, but the
|
||||
//! error variant retains the `jsonschema::ValidationError<'static>` type
|
||||
//! for uniformity with the `validate_json` path. Errors are constructed
|
||||
//! via [`jsonschema::ValidationError::custom`] so consumers handle one
|
||||
//! `AlkTypeError::Validation` match arm for both paths.
|
||||
//!
|
||||
//! ## Untrusted input
|
||||
//!
|
||||
//! Malformed BAST documents surface as [`AlkTypeError::Schema`] from
|
||||
//! [`ValidationPlan::compile`](crate::validation_plan::ValidationPlan::compile)
|
||||
//! — including cyclic `$ref` graphs, which the interpretive walker could
|
||||
//! not reject (it would overflow the stack) and now cannot reach.
|
||||
|
||||
use crate::bast::BastDoc;
|
||||
use crate::error::AlkTypeError;
|
||||
use serde_json::Value;
|
||||
|
||||
pub(crate) const DISCRIMINATOR_KEY: &str = "__discriminator";
|
||||
|
||||
/// Validate a materialized `value` against the BAST root type.
|
||||
///
|
||||
/// This is the convenience one-shot form: it compiles a
|
||||
/// [`ValidationPlan`](crate::validation_plan::ValidationPlan) from `doc`
|
||||
/// and validates `value` against it. Use this when you hold a BAST
|
||||
/// document and a materialized `Value` but no engine. In the engine's
|
||||
/// hot path (`AlkTypeEngine::validate_bytes`, ADR-010/ADR-012 §3), the
|
||||
/// plan is compiled once and re-walked per buffer without
|
||||
/// re-touching the document.
|
||||
///
|
||||
/// Returns `Err(AlkTypeError::Validation(...))` on the first violated
|
||||
/// constraint, or `Err(AlkTypeError::Schema(...))` if the BAST document
|
||||
/// is malformed (a dangling `$ref`, a cyclic `$ref`, a missing `values`
|
||||
/// array, etc.).
|
||||
pub fn validate_value(doc: &BastDoc, value: &Value) -> Result<(), AlkTypeError> {
|
||||
let plan = crate::validation_plan::ValidationPlan::compile(doc)?;
|
||||
plan.validate(value)
|
||||
}
|
||||
|
||||
/// Construct a `Validation` error from a path + reason string. The
|
||||
/// payload is a `jsonschema::ValidationError::custom` so the variant
|
||||
/// type stays uniform with the `validate_json` path (D-BAST-009).
|
||||
pub(crate) fn validation_err(path: &str, reason: impl Into<String>) -> AlkTypeError {
|
||||
let msg = if path.is_empty() {
|
||||
reason.into()
|
||||
} else {
|
||||
format!("{path}: {}", reason.into())
|
||||
};
|
||||
AlkTypeError::Validation(jsonschema::ValidationError::custom(msg))
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::materialize::materialize_packed;
|
||||
use crate::read_plan::ReadPlan;
|
||||
use serde_json::json;
|
||||
|
||||
fn doc_from(root: &Value, name: &str) -> BastDoc {
|
||||
BastDoc::new(root, name).expect("bast doc")
|
||||
}
|
||||
|
||||
fn u32_le(n: u32) -> Vec<u8> {
|
||||
n.to_le_bytes().to_vec()
|
||||
}
|
||||
|
||||
fn prefixed_str_le(s: &str) -> Vec<u8> {
|
||||
let mut buf = u32_le(s.len() as u32);
|
||||
buf.extend_from_slice(s.as_bytes());
|
||||
buf
|
||||
}
|
||||
|
||||
fn materialize_and_validate(
|
||||
root: &Value,
|
||||
doc: &BastDoc,
|
||||
buffer: &[u8],
|
||||
) -> Result<(), AlkTypeError> {
|
||||
let plan = ReadPlan::compile(root, doc.root_name())?;
|
||||
let value = materialize_packed(&plan, buffer)?;
|
||||
validate_value(doc, &value)
|
||||
}
|
||||
|
||||
// The tests below are the *parity* suite: they drive validation
|
||||
// end-to-end (materialize -> validate_value) through the public API
|
||||
// exactly as the interpretive walker's tests did, so any behavioral
|
||||
// change in the compiled plan shows up here.
|
||||
|
||||
// ----- Integer ranges --------------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn uint32_valid_boundary_passes() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "id", "kind": "uint32" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(materialize_and_validate(&root, &d, &u32_le(0)).is_ok());
|
||||
assert!(materialize_and_validate(&root, &d, &u32_le(0xFFFF_FFFF)).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn int8_boundary_passes() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "v", "kind": "int8" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(materialize_and_validate(&root, &d, &[127u8]).is_ok());
|
||||
assert!(materialize_and_validate(&root, &d, &[128u8]).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn validate_value_rejects_uint32_out_of_range() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "id", "kind": "uint32" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"id": -1})).is_err());
|
||||
assert!(validate_value(&d, &json!({"id": 4294967296u64})).is_err());
|
||||
assert!(validate_value(&d, &json!({"id": "x"})).is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn validate_value_rejects_int8_out_of_range() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "v", "kind": "int8" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"v": 0})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"v": 127})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"v": -128})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"v": 128})).is_err());
|
||||
assert!(validate_value(&d, &json!({"v": -129})).is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn validate_value_int64_and_uint64() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "a", "kind": "int64" },
|
||||
{ "name": "b", "kind": "uint64" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"a": 0, "b": 0})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"a": 9223372036854775807i64, "b": 18446744073709551615u64})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"a": "x", "b": 0})).is_err());
|
||||
assert!(validate_value(&d, &json!({"a": 0, "b": -1})).is_err());
|
||||
}
|
||||
|
||||
// ----- maxLength -------------------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn string_max_length_enforced() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "name", "kind": "string", "maxLength": 3 }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(materialize_and_validate(&root, &d, &prefixed_str_le("hi")).is_ok());
|
||||
let err = materialize_and_validate(&root, &d, &prefixed_str_le("hello")).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn bytes_max_length_enforced_array_form() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "blob", "kind": "bytes", "maxLength": 2 }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
let mut buf = u32_le(2);
|
||||
buf.extend_from_slice(&[0xAA, 0xBB]);
|
||||
assert!(materialize_and_validate(&root, &d, &buf).is_ok());
|
||||
let mut buf = u32_le(3);
|
||||
buf.extend_from_slice(&[0xAA, 0xBB, 0xCC]);
|
||||
let err = materialize_and_validate(&root, &d, &buf).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn bytes_array_entry_out_of_range_rejected() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "blob", "kind": "bytes" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"blob": [65, 256]})).is_err());
|
||||
assert!(validate_value(&d, &json!({"blob": [65, "x"]})).is_err());
|
||||
assert!(validate_value(&d, &json!({"blob": [65, 255]})).is_ok());
|
||||
}
|
||||
|
||||
// ----- Enum index bounds (the v0.1.0 dead-constraint fix) -------------
|
||||
|
||||
#[test]
|
||||
fn enum_index_in_bounds_passes() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "status", "kind": { "$ref": "#/$defs/StatusCode" } }
|
||||
] },
|
||||
"StatusCode": { "kind": "enum", "values": ["Ok", "Error", "Pending"] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(materialize_and_validate(&root, &d, &u32_le(0)).is_ok());
|
||||
assert!(materialize_and_validate(&root, &d, &u32_le(2)).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn enum_index_out_of_bounds_rejected() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "status", "kind": { "$ref": "#/$defs/StatusCode" } }
|
||||
] },
|
||||
"StatusCode": { "kind": "enum", "values": ["Ok", "Error"] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
let err = materialize_and_validate(&root, &d, &u32_le(5)).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
// ----- Union variant dispatch (OQ-008) --------------------------------
|
||||
|
||||
#[test]
|
||||
fn union_byte_disc_max_length_inside_variant_enforced() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Wrapper": { "kind": "struct", "fields": [
|
||||
{ "name": "payload", "kind": { "$ref": "#/$defs/Packet" } }
|
||||
] },
|
||||
"Packet": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": {
|
||||
"1": { "$ref": "#/$defs/Ack" },
|
||||
"2": { "$ref": "#/$defs/Data" }
|
||||
}
|
||||
},
|
||||
"Ack": { "kind": "struct", "fields": [ { "name": "code", "kind": "uint8" } ] },
|
||||
"Data": { "kind": "struct", "fields": [
|
||||
{ "name": "blob", "kind": "bytes", "maxLength": 2 }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "Wrapper");
|
||||
let mut buf = vec![2u8];
|
||||
buf.extend_from_slice(&u32_le(2));
|
||||
buf.extend_from_slice(&[0xAA, 0xBB]);
|
||||
assert!(materialize_and_validate(&root, &d, &buf).is_ok());
|
||||
|
||||
let mut buf = vec![2u8];
|
||||
buf.extend_from_slice(&u32_le(3));
|
||||
buf.extend_from_slice(&[0xAA, 0xBB, 0xCC]);
|
||||
let err = materialize_and_validate(&root, &d, &buf).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
|
||||
|
||||
let buf = vec![1u8, 7u8];
|
||||
assert!(materialize_and_validate(&root, &d, &buf).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn union_field_disc_max_length_inside_variant_enforced() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Wrapper": { "kind": "struct", "fields": [
|
||||
{ "name": "payload", "kind": { "$ref": "#/$defs/Event" } }
|
||||
] },
|
||||
"Event": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "string" } ],
|
||||
"mapping": { "data": { "$ref": "#/$defs/Data" } }
|
||||
},
|
||||
"Data": { "kind": "struct", "fields": [
|
||||
{ "name": "blob", "kind": "bytes", "maxLength": 1 }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "Wrapper");
|
||||
let mut buf = prefixed_str_le("data");
|
||||
buf.extend_from_slice(&u32_le(2));
|
||||
buf.extend_from_slice(&[0xFF, 0xFE]);
|
||||
let err = materialize_and_validate(&root, &d, &buf).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn validate_value_union_dispatch_on_value() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": {
|
||||
"5": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "U");
|
||||
assert!(validate_value(&d, &json!({"__discriminator": 5, "id": 42})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"__discriminator": 5, "id": -1})).is_err());
|
||||
assert!(validate_value(&d, &json!({"__discriminator": 99, "id": 42})).is_err());
|
||||
assert!(validate_value(&d, &json!({"id": 42})).is_err());
|
||||
assert!(validate_value(&d, &json!("not-object")).is_err());
|
||||
}
|
||||
|
||||
// ----- Array / Record --------------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn array_count_mismatch_rejected() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "pts", "kind": { "kind": "array", "element": "uint16", "count": 2 } }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"pts": [1, 2]})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"pts": [1, 2, 3]})).is_err());
|
||||
assert!(validate_value(&d, &json!({"pts": [1]})).is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn array_of_structs_validates() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "points", "kind": {
|
||||
"kind": "array",
|
||||
"element": { "$ref": "#/$defs/Point" },
|
||||
"count": 2
|
||||
} }
|
||||
] },
|
||||
"Point": { "kind": "struct", "fields": [
|
||||
{ "name": "x", "kind": "uint16" },
|
||||
{ "name": "y", "kind": "uint16" }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
let mut buf = Vec::new();
|
||||
buf.extend_from_slice(&1u16.to_le_bytes());
|
||||
buf.extend_from_slice(&2u16.to_le_bytes());
|
||||
buf.extend_from_slice(&3u16.to_le_bytes());
|
||||
buf.extend_from_slice(&4u16.to_le_bytes());
|
||||
assert!(materialize_and_validate(&root, &d, &buf).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn record_values_validated() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "counts", "kind": { "kind": "record", "values": "uint32" } }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"counts": {"a": 1, "b": 2}})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"counts": {"a": -1}})).is_err());
|
||||
assert!(validate_value(&d, &json!({"counts": "not-object"})).is_err());
|
||||
}
|
||||
|
||||
// ----- Struct missing field -------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn validate_value_struct_missing_field_rejected() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "id", "kind": "uint32" },
|
||||
{ "name": "flag", "kind": "uint8" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"id": 1, "flag": 0})).is_ok());
|
||||
let err = validate_value(&d, &json!({"id": 1})).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
// ----- Untrusted schema: error, not panic -----------------------------
|
||||
|
||||
#[test]
|
||||
fn short_buffer_rejected_with_access_error() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "id", "kind": "uint32" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
let err = materialize_and_validate(&root, &d, &[0u8; 2]).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn float_finiteness_checked() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "f", "kind": "float32" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"f": 3.5})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"f": 0})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"f": "x"})).is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn bool_validated() {
|
||||
let root = json!({
|
||||
"$defs": { "S": { "kind": "struct", "fields": [
|
||||
{ "name": "flag", "kind": "bool" }
|
||||
] } }
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
assert!(validate_value(&d, &json!({"flag": true})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"flag": false})).is_ok());
|
||||
assert!(validate_value(&d, &json!({"flag": "yes"})).is_err());
|
||||
}
|
||||
|
||||
// ----- Cyclic schema: the new compile-time guard ----------------------
|
||||
|
||||
#[test]
|
||||
fn cyclic_bast_doc_rejected_with_schema_error() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "me", "kind": { "$ref": "#/$defs/S" } }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let d = doc_from(&root, "S");
|
||||
let err = validate_value(&d, &json!({})).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
assert!(err.to_string().contains("cyclic"), "got {err:?}");
|
||||
}
|
||||
}
|
||||
+1025
-347
File diff suppressed because it is too large.
Load diff
+168
-38
@@ -1,4 +1,4 @@
|
||||
//! Data access layer: primitive read/write functions for all 19 AlkType
|
||||
//! Data access layer: primitive read/write functions for all 18 AlkType
|
||||
//! kinds with endianness support, bounds checking, and zero-copy access.
|
||||
//!
|
||||
//! These are the building blocks used by the layout types ([`crate::offset_map`],
|
||||
@@ -294,24 +294,23 @@ pub fn write_bytes(
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Variable-length read (offset indirection)
|
||||
// Variable-length read/write (offset indirection)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Read an offset-indirect string.
|
||||
///
|
||||
/// The 8-byte struct at `buffer[offset..offset+8]` is
|
||||
/// `{ data_offset: u32, data_length: u32 }` (endian-aware). The actual UTF-8
|
||||
/// bytes live in `data_region[data_offset..data_offset+data_length]`. Returns
|
||||
/// a `&'a str` borrowing from `data_region`. Invalid UTF-8 produces
|
||||
/// [`AlkTypeError::Access`].
|
||||
/// bytes live at `buffer[data_offset..data_offset+data_length]` — offsets are
|
||||
/// absolute into the same buffer (safetensors-style). Returns a `&'a str`
|
||||
/// borrowing from `buffer`. Invalid UTF-8 produces [`AlkTypeError::Access`].
|
||||
pub fn read_string_indirect<'a>(
|
||||
buffer: &'a [u8],
|
||||
offset: usize,
|
||||
data_region: &'a [u8],
|
||||
field_path: &str,
|
||||
endian: Endian,
|
||||
) -> Result<&'a str, AlkTypeError> {
|
||||
let bytes = read_bytes_indirect(buffer, offset, data_region, field_path, endian)?;
|
||||
let bytes = read_bytes_indirect(buffer, offset, field_path, endian)?;
|
||||
std::str::from_utf8(bytes).map_err(|e| {
|
||||
access_err(
|
||||
field_path,
|
||||
@@ -324,34 +323,97 @@ pub fn read_string_indirect<'a>(
|
||||
///
|
||||
/// The 8-byte struct at `buffer[offset..offset+8]` is
|
||||
/// `{ data_offset: u32, data_length: u32 }` (endian-aware). Returns a
|
||||
/// `&'a [u8]` slice of `data_region[data_offset..data_offset+data_length]`.
|
||||
/// `&'a [u8]` slice of `buffer[data_offset..data_offset+data_length]` —
|
||||
/// offsets are absolute into the same buffer.
|
||||
pub fn read_bytes_indirect<'a>(
|
||||
buffer: &'a [u8],
|
||||
offset: usize,
|
||||
data_region: &'a [u8],
|
||||
field_path: &str,
|
||||
endian: Endian,
|
||||
) -> Result<&'a [u8], AlkTypeError> {
|
||||
let struct_end = offset
|
||||
.checked_add(8)
|
||||
.ok_or_else(|| access_err(field_path, format!("offset {offset} + 8 overflows usize")))?;
|
||||
check_bounds(buffer.len(), offset, struct_end, field_path)?;
|
||||
let off_bytes: [u8; U32_SIZE] = buffer[offset..offset + U32_SIZE]
|
||||
.try_into()
|
||||
.map_err(|_| access_err(field_path, "internal: try_into failed for data_offset"))?;
|
||||
let len_bytes: [u8; U32_SIZE] = buffer[offset + U32_SIZE..offset + 8]
|
||||
.try_into()
|
||||
.map_err(|_| access_err(field_path, "internal: try_into failed for data_length"))?;
|
||||
let data_offset = u32_from(off_bytes, endian) as usize;
|
||||
let data_length = u32_from(len_bytes, endian) as usize;
|
||||
let data_offset = read_u32(buffer, offset, field_path, endian)? as usize;
|
||||
let len_offset = offset.checked_add(U32_SIZE).ok_or_else(|| {
|
||||
access_err(
|
||||
field_path,
|
||||
format!("offset {offset} + {U32_SIZE} overflows usize"),
|
||||
)
|
||||
})?;
|
||||
let data_length = read_u32(buffer, len_offset, field_path, endian)? as usize;
|
||||
let data_end = data_offset.checked_add(data_length).ok_or_else(|| {
|
||||
access_err(
|
||||
field_path,
|
||||
format!("data_offset {data_offset} + data_length {data_length} overflows usize"),
|
||||
)
|
||||
})?;
|
||||
check_bounds(data_region.len(), data_offset, data_end, field_path)?;
|
||||
Ok(&data_region[data_offset..data_end])
|
||||
check_bounds(buffer.len(), data_offset, data_end, field_path)?;
|
||||
Ok(&buffer[data_offset..data_end])
|
||||
}
|
||||
|
||||
/// Write an offset-indirect string.
|
||||
///
|
||||
/// Writes the `{ data_offset: u32, data_length: u32 }` pair at `offset` and
|
||||
/// the UTF-8 bytes at `data_offset` (absolute into the same buffer). Returns
|
||||
/// the number of bytes written for the pair (`8`).
|
||||
pub fn write_string_indirect(
|
||||
buffer: &mut [u8],
|
||||
offset: usize,
|
||||
data_offset: usize,
|
||||
value: &str,
|
||||
field_path: &str,
|
||||
endian: Endian,
|
||||
) -> Result<usize, AlkTypeError> {
|
||||
write_bytes_indirect(buffer, offset, data_offset, value.as_bytes(), field_path, endian)
|
||||
}
|
||||
|
||||
/// Write offset-indirect raw bytes.
|
||||
///
|
||||
/// Writes the `{ data_offset: u32, data_length: u32 }` pair at `offset` and
|
||||
/// the raw bytes at `data_offset` (absolute into the same buffer). Returns
|
||||
/// the number of bytes written for the pair (`8`).
|
||||
pub fn write_bytes_indirect(
|
||||
buffer: &mut [u8],
|
||||
offset: usize,
|
||||
data_offset: usize,
|
||||
value: &[u8],
|
||||
field_path: &str,
|
||||
endian: Endian,
|
||||
) -> Result<usize, AlkTypeError> {
|
||||
let data_len = value.len();
|
||||
let data_len_u32 = u32::try_from(data_len).map_err(|_| {
|
||||
access_err(
|
||||
field_path,
|
||||
format!("data length {data_len} exceeds u32::MAX (length field width)"),
|
||||
)
|
||||
})?;
|
||||
let data_offset_u32 = u32::try_from(data_offset).map_err(|_| {
|
||||
access_err(
|
||||
field_path,
|
||||
format!("data offset {data_offset} exceeds u32::MAX (offset field width)"),
|
||||
)
|
||||
})?;
|
||||
write_u32(buffer, offset, data_offset_u32, field_path, endian)?;
|
||||
let len_offset = offset.checked_add(U32_SIZE).ok_or_else(|| {
|
||||
access_err(
|
||||
field_path,
|
||||
format!("offset {offset} + {U32_SIZE} overflows usize"),
|
||||
)
|
||||
})?;
|
||||
write_u32(buffer, len_offset, data_len_u32, field_path, endian)?;
|
||||
let data_end = data_offset.checked_add(data_len).ok_or_else(|| {
|
||||
access_err(
|
||||
field_path,
|
||||
format!("data offset {data_offset} + length {data_len} overflows usize"),
|
||||
)
|
||||
})?;
|
||||
check_bounds(buffer.len(), data_offset, data_end, field_path)?;
|
||||
let dest = buffer.get_mut(data_offset..data_end).ok_or_else(|| {
|
||||
access_err(
|
||||
field_path,
|
||||
format!("mutable data slice [{data_offset}..{data_end}) unavailable"),
|
||||
)
|
||||
})?;
|
||||
dest.copy_from_slice(value);
|
||||
Ok(8)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
@@ -583,42 +645,110 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn read_string_indirect_round_trip() {
|
||||
let data_region = b"the quick brown fox";
|
||||
let mut index = [0u8; 8];
|
||||
write_u32(&mut index, 0, 4, "idx.off", LE).unwrap();
|
||||
write_u32(&mut index, 4, 11, "idx.len", LE).unwrap();
|
||||
let s = read_string_indirect(&index, 0, data_region, "msg", LE).unwrap();
|
||||
let data = b"the quick brown fox";
|
||||
let mut full = vec![0u8; 8 + data.len()];
|
||||
write_u32(&mut full, 0, 12, "idx.off", LE).unwrap();
|
||||
write_u32(&mut full, 4, 11, "idx.len", LE).unwrap();
|
||||
full[8..].copy_from_slice(data);
|
||||
let s = read_string_indirect(&full, 0, "msg", LE).unwrap();
|
||||
assert_eq!(s, "quick brown");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_bytes_indirect_round_trip() {
|
||||
let data_region: &[u8] = b"HEADERbody-payloadTAIL";
|
||||
let mut index = [0u8; 8];
|
||||
write_u32(&mut index, 0, 6, "idx.off", BE).unwrap();
|
||||
write_u32(&mut index, 4, 12, "idx.len", BE).unwrap();
|
||||
let bytes = read_bytes_indirect(&index, 0, data_region, "blob", BE).unwrap();
|
||||
let data: &[u8] = b"HEADERbody-payloadTAIL";
|
||||
let mut full = vec![0u8; 8 + data.len()];
|
||||
write_u32(&mut full, 0, 14, "idx.off", BE).unwrap();
|
||||
write_u32(&mut full, 4, 12, "idx.len", BE).unwrap();
|
||||
full[8..].copy_from_slice(data);
|
||||
let bytes = read_bytes_indirect(&full, 0, "blob", BE).unwrap();
|
||||
assert_eq!(bytes, b"body-payload");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn write_bytes_indirect_round_trip() {
|
||||
let mut buf = vec![0u8; 8 + 11];
|
||||
let written = write_bytes_indirect(&mut buf, 0, 8, b"quick brown", "blob", LE).unwrap();
|
||||
assert_eq!(written, 8);
|
||||
assert_eq!(&buf[0..4], &8u32.to_le_bytes());
|
||||
assert_eq!(&buf[4..8], &11u32.to_le_bytes());
|
||||
assert_eq!(&buf[8..19], b"quick brown");
|
||||
let bytes = read_bytes_indirect(&buf, 0, "blob", LE).unwrap();
|
||||
assert_eq!(bytes, b"quick brown");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn write_string_indirect_round_trip() {
|
||||
let mut buf = vec![0u8; 8 + 5];
|
||||
let written = write_string_indirect(&mut buf, 0, 8, "hello", "msg", BE).unwrap();
|
||||
assert_eq!(written, 8);
|
||||
assert_eq!(&buf[0..4], &8u32.to_be_bytes());
|
||||
assert_eq!(&buf[4..8], &5u32.to_be_bytes());
|
||||
assert_eq!(&buf[8..13], b"hello");
|
||||
let s = read_string_indirect(&buf, 0, "msg", BE).unwrap();
|
||||
assert_eq!(s, "hello");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_bytes_indirect_bounds_failure_on_index() {
|
||||
let buf = [0u8; 4];
|
||||
let data_region = b"anything";
|
||||
let err = read_bytes_indirect(&buf, 0, data_region, "blob", LE).unwrap_err();
|
||||
let err = read_bytes_indirect(&buf, 0, "blob", LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_bytes_indirect_bounds_failure_on_data_region() {
|
||||
fn read_bytes_indirect_bounds_failure_on_data() {
|
||||
let mut buf = [0u8; 8];
|
||||
write_u32(&mut buf, 0, 100, "idx.off", LE).unwrap();
|
||||
write_u32(&mut buf, 4, 10, "idx.len", LE).unwrap();
|
||||
let data_region = b"too short";
|
||||
let err = read_bytes_indirect(&buf, 0, data_region, "blob", LE).unwrap_err();
|
||||
let err = read_bytes_indirect(&buf, 0, "blob", LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn write_bytes_indirect_bounds_failure_on_data() {
|
||||
let mut buf = vec![0u8; 8];
|
||||
let err = write_bytes_indirect(&mut buf, 0, 100, b"hello", "blob", LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn m4_write_bytes_indirect_pair_at_nonzero_offset_data_ok() {
|
||||
// The indirect write's pair offset and data offset are
|
||||
// independent; pin the pair landing mid-buffer with the data
|
||||
// elsewhere (review #006 M4 item 3 — the write_bytes_indirect
|
||||
// guard family is the canonical overflow-guard pattern, review
|
||||
// #002 M2, and its nonzero-offset paths were uncovered).
|
||||
let mut buf = vec![0u8; 24];
|
||||
let written = write_bytes_indirect(&mut buf, 4, 16, b"xyz", "blob", BE).unwrap();
|
||||
assert_eq!(written, 8);
|
||||
assert_eq!(&buf[4..8], &16u32.to_be_bytes());
|
||||
assert_eq!(&buf[8..12], &3u32.to_be_bytes());
|
||||
assert_eq!(&buf[16..19], b"xyz");
|
||||
let bytes = read_bytes_indirect(&buf, 4, "blob", BE).unwrap();
|
||||
assert_eq!(bytes, b"xyz");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn m4_write_bytes_indirect_data_length_exceeds_bounds() {
|
||||
// The data-region bounds check in write_bytes_indirect: the pair
|
||||
// fits, but the {data_offset, length} target runs past the buffer
|
||||
// end — the write must refuse without corrupting the pair.
|
||||
let mut buf = vec![0u8; 20];
|
||||
let err = write_bytes_indirect(&mut buf, 0, 16, b"hello", "blob", LE).unwrap_err();
|
||||
match err {
|
||||
AlkTypeError::Access { field_path, reason } => {
|
||||
assert_eq!(field_path, "blob");
|
||||
assert!(reason.contains("bounds"), "reason: {reason}");
|
||||
}
|
||||
other => panic!("expected Access, got {other:?}"),
|
||||
}
|
||||
// The pair was already written (offset+length at 0..8) — that's
|
||||
// fine; the error refuses the data copy only.
|
||||
assert_eq!(&buf[0..4], &16u32.to_le_bytes());
|
||||
assert_eq!(&buf[4..8], &5u32.to_le_bytes());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_at_nonzero_offset() {
|
||||
let mut buf = vec![0u8; 16];
|
||||
|
||||
+1088
-365
File diff suppressed because it is too large.
Load diff
+3
-3
@@ -1,6 +1,6 @@
|
||||
//! Error types for the alktype engine.
|
||||
//!
|
||||
//! Decided in ADR-098: a single `AlkTypeError` enum covers all error
|
||||
//! Decided in ADR-004: a single `AlkTypeError` enum covers all error
|
||||
//! conditions across the engine's three phases (schema parsing, offset
|
||||
//! computation, read/write) plus validation.
|
||||
|
||||
@@ -9,8 +9,8 @@ use std::fmt;
|
||||
/// Errors produced by the alktype engine across all phases.
|
||||
#[derive(Debug)]
|
||||
pub enum AlkTypeError {
|
||||
/// Schema parsing errors — invalid JSON, missing required keywords,
|
||||
/// unknown `AlkType:*` kinds, malformed annotations.
|
||||
/// Schema parsing errors — invalid JSON, missing required properties,
|
||||
/// unknown BAST kinds, malformed annotations.
|
||||
Schema(String),
|
||||
|
||||
/// Offset computation errors — field not found, type not supported
|
||||
|
||||
+1006
-875
File diff suppressed because it is too large.
Load diff
+44
-17
@@ -1,33 +1,52 @@
|
||||
//! alktype: The binary struct engine.
|
||||
//!
|
||||
//! Takes a JSON Schema with `AlkType:*` custom keywords and produces
|
||||
//! Takes a BAST (Binary Abstract Syntax Tree) document and produces
|
||||
//! an offset map, read/write functions, and validation — all driven
|
||||
//! by the schema. The schema is the format definition; the engine is
|
||||
//! generic.
|
||||
//!
|
||||
//! ## Architecture
|
||||
//!
|
||||
//! - **Schema layer** ([`schema`]): AlkType kind detection, annotation
|
||||
//! parsing, `$ref` normalization, endianness.
|
||||
//! - **BAST parser** ([`bast`]): Typed tree over a BAST document —
|
||||
//! `BastDoc`/`BastDef`/`BastStruct`/`BastField`/`BastType`/etc.
|
||||
//! Owns its data (ADR-012 §2a) so consumers (`AlkTypeEngine`,
|
||||
//! `LayoutBuilder`) can store the parsed tree without lifetime
|
||||
//! entanglement; the clone happens once at `BastDoc::new`.
|
||||
//! - **Layout engine** ([`offset_map`], [`layout_builder`],
|
||||
//! [`sequential_reader`]): Two layout modes — aligned static for
|
||||
//! mmap-friendly formats, packed sequential for protocol wire formats.
|
||||
//! [`sequential_reader`], [`read_plan`]): Two layout modes — aligned
|
||||
//! static for mmap-friendly formats, packed sequential for protocol
|
||||
//! wire formats. All consume the BAST typed tree; `read_plan` is the
|
||||
//! packed read-side compiled form (ADR-011).
|
||||
//! - **Data access** ([`data_access`]): Typed read/write at computed
|
||||
//! offsets, zero-copy for fixed-size types.
|
||||
//! - **TUnion dispatch** ([`tunion`]): Byte-offset and field-name
|
||||
//! discriminator dispatch.
|
||||
//! - **Validation** ([`validation`]): Custom keyword validators for all
|
||||
//! 19 `AlkType:*` kinds, delegated to the `jsonschema` crate.
|
||||
//! - **Builder** ([`builder`]): Fluent Rust API for constructing alktype
|
||||
//! JSON Schemas at runtime, producing `serde_json::Value` (ADR-009).
|
||||
//! discriminator dispatch over `BastUnion`.
|
||||
//! - **Validation** ([`validation`], [`bast_validation`],
|
||||
//! [`validation_plan`]): two validators for two paths.
|
||||
//! `validation_plan` is the compiled `ValidationPlan` — a
|
||||
//! compile-once-walk-many constraint tree over the BAST document's
|
||||
//! value-domain constraints, walked by `validate_bytes` per buffer
|
||||
//! without re-touching the document (ADR-012 §3). `bast_validation`
|
||||
//! hosts the one-shot `validate_value` wrapper over the plan.
|
||||
//! `validation` builds a standard `jsonschema::Validator` from a
|
||||
//! consumer-provided JSON Schema for `validate_json` /
|
||||
//! `is_valid_json` (D-BAST-007) — no custom keywords, no BAST
|
||||
//! involvement.
|
||||
//! - **Builder** ([`builder`]): Fluent Rust API for constructing BAST
|
||||
//! documents (binary layout) and standard JSON Schemas (JSON
|
||||
//! validation) at runtime, producing `serde_json::Value` (ADR-009,
|
||||
//! D-BAST-008). `struct_()` → BAST, `object()` → standard JSON Schema.
|
||||
//! - **Materialize** ([`materialize`]): Materialize a `serde_json::Value`
|
||||
//! tree from a binary buffer by walking the schema. Used by
|
||||
//! tree from a binary buffer by walking the BAST typed tree. Used by
|
||||
//! `AlkTypeEngine::validate_bytes` (ADR-010).
|
||||
//! - **Engine** ([`engine`]): `AlkTypeEngine` — the compiled form of a
|
||||
//! schema, combining layout and validation.
|
||||
//! BAST document, combining layout and validation.
|
||||
|
||||
#[macro_use]
|
||||
mod macros;
|
||||
pub mod bast;
|
||||
pub mod bast_meta;
|
||||
pub mod bast_validation;
|
||||
pub mod builder;
|
||||
pub mod data_access;
|
||||
pub mod engine;
|
||||
@@ -35,21 +54,29 @@ pub mod error;
|
||||
pub mod layout_builder;
|
||||
pub mod materialize;
|
||||
pub mod offset_map;
|
||||
pub mod read_plan;
|
||||
pub mod schema;
|
||||
pub mod sequential_reader;
|
||||
pub mod tunion;
|
||||
pub mod validation;
|
||||
pub mod validation_plan;
|
||||
pub(crate) mod walk_guard;
|
||||
|
||||
pub use bast_meta::BAST_META_SCHEMA;
|
||||
pub use bast::{
|
||||
BastArray, BastDef, BastDefKind, BastDiscriminator, BastDoc, BastEnum, BastField, BastRef,
|
||||
BastRecord, BastStruct, BastType, BastUnion,
|
||||
};
|
||||
pub use builder::{Definitions, Discriminator, Schema};
|
||||
pub use engine::{LayoutMode, AlkTypeEngine};
|
||||
pub use error::AlkTypeError;
|
||||
pub use layout_builder::{FieldPosition, LayoutBuilder, PackedLayout};
|
||||
pub use offset_map::{ByteRange, OffsetMap};
|
||||
pub use schema::{
|
||||
get_alktype_kind_loose, get_alktype_kind_loose_enum, inline_union_variant_refs, normalize_refs,
|
||||
parse_align, parse_discriminator, parse_encoding, parse_endian, parse_max_length, resolve_ref,
|
||||
resolve_ref_or_inline, DiscriminatorKind, Endian, AlkTypeKind, VariableEncoding,
|
||||
pub use offset_map::{ByteRange, LeafMeta, OffsetEntry, OffsetMap};
|
||||
pub use read_plan::{
|
||||
CompositePlan, DiscriminatorPlan, FieldPlan, ReadKind, ReadPlan,
|
||||
};
|
||||
pub use schema::{Endian, AlkTypeKind, VariableEncoding};
|
||||
pub use sequential_reader::{FieldValue, SequentialReader};
|
||||
pub use tunion::UnionDispatch;
|
||||
pub use validation::build_validator;
|
||||
pub use validation_plan::{ValidField, ValidNode, ValidVariant, ValidationPlan};
|
||||
+5
-169
@@ -1,173 +1,9 @@
|
||||
//! Macros for generating repetitive code across the 19 AlkType kinds.
|
||||
//! Macros for generating repetitive read/write code across the
|
||||
//! endian-sensitive fixed-size kinds.
|
||||
//!
|
||||
//! These macros eliminate boilerplate in validation, data access, and
|
||||
//! dispatch. Each macro takes a compact specification and generates the
|
||||
//! full implementation, ensuring consistency across all types.
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Validation macros
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Generate a signed integer validator struct and its factory closure.
|
||||
#[macro_export]
|
||||
macro_rules! define_int_validator {
|
||||
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $min:literal, $max:literal) => {
|
||||
struct $validator_struct;
|
||||
impl jsonschema::Keyword for $validator_struct {
|
||||
fn validate<'i>(
|
||||
&self,
|
||||
instance: &'i serde_json::Value,
|
||||
) -> Result<(), jsonschema::ValidationError<'i>> {
|
||||
match instance.as_i64() {
|
||||
Some(n) if ($min..=$max).contains(&n) => Ok(()),
|
||||
_ => Err(jsonschema::ValidationError::custom(concat!(
|
||||
"expected an integer in range [",
|
||||
stringify!($min),
|
||||
", ",
|
||||
stringify!($max),
|
||||
"]"
|
||||
))),
|
||||
}
|
||||
}
|
||||
fn is_valid(&self, instance: &serde_json::Value) -> bool {
|
||||
instance
|
||||
.as_i64()
|
||||
.is_some_and(|n| ($min..=$max).contains(&n))
|
||||
}
|
||||
}
|
||||
|
||||
fn $factory_fn<'a>(
|
||||
_parent: &'a serde_json::Map<String, serde_json::Value>,
|
||||
value: &'a serde_json::Value,
|
||||
_path: jsonschema::paths::Location,
|
||||
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
|
||||
if value.as_bool() == Some(true) {
|
||||
Ok(Box::new($validator_struct))
|
||||
} else {
|
||||
Err(jsonschema::ValidationError::schema(concat!(
|
||||
$keyword,
|
||||
" must be set to true"
|
||||
)))
|
||||
}
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
/// Generate an unsigned integer validator struct and its factory closure.
|
||||
#[macro_export]
|
||||
macro_rules! define_uint_validator {
|
||||
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $max:literal) => {
|
||||
struct $validator_struct;
|
||||
impl jsonschema::Keyword for $validator_struct {
|
||||
fn validate<'i>(
|
||||
&self,
|
||||
instance: &'i serde_json::Value,
|
||||
) -> Result<(), jsonschema::ValidationError<'i>> {
|
||||
match instance.as_u64() {
|
||||
Some(n) if n <= $max => Ok(()),
|
||||
_ => Err(jsonschema::ValidationError::custom(concat!(
|
||||
"expected an unsigned integer in range [0, ",
|
||||
stringify!($max),
|
||||
"]"
|
||||
))),
|
||||
}
|
||||
}
|
||||
fn is_valid(&self, instance: &serde_json::Value) -> bool {
|
||||
instance.as_u64().is_some_and(|n| n <= $max)
|
||||
}
|
||||
}
|
||||
|
||||
fn $factory_fn<'a>(
|
||||
_parent: &'a serde_json::Map<String, serde_json::Value>,
|
||||
value: &'a serde_json::Value,
|
||||
_path: jsonschema::paths::Location,
|
||||
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
|
||||
if value.as_bool() == Some(true) {
|
||||
Ok(Box::new($validator_struct))
|
||||
} else {
|
||||
Err(jsonschema::ValidationError::schema(concat!(
|
||||
$keyword,
|
||||
" must be set to true"
|
||||
)))
|
||||
}
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
/// Generate a float validator struct and its factory closure.
|
||||
#[macro_export]
|
||||
macro_rules! define_float_validator {
|
||||
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $error_msg:literal) => {
|
||||
struct $validator_struct;
|
||||
impl jsonschema::Keyword for $validator_struct {
|
||||
fn validate<'i>(
|
||||
&self,
|
||||
instance: &'i serde_json::Value,
|
||||
) -> Result<(), jsonschema::ValidationError<'i>> {
|
||||
match instance.as_f64() {
|
||||
Some(f) if f.is_finite() => Ok(()),
|
||||
_ => Err(jsonschema::ValidationError::custom($error_msg)),
|
||||
}
|
||||
}
|
||||
fn is_valid(&self, instance: &serde_json::Value) -> bool {
|
||||
instance.as_f64().is_some_and(|f| f.is_finite())
|
||||
}
|
||||
}
|
||||
|
||||
fn $factory_fn<'a>(
|
||||
_parent: &'a serde_json::Map<String, serde_json::Value>,
|
||||
value: &'a serde_json::Value,
|
||||
_path: jsonschema::paths::Location,
|
||||
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
|
||||
if value.as_bool() == Some(true) {
|
||||
Ok(Box::new($validator_struct))
|
||||
} else {
|
||||
Err(jsonschema::ValidationError::schema(concat!(
|
||||
$keyword,
|
||||
" must be set to true"
|
||||
)))
|
||||
}
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
/// Generate a simple type-check validator (object/array/boolean) and its factory.
|
||||
#[macro_export]
|
||||
macro_rules! define_type_validator {
|
||||
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $check_method:ident, $error_msg:literal) => {
|
||||
struct $validator_struct;
|
||||
impl jsonschema::Keyword for $validator_struct {
|
||||
fn validate<'i>(
|
||||
&self,
|
||||
instance: &'i serde_json::Value,
|
||||
) -> Result<(), jsonschema::ValidationError<'i>> {
|
||||
if instance.$check_method() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(jsonschema::ValidationError::custom($error_msg))
|
||||
}
|
||||
}
|
||||
fn is_valid(&self, instance: &serde_json::Value) -> bool {
|
||||
instance.$check_method()
|
||||
}
|
||||
}
|
||||
|
||||
fn $factory_fn<'a>(
|
||||
_parent: &'a serde_json::Map<String, serde_json::Value>,
|
||||
value: &'a serde_json::Value,
|
||||
_path: jsonschema::paths::Location,
|
||||
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
|
||||
if value.as_bool() == Some(true) {
|
||||
Ok(Box::new($validator_struct))
|
||||
} else {
|
||||
Err(jsonschema::ValidationError::schema(concat!(
|
||||
$keyword,
|
||||
" must be set to true"
|
||||
)))
|
||||
}
|
||||
}
|
||||
};
|
||||
}
|
||||
//! These macros eliminate boilerplate in data access. Each macro takes a
|
||||
//! compact specification and generates the full implementation, ensuring
|
||||
//! consistency across all types.
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Data access macros
|
||||
|
||||
+1912
-658
File diff suppressed because it is too large.
Load diff
+1168
-460
File diff suppressed because it is too large.
Load diff
+1659
File diff suppressed because it is too large.
Load diff
+184
-792
File diff suppressed because it is too large.
Load diff
+1156
-837
File diff suppressed because it is too large.
Load diff
+281
-268
@@ -1,18 +1,18 @@
|
||||
//! TUnion discriminator dispatch (ADR-097 §4).
|
||||
//! TUnion discriminator dispatch (ADR-003 §4).
|
||||
//!
|
||||
//! TUnion supports two discriminator kinds: byte-offset (protocol
|
||||
//! dispatch, e.g., SFTP type bytes) and field-name (typedef.ts string
|
||||
//! pattern). This module reads the discriminator value from a byte
|
||||
//! buffer, looks up the variant schema in the union's `mapping`, and
|
||||
//! buffer, looks up the variant type in the union's `mapping`, and
|
||||
//! reports the offset where the variant struct begins.
|
||||
//!
|
||||
//! All reads go through [`crate::data_access`] so bounds checks and
|
||||
//! endianness handling are uniform with the rest of the engine.
|
||||
|
||||
use crate::bast::{BastDiscriminator, BastType, BastUnion};
|
||||
use crate::data_access::{read_enum, read_string, read_u16, read_u32, read_u8};
|
||||
use crate::error::AlkTypeError;
|
||||
use crate::schema::{get_alktype_kind, parse_discriminator, DiscriminatorKind, Endian, AlkTypeKind, DISCRIMINATOR_PATH, U32_SIZE};
|
||||
use serde_json::Value;
|
||||
use crate::schema::{AlkTypeKind, Endian, DISCRIMINATOR_PATH, U32_SIZE};
|
||||
|
||||
const STRING_PREFIX_SIZE: usize = 4;
|
||||
|
||||
@@ -39,21 +39,19 @@ pub struct UnionDispatch {
|
||||
///
|
||||
/// # Errors
|
||||
///
|
||||
/// - [`AlkTypeError::Schema`] if the discriminator annotation is missing
|
||||
/// or malformed, or if the discriminator `type` is not one of
|
||||
/// `AlkType:Uint8` / `AlkType:Uint16` / `AlkType:Uint32`.
|
||||
/// - [`AlkTypeError::Schema`] if the union does not have a byte-offset
|
||||
/// discriminator.
|
||||
/// - [`AlkTypeError::Access`] if the buffer is too short to contain the
|
||||
/// discriminator, or if the read value is not present in the union's
|
||||
/// `mapping`.
|
||||
pub fn read_byte_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &Value,
|
||||
union_node: &BastUnion,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, AlkTypeError> {
|
||||
let disc = parse_discriminator(union_schema)?;
|
||||
let (offset, disc_type) = match disc {
|
||||
DiscriminatorKind::Byte { offset, disc_type } => (offset, disc_type),
|
||||
DiscriminatorKind::Field { .. } => {
|
||||
let (offset, disc_type) = match union_node.discriminator() {
|
||||
BastDiscriminator::Byte { offset, disc_type } => (*offset, *disc_type),
|
||||
BastDiscriminator::Field { .. } => {
|
||||
return Err(AlkTypeError::Schema(
|
||||
"read_byte_discriminator requires a byte-offset discriminator".to_string(),
|
||||
));
|
||||
@@ -75,7 +73,7 @@ pub fn read_byte_discriminator(
|
||||
};
|
||||
|
||||
let key = disc_value.to_string();
|
||||
verify_mapping_key(union_schema, &key, DISCRIMINATOR_PATH, &key)?;
|
||||
verify_mapping_key(union_node, &key, DISCRIMINATOR_PATH, &key)?;
|
||||
|
||||
let variant_offset =
|
||||
offset
|
||||
@@ -105,56 +103,47 @@ pub fn read_byte_discriminator(
|
||||
///
|
||||
/// # Errors
|
||||
///
|
||||
/// - [`AlkTypeError::Schema`] if the discriminator annotation is missing
|
||||
/// or malformed, the discriminator field is not declared in
|
||||
/// `properties`, the field has no `AlkType:*` kind, or the field's
|
||||
/// kind is not one of `AlkType:String` / `AlkType:Uint8` /
|
||||
/// `AlkType:Enum`.
|
||||
/// - [`AlkTypeError::Schema`] if the union does not have a field-name
|
||||
/// discriminator, the discriminator field is not declared in
|
||||
/// `fields`, or the field's kind is not one of `string` / `uint8` /
|
||||
/// `uint16` / `uint32` / `enum` (the same kind set the compiled
|
||||
/// reader's `plan_discriminator_string_value` accepts — review #006
|
||||
/// N1 closed the divergence so both public dispatch paths answer
|
||||
/// identically).
|
||||
/// - [`AlkTypeError::Access`] if the buffer is too short to contain the
|
||||
/// discriminator field, or if the read value is not present in the
|
||||
/// union's `mapping`.
|
||||
pub fn read_field_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &Value,
|
||||
union_node: &BastUnion,
|
||||
disc_field_offset: usize,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, AlkTypeError> {
|
||||
let disc = parse_discriminator(union_schema)?;
|
||||
let name = match disc {
|
||||
DiscriminatorKind::Field { name } => name,
|
||||
DiscriminatorKind::Byte { .. } => {
|
||||
let name = match union_node.discriminator() {
|
||||
BastDiscriminator::Field { name } => name.as_str(),
|
||||
BastDiscriminator::Byte { .. } => {
|
||||
return Err(AlkTypeError::Schema(
|
||||
"read_field_discriminator requires a field-name discriminator".to_string(),
|
||||
));
|
||||
}
|
||||
};
|
||||
|
||||
let field_schema = union_schema
|
||||
.get("properties")
|
||||
.and_then(Value::as_object)
|
||||
.and_then(|props| props.get(&name))
|
||||
.ok_or_else(|| {
|
||||
AlkTypeError::Schema(format!(
|
||||
"discriminator field '{name}' not found in union properties"
|
||||
))
|
||||
})?;
|
||||
let disc_field = union_node.fields().iter().find(|f| f.name() == name).ok_or_else(|| {
|
||||
AlkTypeError::Schema(format!(
|
||||
"discriminator field '{name}' not found in union fields"
|
||||
))
|
||||
})?;
|
||||
|
||||
let kind = get_alktype_kind(field_schema)
|
||||
.and_then(|s| s.parse::<AlkTypeKind>().ok())
|
||||
.ok_or_else(|| {
|
||||
AlkTypeError::Schema(format!(
|
||||
"discriminator field '{name}' has no AlkType:* kind"
|
||||
))
|
||||
})?;
|
||||
let kind = disc_field.ty().alk_kind();
|
||||
|
||||
let (key, discriminator_field_size) = match kind {
|
||||
AlkTypeKind::String => {
|
||||
let s = read_string(buffer, disc_field_offset, &name, endian)?;
|
||||
let s = read_string(buffer, disc_field_offset, name, endian)?;
|
||||
let size =
|
||||
STRING_PREFIX_SIZE
|
||||
.checked_add(s.len())
|
||||
.ok_or_else(|| AlkTypeError::Access {
|
||||
field_path: name.clone(),
|
||||
field_path: name.to_string(),
|
||||
reason: format!(
|
||||
"string prefix {STRING_PREFIX_SIZE} + data length {} overflows usize",
|
||||
s.len()
|
||||
@@ -163,11 +152,19 @@ pub fn read_field_discriminator(
|
||||
(s.to_string(), size)
|
||||
}
|
||||
AlkTypeKind::Uint8 => {
|
||||
let v = read_u8(buffer, disc_field_offset, &name)?;
|
||||
let v = read_u8(buffer, disc_field_offset, name)?;
|
||||
(v.to_string(), 1)
|
||||
}
|
||||
AlkTypeKind::Uint16 => {
|
||||
let v = read_u16(buffer, disc_field_offset, name, endian)?;
|
||||
(v.to_string(), 2)
|
||||
}
|
||||
AlkTypeKind::Uint32 => {
|
||||
let v = read_u32(buffer, disc_field_offset, name, endian)?;
|
||||
(v.to_string(), U32_SIZE)
|
||||
}
|
||||
AlkTypeKind::Enum => {
|
||||
let v = read_enum(buffer, disc_field_offset, &name, endian)?;
|
||||
let v = read_enum(buffer, disc_field_offset, name, endian)?;
|
||||
(v.to_string(), U32_SIZE)
|
||||
}
|
||||
other => {
|
||||
@@ -177,12 +174,12 @@ pub fn read_field_discriminator(
|
||||
}
|
||||
};
|
||||
|
||||
verify_mapping_key(union_schema, &key, &name, &key)?;
|
||||
verify_mapping_key(union_node, &key, name, &key)?;
|
||||
|
||||
let variant_offset = disc_field_offset
|
||||
.checked_add(discriminator_field_size)
|
||||
.ok_or_else(|| AlkTypeError::Access {
|
||||
field_path: name.clone(),
|
||||
field_path: name.to_string(),
|
||||
reason: format!(
|
||||
"disc_field_offset {disc_field_offset} + discriminator_field_size {discriminator_field_size} overflows usize"
|
||||
),
|
||||
@@ -195,63 +192,39 @@ pub fn read_field_discriminator(
|
||||
})
|
||||
}
|
||||
|
||||
/// Look up a variant schema from the union's mapping.
|
||||
/// Look up a variant type from the union's mapping.
|
||||
///
|
||||
/// Returns the variant schema. Inline schemas are returned directly.
|
||||
/// `$ref` pointers of the form `"#/$defs/<name>"` are resolved against
|
||||
/// the `union_schema`'s own `$defs` block (when the union schema is the
|
||||
/// schema root). For nested unions whose `$defs` live on an ancestor,
|
||||
/// the caller (typically `AlkTypeEngine::compile`) is expected to
|
||||
/// resolve refs before reaching this function, or to inline the
|
||||
/// variant schemas into the mapping at load time.
|
||||
/// Returns the variant [`BastType`]. Inline struct/union/enum types are
|
||||
/// returned directly; `$ref` pointers are returned as
|
||||
/// [`BastType::Ref`] — the caller resolves them via
|
||||
/// [`crate::bast::BastDoc::resolve_typeref`] when a concrete definition is needed.
|
||||
///
|
||||
/// # Errors
|
||||
///
|
||||
/// - [`AlkTypeError::Schema`] if the union has no `mapping` object, the
|
||||
/// `key` is not present, a `$ref` is malformed, or a `$ref` cannot be
|
||||
/// resolved against the union schema's own `$defs`.
|
||||
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str) -> Result<&'a Value, AlkTypeError> {
|
||||
let mapping = union_schema
|
||||
.get("mapping")
|
||||
.and_then(Value::as_object)
|
||||
.ok_or_else(|| AlkTypeError::Schema("union is missing 'mapping' object".to_string()))?;
|
||||
|
||||
let variant = mapping
|
||||
.get(key)
|
||||
.ok_or_else(|| AlkTypeError::Schema(format!("unknown mapping key: {key}")))?;
|
||||
|
||||
let ref_str = match variant.get("$ref").and_then(Value::as_str) {
|
||||
Some(r) => r,
|
||||
None => return Ok(variant),
|
||||
};
|
||||
|
||||
let pointer = ref_str
|
||||
.strip_prefix('#')
|
||||
.ok_or_else(|| AlkTypeError::Schema(format!("unsupported $ref form: {ref_str}")))?;
|
||||
|
||||
let resolved = resolve_json_pointer(union_schema, pointer).ok_or_else(|| {
|
||||
AlkTypeError::Schema(format!(
|
||||
"cannot resolve $ref {ref_str} against union schema; ensure refs are inlined or the union schema contains $defs"
|
||||
))
|
||||
})?;
|
||||
Ok(resolved)
|
||||
/// - [`AlkTypeError::Schema`] if the union has no `mapping` entries or
|
||||
/// the `key` is not present.
|
||||
pub fn resolve_variant<'a>(
|
||||
union_node: &'a BastUnion,
|
||||
key: &str,
|
||||
) -> Result<&'a BastType, AlkTypeError> {
|
||||
union_node
|
||||
.variant_for(key)
|
||||
.ok_or_else(|| AlkTypeError::Schema(format!("unknown mapping key: {key}")))
|
||||
}
|
||||
|
||||
/// Get the discriminator size in bytes for a byte-offset discriminator.
|
||||
///
|
||||
/// Returns 1 for `AlkType:Uint8`, 2 for `AlkType:Uint16`, and 4 for
|
||||
/// `AlkType:Uint32`. Field-name discriminators have no fixed size and
|
||||
/// produce a [`AlkTypeError::Schema`].
|
||||
/// Returns 1 for `uint8`, 2 for `uint16`, and 4 for `uint32`.
|
||||
/// Field-name discriminators have no fixed size and produce a
|
||||
/// [`AlkTypeError::Schema`].
|
||||
///
|
||||
/// # Errors
|
||||
///
|
||||
/// - [`AlkTypeError::Schema`] if the discriminator annotation is
|
||||
/// missing/malformed, the discriminator `type` is unsupported, or the
|
||||
/// discriminator is a field-name discriminator.
|
||||
pub fn discriminator_size(union_schema: &Value) -> Result<usize, AlkTypeError> {
|
||||
let disc = parse_discriminator(union_schema)?;
|
||||
match disc {
|
||||
DiscriminatorKind::Byte { disc_type, .. } => match disc_type {
|
||||
/// - [`AlkTypeError::Schema`] if the discriminator is a field-name
|
||||
/// discriminator.
|
||||
pub fn discriminator_size(union_node: &BastUnion) -> Result<usize, AlkTypeError> {
|
||||
match union_node.discriminator() {
|
||||
BastDiscriminator::Byte { disc_type, .. } => match disc_type {
|
||||
AlkTypeKind::Uint8 => Ok(1),
|
||||
AlkTypeKind::Uint16 => Ok(2),
|
||||
AlkTypeKind::Uint32 => Ok(4),
|
||||
@@ -259,24 +232,19 @@ pub fn discriminator_size(union_schema: &Value) -> Result<usize, AlkTypeError> {
|
||||
"unsupported byte discriminator type: {other}"
|
||||
))),
|
||||
},
|
||||
DiscriminatorKind::Field { .. } => Err(AlkTypeError::Schema(
|
||||
BastDiscriminator::Field { .. } => Err(AlkTypeError::Schema(
|
||||
"field-name discriminator has no fixed size".to_string(),
|
||||
)),
|
||||
}
|
||||
}
|
||||
|
||||
fn verify_mapping_key(
|
||||
union_schema: &Value,
|
||||
union_node: &BastUnion,
|
||||
key: &str,
|
||||
field_path: &str,
|
||||
raw_value: &str,
|
||||
) -> Result<(), AlkTypeError> {
|
||||
let in_mapping = union_schema
|
||||
.get("mapping")
|
||||
.and_then(Value::as_object)
|
||||
.map(|m| m.contains_key(key))
|
||||
.unwrap_or(false);
|
||||
if in_mapping {
|
||||
if union_node.variant_for(key).is_some() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(AlkTypeError::Access {
|
||||
@@ -286,84 +254,64 @@ fn verify_mapping_key(
|
||||
}
|
||||
}
|
||||
|
||||
fn resolve_json_pointer<'a>(root: &'a Value, pointer: &str) -> Option<&'a Value> {
|
||||
if pointer.is_empty() {
|
||||
return Some(root);
|
||||
}
|
||||
let trimmed = pointer.strip_prefix('/')?;
|
||||
let mut current = root;
|
||||
for unescaped in trimmed.split('/') {
|
||||
let segment = unescape_json_pointer_token(unescaped)?;
|
||||
current = current.get(&segment)?;
|
||||
}
|
||||
Some(current)
|
||||
}
|
||||
|
||||
fn unescape_json_pointer_token(token: &str) -> Option<String> {
|
||||
let mut out = String::with_capacity(token.len());
|
||||
let mut chars = token.chars();
|
||||
while let Some(c) = chars.next() {
|
||||
match c {
|
||||
'~' => match chars.next() {
|
||||
Some('0') => out.push('~'),
|
||||
Some('1') => out.push('/'),
|
||||
_ => return None,
|
||||
},
|
||||
other => out.push(other),
|
||||
}
|
||||
}
|
||||
Some(out)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::bast::BastDoc;
|
||||
use serde_json::json;
|
||||
|
||||
const LE: Endian = Endian::Little;
|
||||
const BE: Endian = Endian::Big;
|
||||
|
||||
fn byte_union_schema(offset: usize, disc_type: &str) -> Value {
|
||||
fn doc_union(root: &serde_json::Value, name: &str) -> BastUnion {
|
||||
let doc = BastDoc::new(root, name).expect("bast doc");
|
||||
match doc.root_def().kind() {
|
||||
crate::bast::BastDefKind::Union(u) => u.clone(),
|
||||
_ => panic!("root must be a union"),
|
||||
}
|
||||
}
|
||||
|
||||
fn byte_union_root(offset: usize, disc_type: &str) -> serde_json::Value {
|
||||
json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "offset": offset, "type": disc_type},
|
||||
"mapping": {
|
||||
"5": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}},
|
||||
"6": {"AlkType:Struct": true, "properties": {"len": {"AlkType:Uint16": true}}}
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": offset, "type": disc_type },
|
||||
"mapping": {
|
||||
"5": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] },
|
||||
"6": { "kind": "struct", "fields": [ { "name": "len", "kind": "uint16" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
fn field_union_schema(field_name: &str, field_kind: &str) -> Value {
|
||||
let field_schema = match field_kind {
|
||||
"AlkType:Enum" => json!({
|
||||
"AlkType:Enum": true,
|
||||
"enum": ["read", "write"]
|
||||
}),
|
||||
_ => json!({field_kind: true}),
|
||||
};
|
||||
fn field_union_root(field_name: &str, field_kind: &str) -> serde_json::Value {
|
||||
let (key_a, key_b) = match field_kind {
|
||||
"AlkType:String" => ("read", "write"),
|
||||
"string" => ("read", "write"),
|
||||
_ => ("0", "1"),
|
||||
};
|
||||
json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": field_name},
|
||||
"properties": {
|
||||
field_name: field_schema
|
||||
},
|
||||
"mapping": {
|
||||
key_a: {"AlkType:Struct": true, "properties": {"n": {"AlkType:Uint32": true}}},
|
||||
key_b: {"AlkType:Struct": true, "properties": {"m": {"AlkType:Uint16": true}}}
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": field_name },
|
||||
"fields": [ { "name": field_name, "kind": field_kind } ],
|
||||
"mapping": {
|
||||
key_a: { "kind": "struct", "fields": [ { "name": "n", "kind": "uint32" } ] },
|
||||
key_b: { "kind": "struct", "fields": [ { "name": "m", "kind": "uint16" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint8_default_offset() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint8");
|
||||
let root = byte_union_root(0, "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [5u8, 0xAA, 0xBB, 0xCC];
|
||||
let d = read_byte_discriminator(&buf, &schema, LE).expect("read");
|
||||
let d = read_byte_discriminator(&buf, &u, LE).expect("read");
|
||||
assert_eq!(d.key, "5");
|
||||
assert_eq!(d.variant_offset, 1);
|
||||
assert_eq!(d.discriminator_size, 1);
|
||||
@@ -371,19 +319,21 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint8_big_endian() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint8");
|
||||
let root = byte_union_root(0, "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [6u8];
|
||||
let d = read_byte_discriminator(&buf, &schema, BE).expect("read");
|
||||
let d = read_byte_discriminator(&buf, &u, BE).expect("read");
|
||||
assert_eq!(d.key, "6");
|
||||
assert_eq!(d.variant_offset, 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint16_little_endian() {
|
||||
let schema = byte_union_schema(2, "AlkType:Uint16");
|
||||
let root = byte_union_root(2, "uint16");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 4];
|
||||
buf[2..4].copy_from_slice(&5u16.to_le_bytes());
|
||||
let d = read_byte_discriminator(&buf, &schema, LE).expect("read");
|
||||
let d = read_byte_discriminator(&buf, &u, LE).expect("read");
|
||||
assert_eq!(d.key, "5");
|
||||
assert_eq!(d.variant_offset, 4);
|
||||
assert_eq!(d.discriminator_size, 2);
|
||||
@@ -391,19 +341,21 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint16_big_endian() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint16");
|
||||
let root = byte_union_root(0, "uint16");
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [0x00, 0x06, 0xAA, 0xBB];
|
||||
let d = read_byte_discriminator(&buf, &schema, BE).expect("read");
|
||||
let d = read_byte_discriminator(&buf, &u, BE).expect("read");
|
||||
assert_eq!(d.key, "6");
|
||||
assert_eq!(d.variant_offset, 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint32_little_endian() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint32");
|
||||
let root = byte_union_root(0, "uint32");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..4].copy_from_slice(&5u32.to_le_bytes());
|
||||
let d = read_byte_discriminator(&buf, &schema, LE).expect("read");
|
||||
let d = read_byte_discriminator(&buf, &u, LE).expect("read");
|
||||
assert_eq!(d.key, "5");
|
||||
assert_eq!(d.variant_offset, 4);
|
||||
assert_eq!(d.discriminator_size, 4);
|
||||
@@ -411,19 +363,21 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint32_big_endian() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint32");
|
||||
let root = byte_union_root(0, "uint32");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..4].copy_from_slice(&6u32.to_be_bytes());
|
||||
let d = read_byte_discriminator(&buf, &schema, BE).expect("read");
|
||||
let d = read_byte_discriminator(&buf, &u, BE).expect("read");
|
||||
assert_eq!(d.key, "6");
|
||||
assert_eq!(d.variant_offset, 4);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_unknown_value_is_access_error() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint8");
|
||||
let root = byte_union_root(0, "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [99u8];
|
||||
let err = read_byte_discriminator(&buf, &schema, LE).unwrap_err();
|
||||
let err = read_byte_discriminator(&buf, &u, LE).unwrap_err();
|
||||
match err {
|
||||
AlkTypeError::Access { field_path, reason } => {
|
||||
assert_eq!(field_path, DISCRIMINATOR_PATH);
|
||||
@@ -435,29 +389,32 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_buffer_too_short_is_access_error() {
|
||||
let schema = byte_union_schema(4, "AlkType:Uint32");
|
||||
let root = byte_union_root(4, "uint32");
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [0u8; 2];
|
||||
let err = read_byte_discriminator(&buf, &schema, LE).unwrap_err();
|
||||
let err = read_byte_discriminator(&buf, &u, LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_field_kind_is_schema_error() {
|
||||
let schema = field_union_schema("type", "AlkType:String");
|
||||
let root = field_union_root("type", "string");
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [0u8; 16];
|
||||
let err = read_byte_discriminator(&buf, &schema, LE).unwrap_err();
|
||||
let err = read_byte_discriminator(&buf, &u, LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_string() {
|
||||
let schema = field_union_schema("type", "AlkType:String");
|
||||
let root = field_union_root("type", "string");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 32];
|
||||
let value = "read";
|
||||
let len_bytes = (value.len() as u32).to_le_bytes();
|
||||
buf[0..4].copy_from_slice(&len_bytes);
|
||||
buf[4..4 + value.len()].copy_from_slice(value.as_bytes());
|
||||
let d = read_field_discriminator(&buf, &schema, 0, LE).expect("read");
|
||||
let d = read_field_discriminator(&buf, &u, 0, LE).expect("read");
|
||||
assert_eq!(d.key, "read");
|
||||
assert_eq!(d.variant_offset, 4 + value.len());
|
||||
assert_eq!(d.discriminator_size, 4 + value.len());
|
||||
@@ -465,45 +422,109 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_uint8() {
|
||||
let schema = field_union_schema("type", "AlkType:Uint8");
|
||||
let root = field_union_root("type", "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0] = 0;
|
||||
let d = read_field_discriminator(&buf, &schema, 0, LE).expect("read");
|
||||
let d = read_field_discriminator(&buf, &u, 0, LE).expect("read");
|
||||
assert_eq!(d.key, "0");
|
||||
assert_eq!(d.variant_offset, 1);
|
||||
assert_eq!(d.discriminator_size, 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_enum() {
|
||||
let schema = field_union_schema("type", "AlkType:Enum");
|
||||
fn n1_read_field_discriminator_uint16_little_endian() {
|
||||
// N1: tunion now accepts the same field-disc kinds as the plan
|
||||
// reader (string/uint8/uint16/uint32/enum) — uint16/uint32 were
|
||||
// previously rejected with a Schema error, diverging from
|
||||
// `plan_discriminator_string_value`.
|
||||
let root = field_union_root("type", "uint16");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..4].copy_from_slice(&0u32.to_le_bytes());
|
||||
let d = read_field_discriminator(&buf, &schema, 0, LE).expect("read");
|
||||
assert_eq!(d.key, "0");
|
||||
buf[0..2].copy_from_slice(&1u16.to_le_bytes());
|
||||
let d = read_field_discriminator(&buf, &u, 0, LE).expect("read");
|
||||
assert_eq!(d.key, "1");
|
||||
assert_eq!(d.variant_offset, 2);
|
||||
assert_eq!(d.discriminator_size, 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn n1_read_field_discriminator_uint16_big_endian() {
|
||||
let root = field_union_root("type", "uint16");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..2].copy_from_slice(&1u16.to_be_bytes());
|
||||
let d = read_field_discriminator(&buf, &u, 0, BE).expect("read");
|
||||
assert_eq!(d.key, "1");
|
||||
assert_eq!(d.discriminator_size, 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn n1_read_field_discriminator_uint32_little_endian() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "uint32" } ],
|
||||
"mapping": {
|
||||
"1024": { "kind": "struct", "fields": [ { "name": "n", "kind": "uint32" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..4].copy_from_slice(&1024u32.to_le_bytes());
|
||||
let d = read_field_discriminator(&buf, &u, 0, LE).expect("read");
|
||||
assert_eq!(d.key, "1024");
|
||||
assert_eq!(d.variant_offset, 4);
|
||||
assert_eq!(d.discriminator_size, 4);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn n1_read_field_discriminator_uint32_big_endian() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "uint32" } ],
|
||||
"mapping": {
|
||||
"1024": { "kind": "struct", "fields": [ { "name": "n", "kind": "uint32" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..4].copy_from_slice(&1024u32.to_be_bytes());
|
||||
let d = read_field_discriminator(&buf, &u, 0, BE).expect("read");
|
||||
assert_eq!(d.key, "1024");
|
||||
assert_eq!(d.discriminator_size, 4);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_string_big_endian() {
|
||||
let schema = field_union_schema("type", "AlkType:String");
|
||||
let root = field_union_root("type", "string");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 32];
|
||||
let value = "write";
|
||||
let len_bytes = (value.len() as u32).to_be_bytes();
|
||||
buf[0..4].copy_from_slice(&len_bytes);
|
||||
buf[4..4 + value.len()].copy_from_slice(value.as_bytes());
|
||||
let d = read_field_discriminator(&buf, &schema, 0, BE).expect("read");
|
||||
let d = read_field_discriminator(&buf, &u, 0, BE).expect("read");
|
||||
assert_eq!(d.key, "write");
|
||||
assert_eq!(d.variant_offset, 4 + value.len());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_unknown_value_is_access_error() {
|
||||
let schema = field_union_schema("type", "AlkType:Uint8");
|
||||
let root = field_union_root("type", "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0] = 99;
|
||||
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
|
||||
let err = read_field_discriminator(&buf, &u, 0, LE).unwrap_err();
|
||||
match err {
|
||||
AlkTypeError::Access { field_path, reason } => {
|
||||
assert_eq!(field_path, "type");
|
||||
@@ -515,131 +536,123 @@ mod tests {
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_field_not_found_is_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "missing"},
|
||||
"properties": {"other": {"AlkType:Uint8": true}},
|
||||
"mapping": {"5": {"AlkType:Struct": true}}
|
||||
// H3 enforcement: the missing-discriminator-field case is now
|
||||
// rejected at parse (BastUnion::parse), before any dispatch
|
||||
// reader can see the union.
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "missing" },
|
||||
"fields": [ { "name": "other", "kind": "uint8" } ],
|
||||
"mapping": { "5": { "kind": "struct", "fields": [] } }
|
||||
}
|
||||
}
|
||||
});
|
||||
let buf = [0u8; 4];
|
||||
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_no_alktype_kind_is_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "type"},
|
||||
"properties": {"type": {"type": "string"}},
|
||||
"mapping": {"read": {"AlkType:Struct": true}}
|
||||
});
|
||||
let buf = [0u8; 4];
|
||||
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
let err = BastDoc::new(&root, "U").unwrap_err();
|
||||
match err {
|
||||
AlkTypeError::Schema(reason) => {
|
||||
assert!(reason.contains("no field"), "reason: {reason}");
|
||||
}
|
||||
other => panic!("expected Schema error, got {other:?}"),
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_unsupported_kind_is_schema_error() {
|
||||
let schema = field_union_schema("type", "AlkType:Float32");
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "float32" } ],
|
||||
"mapping": { "read": { "kind": "struct", "fields": [] } }
|
||||
}
|
||||
}
|
||||
});
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [0u8; 8];
|
||||
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
|
||||
let err = read_field_discriminator(&buf, &u, 0, LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_byte_kind_is_schema_error() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint8");
|
||||
let root = byte_union_root(0, "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let buf = [5u8];
|
||||
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
|
||||
let err = read_field_discriminator(&buf, &u, 0, LE).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn resolve_variant_inline_schema() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint8");
|
||||
let variant = resolve_variant(&schema, "5").expect("resolve");
|
||||
assert_eq!(
|
||||
variant.get("AlkType:Struct").and_then(Value::as_bool),
|
||||
Some(true)
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn resolve_variant_ref_against_own_defs() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte"},
|
||||
"mapping": {
|
||||
"5": {"$ref": "#/$defs/Read"}
|
||||
},
|
||||
"$defs": {
|
||||
"Read": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}}
|
||||
}
|
||||
});
|
||||
let variant = resolve_variant(&schema, "5").expect("resolve");
|
||||
assert_eq!(
|
||||
variant.get("AlkType:Struct").and_then(Value::as_bool),
|
||||
Some(true)
|
||||
);
|
||||
fn resolve_variant_returns_variant_type() {
|
||||
let root = byte_union_root(0, "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let variant = resolve_variant(&u, "5").expect("resolve");
|
||||
assert!(matches!(variant, BastType::Struct(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn resolve_variant_unknown_key_is_schema_error() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint8");
|
||||
let err = resolve_variant(&schema, "999").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn resolve_variant_missing_mapping_is_schema_error() {
|
||||
let schema = json!({"AlkType:Union": true, "discriminator": {"kind": "byte"}});
|
||||
let err = resolve_variant(&schema, "5").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn resolve_variant_unresolvable_ref_is_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte"},
|
||||
"mapping": {
|
||||
"5": {"$ref": "#/$defs/Read"}
|
||||
}
|
||||
});
|
||||
let err = resolve_variant(&schema, "5").unwrap_err();
|
||||
let root = byte_union_root(0, "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
let err = resolve_variant(&u, "999").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn discriminator_size_uint8() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint8");
|
||||
assert_eq!(discriminator_size(&schema).unwrap(), 1);
|
||||
let root = byte_union_root(0, "uint8");
|
||||
let u = doc_union(&root, "U");
|
||||
assert_eq!(discriminator_size(&u).unwrap(), 1);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn discriminator_size_uint16() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint16");
|
||||
assert_eq!(discriminator_size(&schema).unwrap(), 2);
|
||||
let root = byte_union_root(0, "uint16");
|
||||
let u = doc_union(&root, "U");
|
||||
assert_eq!(discriminator_size(&u).unwrap(), 2);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn discriminator_size_uint32() {
|
||||
let schema = byte_union_schema(0, "AlkType:Uint32");
|
||||
assert_eq!(discriminator_size(&schema).unwrap(), 4);
|
||||
let root = byte_union_root(0, "uint32");
|
||||
let u = doc_union(&root, "U");
|
||||
assert_eq!(discriminator_size(&u).unwrap(), 4);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn discriminator_size_field_kind_is_schema_error() {
|
||||
let schema = field_union_schema("type", "AlkType:String");
|
||||
let err = discriminator_size(&schema).unwrap_err();
|
||||
let root = field_union_root("type", "string");
|
||||
let u = doc_union(&root, "U");
|
||||
let err = discriminator_size(&u).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn discriminator_size_missing_discriminator_is_schema_error() {
|
||||
let schema = json!({"AlkType:Union": true});
|
||||
let err = discriminator_size(&schema).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
fn l2_read_field_discriminator_enum_dispatches_on_index() {
|
||||
// Review #007 L2: the enum arm of read_field_discriminator had
|
||||
// zero executions — N1's parity tests covered uint16/uint32 but
|
||||
// skipped enum, the last arm of the documented kind set.
|
||||
let root = field_union_root("kind", "enum");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..4].copy_from_slice(&0u32.to_le_bytes());
|
||||
let d = read_field_discriminator(&buf, &u, 0, LE).expect("read");
|
||||
assert_eq!(d.key, "0");
|
||||
assert_eq!(d.variant_offset, 4);
|
||||
assert_eq!(d.discriminator_size, 4);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn l2_read_field_discriminator_enum_big_endian() {
|
||||
let root = field_union_root("kind", "enum");
|
||||
let u = doc_union(&root, "U");
|
||||
let mut buf = vec![0u8; 8];
|
||||
buf[0..4].copy_from_slice(&1u32.to_be_bytes());
|
||||
let d = read_field_discriminator(&buf, &u, 0, Endian::Big).expect("read");
|
||||
assert_eq!(d.key, "1");
|
||||
assert_eq!(d.variant_offset, 4);
|
||||
}
|
||||
}
|
||||
+101
-1106
File diff suppressed because it is too large.
Load diff
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,241 @@
|
||||
//! Shared reference-graph guard for the standalone schema walkers.
|
||||
//!
|
||||
//! [`check_ref_graph`] is the one-shot pre-walk `AlkTypeEngine::compile`
|
||||
//! runs before any layout builder, and the one
|
||||
//! [`crate::offset_map::OffsetMap::compute`],
|
||||
//! [`crate::layout_builder::LayoutBuilder::new`], and
|
||||
//! [`crate::materialize::materialize_aligned`] run at their own entry so
|
||||
//! each is untrusted-input-safe when called without the engine. It
|
||||
//! rejects reference graphs that nest deeper than [`MAX_GRAPH_DEPTH`] or
|
||||
//! that contain a `$ref` cycle — the two shapes that would otherwise
|
||||
//! overflow the walkers' unguarded struct/union recursion (review #006
|
||||
//! H2; AGENTS.md §3: a cyclic or adversarially deep schema must produce
|
||||
//! a handleable error, not a stack overflow).
|
||||
//!
|
||||
//! The compiled forms keep their own inline guards (their compile walks
|
||||
//! inline types too, so a graph check alone is not enough for them); the
|
||||
//! error text intentionally mirrors theirs ("cyclic $ref through…",
|
||||
//! "compile depth exceeded…") so downstream matching sees one shape.
|
||||
|
||||
use crate::bast::{BastDefKind, BastDoc, BastStruct, BastType};
|
||||
use crate::error::AlkTypeError;
|
||||
use std::collections::BTreeSet;
|
||||
|
||||
/// Maximum `$ref`-graph depth accepted by [`check_ref_graph`]. Matches
|
||||
/// the plan compilers' `MAX_COMPILE_DEPTH` (128) so every walker rejects
|
||||
/// the same documents.
|
||||
pub(crate) const MAX_GRAPH_DEPTH: usize = 128;
|
||||
|
||||
pub(crate) fn depth_err(path: &str) -> AlkTypeError {
|
||||
AlkTypeError::Schema(format!(
|
||||
"schema walk: compile depth exceeded {MAX_GRAPH_DEPTH} at {path} \
|
||||
(cyclic $ref or adversarially deep nesting)"
|
||||
))
|
||||
}
|
||||
|
||||
pub(crate) fn cycle_err(name: &str, path: &str) -> AlkTypeError {
|
||||
AlkTypeError::Schema(format!(
|
||||
"schema walk: cyclic $ref through {name:?} at {path}"
|
||||
))
|
||||
}
|
||||
|
||||
/// Reject cyclic or over-deep `$ref` graphs before a recursive walker
|
||||
/// sees the document.
|
||||
///
|
||||
/// One walk over the reachable definitions: inline structs recurse into
|
||||
/// their fields; named defs are entered with the path-scoped cycle set
|
||||
/// (a def currently being expanded) and the depth counter. Diamond
|
||||
/// references (two fields `$ref`-ing the same def, neither nested inside
|
||||
/// the other) are allowed — `seen` is removed on exit, so only genuine
|
||||
/// cycles trip it, the same semantics the plan compilers use.
|
||||
pub(crate) fn check_ref_graph(doc: &BastDoc) -> Result<(), AlkTypeError> {
|
||||
let root_def = doc.root_def();
|
||||
let mut seen = BTreeSet::new();
|
||||
let path = root_def.name().to_string();
|
||||
match root_def.kind() {
|
||||
BastDefKind::Struct(s) => check_struct(doc, s, &path, 0, &mut seen),
|
||||
BastDefKind::Union(u) => {
|
||||
check_typeref_list(
|
||||
doc,
|
||||
u.mapping().iter().map(|(_, ty)| ty),
|
||||
&path,
|
||||
0,
|
||||
&mut seen,
|
||||
)?;
|
||||
check_field_list(doc, u.fields(), &path, 0, &mut seen)
|
||||
}
|
||||
BastDefKind::Enum(_) => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
fn check_struct(
|
||||
doc: &BastDoc,
|
||||
s: &BastStruct,
|
||||
path: &str,
|
||||
depth: usize,
|
||||
seen: &mut BTreeSet<String>,
|
||||
) -> Result<(), AlkTypeError> {
|
||||
if depth > MAX_GRAPH_DEPTH {
|
||||
return Err(depth_err(path));
|
||||
}
|
||||
check_field_list(doc, s.fields(), path, depth, seen)
|
||||
}
|
||||
|
||||
fn check_field_list(
|
||||
doc: &BastDoc,
|
||||
fields: &[crate::bast::BastField],
|
||||
path: &str,
|
||||
depth: usize,
|
||||
seen: &mut BTreeSet<String>,
|
||||
) -> Result<(), AlkTypeError> {
|
||||
for field in fields {
|
||||
let field_path = format!("{path}.{}", field.name());
|
||||
check_typeref(doc, field.ty(), &field_path, depth, seen)?;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn check_typeref(
|
||||
doc: &BastDoc,
|
||||
ty: &BastType,
|
||||
path: &str,
|
||||
depth: usize,
|
||||
seen: &mut BTreeSet<String>,
|
||||
) -> Result<(), AlkTypeError> {
|
||||
if depth > MAX_GRAPH_DEPTH {
|
||||
return Err(depth_err(path));
|
||||
}
|
||||
match ty {
|
||||
BastType::Ref(r) => {
|
||||
let name = r.name();
|
||||
if !seen.insert(name.to_string()) {
|
||||
return Err(cycle_err(name, path));
|
||||
}
|
||||
let def = doc.resolve_ref(r)?;
|
||||
let def_path = format!("{path} -> {name}");
|
||||
let out = match def.kind() {
|
||||
BastDefKind::Struct(s) => {
|
||||
check_struct(doc, s, &def_path, depth + 1, seen)
|
||||
}
|
||||
BastDefKind::Union(u) => {
|
||||
check_typeref_list(doc, u.mapping().iter().map(|(_, ty)| ty), &def_path, depth + 1, seen)?;
|
||||
check_field_list(doc, u.fields(), &def_path, depth + 1, seen)
|
||||
}
|
||||
BastDefKind::Enum(_) => Ok(()),
|
||||
};
|
||||
seen.remove(name);
|
||||
out
|
||||
}
|
||||
BastType::Struct(s) => check_struct(doc, s, path, depth + 1, seen),
|
||||
BastType::Union(u) => {
|
||||
check_typeref_list(doc, u.mapping().iter().map(|(_, ty)| ty), path, depth + 1, seen)?;
|
||||
check_field_list(doc, u.fields(), path, depth + 1, seen)
|
||||
}
|
||||
BastType::Array(a) => {
|
||||
check_typeref(doc, a.element(), path, depth + 1, seen)
|
||||
}
|
||||
BastType::Record(r) => check_typeref(doc, r.values(), path, depth + 1, seen),
|
||||
BastType::Primitive(_) | BastType::Enum(_) => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
fn check_typeref_list<'a, I>(
|
||||
doc: &BastDoc,
|
||||
tys: I,
|
||||
path: &str,
|
||||
depth: usize,
|
||||
seen: &mut BTreeSet<String>,
|
||||
) -> Result<(), AlkTypeError>
|
||||
where
|
||||
I: IntoIterator<Item = &'a BastType>,
|
||||
{
|
||||
for ty in tys {
|
||||
check_typeref(doc, ty, path, depth, seen)?;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use serde_json::json;
|
||||
|
||||
#[test]
|
||||
fn self_cycle_rejected() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "me", "kind": { "$ref": "#/$defs/S" } }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "S").expect("cyclic doc parses");
|
||||
let err = check_ref_graph(&doc).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
assert!(err.to_string().contains("cyclic"), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn two_def_cycle_rejected() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"A": { "kind": "struct", "fields": [
|
||||
{ "name": "next", "kind": { "$ref": "#/$defs/B" } }
|
||||
] },
|
||||
"B": { "kind": "struct", "fields": [
|
||||
{ "name": "back", "kind": { "$ref": "#/$defs/A" } }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "A").expect("cyclic doc parses");
|
||||
let err = check_ref_graph(&doc).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn diamond_refs_allowed() {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": { "kind": "struct", "fields": [
|
||||
{ "name": "a", "kind": { "$ref": "#/$defs/Point" } },
|
||||
{ "name": "b", "kind": { "$ref": "#/$defs/Point" } }
|
||||
] },
|
||||
"Point": { "kind": "struct", "fields": [
|
||||
{ "name": "x", "kind": "uint16" }
|
||||
] }
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "S").expect("doc");
|
||||
assert!(check_ref_graph(&doc).is_ok());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn deep_ref_chain_rejected_not_overflow() {
|
||||
// 201 named defs chained by refs — depth 201 exceeds the cap of
|
||||
// 128, but the check itself must complete (bounded stack, clean
|
||||
// error). Note the chain uses distinct defs, so the cycle set
|
||||
// never trips; only the depth cap stops it.
|
||||
let mut defs = serde_json::Map::new();
|
||||
defs.insert(
|
||||
"L200".to_string(),
|
||||
json!({ "kind": "struct", "fields": [ { "name": "v", "kind": "uint8" } ] }),
|
||||
);
|
||||
for i in (0..200).rev() {
|
||||
defs.insert(
|
||||
format!("L{i}"),
|
||||
json!({
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "next", "kind": { "$ref": format!("#/$defs/L{}", i + 1) } } ]
|
||||
}),
|
||||
);
|
||||
}
|
||||
let root = json!({ "$defs": defs });
|
||||
let doc = BastDoc::new(&root, "L0").expect("deep doc parses (no cycle)");
|
||||
let err = check_ref_graph(&doc).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
assert!(
|
||||
err.to_string().contains("depth exceeded"),
|
||||
"expected depth error, got {err:?}"
|
||||
);
|
||||
}
|
||||
}
|
||||
+161
-170
@@ -4,27 +4,34 @@
|
||||
//! accessors, validation convenience methods, and the aligned-mode
|
||||
//! `read_field` / `write_field` round-trip for the fixed-size primitive
|
||||
//! kinds and length-prefixed `String` / `Bytes`.
|
||||
//!
|
||||
//! All schemas are BAST documents (`{ "$defs": { ... } }` with `kind`-
|
||||
//! based vocabulary).
|
||||
|
||||
use alktype::*;
|
||||
use serde_json::json;
|
||||
|
||||
fn mixed_fixed_struct_schema() -> serde_json::Value {
|
||||
fn mixed_fixed_struct_doc() -> serde_json::Value {
|
||||
json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "little",
|
||||
"properties": {
|
||||
"flag": { "AlkType:Uint8": true },
|
||||
"id": { "AlkType:Uint32": true },
|
||||
"score": { "AlkType:Float32": true },
|
||||
"tag": { "AlkType:String": true }
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "little",
|
||||
"fields": [
|
||||
{ "name": "flag", "kind": "uint8" },
|
||||
{ "name": "id", "kind": "uint32" },
|
||||
{ "name": "score", "kind": "float32" },
|
||||
{ "name": "tag", "kind": "string" }
|
||||
]
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn compile_aligned_builds_engine_with_offset_map() -> Result<(), AlkTypeError> {
|
||||
let mut schema = mixed_fixed_struct_schema();
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let doc = mixed_fixed_struct_doc();
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
assert_eq!(engine.mode(), LayoutMode::Aligned);
|
||||
assert!(engine.offset_map().is_some());
|
||||
assert!(engine.layout_builder().is_none());
|
||||
@@ -34,8 +41,8 @@ fn compile_aligned_builds_engine_with_offset_map() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn compile_packed_builds_engine_with_builder_and_reader() -> Result<(), AlkTypeError> {
|
||||
let mut schema = mixed_fixed_struct_schema();
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
let doc = mixed_fixed_struct_doc();
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
|
||||
assert_eq!(engine.mode(), LayoutMode::Packed);
|
||||
assert!(engine.offset_map().is_none());
|
||||
assert!(engine.layout_builder().is_some());
|
||||
@@ -44,147 +51,96 @@ fn compile_packed_builds_engine_with_builder_and_reader() -> Result<(), AlkTypeE
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn compile_normalizes_bare_name_refs() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"child": { "$ref": "Child" }
|
||||
},
|
||||
fn compile_resolves_ref_fields() -> Result<(), AlkTypeError> {
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "child", "kind": { "$ref": "#/$defs/Child" } }
|
||||
]
|
||||
},
|
||||
"Child": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "x": { "AlkType:Uint8": true } }
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "x", "kind": "uint8" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let _engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
assert_eq!(
|
||||
schema["properties"]["child"]["$ref"],
|
||||
json!("#/$defs/Child")
|
||||
);
|
||||
let _engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn compile_leaves_full_pointer_refs_unchanged() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"child": { "$ref": "#/$defs/Child" }
|
||||
},
|
||||
"$defs": {
|
||||
"Child": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "x": { "AlkType:Uint8": true } }
|
||||
}
|
||||
}
|
||||
});
|
||||
let _engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
assert_eq!(
|
||||
schema["properties"]["child"]["$ref"],
|
||||
json!("#/$defs/Child")
|
||||
);
|
||||
Ok(())
|
||||
fn compile_returns_schema_error_when_no_defs() {
|
||||
let doc = json!({ "type": "object", "properties": {} });
|
||||
let err = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn compile_returns_schema_error_when_no_alktype_kind() {
|
||||
let mut schema = json!({ "type": "object", "properties": {} });
|
||||
let err = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned).unwrap_err();
|
||||
fn compile_returns_schema_error_for_missing_root() {
|
||||
let doc = json!({ "$defs": { "Other": { "kind": "struct", "fields": [] } } });
|
||||
let err = AlkTypeEngine::compile(&doc, "Missing", LayoutMode::Aligned, None).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn endian_parsed_from_schema_big() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "big",
|
||||
"properties": { "id": { "AlkType:Uint32": true } }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
|
||||
assert_eq!(engine.endian(), Endian::Big);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn endian_defaults_to_little() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "id": { "AlkType:Uint32": true } }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
|
||||
assert_eq!(engine.endian(), Endian::Little);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn validate_json_accepts_valid_instance() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint32": true, "type": "integer" }
|
||||
},
|
||||
"required": ["id"]
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
assert!(engine.validate_json(&json!({"id": 42})).is_ok());
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn validate_json_rejects_invalid_instance() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint32": true, "type": "integer" }
|
||||
},
|
||||
"required": ["id"]
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let err = engine.validate_json(&json!({"id": -1})).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn is_valid_json_returns_bool() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint32": true, "type": "integer" }
|
||||
},
|
||||
"required": ["id"]
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
assert!(engine.is_valid_json(&json!({"id": 42})));
|
||||
assert!(!engine.is_valid_json(&json!({"id": -1})));
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_write_aligned_round_trips_all_fixed_size_kinds() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "little",
|
||||
"properties": {
|
||||
"i8": { "AlkType:Int8": true },
|
||||
"u8": { "AlkType:Uint8": true },
|
||||
"i16": { "AlkType:Int16": true },
|
||||
"u16": { "AlkType:Uint16": true },
|
||||
"i32": { "AlkType:Int32": true },
|
||||
"u32": { "AlkType:Uint32": true },
|
||||
"i64": { "AlkType:Int64": true },
|
||||
"u64": { "AlkType:Uint64": true },
|
||||
"f32": { "AlkType:Float32": true },
|
||||
"f64": { "AlkType:Float64": true },
|
||||
"b": { "AlkType:Boolean": true },
|
||||
"e": { "AlkType:Enum": true }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "little",
|
||||
"fields": [
|
||||
{ "name": "i8", "kind": "int8" },
|
||||
{ "name": "u8", "kind": "uint8" },
|
||||
{ "name": "i16", "kind": "int16" },
|
||||
{ "name": "u16", "kind": "uint16" },
|
||||
{ "name": "i32", "kind": "int32" },
|
||||
{ "name": "u32", "kind": "uint32" },
|
||||
{ "name": "i64", "kind": "int64" },
|
||||
{ "name": "u64", "kind": "uint64" },
|
||||
{ "name": "f32", "kind": "float32" },
|
||||
{ "name": "f64", "kind": "float64" },
|
||||
{ "name": "b", "kind": "bool" },
|
||||
{ "name": "e", "kind": { "$ref": "#/$defs/E" } }
|
||||
]
|
||||
},
|
||||
"E": { "kind": "enum", "values": ["A", "B", "C"] }
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
|
||||
let mut buffer = vec![0u8; offset_map.total_size()];
|
||||
|
||||
@@ -230,13 +186,15 @@ fn read_write_aligned_round_trips_all_fixed_size_kinds() -> Result<(), AlkTypeEr
|
||||
|
||||
#[test]
|
||||
fn read_write_aligned_round_trips_string() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"name": { "AlkType:String": true }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "name", "kind": "string" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
|
||||
let mut buffer = vec![0u8; offset_map.total_size() + 64];
|
||||
engine.write_field(&mut buffer, "name", &FieldValue::String("hello world"))?;
|
||||
@@ -249,13 +207,15 @@ fn read_write_aligned_round_trips_string() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn read_write_aligned_round_trips_bytes() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"blob": { "AlkType:Bytes": true }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "blob", "kind": "bytes" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
|
||||
let payload = b"the quick brown fox".to_vec();
|
||||
let mut buffer = vec![0u8; offset_map.total_size() + payload.len()];
|
||||
@@ -269,11 +229,15 @@ fn read_write_aligned_round_trips_bytes() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn read_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "id": { "AlkType:Uint32": true } }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
|
||||
let buffer = [0u8; 4];
|
||||
let err = engine.read_field(&buffer, "id").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
@@ -282,11 +246,15 @@ fn read_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError>
|
||||
|
||||
#[test]
|
||||
fn write_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "id": { "AlkType:Uint32": true } }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
|
||||
let mut buffer = [0u8; 4];
|
||||
let err = engine
|
||||
.write_field(&mut buffer, "id", &FieldValue::U32(1))
|
||||
@@ -297,11 +265,15 @@ fn write_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError>
|
||||
|
||||
#[test]
|
||||
fn read_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "id": { "AlkType:Uint32": true } }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let buffer = [0u8; 8];
|
||||
let err = engine.read_field(&buffer, "missing").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Offset { .. }), "got {err:?}");
|
||||
@@ -310,11 +282,15 @@ fn read_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeError
|
||||
|
||||
#[test]
|
||||
fn write_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "id": { "AlkType:Uint32": true } }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let mut buffer = [0u8; 8];
|
||||
let err = engine
|
||||
.write_field(&mut buffer, "missing", &FieldValue::U32(1))
|
||||
@@ -324,30 +300,38 @@ fn write_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeErro
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_returns_access_error_for_composite_types() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"vals": {
|
||||
"AlkType:Array": true,
|
||||
"items": { "AlkType:Uint32": true }
|
||||
fn read_field_returns_error_for_composite_types() -> Result<(), AlkTypeError> {
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "vals", "kind": { "kind": "array", "element": "uint32", "count": 2 } }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let buffer = [0u8; 8];
|
||||
let err = engine.read_field(&buffer, "vals").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
assert!(
|
||||
matches!(err, AlkTypeError::Access { .. } | AlkTypeError::Offset { .. }),
|
||||
"got {err:?}"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn write_field_returns_access_error_for_composite_value() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "id": { "AlkType:Uint32": true } }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let mut buffer = [0u8; 8];
|
||||
let err = engine
|
||||
.write_field(&mut buffer, "id", &FieldValue::Struct { start: 0, end: 4 })
|
||||
@@ -358,19 +342,26 @@ fn write_field_returns_access_error_for_composite_value() -> Result<(), AlkTypeE
|
||||
|
||||
#[test]
|
||||
fn read_field_aligned_reads_nested_struct_byte_range() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"header": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"version": { "AlkType:Uint8": true },
|
||||
"magic": { "AlkType:Uint32": true }
|
||||
}
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{
|
||||
"name": "header",
|
||||
"kind": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "version", "kind": "uint8" },
|
||||
{ "name": "magic", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let offset_map = engine.offset_map().expect("aligned mode");
|
||||
let mut buffer = vec![0u8; offset_map.total_size()];
|
||||
|
||||
@@ -386,4 +377,4 @@ fn read_field_aligned_reads_nested_struct_byte_range() -> Result<(), AlkTypeErro
|
||||
FieldValue::U32(0xCAFEBABE)
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
+165
-138
@@ -2,10 +2,13 @@
|
||||
//!
|
||||
//! Exercises the `AlkTypeError` variants across the crate:
|
||||
//! `Access` (buffer too short, invalid UTF-8, invalid boolean byte,
|
||||
//! unknown discriminator value), `Schema` (missing AlkType kind,
|
||||
//! malformed discriminator annotation), and `Offset` (missing
|
||||
//! variable-length field size in `LayoutBuilder::build`).
|
||||
//! unknown discriminator value), `Schema` (missing root, malformed
|
||||
//! discriminator annotation), and `Offset` (missing variable-length
|
||||
//! field size in `LayoutBuilder::build`).
|
||||
//!
|
||||
//! All schemas are BAST documents.
|
||||
|
||||
use alktype::bast::BastDoc;
|
||||
use alktype::data_access;
|
||||
use alktype::tunion;
|
||||
use alktype::*;
|
||||
@@ -149,128 +152,145 @@ fn write_string_buffer_too_short_returns_access_error() {
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn compile_missing_alktype_kind_returns_schema_error() {
|
||||
let mut schema = json!({ "type": "object", "properties": {} });
|
||||
let err = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned).unwrap_err();
|
||||
fn compile_missing_defs_returns_schema_error() {
|
||||
let doc = json!({ "type": "object", "properties": {} });
|
||||
let err = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn offset_map_compute_missing_alktype_kind_returns_schema_error() {
|
||||
let schema = json!({ "type": "object", "properties": {} });
|
||||
let err = OffsetMap::compute(&schema).unwrap_err();
|
||||
fn compile_missing_root_returns_schema_error() {
|
||||
let doc = json!({ "$defs": { "Other": { "kind": "struct", "fields": [] } } });
|
||||
let err = AlkTypeEngine::compile(&doc, "Missing", LayoutMode::Aligned, None).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn offset_map_compute_non_struct_top_level_returns_schema_error() {
|
||||
let schema = json!({ "AlkType:Uint32": true });
|
||||
let err = OffsetMap::compute(&schema).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": { "1": { "$ref": "#/$defs/A" } }
|
||||
},
|
||||
"A": { "kind": "struct", "fields": [] }
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "U").expect("doc");
|
||||
let err = OffsetMap::compute(&doc).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn layout_builder_new_missing_alktype_kind_returns_schema_error() {
|
||||
let schema = json!({ "type": "object", "properties": {} });
|
||||
let err = LayoutBuilder::new(&schema).unwrap_err();
|
||||
fn layout_builder_new_missing_root_returns_schema_error() {
|
||||
let root = json!({ "$defs": { "Other": { "kind": "struct", "fields": [] } } });
|
||||
let err = LayoutBuilder::new(&root, "Missing").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn layout_builder_new_non_struct_top_level_returns_schema_error() {
|
||||
let schema = json!({ "AlkType:Uint32": true });
|
||||
let err = LayoutBuilder::new(&schema).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_discriminator_missing_returns_schema_error() {
|
||||
let schema = json!({"AlkType:Union": true});
|
||||
let err = parse_discriminator(&schema).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_discriminator_field_missing_name_returns_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field"}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": { "1": { "$ref": "#/$defs/A" } }
|
||||
},
|
||||
"A": { "kind": "struct", "fields": [] }
|
||||
}
|
||||
});
|
||||
let err = parse_discriminator(&schema).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_discriminator_unknown_kind_returns_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "magic"}
|
||||
});
|
||||
let err = parse_discriminator(&schema).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_discriminator_byte_invalid_type_returns_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Float32"}
|
||||
});
|
||||
let err = parse_discriminator(&schema).unwrap_err();
|
||||
let err = LayoutBuilder::new(&root, "U").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
|
||||
"mapping": {"5": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "type": "uint8" },
|
||||
"mapping": {
|
||||
"5": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "U")?;
|
||||
let union_def = match doc.root_def().kind() {
|
||||
alktype::bast::BastDefKind::Union(u) => u,
|
||||
_ => unreachable!(),
|
||||
};
|
||||
let buffer = [99u8, 0x00, 0x00];
|
||||
let err = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Little).unwrap_err();
|
||||
let err = tunion::read_byte_discriminator(&buffer, union_def, Endian::Little).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_buffer_too_short_returns_access_error() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "offset": 4, "type": "AlkType:Uint32"},
|
||||
"mapping": {"5": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 4, "type": "uint32" },
|
||||
"mapping": {
|
||||
"5": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "U")?;
|
||||
let union_def = match doc.root_def().kind() {
|
||||
alktype::bast::BastDefKind::Union(u) => u,
|
||||
_ => unreachable!(),
|
||||
};
|
||||
let buffer = [0u8; 2];
|
||||
let err = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Little).unwrap_err();
|
||||
let err = tunion::read_byte_discriminator(&buffer, union_def, Endian::Little).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "type"},
|
||||
"properties": {"type": {"AlkType:Uint8": true}},
|
||||
"mapping": {"0": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "uint8" } ],
|
||||
"mapping": {
|
||||
"0": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "U")?;
|
||||
let union_def = match doc.root_def().kind() {
|
||||
alktype::bast::BastDefKind::Union(u) => u,
|
||||
_ => unreachable!(),
|
||||
};
|
||||
let mut buffer = vec![0u8; 8];
|
||||
buffer[0] = 99;
|
||||
let err =
|
||||
tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little).unwrap_err();
|
||||
tunion::read_field_discriminator(&buffer, union_def, 0, Endian::Little).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn layout_builder_missing_var_size_returns_offset_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"name": { "AlkType:String": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "name", "kind": "string" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let builder = LayoutBuilder::new(&schema).expect("builder");
|
||||
let builder = LayoutBuilder::new(&root, "S").expect("builder");
|
||||
let empty: HashMap<String, usize> = HashMap::new();
|
||||
let err = builder.build(&empty).unwrap_err();
|
||||
match err {
|
||||
@@ -285,39 +305,28 @@ fn layout_builder_missing_var_size_returns_offset_error() {
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn layout_builder_missing_array_data_size_returns_offset_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"vals": {
|
||||
"AlkType:Array": true,
|
||||
"items": { "AlkType:Uint32": true }
|
||||
}
|
||||
}
|
||||
});
|
||||
let builder = LayoutBuilder::new(&schema).expect("builder");
|
||||
let empty: HashMap<String, usize> = HashMap::new();
|
||||
let err = builder.build(&empty).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Offset { .. }), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn layout_builder_missing_discriminator_value_returns_offset_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"payload": {
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
|
||||
"mapping": {"5": {"$ref": "#/$defs/Read"}}
|
||||
}
|
||||
},
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Read": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "payload", "kind": { "$ref": "#/$defs/Packet" } }
|
||||
]
|
||||
},
|
||||
"Packet": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "type": "uint8" },
|
||||
"mapping": { "5": { "$ref": "#/$defs/Read" } }
|
||||
},
|
||||
"Read": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "x", "kind": "uint8" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let builder = LayoutBuilder::new(&schema).expect("builder");
|
||||
let builder = LayoutBuilder::new(&root, "S").expect("builder");
|
||||
let empty: HashMap<String, usize> = HashMap::new();
|
||||
let err = builder.build(&empty).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Offset { .. }), "got {err:?}");
|
||||
@@ -325,20 +334,26 @@ fn layout_builder_missing_discriminator_value_returns_offset_error() {
|
||||
|
||||
#[test]
|
||||
fn layout_builder_unknown_discriminator_value_returns_offset_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"payload": {
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
|
||||
"mapping": {"5": {"$ref": "#/$defs/Read"}}
|
||||
}
|
||||
},
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Read": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "payload", "kind": { "$ref": "#/$defs/Packet" } }
|
||||
]
|
||||
},
|
||||
"Packet": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "type": "uint8" },
|
||||
"mapping": { "5": { "$ref": "#/$defs/Read" } }
|
||||
},
|
||||
"Read": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "x", "kind": "uint8" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let builder = LayoutBuilder::new(&schema).expect("builder");
|
||||
let builder = LayoutBuilder::new(&root, "S").expect("builder");
|
||||
let mut vs = HashMap::new();
|
||||
vs.insert("payload.__discriminator".to_string(), 99);
|
||||
let err = builder.build(&vs).unwrap_err();
|
||||
@@ -352,64 +367,76 @@ fn layout_builder_unknown_discriminator_value_returns_offset_error() {
|
||||
|
||||
#[test]
|
||||
fn sequential_reader_buffer_too_short_returns_access_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint32": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "id", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let buffer = [0u8; 2];
|
||||
let mut reader = SequentialReader::new(&schema).unwrap();
|
||||
let plan = ReadPlan::compile(&root, "S").unwrap();
|
||||
let mut reader = SequentialReader::new(std::sync::Arc::new(plan));
|
||||
let err = reader.read_next(&buffer).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sequential_reader_unknown_field_returns_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": { "a": { "AlkType:Uint8": true } }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "a", "kind": "uint8" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let buffer = [0u8; 4];
|
||||
let mut reader = SequentialReader::new(&schema).unwrap();
|
||||
let plan = ReadPlan::compile(&root, "S").unwrap();
|
||||
let mut reader = SequentialReader::new(std::sync::Arc::new(plan));
|
||||
let err = reader.read_field(&buffer, "missing").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn sequential_reader_new_non_struct_returns_schema_error() {
|
||||
let schema = json!({ "AlkType:Uint32": true });
|
||||
let err = SequentialReader::new(&schema).unwrap_err();
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": { "1": { "$ref": "#/$defs/A" } }
|
||||
},
|
||||
"A": { "kind": "struct", "fields": [] }
|
||||
}
|
||||
});
|
||||
let err = ReadPlan::compile(&root, "U").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_string_indirect_data_region_too_short_returns_access_error() {
|
||||
fn read_string_indirect_data_too_short_returns_access_error() {
|
||||
let mut index = [0u8; 8];
|
||||
let _ = data_access::write_u32(&mut index, 0, 100, "idx.off", Endian::Little);
|
||||
let _ = data_access::write_u32(&mut index, 4, 10, "idx.len", Endian::Little);
|
||||
let data_region = b"too short";
|
||||
let err = data_access::read_bytes_indirect(&index, 0, data_region, "blob", Endian::Little)
|
||||
.unwrap_err();
|
||||
let err = data_access::read_bytes_indirect(&index, 0, "blob", Endian::Little).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_bytes_indirect_index_too_short_returns_access_error() {
|
||||
let buffer = [0u8; 4];
|
||||
let data_region = b"anything";
|
||||
let err = data_access::read_bytes_indirect(&buffer, 0, data_region, "blob", Endian::Little)
|
||||
.unwrap_err();
|
||||
let err = data_access::read_bytes_indirect(&buffer, 0, "blob", Endian::Little).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_string_indirect_invalid_utf8_returns_access_error() {
|
||||
let data_region: &[u8] = &[0xFF, 0xFE, 0xFD];
|
||||
let mut index = [0u8; 8];
|
||||
let _ = data_access::write_u32(&mut index, 0, 0, "idx.off", Endian::Little);
|
||||
let _ = data_access::write_u32(&mut index, 4, 3, "idx.len", Endian::Little);
|
||||
let err = data_access::read_string_indirect(&index, 0, data_region, "name", Endian::Little)
|
||||
.unwrap_err();
|
||||
let mut buf = vec![0u8; 8 + 3];
|
||||
let _ = data_access::write_u32(&mut buf, 0, 8, "idx.off", Endian::Little);
|
||||
let _ = data_access::write_u32(&mut buf, 4, 3, "idx.len", Endian::Little);
|
||||
buf[8..11].copy_from_slice(&[0xFF, 0xFE, 0xFD]);
|
||||
let err = data_access::read_string_indirect(&buf, 0, "name", Endian::Little).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
}
|
||||
}
|
||||
+344
-167
@@ -8,7 +8,12 @@
|
||||
//! walks. Each test writes values to a buffer at computed offsets and
|
||||
//! reads them back, asserting both the values and (where applicable)
|
||||
//! the byte positions.
|
||||
//!
|
||||
//! All schemas are BAST documents (`{ "$defs": { ... } }` with `kind`-
|
||||
//! based vocabulary). The root type name is passed to `OffsetMap::compute`
|
||||
//! / `LayoutBuilder::new` / `SequentialReader::new` / `AlkTypeEngine::compile`.
|
||||
|
||||
use alktype::bast::BastDoc;
|
||||
use alktype::data_access;
|
||||
use alktype::tunion;
|
||||
use alktype::*;
|
||||
@@ -21,42 +26,47 @@ fn var_sizes(pairs: &[(&str, usize)]) -> HashMap<String, usize> {
|
||||
|
||||
#[test]
|
||||
fn fixed_size_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint32": true },
|
||||
"score": { "AlkType:Float32": true },
|
||||
"flag": { "AlkType:Uint8": true },
|
||||
"count": { "AlkType:Uint16": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "id", "kind": "uint32" },
|
||||
{ "name": "score", "kind": "float32" },
|
||||
{ "name": "flag", "kind": "uint8" },
|
||||
{ "name": "count", "kind": "uint16" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let offset_map = OffsetMap::compute(&schema)?;
|
||||
let doc = BastDoc::new(&root, "S")?;
|
||||
let offset_map = OffsetMap::compute(&doc)?;
|
||||
let mut buffer = vec![0u8; offset_map.total_size()];
|
||||
|
||||
let id_range = offset_map.get("id").expect("id range");
|
||||
data_access::write_u32(&mut buffer, id_range.start, 42, "id", Endian::Little)?;
|
||||
data_access::write_u32(&mut buffer, id_range.start(), 42, "id", Endian::Little)?;
|
||||
let score_range = offset_map.get("score").expect("score range");
|
||||
data_access::write_f32(&mut buffer, score_range.start, 1.5, "score", Endian::Little)?;
|
||||
data_access::write_f32(&mut buffer, score_range.start(), 1.5, "score", Endian::Little)?;
|
||||
let flag_range = offset_map.get("flag").expect("flag range");
|
||||
data_access::write_u8(&mut buffer, flag_range.start, 1, "flag")?;
|
||||
data_access::write_u8(&mut buffer, flag_range.start(), 1, "flag")?;
|
||||
let count_range = offset_map.get("count").expect("count range");
|
||||
data_access::write_u16(
|
||||
&mut buffer,
|
||||
count_range.start,
|
||||
count_range.start(),
|
||||
1000,
|
||||
"count",
|
||||
Endian::Little,
|
||||
)?;
|
||||
|
||||
assert_eq!(
|
||||
data_access::read_u32(&buffer, id_range.start, "id", Endian::Little)?,
|
||||
data_access::read_u32(&buffer, id_range.start(), "id", Endian::Little)?,
|
||||
42
|
||||
);
|
||||
let score = data_access::read_f32(&buffer, score_range.start, "score", Endian::Little)?;
|
||||
let score = data_access::read_f32(&buffer, score_range.start(), "score", Endian::Little)?;
|
||||
assert!((score - 1.5).abs() < 0.001, "score: {score}");
|
||||
assert_eq!(data_access::read_u8(&buffer, flag_range.start, "flag")?, 1);
|
||||
assert_eq!(data_access::read_u8(&buffer, flag_range.start(), "flag")?, 1);
|
||||
assert_eq!(
|
||||
data_access::read_u16(&buffer, count_range.start, "count", Endian::Little)?,
|
||||
data_access::read_u16(&buffer, count_range.start(), "count", Endian::Little)?,
|
||||
1000
|
||||
);
|
||||
Ok(())
|
||||
@@ -64,16 +74,20 @@ fn fixed_size_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn fixed_size_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "little",
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint32": true },
|
||||
"score": { "AlkType:Float32": true },
|
||||
"flag": { "AlkType:Uint8": true }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "little",
|
||||
"fields": [
|
||||
{ "name": "id", "kind": "uint32" },
|
||||
{ "name": "score", "kind": "float32" },
|
||||
{ "name": "flag", "kind": "uint8" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
|
||||
let mut buffer = vec![0u8; offset_map.total_size()];
|
||||
|
||||
@@ -107,13 +121,15 @@ fn string_round_trip_via_data_access() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn string_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"name": { "AlkType:String": true }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "name", "kind": "string" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
|
||||
let mut buffer = vec![0u8; offset_map.total_size() + 64];
|
||||
engine.write_field(&mut buffer, "name", &FieldValue::String("hello"))?;
|
||||
@@ -141,48 +157,56 @@ fn bytes_round_trip_via_data_access() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn nested_struct_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"header": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"version": { "AlkType:Uint32": true },
|
||||
"magic": { "AlkType:Uint32": true }
|
||||
}
|
||||
},
|
||||
"payload": { "AlkType:Bytes": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{
|
||||
"name": "header",
|
||||
"kind": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "version", "kind": "uint32" },
|
||||
{ "name": "magic", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
},
|
||||
{ "name": "payload", "kind": "bytes" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let offset_map = OffsetMap::compute(&schema)?;
|
||||
let doc = BastDoc::new(&root, "S")?;
|
||||
let offset_map = OffsetMap::compute(&doc)?;
|
||||
|
||||
let header_version = offset_map.get("header.version").expect("header.version");
|
||||
let header_magic = offset_map.get("header.magic").expect("header.magic");
|
||||
let payload_prefix = offset_map.get("payload").expect("payload");
|
||||
|
||||
assert_eq!(header_version.start, 0);
|
||||
assert_eq!(header_magic.start, 4);
|
||||
assert_eq!(payload_prefix.start, 8);
|
||||
assert_eq!(header_version.start(), 0);
|
||||
assert_eq!(header_magic.start(), 4);
|
||||
assert_eq!(payload_prefix.start(), 8);
|
||||
|
||||
let data = b"body-data".to_vec();
|
||||
let mut buffer = vec![0u8; offset_map.total_size() + data.len()];
|
||||
data_access::write_u32(
|
||||
&mut buffer,
|
||||
header_version.start,
|
||||
header_version.start(),
|
||||
1,
|
||||
"header.version",
|
||||
Endian::Little,
|
||||
)?;
|
||||
data_access::write_u32(
|
||||
&mut buffer,
|
||||
header_magic.start,
|
||||
header_magic.start(),
|
||||
0xCAFEBABE,
|
||||
"header.magic",
|
||||
Endian::Little,
|
||||
)?;
|
||||
data_access::write_bytes(
|
||||
&mut buffer,
|
||||
payload_prefix.start,
|
||||
payload_prefix.start(),
|
||||
&data,
|
||||
"payload",
|
||||
Endian::Little,
|
||||
@@ -191,18 +215,18 @@ fn nested_struct_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
|
||||
assert_eq!(
|
||||
data_access::read_u32(
|
||||
&buffer,
|
||||
header_version.start,
|
||||
header_version.start(),
|
||||
"header.version",
|
||||
Endian::Little
|
||||
)?,
|
||||
1
|
||||
);
|
||||
assert_eq!(
|
||||
data_access::read_u32(&buffer, header_magic.start, "header.magic", Endian::Little)?,
|
||||
data_access::read_u32(&buffer, header_magic.start(), "header.magic", Endian::Little)?,
|
||||
0xCAFEBABE
|
||||
);
|
||||
assert_eq!(
|
||||
data_access::read_bytes(&buffer, payload_prefix.start, "payload", Endian::Little)?,
|
||||
data_access::read_bytes(&buffer, payload_prefix.start(), "payload", Endian::Little)?,
|
||||
&data[..]
|
||||
);
|
||||
Ok(())
|
||||
@@ -210,25 +234,32 @@ fn nested_struct_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn nested_struct_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
|
||||
let mut schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"header": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"version": { "AlkType:Uint8": true },
|
||||
"flags": { "AlkType:Uint8": true }
|
||||
}
|
||||
},
|
||||
"payload_len": { "AlkType:Uint32": true }
|
||||
let doc = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{
|
||||
"name": "header",
|
||||
"kind": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "version", "kind": "uint8" },
|
||||
{ "name": "flags", "kind": "uint8" }
|
||||
]
|
||||
}
|
||||
},
|
||||
{ "name": "payload_len", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
|
||||
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
|
||||
let offset_map = engine.offset_map().expect("aligned mode");
|
||||
|
||||
assert_eq!(offset_map.get("header.version").unwrap().start, 0);
|
||||
assert_eq!(offset_map.get("header.flags").unwrap().start, 1);
|
||||
assert_eq!(offset_map.get("payload_len").unwrap().start, 4);
|
||||
assert_eq!(offset_map.get("header.version").unwrap().start(), 0);
|
||||
assert_eq!(offset_map.get("header.flags").unwrap().start(), 1);
|
||||
assert_eq!(offset_map.get("payload_len").unwrap().start(), 4);
|
||||
|
||||
let mut buffer = vec![0u8; offset_map.total_size()];
|
||||
engine.write_field(&mut buffer, "header.version", &FieldValue::U8(1))?;
|
||||
@@ -252,67 +283,76 @@ fn nested_struct_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn big_endian_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "big",
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint32": true },
|
||||
"offset": { "AlkType:Float64": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "big",
|
||||
"fields": [
|
||||
{ "name": "id", "kind": "uint32" },
|
||||
{ "name": "offset", "kind": "float64" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let offset_map = OffsetMap::compute(&schema)?;
|
||||
let endian = Endian::from_schema(&schema);
|
||||
assert_eq!(endian, Endian::Big);
|
||||
let doc = BastDoc::new(&root, "S")?;
|
||||
let offset_map = OffsetMap::compute(&doc)?;
|
||||
let endian = Endian::Big;
|
||||
|
||||
let id_range = offset_map.get("id").expect("id");
|
||||
let offset_range = offset_map.get("offset").expect("offset");
|
||||
|
||||
assert_eq!(id_range.start, 0);
|
||||
assert_eq!(offset_range.start, 8);
|
||||
assert_eq!(id_range.start(), 0);
|
||||
assert_eq!(offset_range.start(), 8);
|
||||
|
||||
let value: f64 = std::f64::consts::PI;
|
||||
let mut buffer = vec![0u8; offset_map.total_size()];
|
||||
data_access::write_u32(&mut buffer, id_range.start, 0x01020304, "id", endian)?;
|
||||
data_access::write_f64(&mut buffer, offset_range.start, value, "offset", endian)?;
|
||||
data_access::write_u32(&mut buffer, id_range.start(), 0x01020304, "id", endian)?;
|
||||
data_access::write_f64(&mut buffer, offset_range.start(), value, "offset", endian)?;
|
||||
|
||||
assert_eq!(&buffer[0..4], &[0x01, 0x02, 0x03, 0x04]);
|
||||
assert_eq!(&buffer[4..8], &[0x00, 0x00, 0x00, 0x00]);
|
||||
assert_eq!(&buffer[8..16], value.to_be_bytes());
|
||||
|
||||
assert_eq!(
|
||||
data_access::read_u32(&buffer, id_range.start, "id", endian)?,
|
||||
data_access::read_u32(&buffer, id_range.start(), "id", endian)?,
|
||||
0x01020304
|
||||
);
|
||||
let read = data_access::read_f64(&buffer, offset_range.start, "offset", endian)?;
|
||||
let read = data_access::read_f64(&buffer, offset_range.start(), "offset", endian)?;
|
||||
assert!((read - value).abs() < 1e-12);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn alignment_padding_round_trip_u8_then_u32() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"flag": { "AlkType:Uint8": true },
|
||||
"id": { "AlkType:Uint32": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "flag", "kind": "uint8" },
|
||||
{ "name": "id", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let offset_map = OffsetMap::compute(&schema)?;
|
||||
let doc = BastDoc::new(&root, "S")?;
|
||||
let offset_map = OffsetMap::compute(&doc)?;
|
||||
|
||||
let flag_range = offset_map.get("flag").expect("flag");
|
||||
let id_range = offset_map.get("id").expect("id");
|
||||
|
||||
assert_eq!(flag_range.start, 0);
|
||||
assert_eq!(flag_range.end, 1);
|
||||
assert_eq!(id_range.start, 4);
|
||||
assert_eq!(id_range.end, 8);
|
||||
assert_eq!(flag_range.start(), 0);
|
||||
assert_eq!(flag_range.end(), 1);
|
||||
assert_eq!(id_range.start(), 4);
|
||||
assert_eq!(id_range.end(), 8);
|
||||
assert_eq!(offset_map.total_size(), 8);
|
||||
|
||||
let mut buffer = vec![0u8; offset_map.total_size()];
|
||||
data_access::write_u8(&mut buffer, flag_range.start, 0xAB, "flag")?;
|
||||
data_access::write_u8(&mut buffer, flag_range.start(), 0xAB, "flag")?;
|
||||
data_access::write_u32(
|
||||
&mut buffer,
|
||||
id_range.start,
|
||||
id_range.start(),
|
||||
0x01020304,
|
||||
"id",
|
||||
Endian::Little,
|
||||
@@ -323,11 +363,11 @@ fn alignment_padding_round_trip_u8_then_u32() -> Result<(), AlkTypeError> {
|
||||
assert_eq!(&buffer[4..8], 0x01020304u32.to_le_bytes());
|
||||
|
||||
assert_eq!(
|
||||
data_access::read_u8(&buffer, flag_range.start, "flag")?,
|
||||
data_access::read_u8(&buffer, flag_range.start(), "flag")?,
|
||||
0xAB
|
||||
);
|
||||
assert_eq!(
|
||||
data_access::read_u32(&buffer, id_range.start, "id", Endian::Little)?,
|
||||
data_access::read_u32(&buffer, id_range.start(), "id", Endian::Little)?,
|
||||
0x01020304
|
||||
);
|
||||
Ok(())
|
||||
@@ -335,16 +375,20 @@ fn alignment_padding_round_trip_u8_then_u32() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn packed_layout_round_trip_via_layout_builder() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "little",
|
||||
"properties": {
|
||||
"flag": { "AlkType:Uint8": true },
|
||||
"id": { "AlkType:Uint32": true },
|
||||
"payload": { "AlkType:String": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "little",
|
||||
"fields": [
|
||||
{ "name": "flag", "kind": "uint8" },
|
||||
{ "name": "id", "kind": "uint32" },
|
||||
{ "name": "payload", "kind": "string" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let builder = LayoutBuilder::new(&schema)?;
|
||||
let builder = LayoutBuilder::new(&root, "S")?;
|
||||
let layout = builder.build(&var_sizes(&[("payload", 10)]))?;
|
||||
|
||||
let flag_pos = layout.get("flag").expect("flag");
|
||||
@@ -392,16 +436,20 @@ fn packed_layout_round_trip_via_layout_builder() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn sequential_reader_round_trip_packed_buffer() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "little",
|
||||
"properties": {
|
||||
"id": { "AlkType:Uint8": true },
|
||||
"name": { "AlkType:String": true },
|
||||
"tail": { "AlkType:Uint8": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "little",
|
||||
"fields": [
|
||||
{ "name": "id", "kind": "uint8" },
|
||||
{ "name": "name", "kind": "string" },
|
||||
{ "name": "tail", "kind": "uint8" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let builder = LayoutBuilder::new(&schema)?;
|
||||
let builder = LayoutBuilder::new(&root, "S")?;
|
||||
let payload = "hello";
|
||||
let layout = builder.build(&var_sizes(&[("name", payload.len())]))?;
|
||||
|
||||
@@ -411,7 +459,8 @@ fn sequential_reader_round_trip_packed_buffer() -> Result<(), AlkTypeError> {
|
||||
let after = 1 + 4 + payload.len();
|
||||
data_access::write_u8(&mut buffer, after, 99, "tail")?;
|
||||
|
||||
let mut reader = SequentialReader::new(&schema)?;
|
||||
let plan = ReadPlan::compile(&root, "S")?;
|
||||
let mut reader = SequentialReader::new(std::sync::Arc::new(plan));
|
||||
assert_eq!(reader.endian(), Endian::Little);
|
||||
assert_eq!(reader.position(), 0);
|
||||
|
||||
@@ -436,13 +485,17 @@ fn sequential_reader_round_trip_packed_buffer() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn sequential_reader_read_field_walks_preceding_fields() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Struct": true,
|
||||
"endian": "little",
|
||||
"properties": {
|
||||
"a": { "AlkType:Uint8": true },
|
||||
"b": { "AlkType:Uint32": true },
|
||||
"c": { "AlkType:Uint8": true }
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "little",
|
||||
"fields": [
|
||||
{ "name": "a", "kind": "uint8" },
|
||||
{ "name": "b", "kind": "uint32" },
|
||||
{ "name": "c", "kind": "uint8" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let mut buffer = vec![0u8; 16];
|
||||
@@ -450,7 +503,8 @@ fn sequential_reader_read_field_walks_preceding_fields() -> Result<(), AlkTypeEr
|
||||
data_access::write_u32(&mut buffer, 1, 0xDEADBEEF, "b", Endian::Little)?;
|
||||
data_access::write_u8(&mut buffer, 5, 9, "c")?;
|
||||
|
||||
let mut reader = SequentialReader::new(&schema)?;
|
||||
let plan = ReadPlan::compile(&root, "S")?;
|
||||
let mut reader = SequentialReader::new(std::sync::Arc::new(plan));
|
||||
let value = reader.read_field(&buffer, "c")?;
|
||||
assert_eq!(value, FieldValue::U8(9));
|
||||
assert_eq!(reader.position(), 6);
|
||||
@@ -463,74 +517,197 @@ fn sequential_reader_read_field_walks_preceding_fields() -> Result<(), AlkTypeEr
|
||||
|
||||
#[test]
|
||||
fn tunion_byte_offset_discriminator_dispatch() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "AlkType:Uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" }
|
||||
},
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Read": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"handle": { "AlkType:Uint32": true },
|
||||
"length": { "AlkType:Uint32": true }
|
||||
"Packet": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
|
||||
"mapping": {
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
},
|
||||
"Read": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" }
|
||||
]
|
||||
},
|
||||
"Write": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"handle": { "AlkType:Uint32": true },
|
||||
"length": { "AlkType:Uint32": true },
|
||||
"data": { "AlkType:Uint32": true }
|
||||
}
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" },
|
||||
{ "name": "data", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let doc = BastDoc::new(&root, "Packet")?;
|
||||
let union_def = match doc.root_def().kind() {
|
||||
alktype::bast::BastDefKind::Union(u) => u,
|
||||
_ => unreachable!(),
|
||||
};
|
||||
let mut buffer = vec![0u8; 32];
|
||||
buffer[0] = 5;
|
||||
data_access::write_u32(&mut buffer, 1, 0x01020304, "Read.handle", Endian::Big)?;
|
||||
data_access::write_u32(&mut buffer, 5, 4096, "Read.length", Endian::Big)?;
|
||||
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big)?;
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, union_def, Endian::Big)?;
|
||||
assert_eq!(dispatch.key, "5");
|
||||
assert_eq!(dispatch.variant_offset, 1);
|
||||
assert_eq!(dispatch.discriminator_size, 1);
|
||||
|
||||
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
|
||||
assert_eq!(
|
||||
variant
|
||||
.get("AlkType:Struct")
|
||||
.and_then(serde_json::Value::as_bool),
|
||||
Some(true)
|
||||
);
|
||||
let variant = tunion::resolve_variant(union_def, &dispatch.key)?;
|
||||
match variant {
|
||||
alktype::bast::BastType::Ref(r) => assert_eq!(r.name(), "Read"),
|
||||
other => panic!("expected Ref to Read, got {other:?}"),
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn tunion_byte_offset_discriminator_size_lookup() -> Result<(), AlkTypeError> {
|
||||
let u8_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
|
||||
"mapping": {}
|
||||
});
|
||||
let u16_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint16"},
|
||||
"mapping": {}
|
||||
});
|
||||
let u32_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint32"},
|
||||
"mapping": {}
|
||||
});
|
||||
assert_eq!(tunion::discriminator_size(&u8_schema)?, 1);
|
||||
assert_eq!(tunion::discriminator_size(&u16_schema)?, 2);
|
||||
assert_eq!(tunion::discriminator_size(&u32_schema)?, 4);
|
||||
fn union_with(disc_type: &str) -> serde_json::Value {
|
||||
json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "type": disc_type },
|
||||
"mapping": { "1": { "$ref": "#/$defs/A" } }
|
||||
},
|
||||
"A": { "kind": "struct", "fields": [] }
|
||||
}
|
||||
})
|
||||
}
|
||||
let u8_root = union_with("uint8");
|
||||
let u16_root = union_with("uint16");
|
||||
let u32_root = union_with("uint32");
|
||||
let u8_doc = BastDoc::new(&u8_root, "U")?;
|
||||
let u16_doc = BastDoc::new(&u16_root, "U")?;
|
||||
let u32_doc = BastDoc::new(&u32_root, "U")?;
|
||||
let u8_union = match u8_doc.root_def().kind() {
|
||||
alktype::bast::BastDefKind::Union(u) => u,
|
||||
_ => unreachable!(),
|
||||
};
|
||||
let u16_union = match u16_doc.root_def().kind() {
|
||||
alktype::bast::BastDefKind::Union(u) => u,
|
||||
_ => unreachable!(),
|
||||
};
|
||||
let u32_union = match u32_doc.root_def().kind() {
|
||||
alktype::bast::BastDefKind::Union(u) => u,
|
||||
_ => unreachable!(),
|
||||
};
|
||||
assert_eq!(tunion::discriminator_size(u8_union)?, 1);
|
||||
assert_eq!(tunion::discriminator_size(u16_union)?, 2);
|
||||
assert_eq!(tunion::discriminator_size(u32_union)?, 4);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// L6 (review #006): the single roundtrip test through a field-disc union.
|
||||
// Write with LayoutBuilder → read with SequentialReader → materialize →
|
||||
// validate_bytes on the engine — one schema, all three packed-mode
|
||||
// consumers, so the H3 wire-convention split can never reappear silently.
|
||||
//
|
||||
// Fixture shape: the union's `fields` carry the discriminator (`type`,
|
||||
// uint8) *and* a second shared field (`seq`, uint32); the variant
|
||||
// (`Read`) declares only its own field (`handle`) — it does NOT
|
||||
// re-declare the discriminator or any shared field (forbidden since the
|
||||
// ADR-011 addendum). The disc field is a uint8, so the mapping keys are
|
||||
// stringified integers ("1" = read). The builder lays out
|
||||
// shared-then-variant:
|
||||
// type@0 (1B), seq@1 (4B), handle@5 (4B) — total 9.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
#[test]
|
||||
fn field_disc_union_roundtrip_build_read_materialize_validate() -> Result<(), AlkTypeError> {
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"S": {
|
||||
"kind": "struct",
|
||||
"endian": "little",
|
||||
"fields": [
|
||||
{ "name": "event", "kind": { "$ref": "#/$defs/Event" } },
|
||||
{ "name": "trailer", "kind": "uint8" }
|
||||
]
|
||||
},
|
||||
"Event": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [
|
||||
{ "name": "type", "kind": "uint8" },
|
||||
{ "name": "seq", "kind": "uint32" }
|
||||
],
|
||||
"mapping": {
|
||||
"1": { "$ref": "#/$defs/Read" },
|
||||
"2": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
},
|
||||
"Read": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "handle", "kind": "uint32" } ]
|
||||
},
|
||||
"Write": {
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "handle", "kind": "uint32" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
// --- Write side: LayoutBuilder (shared-then-variant layout) --------
|
||||
let builder = LayoutBuilder::new(&root, "S")?;
|
||||
let layout = builder.build(&var_sizes(&[("event.__variant", 0)]))?;
|
||||
assert_eq!(layout.total_size(), 9 + 1, "shared(1+4) + variant(4) + trailer(1)");
|
||||
|
||||
let handle_pos = layout.get("event.handle").expect("event.handle");
|
||||
assert_eq!(handle_pos.offset, 5, "variant fields start after shared (type@0, seq@1)");
|
||||
let disc_pos = layout.get("event.type").expect("event.type");
|
||||
assert_eq!(disc_pos.offset, 0);
|
||||
let seq_pos = layout.get("event.seq").expect("event.seq");
|
||||
assert_eq!(seq_pos.offset, 1);
|
||||
|
||||
let mut buffer = vec![0u8; layout.total_size()];
|
||||
data_access::write_u8(&mut buffer, 0, 1, "event.type")?; // mapping key "1"
|
||||
data_access::write_u32(&mut buffer, 1, 77, "event.seq", Endian::Little)?;
|
||||
data_access::write_u32(&mut buffer, 5, 4242, "event.handle", Endian::Little)?;
|
||||
data_access::write_u8(&mut buffer, 9, 55, "trailer")?;
|
||||
|
||||
// --- Engine (packed) + reader + materializer + validator -----------
|
||||
let engine = AlkTypeEngine::compile(&root, "S", LayoutMode::Packed, None)?;
|
||||
|
||||
// validate_bytes (materializer + ValidationPlan) accepts the buffer.
|
||||
engine.validate_bytes(&buffer)?;
|
||||
|
||||
// SequentialReader: the union field reports the disc value and the
|
||||
// variant start (after the shared walk).
|
||||
let mut reader = engine.sequential_reader().expect("packed mode has reader");
|
||||
let (name, value) = reader.read_next(&buffer)?.expect("event");
|
||||
assert_eq!(name, "event");
|
||||
let (disc, variant_start) = match &value {
|
||||
FieldValue::Union { discriminator, variant_start } => (discriminator.clone(), *variant_start),
|
||||
other => panic!("expected Union, got {other:?}"),
|
||||
};
|
||||
assert_eq!(disc, "1", "uint8 disc value 1 stringifies to the mapping key");
|
||||
assert_eq!(variant_start, 5, "variant starts after the shared walk");
|
||||
let (name, value) = reader.read_next(&buffer)?.expect("trailer");
|
||||
assert_eq!(name, "trailer");
|
||||
assert_eq!(value, FieldValue::U8(55));
|
||||
|
||||
// materialize_packed through the engine's plan: shared fields land
|
||||
// in the object, the variant's handle flattens in.
|
||||
let plan = engine_sequential_plan(&root, "S")?;
|
||||
let value = alktype::materialize::materialize_packed(&plan, &buffer)?;
|
||||
assert_eq!(value["event"]["type"], json!(1));
|
||||
assert_eq!(value["event"]["seq"], json!(77));
|
||||
assert_eq!(value["event"]["handle"], json!(4242));
|
||||
assert_eq!(value["event"]["__discriminator"], json!("1"));
|
||||
assert_eq!(value["trailer"], json!(55));
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn engine_sequential_plan(root: &serde_json::Value, name: &str) -> Result<std::sync::Arc<ReadPlan>, AlkTypeError> {
|
||||
Ok(std::sync::Arc::new(ReadPlan::compile(root, name)?))
|
||||
}
|
||||
+186
-165
@@ -6,103 +6,116 @@
|
||||
//! correct mapping key and variant offset, that `resolve_variant`
|
||||
//! follows `$ref` pointers, and that `discriminator_size` reports the
|
||||
//! right fixed sizes.
|
||||
//!
|
||||
//! All schemas are BAST documents.
|
||||
|
||||
use alktype::bast::{BastDefKind, BastDoc, BastType, BastUnion};
|
||||
use alktype::data_access;
|
||||
use alktype::tunion;
|
||||
use alktype::{Endian, AlkTypeError};
|
||||
use serde_json::json;
|
||||
|
||||
fn sftp_like_byte_union() -> serde_json::Value {
|
||||
fn sftp_like_byte_union_doc() -> serde_json::Value {
|
||||
json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "AlkType:Uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" }
|
||||
},
|
||||
"$defs": {
|
||||
"Read": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"handle": { "AlkType:Uint32": true },
|
||||
"length": { "AlkType:Uint32": true }
|
||||
"Packet": {
|
||||
"kind": "union",
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
},
|
||||
"Read": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" }
|
||||
]
|
||||
},
|
||||
"Write": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"handle": { "AlkType:Uint32": true },
|
||||
"length": { "AlkType:Uint32": true },
|
||||
"data": { "AlkType:Uint32": true }
|
||||
}
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" },
|
||||
{ "name": "data", "kind": "uint32" }
|
||||
]
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
fn union_of(root: &serde_json::Value, name: &str) -> BastUnion {
|
||||
let doc = BastDoc::new(root, name).expect("bast doc");
|
||||
match doc.root_def().kind() {
|
||||
BastDefKind::Union(u) => u.clone(),
|
||||
_ => panic!("root must be a union"),
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint8_dispatches_to_read() -> Result<(), AlkTypeError> {
|
||||
let union_schema = sftp_like_byte_union();
|
||||
let root = sftp_like_byte_union_doc();
|
||||
let union_def = union_of(&root, "Packet");
|
||||
let mut buffer = vec![0u8; 16];
|
||||
buffer[0] = 5;
|
||||
data_access::write_u32(&mut buffer, 1, 0x01020304, "Read.handle", Endian::Big)?;
|
||||
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big)?;
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big)?;
|
||||
assert_eq!(dispatch.key, "5");
|
||||
assert_eq!(dispatch.variant_offset, 1);
|
||||
assert_eq!(dispatch.discriminator_size, 1);
|
||||
|
||||
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
|
||||
assert_eq!(
|
||||
variant
|
||||
.get("AlkType:Struct")
|
||||
.and_then(serde_json::Value::as_bool),
|
||||
Some(true)
|
||||
);
|
||||
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
|
||||
match variant {
|
||||
BastType::Ref(r) => assert_eq!(r.name(), "Read"),
|
||||
other => panic!("expected Ref to Read, got {other:?}"),
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint8_dispatches_to_write() -> Result<(), AlkTypeError> {
|
||||
let union_schema = sftp_like_byte_union();
|
||||
let root = sftp_like_byte_union_doc();
|
||||
let union_def = union_of(&root, "Packet");
|
||||
let mut buffer = vec![0u8; 16];
|
||||
buffer[0] = 6;
|
||||
data_access::write_u32(&mut buffer, 1, 0xDEADBEEF, "Write.handle", Endian::Big)?;
|
||||
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big)?;
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big)?;
|
||||
assert_eq!(dispatch.key, "6");
|
||||
assert_eq!(dispatch.variant_offset, 1);
|
||||
assert_eq!(dispatch.discriminator_size, 1);
|
||||
|
||||
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
|
||||
let props = variant
|
||||
.get("properties")
|
||||
.and_then(serde_json::Value::as_object)
|
||||
.expect("variant has properties");
|
||||
assert!(props.contains_key("data"));
|
||||
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
|
||||
match variant {
|
||||
BastType::Ref(r) => assert_eq!(r.name(), "Write"),
|
||||
other => panic!("expected Ref to Write, got {other:?}"),
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint16_little_endian() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 2,
|
||||
"type": "AlkType:Uint16"
|
||||
},
|
||||
"mapping": {
|
||||
"5": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 2, "type": "uint16" },
|
||||
"mapping": {
|
||||
"5": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let union_def = union_of(&root, "U");
|
||||
let mut buffer = vec![0u8; 16];
|
||||
buffer[2..4].copy_from_slice(&5u16.to_le_bytes());
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &schema, Endian::Little)?;
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Little)?;
|
||||
assert_eq!(dispatch.key, "5");
|
||||
assert_eq!(dispatch.variant_offset, 4);
|
||||
assert_eq!(dispatch.discriminator_size, 2);
|
||||
@@ -111,20 +124,21 @@ fn read_byte_discriminator_uint16_little_endian() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_uint32_big_endian() -> Result<(), AlkTypeError> {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "AlkType:Uint32"
|
||||
},
|
||||
"mapping": {
|
||||
"101": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "offset": 0, "type": "uint32" },
|
||||
"mapping": {
|
||||
"101": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let union_def = union_of(&root, "U");
|
||||
let mut buffer = vec![0u8; 16];
|
||||
buffer[0..4].copy_from_slice(&101u32.to_be_bytes());
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &schema, Endian::Big)?;
|
||||
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big)?;
|
||||
assert_eq!(dispatch.key, "101");
|
||||
assert_eq!(dispatch.variant_offset, 4);
|
||||
assert_eq!(dispatch.discriminator_size, 4);
|
||||
@@ -133,191 +147,198 @@ fn read_byte_discriminator_uint32_big_endian() -> Result<(), AlkTypeError> {
|
||||
|
||||
#[test]
|
||||
fn read_byte_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
|
||||
let union_schema = sftp_like_byte_union();
|
||||
let root = sftp_like_byte_union_doc();
|
||||
let union_def = union_of(&root, "Packet");
|
||||
let buffer = [99u8, 0x00, 0x00, 0x00];
|
||||
let err = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big).unwrap_err();
|
||||
let err = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_string_dispatches_to_read() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "type"},
|
||||
"properties": {
|
||||
"type": { "AlkType:String": true }
|
||||
},
|
||||
"mapping": {
|
||||
"read": {"$ref": "#/$defs/Read"},
|
||||
"write": {"$ref": "#/$defs/Write"}
|
||||
},
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Read": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"handle": { "AlkType:Uint32": true },
|
||||
"length": { "AlkType:Uint32": true }
|
||||
"Event": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "string" } ],
|
||||
"mapping": {
|
||||
"read": { "$ref": "#/$defs/Read" },
|
||||
"write": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
},
|
||||
"Read": {
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "length", "kind": "uint32" }
|
||||
]
|
||||
},
|
||||
"Write": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {
|
||||
"handle": { "AlkType:Uint32": true },
|
||||
"data": { "AlkType:Bytes": true }
|
||||
}
|
||||
"kind": "struct",
|
||||
"fields": [
|
||||
{ "name": "handle", "kind": "uint32" },
|
||||
{ "name": "data", "kind": "bytes" }
|
||||
]
|
||||
}
|
||||
}
|
||||
});
|
||||
let union_def = union_of(&root, "Event");
|
||||
let value = "read";
|
||||
let mut buffer = vec![0u8; 32];
|
||||
data_access::write_string(&mut buffer, 0, value, "type", Endian::Little)?;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
|
||||
assert_eq!(dispatch.key, "read");
|
||||
assert_eq!(dispatch.variant_offset, 4 + value.len());
|
||||
assert_eq!(dispatch.discriminator_size, 4 + value.len());
|
||||
|
||||
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
|
||||
assert_eq!(
|
||||
variant
|
||||
.get("AlkType:Struct")
|
||||
.and_then(serde_json::Value::as_bool),
|
||||
Some(true)
|
||||
);
|
||||
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
|
||||
match variant {
|
||||
BastType::Ref(r) => assert_eq!(r.name(), "Read"),
|
||||
other => panic!("expected Ref to Read, got {other:?}"),
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_string_dispatches_to_write() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "type"},
|
||||
"properties": {
|
||||
"type": { "AlkType:String": true }
|
||||
},
|
||||
"mapping": {
|
||||
"read": {"$ref": "#/$defs/Read"},
|
||||
"write": {"$ref": "#/$defs/Write"}
|
||||
},
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Event": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "string" } ],
|
||||
"mapping": {
|
||||
"read": { "$ref": "#/$defs/Read" },
|
||||
"write": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
},
|
||||
"Read": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {"x": {"AlkType:Uint8": true}}
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "x", "kind": "uint8" } ]
|
||||
},
|
||||
"Write": {
|
||||
"AlkType:Struct": true,
|
||||
"properties": {"y": {"AlkType:Uint16": true}}
|
||||
"kind": "struct",
|
||||
"fields": [ { "name": "y", "kind": "uint16" } ]
|
||||
}
|
||||
}
|
||||
});
|
||||
let union_def = union_of(&root, "Event");
|
||||
let value = "write";
|
||||
let mut buffer = vec![0u8; 32];
|
||||
data_access::write_string(&mut buffer, 0, value, "type", Endian::Little)?;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
|
||||
assert_eq!(dispatch.key, "write");
|
||||
assert_eq!(dispatch.variant_offset, 4 + value.len());
|
||||
|
||||
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
|
||||
let props = variant
|
||||
.get("properties")
|
||||
.and_then(serde_json::Value::as_object)
|
||||
.expect("variant has properties");
|
||||
assert!(props.contains_key("y"));
|
||||
assert!(!props.contains_key("x"));
|
||||
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
|
||||
match variant {
|
||||
BastType::Ref(r) => assert_eq!(r.name(), "Write"),
|
||||
other => panic!("expected Ref to Write, got {other:?}"),
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_uint8_field() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "tag"},
|
||||
"properties": {
|
||||
"tag": { "AlkType:Uint8": true }
|
||||
},
|
||||
"mapping": {
|
||||
"0": {"AlkType:Struct": true, "properties": {"a": {"AlkType:Uint32": true}}},
|
||||
"1": {"AlkType:Struct": true, "properties": {"b": {"AlkType:Uint16": true}}}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Event": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "tag" },
|
||||
"fields": [ { "name": "tag", "kind": "uint8" } ],
|
||||
"mapping": {
|
||||
"0": { "kind": "struct", "fields": [ { "name": "a", "kind": "uint32" } ] },
|
||||
"1": { "kind": "struct", "fields": [ { "name": "b", "kind": "uint16" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let union_def = union_of(&root, "Event");
|
||||
let mut buffer = vec![0u8; 8];
|
||||
buffer[0] = 0;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
|
||||
assert_eq!(dispatch.key, "0");
|
||||
assert_eq!(dispatch.variant_offset, 1);
|
||||
assert_eq!(dispatch.discriminator_size, 1);
|
||||
|
||||
buffer[0] = 1;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
|
||||
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
|
||||
assert_eq!(dispatch.key, "1");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn read_field_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
|
||||
let union_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "tag"},
|
||||
"properties": {
|
||||
"tag": { "AlkType:Uint8": true }
|
||||
},
|
||||
"mapping": {
|
||||
"0": {"AlkType:Struct": true, "properties": {"a": {"AlkType:Uint32": true}}}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"Event": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "tag" },
|
||||
"fields": [ { "name": "tag", "kind": "uint8" } ],
|
||||
"mapping": {
|
||||
"0": { "kind": "struct", "fields": [ { "name": "a", "kind": "uint32" } ] }
|
||||
}
|
||||
}
|
||||
}
|
||||
});
|
||||
let union_def = union_of(&root, "Event");
|
||||
let mut buffer = vec![0u8; 8];
|
||||
buffer[0] = 99;
|
||||
let err =
|
||||
tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little).unwrap_err();
|
||||
tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn discriminator_size_returns_correct_values() -> Result<(), AlkTypeError> {
|
||||
let u8_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
|
||||
"mapping": {}
|
||||
});
|
||||
let u16_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint16"},
|
||||
"mapping": {}
|
||||
});
|
||||
let u32_schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "byte", "type": "AlkType:Uint32"},
|
||||
"mapping": {}
|
||||
});
|
||||
assert_eq!(tunion::discriminator_size(&u8_schema)?, 1);
|
||||
assert_eq!(tunion::discriminator_size(&u16_schema)?, 2);
|
||||
assert_eq!(tunion::discriminator_size(&u32_schema)?, 4);
|
||||
fn union_with(disc_type: &str) -> serde_json::Value {
|
||||
json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "byte", "type": disc_type },
|
||||
"mapping": { "1": { "$ref": "#/$defs/A" } }
|
||||
},
|
||||
"A": { "kind": "struct", "fields": [] }
|
||||
}
|
||||
})
|
||||
}
|
||||
let u8_root = union_with("uint8");
|
||||
let u16_root = union_with("uint16");
|
||||
let u32_root = union_with("uint32");
|
||||
let u8_union = union_of(&u8_root, "U");
|
||||
let u16_union = union_of(&u16_root, "U");
|
||||
let u32_union = union_of(&u32_root, "U");
|
||||
assert_eq!(tunion::discriminator_size(&u8_union)?, 1);
|
||||
assert_eq!(tunion::discriminator_size(&u16_union)?, 2);
|
||||
assert_eq!(tunion::discriminator_size(&u32_union)?, 4);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn discriminator_size_field_kind_returns_schema_error() {
|
||||
let schema = json!({
|
||||
"AlkType:Union": true,
|
||||
"discriminator": {"kind": "field", "name": "type"},
|
||||
"properties": {"type": {"AlkType:Uint8": true}},
|
||||
"mapping": {}
|
||||
let root = json!({
|
||||
"$defs": {
|
||||
"U": {
|
||||
"kind": "union",
|
||||
"discriminator": { "kind": "field", "name": "type" },
|
||||
"fields": [ { "name": "type", "kind": "uint8" } ],
|
||||
"mapping": { "0": { "kind": "struct", "fields": [] } }
|
||||
}
|
||||
}
|
||||
});
|
||||
let err = tunion::discriminator_size(&schema).unwrap_err();
|
||||
let union_def = union_of(&root, "U");
|
||||
let err = tunion::discriminator_size(&union_def).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn resolve_variant_returns_schema_error_for_unknown_key() {
|
||||
let union_schema = sftp_like_byte_union();
|
||||
let err = tunion::resolve_variant(&union_schema, "999").unwrap_err();
|
||||
let root = sftp_like_byte_union_doc();
|
||||
let union_def = union_of(&root, "Packet");
|
||||
let err = tunion::resolve_variant(&union_def, "999").unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parse_discriminator_missing_returns_schema_error() {
|
||||
let schema = json!({"AlkType:Union": true});
|
||||
let err = alktype::parse_discriminator(&schema).unwrap_err();
|
||||
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
|
||||
}
|
||||
}
|
||||
Reference in new issue
Block a user