26 Commits
Author SHA1 Message Date
glm-5.2 cab493206c Release v0.2.0: BAST pivot
Bump version to 0.2.0, exclude AGENTS.md from the published crate, and
add a CHANGELOG.md covering the breaking BAST pivot (schema format,
compile signature, validation split) plus bug fixes vs v0.1.0.

Verification:
- cargo test --release: all tests pass
- cargo clippy --all-targets -- -D warnings: clean
- cargo publish --dry-run --allow-dirty: packages as v0.2.0, no collision
- AGENTS.md no longer in cargo package --list; CHANGELOG.md included
2026-08-17 05:47:25 +00:00
glm-5.2 82fec45bc6 Sync docs to 18 BAST kinds (Timestamp removal fallout)
- README: 19 -> 18 kinds, drop the timestamp row from the kinds table,
  drop 'timestamp shape' from the validate_bytes constraint list, remove
  the residual 'upcoming alkcall crate' sentence from the crate
  independence section (alkcall exists now), fix two '19 kinds' refs in
  the documentation pointer list and schema-layer link.
- src/data_access.rs: module doc comment still said 'all 19 AlkType
  kinds' -> 18.
- bast-format.md (normative): 19 -> 18 AlkTypeKind enum variants.
- layout-engine.md: cross-reference to 'the 19 AlkType kinds' -> 18.
- ADR-BAST (bast-bast-format.md): the Decision section claimed the post-
  pivot enum has '19 unchanged' variants; now 18, with the wording
  adjusted so it no longer says 'unchanged' across the pivot.
- ADR-VAL-SPLIT: drop 'timestamp shape' from the value-domain constraint
  list and the validator-arm table row (validate_timestamp no longer
  exists).

Left as historically accurate (describe the v0.1.0 pre-pivot state):
ADR-003/004/006 Context mentions of AlkType:Timestamp, ADR-005 'engine
now has 19', and the '19 jsonschema::Keyword factories' references in
the What-is-removed sections of ADR-BAST and ADR-VAL-SPLIT.

Verification: cargo test --release (407 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean), cargo build --target
wasm32-unknown-unknown --release (clean).
2026-08-16 09:45:01 +00:00
deepseek-v4-pro 510553d800 Remove the Timestamp kind
The Timestamp kind was a residual from an early research reference. It
was byte-identical to String everywhere (length-prefixed UTF-8) and its
only distinguishing behavior was a hand-rolled non-strict RFC 3339 check
that the docs admitted was incomplete (Feb 31 passes, seconds range
unchecked, no leap seconds). JSON-level timestamp validation is
jsonschema's job (format: date-time on the validate_json path), not
alktype's.

Removes the AlkTypeKind::Timestamp variant, its to_bast_str/from_bast_str
mapping, the builder's Schema::timestamp() constructor, the
validate_timestamp/is_rfc3339_timestamp validator arms, and the
materializer/reader/engine timestamp arms. Updates the meta-schema
primitive enum (14 -> 13), the spec docs (bast-format.md, schema-layer.md,
builder.md, validation.md, overview.md, data-access.md, README.md), and
the kind-count references (19 -> 18).

Verification: cargo test --release (407 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean), cargo build --target
wasm32-unknown-unknown --release (clean).
2026-08-16 09:12:33 +00:00
deepseek-v4-pro 62ed009281 Resolve review #003 nits N1, N3, N4
- N1: number_from_f64 now returns AlkTypeError::Access for non-finite
  floats instead of silently materializing Value::Null. A NaN/Inf in the
  buffer surfaces as a clear access error rather than a misleading
  'expected a number' validation error.
- N3: drop the dead Value::String arm from check_bytes; the materializer
  only ever emits bytes as an array of u8. Update the validation-model
  docs to match.
- N4: fix stale ADR references in doc comments (ADR-096 -> ADR-002,
  ADR-097 -> ADR-003, ADR-098 -> ADR-004, ADR-101 -> ADR-007).

N2 (BastType::alk_kind returning Struct for ) left as documented;
every current caller resolves the ref first, so forcing Option through
the call sites is churn without benefit.

Verification: cargo test --release (411 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean).
2026-08-16 08:59:41 +00:00
deepseek-v4-pro 230345a867 Validate BAST documents against the meta-schema at compile time
Resolves review #003 finding M3.

- Add bast_meta::validate_bast_doc and call it from AlkTypeEngine::compile
  before parsing. Malformed annotations (unknown endian/encoding strings,
  non-integer align/maxLength, missing required properties) now surface as
  AlkTypeError::Schema instead of being silently tolerated by the parser.
- Align the meta-schema TypeRef with the parser: allow inline struct/union/
  enum as TypeRefs (the parser and builder already accepted them; the
  meta-schema and spec did not). Update bast-format.md TypeRef table and
  the BastType doc comment to the seven-form vocabulary.

Verification: cargo test --release (410 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean), cargo build --target
wasm32-unknown-unknown --release (clean).
2026-08-16 08:54:57 +00:00
deepseek-v4-pro e5c7cc1ca2 Fix offset-indirect, field-level endian, and aligned materialization
Resolves review #003 findings M1, M2, L1, and L2.

- offset-indirect (L1 + M2 arm): rework read_*_indirect to same-buffer
  absolute offsets (safetensors-style, per the metatensor model) and add
  write_*_indirect. Wire the encoding dispatch into engine.read_field/
  write_field and the aligned materializer. Previously the encoding was
  laid out but never read back.
- field-level endian override (M1): thread field.effective_endian through
  sequential_reader, engine.read_field/write_field, and
  materialize_struct_aligned. Previously a per-field endian override was
  silently ignored, misreading multi-byte values.
- aligned-mode materialization (M2): materialize fixed-size arrays via
  their vals[i] offset-map entries and maxLength reservations as
  zero-padded fixed-size slices (trailing NULs trimmed). Previously both
  read garbage or errored.
- dead endian param (L2): drop the ignored endian argument from
  materialize_packed/materialize_aligned; endianness is read from the
  root struct.

Verification: cargo test --release (404 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo build --target wasm32-unknown-unknown
--release (clean), cargo llvm-cov --release (89.60% lines).
2026-08-16 08:38:26 +00:00
deepseek-v4-pro ec73440c19 Add post-BAST-pivot code review (#003)
Covers the whole crate after the BAST pivot: correctness, code smell,
panic safety, and coverage (cargo-llvm-cov). 3 Medium findings (field-
level endian override ignored, aligned-mode validate_bytes broken for
arrays/maxLength/offset-indirect, meta-schema never applied at compile
time), 2 Low (offset-indirect dead code, dead endian param), 4 Nits.
Timestamp removal recorded as a publisher decision, tracked separately.

Verification: cargo test --release (389 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo llvm-cov --release (90.14% lines / 86.68%
functions).
2026-08-15 14:43:09 +00:00
glm-5.2 562284faf4 Rewrite README for BAST pivot
The README still described the v0.1.0 custom-keyword JSON Schema format
(`AlkType:*` kinds, `compile(&mut schema, mode)`, single jsonschema
validator). Rewrite it for the BAST-era shipped API:

- What-it-is table: two validation specs (bytes via BAST-native, JSON
  via consumer-provided JSON Schema) instead of one.
- Usage example: `Definitions::new().build_doc(ChunkHeader, ...)` +
  `AlkTypeEngine::compile(&doc, ChunkHeader, LayoutMode::Packed,
  None)`; added a json!-literal variant showing the BAST document
  shape.
- The 19 kinds table: lowercase BAST `kind` strings instead of
  `AlkType:*` keywords; note the enum index bounds fix.
- Two layout modes: updated `compile` signature.
- Variable-length handling: BAST field-level `encoding` annotation.
- Union discriminators: `kind: union`, lazy variant $ref
  resolution, D-BAST-005 fields requirement.
- Endianness: struct-level (was 'top-level schema').
- Validation: two validators (ADR-VAL-SPLIT), BAST-native for bytes,
  standard jsonschema for JSON; uniform AlkTypeError::Validation
  payload (D-BAST-009); enum bounds fix noted.
- New 'BAST document shape' section: $defs required, root_name
  parameter, $ref restricted to #/$defs/<name>, BAST_META_SCHEMA
  re-export.
- Untrusted input: BAST document (was 'schema'); BAST parser
  preserves the no-panic invariant.
- Documentation index: updated for the new/rewritten docs and the
  ADR-BAST / ADR-VAL-SPLIT additions.

Verified both code examples compile and pass against the shipped API
(via a throwaway integration test, since removed).
2026-08-15 14:13:34 +00:00
glm-5.2 62270b03ca Sync architecture docs and ADRs to BAST pivot (steps 9-10)
Step 9 (convert tests to BAST format) was a no-op: steps 4-8 converted
the tests as they went. The only remaining  reference in
src/tests was the intentional  rejection test at
src/schema.rs:462 (asserting the old keyword form is rejected). Full
suite passes: 389 tests (312 lib + 77 integration).

Step 10 (sync architecture docs and ADRs):

Descriptive docs rewritten/updated for BAST:
- schema-layer.md: rewritten for the BAST parser (BastDoc/BastDef/
  BastType typed tree, AlkTypeKind enum with to_bast_str/from_bast_str,
  what was removed). Points at bast-format.md for the normative format.
- validation.md: rewritten for the two-validator model
  (bast_validation for validate_bytes, standard jsonschema for
  validate_json). Documents the repurposed build_validator, the
  AlkTypeError::Validation uniform payload (D-BAST-009), and what is
  removed.
- builder.md: updated all output examples to BAST JSON
  (struct_() -> { kind: struct, fields: [...] }; object() -> standard
  JSON Schema). Documents build_doc, count(), and the field-name union
  fields requirement (D-BAST-005).
- overview.md: updated for BAST (what/why, schema-is-the-format table,
  dependencies, architecture pointers, design decisions table).
- README.md (architecture index): updated document table, ADR table
  (new ADR-BAST + ADR-VAL-SPLIT, superseded ADR-001), OQ table
  (OQ-007/OQ-008 resolutions updated for BAST-native validator), and
  key design principles (#1, #2, #7, #10 reworded for BAST).
- data-access.md: updated tunion function signatures to BastUnion and
  the variant resolution to return BastType (resolve_typeref for refs).
- layout-engine.md: updated construct signatures
  (LayoutBuilder::new(bast_doc, root_name), OffsetMap::compute(&doc),
  SequentialReader::new(bast_doc, root_name)), the recursive-walk
  description (BAST typed tree), and composite-kind headings
  (TStruct/TUnion/TArray -> struct/union/array). Added D-BAST-004
  note on array count requirement.

New ADRs:
- ADR-BAST (bast-bast-format.md): the BAST format, meta-schema,
  //kind vocabulary, design principles, what is removed, the
  enum index bounds bug fix. Supersedes ADR-001's format-specific
  content; records D-BAST-001..009.
- ADR-VAL-SPLIT (val-split-two-validator-model.md): the two-validator
  model (BAST-native for validate_bytes, standard jsonschema for
  validate_json), the repurposed build_validator, the uniform
  AlkTypeError::Validation payload. Refines ADR-004's validation
  strategy and ADR-010's validation step; records D-BAST-006/007/009.

Amended ADRs (supersession/amendment notes added; original decision
text preserved as historical record):
- ADR-001: format-specific content superseded by ADR-BAST;
  purpose/scope and schema-is-the-format principle retained.
- ADR-002: unchanged under the pivot; one-line note that the input
  format changed but the modes didn't.
- ADR-003: annotation semantics retained; annotation location moved
  to BAST type-level properties (amended by ADR-BAST).
- ADR-004: AlkTypeError enum retained (D-BAST-009); validation
  strategy section refined by ADR-VAL-SPLIT.
- ADR-009: builder API surface retained; build() output format
  amended to BAST / standard JSON Schema by ADR-BAST (D-BAST-008).
- ADR-010: validate_bytes two-step concept retained; validation step
  amended to the BAST-native validator by ADR-VAL-SPLIT.

Other:
- Cargo.toml description: JSON Schema with AlkType:* custom keywords
  -> BAST document.
- bast-pivot.md research record: status draft -> implemented, with a
  pointer to the ADRs that superseded its decisions.
- bast-implementation.md plan: status draft -> complete, with a note
  that step 9 was a no-op and step 10 is this commit.
- open-questions.md: OQ-006/OQ-007/OQ-008 resolutions updated for the
  BAST-native validator.
- questions/008-unionvalidator-variant-dispatch.md: added a
  post-BAST-pivot note pointing to the current bast_validation
  implementation; v0.1.0 resolution text preserved as historical
  record.

Verification:
- cargo test --release: 389 pass (312 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cross-reference check: every relative link in the new/updated docs
  resolves (verified by script).
2026-08-15 14:03:21 +00:00
glm-5.2 54fd112fde Remove v0.1.0 custom-keyword machinery (step 8)
The BAST parser (step 3) and BAST-native validator (step 5) replaced
the v0.1.0 custom-keyword accessor layer; step 7 moved the builder to
BAST output. This step removes the now-dead code:

Removed from src/schema.rs:
- get_alktype_kind / get_alktype_kind_enum /
  get_alktype_kind_loose / get_alktype_kind_loose_enum
  (replaced by the BAST parser's kind dispatch)
- normalize_refs / inline_union_variant_refs + helpers
  (BAST refs are always #/$defs/<name>; resolution is a single
  hash lookup, variant refs resolve lazily)
- parse_encoding / parse_align / parse_max_length / parse_endian
  (bast.rs has its own BAST-property-form copies)
- parse_discriminator + DiscriminatorKind
  (replaced by bast::BastDiscriminator; builder has its own
  Discriminator enum)
- resolve_ref / resolve_ref_or_inline
  (replaced by BastDoc::lookup_def / resolve_typeref)
- FromStr impl, as_str, Endian::from_schema, ALKTYPE_PREFIX,
  BYTE_DISCRIMINATOR_TYPES, and the associated unit tests

Kept: AlkTypeKind enum + methods (type_size, natural_alignment,
is_fixed_size, needs_endian, is_composite, is_variable_length,
to_bast_str, from_bast_str), Display (now backed by to_bast_str),
Endian, VariableEncoding, U32_SIZE, DISCRIMINATOR_PATH.

src/lib.rs: dropped the 13 schema::* helper re-exports and
DiscriminatorKind from the public surface; kept Endian, AlkTypeKind,
VariableEncoding.

Doc/comment updates: bast.rs, builder.rs, engine.rs, error.rs —
removed references to the deleted functions and the AlkType:*
keyword form.

The jsonschema crate remains a dependency (validate_json path +
BAST meta-schema validation); build_validator was already repurposed
in step 6 (no custom keywords).

Verification:
- cargo test --release: 389 pass (312 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cargo build --target wasm32-unknown-unknown --release: clean
2026-08-15 13:39:05 +00:00
glm-5.2 45f3336201 Builder API produces BAST JSON (step 7)
- Schema internals: flat Map<String, Value> → Repr enum distinguishing
  BAST primitives (bare TypeRef strings), BAST composites (struct/
  union/enum/array/record objects), $ref, standard JSON Schema
  objects, and raw adopted values. Annotations (endian/align/encoding/
  maxLength) stored on the Schema and placed correctly by build()
  (struct-level) or field() (field-level).

- Primitive constructors (uint32, string, bytes, etc.) now produce
  bare BAST kind strings ("uint32") instead of {"AlkType:Uint32":true}.

- struct_().field(...) produces {"kind":"struct","fields":[{name,kind,...annos}]}
  with an ordered fields array (BAST design principle #4 — no reliance
  on preserve_order for field order).

- union_() emits {"kind":"union","discriminator":{...},"mapping":{...}}
  with BAST kind strings in the discriminator ("uint8" not
  "AlkType:Uint8"). Field-name discriminator unions emit the fields
  array; byte-offset unions omit it.

- enum_of() produces {"kind":"enum","values":[...]}.

- array_of(element).count(n) produces {"kind":"array","element":...,"count":n}.
  .count() is an additive method (D-BAST-004 requires count; the
  array_of signature is unchanged per the semver contract).

- record_of(values) produces {"kind":"record","values":...}.

- object()/string_()/etc. unchanged — standard JSON Schema output.

- encoding() now stores the value on the Schema; field() extracts it
  as a field-level property. No more keyword-object duality.

- max_length() on standard types emits the standard keyword; on BAST
  types it is extracted by field() as a field-level constraint.

- Definitions::build_doc(root_name, root) — new additive method
  producing a complete BAST document with the root type inside $defs
  (where BAST requires it). The old build()/merge_into() remain for
  backward compatibility.

- All builder tests updated to expect BAST output shapes. New tests:
  field annotations, field-name union with fields, array with count,
  ref element, encoding emission, Definitions::build_doc round-trip
  through compile(), SFTP-style union compile, array/record compile.

- Module doc comments updated (builder.rs, lib.rs).

Verification:
- cargo test --release: 427 pass (350 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean (no warnings)
- cargo build --target wasm32-unknown-unknown --release: clean
2026-08-15 13:09:18 +00:00
glm-5.2 ba7f8e1bad validate_json against consumer-provided JSON Schema (step 6)
- AlkTypeEngine::compile gains a 4th param json_schema: Option<&Value>.
  When Some, a standard jsonschema::Validator is built from the
  consumer-provided JSON Schema and stored for the JSON-validation
  path. When None, validate_json returns AlkTypeError::Schema and
  is_valid_json returns false (D-BAST-007).

- validate_json / is_valid_json signatures unchanged (per semver
  contract). Behavior: they now validate against the consumer JSON
  Schema, not a custom-keyword validator built from the alktype
  schema. The BAST document is not involved in this path.

- build_validator repurposed (deferred decision #2): same signature,
  now builds a standard jsonschema::Validator with no custom keywords.
  Behavioral break, not a type break. Re-export kept.

- Removed the 19 jsonschema::Keyword implementations and the 4
  define_*_validator! macros (dead on the bytes path since step 5,
  now dead on the JSON path too). The is_rfc3339_timestamp helper
  lives on in bast_validation.rs (already copied there in step 5).

- All compile call sites updated to pass None for json_schema (the
  layout/read/write/validate_bytes tests don't need JSON validation).

- New tests: validate_json accepts/rejects against consumer JSON
  Schema, returns Schema error when no JSON Schema supplied,
  is_valid_json false when no schema, independence from BAST doc,
  malformed JSON Schema build error, nested object JSON Schema.

Verification:
- cargo test --release: 409 pass (332 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo doc --no-deps: clean
2026-08-15 12:53:35 +00:00
glm-5.2 f853dafaf1 Add BAST-native validator for validate_bytes (step 5)
Replace the jsonschema custom-keyword validator on the bytes path with
a recursive walker over the BAST typed tree. The materializer already
guarantees structural correctness; the validator enforces only the
value-domain constraints expressed in the BAST document.

- New `src/bast_validation.rs`: `validate_value` dispatches on
  `BastType`, resolving `$ref`s lazily via `BastDoc::resolve_typeref`.
  Constraint arms: integer ranges (Int8..Uint64), float finiteness,
  string/bytes `maxLength` (from `BastField::max_length`), RFC 3339
  timestamp shape, enum index bounds, union variant dispatch (recurses
  into the variant, recovering OQ-008 per-variant constraints), struct
  field presence, array count, record values.
- `engine::validate_bytes` now calls `bast_validation::validate_value`
  instead of `self.validator.validate`. The `validator` field is still
  built and used by `validate_json`/`is_valid_json` (step 6 reworks
  those).
- Enum index bounds check (`idx < values.len()`) fixes the v0.1.0 dead
  constraint: the built-in `enum` keyword checked string membership, but
  the materializer emits a numeric index that never matched.
- Errors constructed via `jsonschema::ValidationError::custom` so
  `AlkTypeError::Validation` keeps its payload type uniform with the
  `validate_json` path (D-BAST-009).
- `lib.rs`: add `pub mod bast_validation;` (engine-internal, not
  re-exported in the `pub use` block) and update the module doc.

Verification:
- cargo test --release: 446 pass (369 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo doc --no-deps: clean
2026-08-15 12:44:31 +00:00
glm-5.2 04573e1d86 Wire compile() to BAST document + root name (step 4)
Step 4 of the BAST pivot: the layout engines, materializer, tunion
dispatch, and engine now consume the BAST typed tree (BastDoc/
BastStruct/BastField/BastType/...) instead of walking raw JSON with
get_alktype_kind*.

Breaking changes (per the pivot plan's semver contract):
- AlkTypeEngine::compile signature:
    compile(schema: &mut Value, mode)
    -> compile(bast_doc: &Value, root_name: &str, mode)
  Drops &mut (BAST needs no in-place normalize_refs); adds required
  root_name (D-BAST-001); input is a BAST document, not a custom-keyword
  JSON Schema.
- OffsetMap::compute, LayoutBuilder::new, SequentialReader::new now take
  a BAST document (&Value) + root_name (or &BastDoc) instead of a
  v0.1.0 schema.
- tunion::read_byte_discriminator / read_field_discriminator /
  resolve_variant / discriminator_size now take &BastUnion instead of
  &Value.
- materialize::materialize_packed / materialize_aligned now take
  &BastDoc instead of &Value.

Key design points:
- The engine stores a clone of the BAST Value + root_name so
  sequential_reader() and read_field() can re-parse the typed tree on
  demand without lifetime entanglement with the caller's Value.
-  resolution is a single hash lookup via BastDoc::resolve_typeref;
  no normalize_refs, no inline_union_variant_refs.
- bast.rs gains BastField::synthetic() (pub(crate)) for constructing
  synthetic fields wrapping TypeRefs (array elements, record values,
  union variants — these aren't fields and carry no field annotations).
- The v0.1.0 schema.rs helpers and validation.rs custom-keyword
  validators remain defined (step 8 removes them). build_validator still
  runs on the BAST doc — with no AlkType:* keywords present, the custom
  factories don't trigger and jsonschema performs structural validation
  only. The validate_bytes value-constraint enforcement (maxLength, enum
  bounds) is step 5's concern (the BAST-native validator).

Tests:
- All engine, layout, materialize, tunion, and integration tests
  converted to BAST format (kind/fields vocabulary, /
  composition). Expected validation outcomes for the layout path are
  identical; the maxLength/enum-bounds validate_bytes tests are step 5's
  regression target.
- builder.rs::builder_chunk_header_compiles_in_packed_mode uses a
  hand-written BAST doc (the builder still emits v0.1.0 format; step 7
  converts it).

Verification:
- cargo test --release: 425 pass (348 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo doc --no-deps: clean
2026-08-15 12:29:27 +00:00
glm-5.2 f2f9c0326c Add BAST document parser (typed tree, step 3)
- New src/bast.rs: typed surface over a BAST document — BastDoc,
  BastDef, BastDefKind, BastStruct, BastField, BastUnion,
  BastDiscriminator, BastEnum, BastType, BastRef, BastArray,
  BastRecord. Borrows from the source Value (no clone of the tree).
- $ref resolution is a single hash lookup against $defs
  (#/$defs/<name> only); union variant refs resolved lazily via
  BastDoc::resolve_typeref / resolve_ref — replaces
  normalize_refs + inline_union_variant_refs (those stay for now;
  step 8 removes them).
- Untrusted-input safe: every walk returns AlkTypeError::Schema on
  a malformed document, never panic/unwrap (AGENTS.md §3).
  Overflow-safe usize parsing via try_from (AGENTS.md §4).
- D-BAST-005 enforced: fields array only valid with field-name
  discriminators; required for them.
- Additive only: new pub mod bast + re-exports in lib.rs. No
  existing re-exports removed (those go in step 8). 47 new tests.

Verification:
- cargo test --release: 465 pass (379 lib + 86 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cargo build --target wasm32-unknown-unknown --release: clean
2026-08-15 11:56:03 +00:00
glm-5.2 29134789a9 Embed BAST v1 meta-schema as BAST_META_SCHEMA
Step 2 of the BAST pivot. Adds the normative BAST meta-schema
(Draft 2020-12 JSON Schema) as a public serde_json::Value, embedded
at compile time and available for validating BAST document
well-formedness. Copied verbatim from docs/architecture/bast-format.md
§The Meta-Schema.

- New src/bast_meta.rs: BAST_META_SCHEMA static, built lazily via
  LazyLock (the json! macro allocates, so it can't be a const;
  parsed once, reused as &'static Value thereafter).
- src/lib.rs: pub mod bast_meta + re-export BAST_META_SCHEMA.
  Additive public surface.

The meta-schema validates structure (correct $defs shape, known
kind strings, required properties, no additional properties,
$ref restricted to #/$defs/<name>). Value-domain constraints
(maxLength, enum index bounds) are enforced by the BAST-native
validator (step 5), not this meta-schema.

Verification:
- cargo test --release: 418 tests pass (332 lib + 86 integration);
  16 new unit tests exercise the meta-schema against valid and
  invalid BAST documents (missing $defs, unknown kind, additional
  properties, byte-discriminator union, enum, array+count, record,
  $ref, malformed ref, empty enum).
- cargo clippy --all-targets -- -D warnings: clean.
- cargo build --target wasm32-unknown-unknown --release: clean
  (LazyLock + serde_json::json! macro are wasm-safe).
2026-08-15 11:37:14 +00:00
glm-5.2 66ab9d7d93 Add AlkTypeKind::to_bast_str/from_bast_str for BAST kind strings
Step 1 of the BAST pivot. Adds the lowercase-string mapping
("uint32" <-> AlkTypeKind::Uint32) that the BAST parser and
validator dispatch on (D-BAST-002).

- to_bast_str(self) -> &'static str: returns the lowercase BAST
  string for all 19 variants (14 primitives + struct/union/array/
  record/enum). Boolean -> "bool", distinct from the PascalCase
  variant name.
- from_bast_str(s) -> Result<AlkTypeKind, AlkTypeError>: inverse of
  to_bast_str; returns AlkTypeError::Schema for unknown strings.

These are additive inherent methods on the already-re-exported enum.
The existing FromStr impl (parsing the v0.1.0 "AlkType:Uint32"
keyword form) is unchanged and removed in step 8. The two surfaces
are deliberately distinct: from_bast_str rejects "AlkType:Uint32"
and FromStr rejects "uint32".

Verification:
- cargo test --release: 402 tests pass (316 lib + 86 integration);
  6 new unit tests cover both directions, round-trip, and the
  from_bast_str/from_str distinctness.
- cargo clippy --all-targets -- -D warnings: clean.
2026-08-15 11:29:34 +00:00
glm-5.2 f5f52c61e8 Decompose BAST pivot doc into normative spec + implementation plan
The bast-pivot.md research doc had grown to 1477 lines (~58KB) through
iterative editing, pushing its most actionable content (D-BAST
decisions, POC result, migration steps) past the 50KB Read tool cap.
Agents peeking at the truncated file landed in duplicated/out-of-order
sections. Decompose into three readable-sized files with distinct roles:

- docs/architecture/bast-format.md (28KB, new): the normative BAST
  format spec -- meta-schema, TypeRef, examples, validation model.
  Grounded in the POC and D-BAST-001..009. Stable and safe to write
  now; schema-layer.md/validation.md stay describing current code and
  are rewritten post-implementation (per AGENTS.md ADR-grounding rule).
- docs/plans/bast-implementation.md (31KB, new): the execution entry
  point -- ordered 10-step plan with per-step goal/files/spec-ref/
  verification, the public-API semver contract table up front as a
  scope-creep guardrail, and the ADR-sync checklist at the end. Each
  step links to the specific bast-format.md section and D-BAST anchor.
- docs/research/bast-pivot.md (28KB, trimmed): now the research record
  only -- Summary, Motivation, POC scope/result, Decisions, Risks,
  References. The normative format spec, what-changes tables,
  validator-split details, and migration steps moved to the two new
  docs; pointers added. 1155 lines removed, 216 added.
- docs/architecture/README.md: index updated to list bast-format.md
  and the two in-progress pivot docs, with notes on schema-layer.md
  and validation.md being rewritten when the pivot lands.

All three files are under the 50KB Read cap, so an implementing agent
gets the whole document in one call. Cross-reference anchors verified
to resolve. No code changes; cargo test --release (396 tests) green.

Verification: cargo test --release (310 crate + 86 integration, all pass).
2026-08-15 10:59:14 +00:00
glm-5.2 5796d1c22f Add semver and ADR impact mapping for the BAST pivot
Scope-creep guardrail for the public API during implementation. Maps
every item re-exported from src/lib.rs to a class (breaking / additive /
unchanged) with the specific change, and every ADR (001-010) to an
action (supersede / amend / unchanged) with the reason.

Net breaking: compile (signature), validate_json/is_valid_json
(contract), Schema::build/Definitions::build (output format),
build_validator (signature or removal), and the ~13 schema::* helper
re-exports. Net additive: BAST parser, BAST-native validator,
AlkTypeKind::from_str/to_str. Net unchanged: the entire layout +
data-access + materialize + tunion layer, AlkTypeError (D-BAST-009),
the Discriminator builder, AlkTypeKind variants.

Flags three small decisions deferred to their implementation steps
(validate_json JSON Schema source, build_validator fate, schema::*
re-export retention) so they don't become drive-by semver changes.

This is a living guide — it may shift slightly during implementation,
but capturing the contract now prevents public-surface drift. Doc-only.
2026-08-15 10:29:24 +00:00
glm-5.2 89f05850f2 Resolve OQ-BAST-001: keep Validation(ValidationError<'static>) (D-BAST-009)
Converts the open question into a closed decision. Rationale: consumer
ergonomics on the combined validate_json + validate_bytes path — one
uniform payload type means one match arm downstream. Option 2
(Validation(String)) would force validate_json to flatten its
structured errors to a String, losing information on the richer path to
accommodate the less rich one. The no_std/minimal-build angle that
option 2 was meant to enable is moot: validate_json requires jsonschema
regardless, so a bytes-only no_std build already has to give up
validate_json as a separate larger decision; dropping the type from one
error variant doesn't unlock it.

Updates Phase 1 step 5 to reference D-BAST-009 for the error
construction pattern, and rewrites POC Result observation 4 from
'deferred decision' to 'decided — see D-BAST-009'.

No semver-relevant change to the Validation variant. Doc-only.
2026-08-15 10:13:47 +00:00
glm-5.2 e77268c951 Record BAST validator POC result; track OQ-BAST-001 error-payload decision
Adds the POC Result section (POC on branch bast-validator-poc, commit
f371fe4 — 20/20 tests, full 416-test suite green, clippy/wasm/doc clean).
Hypothesis confirmed: a recursive walker over the BAST type tree fully
replaces the 19 custom keyword validators on the validate_bytes path,
recovers OQ-008 union variant dispatch, and fixes the enum-membership
dead constraint. The POC code is reference scaffolding on the branch
and is not merged to main — it is superseded by Phase 1 step 5.

Marks Phase 1 step 2 and the POC Scope section as done with pointers to
the result section.

Elevates the deferred AlkTypeError::Validation payload-shape question
to OQ-BAST-001: keep jsonschema::ValidationError<'static> (POC choice,
simplest, dependency stays) vs introduce Validation(String) (drops
jsonschema from the error type; semver-relevant public-API change).
Decision belongs to the production refactor.

Verification: doc-only change, no code touched.
2026-08-15 09:38:53 +00:00
glm-5.2 19f8162f1a Refine BAST pivot: validation model, POC scope, resolve open questions
- Rewrite Validator Split around BAST-native validator for
  validate_bytes (walks BAST, checks value-domain constraints, no
  external JSON Schema needed); validate_json uses standard
  jsonschema::Validator from consumer-provided JSON Schema
- Add targeted POC: BAST-native validator replacing 19 custom keyword
  validators on the bytes path, verified via existing test suite
- Fix meta-schema: require count on arrays (variable-element arrays
  deferred per OQ-001), add optional fields array to UnionDef for
  field-name discriminators
- Remove lying no-count array example, replace with deferred note
- Document dead enum constraint (materialized index never matches
  string-membered enum); BAST-native validator fixes it via index
  bounds check
- Resolve all 7 OQs + 3 spec gaps as D-BAST-001 through D-BAST-008
- Update Migration Path, Risks table, Engine internals to reflect
  BAST-native validator
- Clarify 'no pocs needed' was an overcorrection: layout swap needs
  no POC, but the validation model does

Verification: docs-only change, no code affected
2026-08-15 09:05:55 +00:00
deepseek-v4-pro 82bc8f29c0 Clean up BAST pivot: drop POCs, remove recursion, add spec gaps
- Remove the four proposed POCs: they were implementation smoke tests,
  not de-risking probes. The pivot is a backend swap on a proven layout
  engine; byte-identity is already proven and the layout code is
  unchanged, so there is nothing empirical left to de-risk.
- Remove the recursive TreeNode example and the recursion mention in
  design principle 2: recursion is not a binary-layout concern and the
  engine has no cycle detection.
- Add a Spec Gaps section: validate_bytes semantics after keyword
  validator removal (UnionValidator variant dispatch regression),
  arrays of variable-length elements (engine rejects them), and
  field-name discriminator unions (meta-schema cannot express them).
- Reframe the engine change as an accessor-layer refactor: the
  walkers' (kind, field list, annotations) reads change; everything
  beneath them carries over unchanged.
2026-08-15 07:46:03 +00:00
deepseek-v4-pro 37bd5d7b7d Fix BAST pivot: jsonschema remains a direct dependency
Correct the validator split section to clarify that jsonschema is
not removed — it remains the JSON Schema validator for both paths
(BAST meta-schema validation and standard JSON payload validation).
Only the custom keyword registration path is removed. Also fix the
build_validator and validate_json entries in the engine changes
table to reflect that they are repurposed, not removed.
2026-08-14 16:06:26 +00:00
deepseek-v4-pro 004d52d505 Add BAST pivot research document
Proposes replacing custom JSON Schema keywords with a standalone
kind-based JSON format (BAST) using / for composition.
Covers format design, meta-schema, engine changes, validator split,
codegen future, ABI adapter potential, migration path, 7 open
questions, and 4 proposed POCs.
2026-08-14 15:55:17 +00:00
glm-5.2 5a4ed9e8e3 Exclude internal artifacts from crates.io package
Trim the published package from 64 → 52 files (865.7KiB → 715.3KiB,
202.8KiB → 153.0KiB compressed) by excluding internal-only files:
.opencode/ agent configs, docs/reviews/, docs/research/, and
docs/sdd_process.md. Keeps docs/architecture/ ADRs and OQs, which
document the public design.

Verification:
- cargo test --release: 396 tests pass
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo publish --dry-run --allow-dirty: clean (52 files, 153.0KiB)
2026-08-11 09:45:15 +00:00
45 changed files with 12783 additions and 7648 deletions

No files matched your search

+119
View File
@@ -0,0 +1,119 @@
# Changelog
All notable changes to this crate are documented here. The format is
based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
this crate adheres to [Semantic Versioning](https://semver.org/).
## [0.2.0] - 2026-08-17
A breaking release that replaces the v0.1.0 `AlkType:*` custom-keyword
JSON Schema format with BAST (Binary Abstract Syntax Tree) — a JSON
document that describes binary layouts using a `kind`-based vocabulary
with `$defs`/`$ref` for composition. BAST is itself a valid JSON Schema
instance (it has a meta-schema), making it self-validating,
editor-friendly, and trivially consumable from any language with a JSON
parser. The engine works the same way as before: compile a document
once into an `AlkTypeEngine`, then read/write fields at computed offsets
and validate bytes/JSON. The pivot was made now because v0.1.0 has no
real consumers (≈15 crates.io downloads, mostly bots/scanners), so the
custom-keyword wart could be removed cleanly.
### Breaking changes
- **Schema format.** The v0.1.0 `AlkType:*` custom-keyword JSON Schema
format (`{ "AlkType:Struct": true, "fields": [...] }`) is removed.
Schemas are now BAST documents:
`{ "$defs": { "<TypeName>": { "kind": "struct", "fields": [...] } } }`.
The `kind`-based vocabulary covers 18 binary kinds (integers, floats,
bytes, string, struct, union, enum, array, etc.).
- **`AlkTypeEngine::compile` signature.** Now takes
`(bast_doc: &Value, root_name: &str, mode: LayoutMode, json_schema: Option<&Value>)`.
The root type name is a required parameter — it selects which `$defs`
entry is the top-level type (previously the root was implicit from the
single top-level schema object).
- **Builder API output.** `Definitions`/`Schema`/`Discriminator` now
produce BAST JSON via `Definitions::build_doc(name, schema)`. The
builder method names are unchanged; only the emitted JSON shape
changed. `Schema::struct_()` produces a BAST struct;
`Schema::object()` produces a standard JSON Schema (for the
`validate_json` path).
- **Validation split.** v0.1.0 used a single `jsonschema` validator
with 19 custom `AlkType:*` keywords for both bytes and JSON
validation. 0.2.0 splits this into two independent paths:
- `validate_bytes` uses a new BAST-native validator
(`bast_validation`) — a recursive walker over the BAST type tree.
- `validate_json` / `is_valid_json` use a standard
`jsonschema::Validator` built from a consumer-provided JSON Schema
(passed to `compile` as the `json_schema` parameter). No custom
keywords; BAST is not involved — BAST describes bytes, not JSON
shape.
Both paths return `AlkTypeError::Validation` with a uniform
`jsonschema::ValidationError<'static>` payload.
- **Removed.** The v0.1.0 custom-keyword accessor layer
(`AlkTypeKind::FromStr`, `parse_*`, `resolve_ref*`,
`DiscriminatorKind`) is removed. The BAST parser (`bast` module)
exposes a cleaner typed surface (`BastDoc`/`BastDef`/`BastStruct`/
`BastField`/`BastType`/etc.) that borrows from the source
`serde_json::Value` without cloning field data.
- **Public module surface.** New public modules: `bast`, `bast_meta`,
`bast_validation`, `builder`, `materialize`. The `schema` module is
retained but now holds only `Endian`/`AlkTypeKind`/`VariableEncoding`
(the binary-kind vocabulary); the v0.1.0 custom-keyword machinery is
gone.
### Additions
- **BAST meta-schema.** Embedded in the crate as
`BAST_META_SCHEMA` (re-exported from the crate root) and published at
`https://alk.dev/bast/v1/schema`. BAST documents are validated against
it at compile time (`AlkTypeEngine::compile` calls
`validate_bast_doc` before parsing).
- **`materialize` module.** Materializes a `serde_json::Value` tree from
a binary buffer by walking the BAST typed tree. Used by
`AlkTypeEngine::validate_bytes` (ADR-010).
- **Builder for JSON Schemas.** `Schema::object()` produces a standard
JSON Schema object (for the `validate_json` path), complementing
`Schema::struct_()` which produces a BAST struct (for the bytes path).
One builder, two output shapes — the method name selects which.
### Bug fixes vs v0.1.0
- **Enum index bounds are now checked.** The v0.1.0 validator had a
dead constraint: enum variant indices were never bounds-checked
against `values.len()`. The BAST-native validator enforces it
(`validate_enum` checks `idx < values.len()`).
- **Offset-indirect, field-level endian, and aligned materialization**
bugs found during review #003 are fixed.
### Non-breaking improvements
- `$ref` is restricted to `#/$defs/<name>` — one hash lookup, no
`normalize_refs` pass (the v0.1.0 engine needed one).
- Schemas remain untrusted input: every engine path that walks a BAST
document returns `Err` on a malformed document, never `panic!`/
`unreachable!`. Overflow-safe arithmetic (`checked_add`,
`usize::try_from`) on all offset/count casts.
- Still two dependencies (`jsonschema` with `default-features = false`,
`serde_json` with `preserve_order`), no `async`, no `unsafe`, no
platform deps, no feature flags. Compiles to
`wasm32-unknown-unknown`.
### Upgrade notes
There is no migration path from v0.1.0 `AlkType:*` schemas — the format
is incompatible. Rewrite schemas as BAST documents (the `builder` API
produces them; see the README usage example) and update `compile` calls
to pass the root type name and the optional JSON Schema. The
read/write/validate API surface (`read_field`, `write_field`,
`sequential_reader`, `validate_bytes`, `validate_json`,
`is_valid_json`) is unchanged.
## [0.1.0] - 2025-11-10
Initial crates.io release. Custom-keyword JSON Schema format
(`AlkType:*`), single `jsonschema` validator for both bytes and JSON,
`AlkTypeEngine` with packed/aligned layout modes, builder API producing
`serde_json::Value`.
[0.2.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.2.0
[0.1.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.1.0
Generated
+1 -1
View File
@@ -27,7 +27,7 @@ dependencies = [
[[package]]
name = "alktype"
version = "0.1.0"
version = "0.2.0"
dependencies = [
"jsonschema",
"serde_json",
+3 -2
View File
@@ -1,14 +1,15 @@
[package]
name = "alktype"
version = "0.1.0"
version = "0.2.0"
edition = "2021"
rust-version = "1.85"
license = "MIT OR Apache-2.0"
description = "Binary struct engine: takes a JSON Schema with AlkType:* custom keywords and produces an offset map, read/write functions, and validation"
description = "Binary struct engine: takes a BAST (Binary Abstract Syntax Tree) document and produces an offset map, read/write functions, and validation"
repository = "https://git.alk.dev/alkdev/alktype"
readme = "README.md"
keywords = ["binary", "jsonschema", "wire-format", "serialization", "layout"]
categories = ["encoding", "data-structures", "parsing"]
exclude = [".opencode/", "docs/reviews/", "docs/research/", "docs/sdd_process.md", "Cargo.lock", "AGENTS.md"]
[lib]
name = "alktype"
+168 -91
View File
@@ -1,51 +1,61 @@
# alktype
The binary struct engine: a small Rust crate that takes a JSON Schema
with `AlkType:*` custom keywords and produces an offset map, read/write
The binary struct engine: a small Rust crate that takes a BAST (Binary
Abstract Syntax Tree) document and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
`alktype` is a standalone crate with **two dependencies**: `jsonschema`
(for validation) and `serde_json` (for schema parsing). No tokio, no
platform deps, no `unsafe`. Compiles to `wasm32-unknown-unknown`.
(for JSON validation and BAST meta-schema validation) and `serde_json`
(for BAST document parsing). No tokio, no platform deps, no `unsafe`.
Compiles to `wasm32-unknown-unknown`.
## What it is
A JSON Schema annotated with `AlkType:*` custom keywords serves three
roles simultaneously:
BAST is a JSON document that describes binary data layouts using a
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
itself a valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser. See
[`docs/architecture/bast-format.md`](docs/architecture/bast-format.md)
for the normative format spec.
A BAST document serves three roles simultaneously:
| Role | Mechanism | When |
|------|-----------|------|
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
| **Validation spec (bytes)** | BAST-native validator (recursive walker over the BAST type tree) | Access time (`validate_bytes`) |
| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) |
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map / packed layout) |
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
No separate format definition, no separate parser, no separate
validator. The schema is the single source of truth for the binary
format. Adding a new field to a protocol is adding a property to the
schema JSON — the engine computes the new offsets automatically.
validator. The BAST document is the single source of truth for the
binary format. Adding a new field to a protocol is adding an entry to
the BAST `fields` array — the engine computes the new offsets
automatically.
This is the same principle as `#[repr(C)]` struct field access, but at
runtime from a portable JSON Schema instead of at compile time from
language-specific annotations. The schema is the ABI contract.
runtime from a portable JSON document instead of at compile time from
language-specific annotations. The BAST document is the ABI contract.
## Usage
Build the schema with the fluent Rust builder (ADR-009), compile it
once into an [`AlkTypeEngine`], then read/write fields at computed
Build the BAST document with the fluent Rust builder (ADR-009), compile
it once into an [`AlkTypeEngine`], then read/write fields at computed
offsets:
```rust
use alktype::{AlkTypeEngine, Endian, LayoutMode, Schema, FieldValue};
use alktype::{AlkTypeEngine, Definitions, Endian, LayoutMode, Schema, FieldValue};
// Channels' 8-byte chunk header: big-endian, packed mode.
let mut schema = Schema::struct_()
let doc = Definitions::new().build_doc("ChunkHeader", Schema::struct_()
.endian(Endian::Big)
.field("channel_id", Schema::uint32())
.field("length", Schema::uint32())
.build();
.field("length", Schema::uint32()));
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
// `json_schema: None` — no JSON-validation path needed for a binary-only schema.
let engine = AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)?;
// Write a frame into a buffer. For fixed-size structs, the byte
// positions are a direct read off the layout — channel_id at 0,
@@ -55,8 +65,8 @@ let mut buf = vec![0u8; 8];
alktype::data_access::write_u32(&mut buf, 0, 42, "channel_id", Endian::Big)?;
alktype::data_access::write_u32(&mut buf, 4, 7, "length", Endian::Big)?;
// Validate the bytes against the schema in one call.
engine.validate_bytes(&buf)?; // materializes a Value, then validates
// Validate the bytes against the BAST document in one call.
engine.validate_bytes(&buf)?; // materializes a Value, then runs the BAST-native validator
// Read the frame back sequentially (packed mode is sequential by
// construction — variable-length fields shift subsequent fields).
@@ -67,43 +77,64 @@ assert_eq!(value, FieldValue::U32(42));
# Ok::<(), alktype::AlkTypeError>(())
```
Schemas may also be authored as plain `serde_json::json!{...}` literals
and passed directly to `AlkTypeEngine::compile` — the builder is a
construction convenience, not a requirement.
BAST documents may also be authored as plain `serde_json::json!{...}`
literals and passed directly to `AlkTypeEngine::compile` — the builder
is a construction convenience, not a requirement.
## The 19 `AlkType:*` kinds
```rust
use alktype::{AlkTypeEngine, LayoutMode};
use serde_json::json;
| Kind | Rust type | Size | Notes |
|------|-----------|-----:|-------|
| `AlkType:Int8` | `i8` | 1 | |
| `AlkType:Int16` | `i16` | 2 | endian-sensitive |
| `AlkType:Int32` | `i32` | 4 | endian-sensitive |
| `AlkType:Int64` | `i64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `AlkType:Uint8` | `u8` | 1 | |
| `AlkType:Uint16` | `u16` | 2 | endian-sensitive |
| `AlkType:Uint32` | `u32` | 4 | endian-sensitive; also the enum/string/bytes length-prefix width |
| `AlkType:Uint64` | `u64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `AlkType:Float32` | `f32` | 4 | endian-sensitive; NaN/inf rejected by validator |
| `AlkType:Float64` | `f64` | 8 | endian-sensitive; NaN/inf rejected by validator |
| `AlkType:Boolean` | `bool` | 1 | |
| `AlkType:Enum` | `u32` index | 4 | index into the schema's `"enum"` array |
| `AlkType:String` | length-prefixed UTF-8 | 4 + N | `[length: u32][bytes]` by default |
| `AlkType:Bytes` | length-prefixed raw bytes | 4 + N | `[length: u32][bytes]` by default |
| `AlkType:Timestamp` | length-prefixed RFC 3339 | 4 + N | non-strict string check (see inline docs) |
| `AlkType:Struct` | record of fields | composite | nested; field paths are dotted (`"header.version"`) |
| `AlkType:Union` | tagged union | composite | byte-offset or field-name discriminator |
| `AlkType:Array` | repeated element | composite | fixed-size elements with stride, or variable count |
| `AlkType:Record` | string-keyed map | composite | `[count: u32][key, value]...` |
let doc = json!({
"$defs": {
"ChunkHeader": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "channel_id", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
}
}
});
let engine = AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)?;
# Ok::<(), alktype::AlkTypeError>(())
```
The engine recognizes a kind when the schema object has a key starting
with `AlkType:` whose value is `true` (the boolean shorthand) or an
annotation object (e.g. `{ "AlkType:String": { "encoding": "offset-indirect" } }`).
## The 18 BAST kinds
| `kind` | Rust type | Size | Notes |
|--------|-----------|-----:|-------|
| `int8` | `i8` | 1 | |
| `int16` | `i16` | 2 | endian-sensitive |
| `int32` | `i32` | 4 | endian-sensitive |
| `int64` | `i64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `uint8` | `u8` | 1 | |
| `uint16` | `u16` | 2 | endian-sensitive |
| `uint32` | `u32` | 4 | endian-sensitive; also the enum/string/bytes length-prefix width |
| `uint64` | `u64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `float32` | `f32` | 4 | endian-sensitive; NaN/inf rejected by validator |
| `float64` | `f64` | 8 | endian-sensitive; NaN/inf rejected by validator |
| `bool` | `bool` | 1 | `0x00`=false, `0x01`=true |
| `enum` | `u32` index | 4 | index into the `values` array; bounds-checked by the BAST-native validator |
| `string` | length-prefixed UTF-8 | 4 + N | `[length: u32][bytes]` by default |
| `bytes` | length-prefixed raw bytes | 4 + N | `[length: u32][bytes]` by default |
| `struct` | record of fields | composite | nested; field paths are dotted (`"header.version"`) |
| `union` | tagged union | composite | byte-offset or field-name discriminator |
| `array` | repeated element | composite | fixed-size elements with stride; `count` required in v1 (D-BAST-004) |
| `record` | string-keyed map | composite | `[count: u32][key, value]...` |
The 18 kinds map to the `AlkTypeKind` Rust enum. `AlkTypeKind::to_bast_str`/
`from_bast_str` convert between the enum and the lowercase BAST strings
(D-BAST-002). Only `struct`, `union`, and `enum` can appear as named
`$defs` entries; primitives, arrays, and records appear as field/element/
value types via [TypeRef](docs/architecture/bast-format.md#typeref).
## Two layout modes
The consumer selects the layout mode at engine construction time via
`AlkTypeEngine::compile(schema, mode)`. The same schema can be compiled
in either mode. Decided in ADR-002.
`AlkTypeEngine::compile(bast_doc, root_name, mode, json_schema)`. The
same BAST document can be compiled in either mode. Decided in ADR-002.
| Mode | Use case | Read API | Write API |
|------|----------|----------|-----------|
@@ -118,60 +149,100 @@ in either mode. Decided in ADR-002.
- **Aligned mode**: a 4-byte length prefix sits at a known offset; the
variable data is not part of the static layout. Offset indirection
(the metatensor blob pattern: `{offset, length}` pointing into a
separate data region) is opt-in via the `encoding` annotation.
separate data region) is opt-in via the field-level `encoding`
annotation. Fixed-size reservation via `maxLength` is also supported.
### TUnion discriminators
### Union discriminators
`AlkType:Union` supports two discriminator kinds (ADR-003):
`kind: "union"` supports two discriminator kinds (ADR-003):
- **Byte-offset** — a fixed-size integer at a known byte offset. The
SFTP `Packet` pattern: byte 0 is the type byte, bytes 1..N are the
variant struct. Mapping keys are stringified integers.
- **Field-name** — a named field within the struct. The TypeBox
- **Byte-offset** — a fixed-size integer (`uint8`/`uint16`/`uint32`) at
a known byte offset. The SFTP `Packet` pattern: byte 0 is the type
byte, bytes 1..N are the variant struct. Mapping keys are stringified
integers.
- **Field-name** — a named field within the union. The TypeBox
`typedef.ts` pattern. Mapping keys are string values matching the
discriminator field's value.
discriminator field's value. The `fields` array declares the
discriminator field (D-BAST-005).
Variant `$ref`s are resolved lazily — no compile-time inlining step.
## Endianness
Per-schema, default little-endian. Set `"endian": "big"` on the
top-level schema (or via `Schema::endian(Endian::Big)`) and the engine
byte-swaps every multi-byte read/write accordingly. SFTP consumers
specify big-endian; channels' chunk header is big-endian.
Per-schema, default little-endian. Set `"endian": "big"` on the root
struct (or via `Schema::endian(Endian::Big)`) and the engine byte-swaps
every multi-byte read/write accordingly. Field-level `endian` overrides
the struct default. SFTP consumers specify big-endian; channels' chunk
header is big-endian.
## Validation
Two entry points on [`AlkTypeEngine`], one underlying `jsonschema`
validator (ADR-010):
Two entry points on [`AlkTypeEngine`], two validators for two input
types (ADR-VAL-SPLIT):
- `validate_bytes(&[u8])` — for raw byte buffers (channels' chunk
header, SFTP packets). Materializes a `Value` tree from the bytes via
the layout engine, then runs the **BAST-native validator** — a
recursive walker over the BAST type tree that checks the value-domain
constraints the materializer doesn't (integer ranges, `maxLength`,
enum index bounds, union variant constraints). No
`jsonschema` involvement; the BAST document is the complete
validation spec for bytes (D-BAST-006).
- `validate_json(&Value)` / `is_valid_json(&Value)` — for already-parsed
JSON (call's payload schemas).
- `validate_bytes(&[u8])` — materializes a `Value` tree from the bytes
via the layout engine, then validates that `Value`. Single-call binary
buffer validation.
JSON (call's `OperationSpec.input_schema` payloads). Validates
against a **standard `jsonschema::Validator`** compiled at
`AlkTypeEngine::compile` time from a consumer-provided JSON Schema
(the `json_schema: Option<&Value>` parameter). BAST is not involved —
BAST describes bytes, not JSON shape (D-BAST-007).
The validator is compiled once at load time; access-time validation is
a fast `is_valid()` check. High-throughput paths can skip validation;
security-sensitive paths can validate every frame.
Both paths return `AlkTypeError::Validation(jsonschema::ValidationError<'static>)`
— one uniform payload, one match arm (D-BAST-009).
Validation is opt-in per operation. High-throughput paths can skip it;
security-sensitive paths can validate every frame. The BAST-native
validator also fixes a v0.1.0 dead constraint: enum index bounds are
now checked (the materializer emits a numeric index; the validator
checks it against `values.len()`).
## BAST document shape
Every BAST document has the same top-level shape:
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
- The `$defs` block is **required** (D-BAST-003).
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001)
— it selects which `$defs` entry is the top-level type.
- `$ref` is restricted to `#/$defs/<name>` — one hash lookup, no
normalization pass.
The BAST meta-schema is embedded in the crate as `BAST_META_SCHEMA`
(re-exported from the crate root) and published at
`https://alk.dev/bast/v1/schema`. Consumers can validate a BAST
document's structure with any JSON Schema validator. See
[`docs/architecture/bast-format.md`](docs/architecture/bast-format.md)
for the full spec.
## Crate independence
`alktype` does **not** depend on any application or networking crate.
It defines its own types (`AlkTypeError`, `AlkTypeEngine`, `FieldValue`,
etc.) and is usable in contexts where networking doesn't exist — CLI
tools, test harnesses, schema-building utilities, and WASM targets. The
upcoming `alkcall` crate (the `alknet-call` + `alknet-channels`
unification) depends on `alktype` for both binary layout and JSON
payload schemas; `alktype` knows nothing about `alkcall`.
tools, test harnesses, schema-building utilities, and WASM targets.
## Schemas as untrusted input
The crate treats schemas as untrusted input. A malformed schema
returns `AlkTypeError::Schema` / `AlkTypeError::Offset` from any
engine path — never a panic. This matters for hub/spoke topologies
where the remote peer provides the schema (e.g. `alkcall` accepting an
`OperationSpec` from an arbitrary internet peer). All `unreachable!()`
sites in production code were converted to `Err` ahead of v0.1.0
(review #002, L2).
The crate treats BAST documents as untrusted input. A malformed
document returns `AlkTypeError::Schema` from any engine path — never a
panic. This matters for hub/spoke topologies where the remote peer
provides the schema (e.g. `alkcall` accepting an `OperationSpec` from
an arbitrary internet peer). Every `unreachable!()` site in production
code was converted to `Err` ahead of v0.1.0 (review #002, L2); the
BAST parser preserves this invariant — overflow-safe arithmetic
(`checked_add`, `usize::try_from`) on all offset/count casts.
## Documentation
@@ -179,20 +250,26 @@ Architecture documentation lives under [`docs/architecture/`](docs/architecture/
- [Overview](docs/architecture/overview.md) — purpose, "schema is the
format" principle, dependencies, consumers, scope boundaries
- [Schema layer](docs/architecture/schema-layer.md) — the 19 kinds,
jsonschema custom keyword integration, schema annotations
- [BAST format](docs/architecture/bast-format.md) — **normative format
spec**: meta-schema, TypeRef, TypeDef shapes, validation model
- [Schema layer](docs/architecture/schema-layer.md) — the BAST parser
(`BastDoc`/`BastDef`/`BastType` typed tree), the 18 kinds, the
`AlkTypeKind` enum
- [Layout engine](docs/architecture/layout-engine.md) — offset
computation, the two layout modes, alignment, endianness
- [Data access](docs/architecture/data-access.md) — read/write
functions, TUnion dispatch, field paths, zero-copy access
- [Validation](docs/architecture/validation.md) — custom keyword
validators, `AlkTypeError`, load-time vs access-time validation
- [Validation](docs/architecture/validation.md) — the two-validator
model, `AlkTypeError`, load-time vs access-time validation
- [Builder](docs/architecture/builder.md) — fluent Rust API for
constructing alktype JSON Schemas at runtime
constructing BAST documents and standard JSON Schemas at runtime
- [Architecture decisions (ADRs)](docs/architecture/decisions/) —
purpose/scope, two layout modes, schema annotations, error handling,
int64/uint64 kinds, packed-mode read factory, TUnion in aligned mode,
builder API, `validate_bytes`
purpose/scope (ADR-001), BAST format (ADR-BAST), two-validator model
(ADR-VAL-SPLIT), two layout modes (ADR-002), schema annotations
(ADR-003), error handling (ADR-004), int64/uint64 kinds (ADR-005),
non-final inline variable fields (ADR-006), packed-mode read factory
(ADR-007), TUnion in aligned mode (ADR-008), builder API (ADR-009),
`validate_bytes` (ADR-010)
## License
+73 -57
View File
@@ -1,12 +1,12 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-15
---
# alktype
The binary struct engine: a small Rust crate that takes a JSON Schema
with `AlkType:*` custom keywords and produces an offset map, read/write
The binary struct engine: a small Rust crate that takes a BAST (Binary
Abstract Syntax Tree) document and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
@@ -14,60 +14,73 @@ format definition; the engine is generic.
| Document | Status | Description |
|----------|--------|-------------|
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [schema-layer.md](schema-layer.md) | draft | The 19 `AlkType:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [overview.md](overview.md) | accepted | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [`bast-format.md`](bast-format.md) | accepted | **Normative BAST format specification.** Meta-schema, TypeRef, TypeDef shapes (Struct/Union/Enum/FieldDef), examples, validation model. The format the engine consumes. |
| [schema-layer.md](schema-layer.md) | accepted | The BAST parser (`src/bast.rs`) — the typed tree (`BastDoc`/`BastDef`/`BastType`/…) every engine module walks, the 18 BAST kinds, the `AlkTypeKind` enum, and the foundational annotation types. |
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine`; `validate_bytes` for binary buffers (ADR-010) |
| [builder.md](builder.md) | draft | Fluent Rust API for constructing alktype JSON Schemas at runtime, producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema (ADR-009) |
| [validation.md](validation.md) | accepted | The two-validator model (BAST-native for `validate_bytes`, standard `jsonschema` for `validate_json`), `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine` as the compiled form of a BAST document (ADR-010, ADR-VAL-SPLIT). |
| [builder.md](builder.md) | accepted | Fluent Rust API for constructing BAST documents (`struct_()`) and standard JSON Schemas (`object()`) at runtime, producing `serde_json::Value` (ADR-009, D-BAST-008). |
### In-progress work
| Document | Status | Description |
|----------|--------|-------------|
| [BAST pivot — research record](../research/bast-pivot.md) | accepted | Motivation, POC scope and result, decisions D-BAST-001..009, risks for the BAST format pivot. Implemented in steps 1–10. |
| [BAST pivot — implementation plan](../plans/bast-implementation.md) | accepted | Ordered implementation steps, the public-API semver contract, and the ADR-sync checklist for the BAST pivot. Steps 1–10 complete. |
## Applicable ADRs
| ADR | Title | Relevance |
|-----|-------|-----------|
| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` |
| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations |
| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors |
| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries. *Format-specific content superseded by ADR-BAST; purpose/scope retained.* |
| [BAST](decisions/bast-bast-format.md) | BAST (Binary Abstract Syntax Tree) as the Schema Format | The BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary. Supersedes ADR-001's format-specific content; records D-BAST-001..009. |
| [VAL-SPLIT](decisions/val-split-two-validator-model.md) | Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON | `validate_bytes` uses the BAST-native validator; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema. Records D-BAST-006/007/009. |
| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` (format-agnostic — input format changed, modes didn't) |
| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Annotation *semantics* (carry forward unchanged); annotation *location* moved to BAST type-level properties under the pivot |
| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum (shape unchanged, D-BAST-009); load-time build, access-time check; field-path-carrying errors. *Validation-strategy section refined by ADR-VAL-SPLIT.* |
| [005](decisions/005-int64-uint64-first-class-kinds.md) | Int64/Uint64 as First-Class Kinds | 64-bit integers (SFTP offsets, metatensor data_offsets); JSON precision caveat |
| [006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode | Prevents silent data corruption (inline variable data clobbering subsequent fields) |
| [007](decisions/007-packed-mode-read-factory.md) | Packed-Mode Read API — Engine as SequentialReader Factory | `engine.sequential_reader()` returns an owned reader, not a reference |
| [008](decisions/008-reject-tunion-in-aligned-mode.md) | Reject TUnion in Aligned Mode for v1 | Unions are the protocol pattern; aligned-mode union semantics were broken |
| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate; two methods on one struct, not a trait |
| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003. *Output format amended to BAST / standard JSON Schema by ADR-BAST.* |
| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate. *Validation step amended to the BAST-native validator by ADR-VAL-SPLIT.* |
## Relevant Open Questions
| OQ | Title | Status | Relevance |
|----|-------|--------|-----------|
| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it |
| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it. BAST arrays require `count` in v1 (D-BAST-004), aligning with this deferral. |
| OQ-002 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case |
| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Resolved in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) |
| OQ-004 | `Discriminator::Field` name — `&str` or `String` | open | Builder API ownership question; resolve before the SFTP Packet POC's field-name discriminator path |
| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | open | Blocks the SFTP Packet `validate_bytes` POC (next round) |
| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | open | Documentation fix in builder.md; the engine requires `AlkType:Struct` at the top level |
| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | open | Blocks the SFTP use case for `validate_bytes` (binary `handle`/`data` fields) |
| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Shipped in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) |
| OQ-004 | `Discriminator::Field` name — `&str` or `String` | resolved | `String`, for ownership simplicity |
| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | resolved | Both kinds return `{ "__discriminator": <value>, ...variant-fields }` |
| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | resolved | [builder.md](builder.md) Example 3 wraps the union in a `Schema::struct_().field("payload", ...)` |
| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | resolved | Array of u8: materializer produces `Value::Array` of `Value::Number`; BAST-native validator accepts both `Value::String` and `Value::Array` |
| OQ-008 | `UnionValidator` variant dispatch | resolved | BAST-native validator recurses into the selected variant's BAST definition on `__discriminator` lookup — no custom keywords, no `inline_union_variant_refs` |
## Key Design Principles
1. **The schema is the format.** A JSON Schema with `AlkType:*` custom
keywords is both the validation spec and the layout spec. No separate
format definition, no separate parser, no separate validator. One
schema, three uses: validate, compute offsets, access data. See
[overview.md](overview.md) and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
1. **The schema is the format.** A BAST document is both the layout
spec and the validation spec for bytes. No separate format
definition, no separate parser, no separate validator. One schema,
three uses: validate, compute offsets, access data. See
[overview.md](overview.md), [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md),
and [ADR-BAST](decisions/bast-bast-format.md).
2. **jsonschema is the validation engine, not a custom engine.** The
`jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with
custom keyword support. The novel code is the offset computation, not
the validation. This eliminates ~14,000 lines of hand-rolled schema
engines (typebox-rs, the @alkdev/alktype prototype). See [schema-layer.md](schema-layer.md)
and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
2. **BAST is a JSON Schema dialect, not a custom format.** A BAST
document is valid JSON conforming to the BAST meta-schema (a
standard Draft 2020-12 JSON Schema). Any JSON Schema validator can
check whether a BAST document is well-formed; editors with JSON
Schema support provide autocomplete for free. See
[`bast-format.md`](bast-format.md) and
[ADR-BAST](decisions/bast-bast-format.md).
3. **Two layout modes for two use cases.** Packed sequential
(`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP,
channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats
(metatensor). The consumer selects the mode; the schema is the same.
See [layout-engine.md](layout-engine.md) and
(metatensor). The consumer selects the mode; the BAST document is the
same. See [layout-engine.md](layout-engine.md) and
[ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md).
4. **Variable-length types default to inline length-prefixing.**
@@ -84,38 +97,41 @@ format definition; the engine is generic.
[ADR-003](decisions/003-schema-annotations.md).
6. **Endianness is per-schema, default little-endian.** The engine reads
the `"endian"` annotation and byte-swaps accordingly. SFTP consumers
specify `"endian": "big"`. See [layout-engine.md](layout-engine.md)
and [ADR-003](decisions/003-schema-annotations.md).
the struct-level `"endian"` annotation and byte-swaps accordingly.
SFTP consumers specify `"endian": "big"`. See
[layout-engine.md](layout-engine.md) and
[ADR-003](decisions/003-schema-annotations.md).
7. **Validation is opt-in, built once at load time.** The jsonschema
validator is compiled once at schema load time. Access-time validation
is a fast `is_valid()` check. High-throughput paths can skip
validation; security-sensitive paths can validate every frame. See
7. **Two validators for two input types.** `validate_bytes(&[u8])` uses
the BAST-native validator (a recursive walker over the BAST type
tree — no `jsonschema` involvement). `validate_json(&Value)` uses a
standard `jsonschema::Validator` from a consumer-provided JSON Schema
(BAST is not involved — BAST describes bytes, not JSON shape). One
`AlkTypeError::Validation` variant covers both (D-BAST-009). See
[validation.md](validation.md) and
[ADR-004](decisions/004-error-handling-validation-strategy.md).
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
8. **Not a serialization framework.** The alktype engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use alktype. See [overview.md](overview.md) and
computed offsets — no reflection, no dynamic dispatch per field. For
JSON data, use serde. For binary data with a known BAST document, use
alktype. See [overview.md](overview.md) and
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
9. **Schemas can be built at runtime from Rust (v0.1.0).** A fluent
builder API produces `serde_json::Value` for both AlkType-kind
schemas and standard JSON Schema, covering alkcall's two roles
9. **Schemas can be built at runtime from Rust.** A fluent builder API
produces `serde_json::Value` for both BAST documents (`struct_()`) and
standard JSON Schemas (`object()`), covering alkcall's two roles
(binary layout + JSON payloads) from one module. The builder is
additive — consumers with static schemas continue to load JSON.
See [builder.md](builder.md) and [ADR-009](decisions/009-builder-api.md).
additive — consumers with static BAST documents continue to load
JSON. See [builder.md](builder.md) and
[ADR-009](decisions/009-builder-api.md).
10. **Two validation entry points, one engine (v0.1.0).**
`validate_json(&Value)` for already-parsed JSON (call's payloads);
`validate_bytes(&[u8])` for binary buffers (channels' chunk header).
Same underlying `jsonschema` validator; the bytes path materializes
a `Value` tree via the layout engine, then validates. See
[validation.md](validation.md) and
[ADR-010](decisions/010-generalized-validation-validate-bytes.md).
10. **Two validation entry points, one engine.** `validate_json(&Value)`
for already-parsed JSON (call's payloads); `validate_bytes(&[u8])`
for binary buffers (channels' chunk header). Different validators,
one `AlkTypeError::Validation` variant. See [validation.md](validation.md),
[ADR-010](decisions/010-generalized-validation-validate-bytes.md),
and [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
## References
@@ -135,4 +151,4 @@ format definition; the engine is generic.
> **Note**: The research findings, POC code, and prior-attempt paths above
> refer to the parent `@alkdev/alknet` workspace where this crate originated.
> They are preserved here as historical context for the architectural
> decisions; the artifacts themselves are not part of this standalone repo.
> decisions; the artifacts themselves are not part of this standalone repo.
+684
View File
@@ -0,0 +1,684 @@
---
status: draft
last_updated: 2026-08-15
---
# alktype — BAST Format
**BAST** (Binary Abstract Syntax Tree) is alktype's schema format: a
JSON document that describes binary data layouts using a `kind`-based
vocabulary with `$defs`/`$ref` for composition. BAST replaces the
v0.1.0 `AlkType:*` custom-keyword JSON Schema format.
This document is the **normative format specification**. It is grounded
in the POC on branch `bast-validator-poc` (commit `f371fe4`) and the
decisions D-BAST-001 through D-BAST-009 in
[the pivot research record](../research/bast-pivot.md#decisions). The
implementation plan is
[`docs/plans/bast-implementation.md`](../plans/bast-implementation.md).
Until the BAST pivot lands in code, [`schema-layer.md`](schema-layer.md)
describes the *current* (custom-keyword) schema layer. This document
describes the *target* (BAST) schema layer. They coexist temporarily;
the implementation plan's final step retires `schema-layer.md`'s
custom-keyword content.
## Design Principles
1. **BAST is a JSON Schema instance.** A BAST document is valid JSON
that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON
Schema). Any JSON Schema validator can check whether a BAST document
is well-formed; editors with JSON Schema support provide autocomplete
and inline validation for free.
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
top-level `$defs` block. `$ref` handles cross-references and union
variant references — the same pattern as JSON Schema's own `$defs`
and TypeBox's `Type.Module`. No custom reference resolution mechanism.
3. **`kind`-based vocabulary.** Every type has a `kind` field whose
value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.).
This replaces the `AlkType:*` custom-keyword pattern with a flat,
easily-matched string. The 18 `AlkTypeKind` enum variants are
unchanged; `AlkTypeKind::from_str`/`to_str` map between the enum and
the lowercase BAST strings (D-BAST-002).
4. **Order is explicit.** Struct fields are an ordered array, not an
object with `properties`. Field order is unambiguous — no reliance on
`serde_json`'s `preserve_order` for correctness — and matches the
mental model of binary layouts.
5. **Annotations are type-level properties.** Endianness, alignment,
encoding, and discriminators are properties of the type definition
or field, not custom keywords on a separate schema object. Their
*semantics* carry forward unchanged from ADR-003; only their
*location* moves.
## Document Shape
Every BAST document has the same top-level shape:
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
- The `$defs` block is **required** (D-BAST-003). Single-type documents
are a special case with one entry. A bare struct at the top level
would be a different shape with different parsing logic and no home
for additional definitions — rejected.
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode)` (D-BAST-001). It
selects which `$defs` entry is the top-level type. Convention (first
entry) is fragile and depends on JSON key order; a `$root` marker is
redundant with an explicit parameter.
## The Meta-Schema
The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that
validates the *structure* of BAST documents (is it well-formed?). It
lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded
in the crate for offline use. A different validator — the BAST-native
validator (see [Validation Model](#validation-model) below) — validates
*binary data* against a BAST document (are the bytes a valid instance?).
These are different validators for different inputs.
```json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://alk.dev/bast/v1/schema",
"title": "Binary Abstract Syntax Tree (BAST) v1",
"description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
"type": "object",
"properties": {
"$defs": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeDef" }
}
},
"required": ["$defs"],
"$defs": {
"TypeDef": {
"oneOf": [
{ "$ref": "#/$defs/StructDef" },
{ "$ref": "#/$defs/UnionDef" },
{ "$ref": "#/$defs/EnumDef" }
]
},
"StructDef": {
"type": "object",
"properties": {
"kind": { "const": "struct" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
}
},
"required": ["kind", "fields"],
"additionalProperties": false
},
"FieldDef": {
"type": "object",
"properties": {
"name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
"kind": { "$ref": "#/$defs/TypeRef" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"encoding": { "enum": ["length-prefixed", "offset-indirect"] },
"maxLength": { "type": "integer", "minimum": 0 }
},
"required": ["name", "kind"],
"additionalProperties": false
},
"TypeRef": {
"oneOf": [
{
"description": "Primitive type",
"type": "string",
"enum": [
"int8", "int16", "int32", "int64",
"uint8", "uint16", "uint32", "uint64",
"float32", "float64",
"bool", "string", "bytes"
]
},
{
"description": "Reference to a named $defs entry",
"type": "object",
"properties": {
"$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
},
"required": ["$ref"],
"additionalProperties": false
},
{
"description": "Array type (fixed-size only in v1 — count is required)",
"type": "object",
"properties": {
"kind": { "const": "array" },
"element": { "$ref": "#/$defs/TypeRef" },
"count": { "type": "integer", "minimum": 0 }
},
"required": ["kind", "element", "count"],
"additionalProperties": false
},
{
"description": "Record (string-keyed map) type",
"type": "object",
"properties": {
"kind": { "const": "record" },
"values": { "$ref": "#/$defs/TypeRef" }
},
"required": ["kind", "values"],
"additionalProperties": false
}
]
},
"UnionDef": {
"type": "object",
"properties": {
"kind": { "const": "union" },
"endian": { "enum": ["little", "big"] },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
},
"discriminator": {
"oneOf": [
{
"type": "object",
"properties": {
"kind": { "const": "byte" },
"offset": { "type": "integer", "minimum": 0 },
"type": { "enum": ["uint8", "uint16", "uint32"] }
},
"required": ["kind", "offset", "type"],
"additionalProperties": false
},
{
"type": "object",
"properties": {
"kind": { "const": "field" },
"name": { "type": "string" }
},
"required": ["kind", "name"],
"additionalProperties": false
}
]
},
"mapping": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeRef" }
}
},
"required": ["kind", "discriminator", "mapping"],
"additionalProperties": false
},
"EnumDef": {
"type": "object",
"properties": {
"kind": { "const": "enum" },
"values": {
"type": "array",
"items": { "type": "string" },
"minItems": 1
}
},
"required": ["kind", "values"],
"additionalProperties": false
}
}
}
```
## Type Definitions
### Struct
```json
{
"kind": "struct",
"endian": "big",
"align": 256,
"fields": [
{ "name": "channel_id", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
}
```
- `kind` (required): `"struct"`.
- `endian` (optional): `"little"` (default) or `"big"`. Sets the default
for all fields; field-level `endian` overrides.
- `align` (optional): struct-level alignment, only meaningful in aligned
static mode (ADR-002/003).
- `fields` (required): ordered array of [FieldDef](#fielddef). Array
position is field order — no reliance on JSON key order.
### FieldDef
```json
{ "name": "handle", "kind": "string", "encoding": "offset-indirect", "maxLength": 256 }
```
- `name` (required): identifier, `^[a-zA-Z_][a-zA-Z0-9_]*$`.
- `kind` (required): a [TypeRef](#typeref) — primitive string, `$ref`
object, array object, or record object.
- `endian` (optional): overrides the struct/union default for this field.
- `align` (optional): field-level alignment (aligned mode only).
- `encoding` (optional): `"length-prefixed"` (default) or
`"offset-indirect"`. See [Variable-length encoding](#variable-length-encoding).
- `maxLength` (optional): byte-length cap. See
[Variable-length encoding](#variable-length-encoding).
### TypeRef
`TypeRef` is the central mechanism for referencing types. Seven forms:
| Form | Example | Meaning |
|------|---------|---------|
| Primitive string | `"uint32"` | A built-in primitive (see [Primitives](#primitives)) |
| `$ref` object | `{ "$ref": "#/$defs/Read" }` | Reference to a named `$defs` entry |
| Array object | `{ "kind": "array", "element": "uint32", "count": 3 }` | Fixed-size array |
| Record object | `{ "kind": "record", "values": "string" }` | String-keyed map |
| Inline struct | `{ "kind": "struct", "fields": [...] }` | Anonymous struct |
| Inline union | `{ "kind": "union", ... }` | Anonymous union |
| Inline enum | `{ "kind": "enum", "values": [...] }` | Anonymous enum |
The `$ref` form uses standard JSON Pointer syntax **restricted to
`#/$defs/<name>`** — no external references, no fragment-only pointers,
no bare names. The restriction keeps resolution a single hash lookup
and eliminates the `normalize_refs` step the v0.1.0 engine needed for
TypeBox's bare-name refs.
Arrays and records are inline type constructors, not top-level `$defs`
entries. Complex element types use nested `$ref`:
```json
{ "kind": "array", "element": { "$ref": "#/$defs/ComplexElement" }, "count": 4 }
```
Inline `struct`/`union`/`enum` TypeRefs are anonymous composites — a
field, array element, record value, or union variant whose type is
declared inline rather than named in `$defs`. They are structurally
identical to their named counterparts (same `StructDef`/`UnionDef`/
`EnumDef` shape); only the reference mechanism differs. Named composites
are preferred for reuse and for `$ref`-based dispatch; inline composites
are convenient for one-off nested types.
### Primitives
The 13 primitive `kind` strings map to the unchanged `AlkTypeKind`
variants (D-BAST-002 — lowercase strings, PascalCase enum variants):
| BAST kind | `AlkTypeKind` | Rust type | Size | Category |
|-----------|---------------|-----------|------|----------|
| `int8` | `Int8` | `i8` | 1 | fixed |
| `int16` | `Int16` | `i16` | 2 | fixed |
| `int32` | `Int32` | `i32` | 4 | fixed |
| `int64` | `Int64` | `i64` | 8 | fixed |
| `uint8` | `Uint8` | `u8` | 1 | fixed |
| `uint16` | `Uint16` | `u16` | 2 | fixed |
| `uint32` | `Uint32` | `u32` | 4 | fixed |
| `uint64` | `Uint64` | `u64` | 8 | fixed |
| `float32` | `Float32` | `f32` | 4 | fixed |
| `float64` | `Float64` | `f64` | 8 | fixed |
| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed |
| `string` | `String` | length-prefixed UTF-8 | variable | variable |
| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable |
`int64`/`uint64` are alktype additions (not in TypeBox's `typedef.ts`),
required by SFTP `offset: u64` and metatensor `data_offsets`. JSON
precision caveat per ADR-005 applies: integers beyond `2^53` lose
precision in `serde_json::Value::Number`; the binary path is exact.
### Enum
```json
{
"kind": "enum",
"values": ["Ok", "PermissionDenied", "NoSuchFile", "Failure"]
}
```
- `kind` (required): `"enum"`.
- `values` (required): non-empty array of strings, in declaration order.
- Binary representation: a `u32` index into `values` (0-based), encoded
per the struct's endianness. This is a deliberate deviation from
TypeBox's string enum in favor of binary efficiency — a `u32` index is
compact, fixed-size, and sufficient for any realistic enum.
**Bug fix vs v0.1.0:** The v0.1.0 engine has a dead constraint on the
bytes path — the built-in `enum` keyword checks string membership, but
the materializer emits `Value::Number(index)`, which never matches. The
BAST-native validator (see [Validation Model](#validation-model)) checks
the materialized index against `values.len()` bounds, fixing this.
### Union
```json
{
"kind": "union",
"endian": "big",
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "uint8"
},
"mapping": {
"1": { "$ref": "#/$defs/Init" },
"3": { "$ref": "#/$defs/Open" },
"5": { "$ref": "#/$defs/Read" }
}
}
```
- `kind` (required): `"union"`.
- `endian` (optional): default endianness for variant fields.
- `discriminator` (required): one of:
- **Byte-offset**: `{ "kind": "byte", "offset": <N>, "type": "uint8"|"uint16"|"uint32" }`.
The discriminator byte is at `offset`; the variant struct starts at
`offset + discriminator_size`. Mapping keys are stringified integers.
- **Field-name**: `{ "kind": "field", "name": "<field>" }`. The
discriminator is a length-prefixed string field; mapping keys are
string values matching the field's value. The union's `fields` array
(optional, only valid with field-name discriminators per D-BAST-005)
provides the field definitions including the discriminator field.
- `mapping` (required): object mapping discriminator values to
[TypeRef](#typeref) entries (typically `$ref` to `$defs` variants).
Variant `$ref`s are resolved **lazily** by the materializer and
validator — no `inline_union_variant_refs` compile step (removed under
BAST). The validator recurses into the selected variant's BAST
definition on `__discriminator` lookup, recovering the OQ-008
per-variant constraint enforcement (e.g., `maxLength` on a `bytes`
field inside a variant) without custom keywords.
### Array
```json
{ "kind": "array", "element": "float32", "count": 3 }
```
- `kind` (required): `"array"`.
- `element` (required): a [TypeRef](#typeref).
- `count` (required in v1): the fixed element count. **Arrays of
variable-length elements without a `count` are not supported in v1**
(D-BAST-004, aligning with OQ-001). The meta-schema enforces this:
`count` is in `required`. Variable-length collections are available
via `record` instead.
For fixed-size elements with a known count, the array size is
`element_size × count`. For variable-length elements (e.g.,
`"element": "string"`) with a known count, each element carries its own
length prefix — the array is count-prefixed in the sense that the count
is known at schema time, but the total byte size is not.
### Record
```json
{ "kind": "record", "values": "string" }
```
- `kind` (required): `"record"`.
- `values` (required): a [TypeRef](#typeref) for the value type.
Binary layout: a count-prefixed sequence of `(key, value)` pairs —
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
Each key is a length-prefixed UTF-8 string. Each value is encoded per
its `values` type. There is no separate `value_len` prefix — the value's
size is determined by its kind (fixed-size kinds have a known size;
variable-length kinds carry their own length prefix). The count and
key-length prefixes respect the struct's endianness.
## Variable-Length Encoding
The three strategies from ADR-003 carry forward with the same semantics,
expressed as field-level properties instead of keyword-value objects:
| Strategy | BAST syntax | Behavior |
|----------|------------|----------|
| Inline length-prefixed (default) | `{ "name": "handle", "kind": "string" }` | `[u32 length][data]` |
| Fixed-size reservation | `{ "name": "name", "kind": "string", "maxLength": 256 }` | Reserve `maxLength` bytes (aligned mode); validation constraint (packed mode) |
| Offset indirection | `{ "name": "blob", "kind": "bytes", "encoding": "offset-indirect" }` | `{offset: u32, length: u32}` pointing to separate data region |
**Default strategy selection (unchanged from v0.1.0):**
- **Packed sequential mode:** always inline length-prefixing.
`maxLength` is a validation constraint only.
- **Aligned static mode:** fixed-size reservation if `maxLength` is
declared; offset indirection if `"encoding": "offset-indirect"` is
declared; inline length-prefixing otherwise.
**Length prefix endianness:** The 4-byte length prefix (strategies 1
and 3) respects the effective endianness (struct default or field
override). In little-endian mode, `u32::from_le_bytes`; in big-endian
mode, `u32::from_be_bytes`. Ensures SFTP consumers (big-endian) have
consistent byte order for field values and length prefixes.
Applies to all variable-length types: `string`, `bytes`,
`record`, and arrays of variable-length elements.
## Endianness
Struct-level or union-level property with per-field override (same
semantics as ADR-003):
- Struct/union-level `"endian"` sets the default for all fields.
- Field-level `"endian"` overrides the struct/union default.
- Default is `"little"` when neither is specified.
- The length prefix for variable-length fields respects the effective
endianness.
```json
{
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "id", "kind": "uint32" },
{ "name": "handle", "kind": "string" },
{ "name": "crc", "kind": "uint32", "endian": "little" }
]
}
```
## Alignment
Struct-level or field-level property, only meaningful in aligned static
mode (same as ADR-003):
```json
{
"kind": "struct",
"align": 256,
"fields": [
{ "name": "header", "kind": { "$ref": "#/$defs/Header" } },
{ "name": "weight", "kind": "float32", "align": 16 }
]
}
```
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/
f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length
prefix), 1 for struct/union/array. Unchanged from v0.1.0.
- Ignored in packed sequential mode.
## Validation Model
BAST separates two concerns that the v0.1.0 format conflates, and in
doing so reveals that the engine has **two distinct validation paths**
with different inputs and guarantees. This is the validator split,
decided in D-BAST-006, D-BAST-007, and D-BAST-009. See
[`validation.md`](validation.md) for the current (pre-pivot) validation
layer; this section specifies the target model.
### Two validators, two inputs
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
**`validate_bytes` — bytes in, BAST is the validator.** The materializer
produces a `Value` tree from bytes. By construction, this `Value` is
*structurally correct*: all declared fields are present (the
materializer iterates the field list), types are correct (`read_u32`
produces `Value::Number`), bounds are checked (via `data_access::
check_bounds`), UTF-8 is valid (via `from_utf8`), the discriminator is
in the mapping, and the boolean byte is 0 or 1. What the materializer
does NOT check — and what the 19 v0.1.0 custom keyword validators check
afterward — are **value-domain constraints expressed in the BAST
document**. The BAST-native validator is a recursive walker over the
BAST type tree that checks exactly these:
| Constraint | Validator arm |
|------------|---------------|
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` with `as_i64`/`as_u64` + range check |
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `validate_float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `check_string` reads the field-level `maxLength` |
| Bytes `maxLength` (array length) | `check_bytes` accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `validate_enum` checks `idx < values.len()` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `validate_union` reads `__discriminator`, looks up the variant, recurses via `validate_typeref` |
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
| Record values | `validate_record` recurses into each value's `values` type |
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
No external JSON Schema is required for `validate_bytes`. The BAST
document is the complete specification of the binary format — it
describes both the layout (how to read) and the constraints (what
values are valid). This is the "schema is the format" principle from
ADR-001, now fully realized.
An optional external JSON Schema can be layered on top for constraints
BAST doesn't express (cross-field consistency, regex patterns on string
content). This is additive, not load-bearing.
**`validate_json` — JSON in, JSON Schema is the validator.** The
consumer provides a JSON `Value` (e.g., an incoming JSON-RPC request).
The BAST document is irrelevant — BAST describes bytes, not JSON shape.
The right validator for a JSON value is a standard
`jsonschema::Validator` built from a standard JSON Schema document the
consumer provides. This is the path alkcall uses for its `OperationSpec`
JSON validation. No custom keywords; BAST is not involved.
### What is removed
Under the BAST pivot, the v0.1.0 validation machinery is removed from
the `validate_bytes` path:
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
factories) — replaced by the BAST-native validator (~250 lines, a
flat match with no factories, no trait objects, no sub-validator
pre-computation).
- `inline_union_variant_refs()` — union variant refs are resolved lazily
by the validator and materializer.
- `build_validator()`'s custom-keyword path — repurposed or removed (see
the implementation plan's step 6 for the decision on its fate).
The `jsonschema` crate **remains a direct dependency** for
`validate_json` and for validating BAST documents against the BAST
meta-schema. The only thing removed is the custom keyword integration
path. The `validate_bytes` path no longer touches `jsonschema` — a
small wasm binary-size win in addition to the architecture
simplification.
### `AlkTypeError::Validation` payload shape
**Decided (D-BAST-009):** Keep
`Validation(jsonschema::ValidationError<'static>)`.
The `validate_bytes` path no longer uses `jsonschema`, so its error
payload is constructed via `jsonschema::ValidationError::custom` purely
to keep the variant's type unchanged. The rationale is consumer
ergonomics on the *combined* path: consumers like alkcall use both
`validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary
channels) and handle `AlkTypeError::Validation` in one place. A single
uniform payload type means one match arm covers both sources. The
alternative (`Validation(String)`) would force `validate_json` to
flatten its structured errors (instance path, schema path, keyword) to
a `String` via `Display` — the more information-rich path loses data to
accommodate the less rich one. That is the wrong direction.
The `no_std`/minimal-build angle (OQ-002) that the alternative was
meant to enable is moot: `validate_json` requires `jsonschema`
regardless, so a bytes-only `no_std` build already has to give up
`validate_json` as a separate, larger decision. The right place to
revisit is when/if OQ-002 is actually pursued.
## Relationship to JSON Schema and TypeBox
### BAST is a JSON Schema dialect
BAST is a specific JSON Schema instance format — like how JSON Schema
itself is a JSON document conforming to the JSON Schema meta-schema.
BAST documents conform to the BAST meta-schema. The entire JSON Schema
tooling ecosystem works with BAST:
- **Validation:** `jsonschema::options().build(&bast_meta_schema)?.validate(&bast_doc)`
- **Editors:** VSCode with `$schema` pointing to the BAST meta-schema URL
- **Documentation:** JSON Schema generators produce human-readable docs
from the meta-schema
### TypeBox interop
TypeBox's `Type.Module({...})` pattern maps naturally to BAST's `$defs`
structure. A TypeBox module defining binary types can serialize to BAST
JSON. The relationship:
- TypeBox → BAST JSON → alktype engine (binary layout)
- TypeBox → standard JSON Schema → jsonschema (JSON validation)
Same TypeBox source, two output formats, two validators.
### Not a replacement for JSON Schema
BAST does not replace JSON Schema for JSON data validation. A BAST
document cannot validate a JSON payload — it describes binary data
layouts and value-domain constraints for bytes. For JSON validation,
consumers use standard JSON Schema documents (which may be derived from
BAST via future codegen, or authored separately). The `validate_json`
path accepts a consumer-provided JSON Schema and uses a standard
`jsonschema::Validator` — BAST is not involved.
This is the split: `validate_bytes` is BAST-native (the BAST document
is both the layout spec and the validation spec for bytes);
`validate_json` is JSON-Schema-native (a standard JSON Schema is the
validation spec for JSON values). One crate, two validators, two input
types.
## Decisions
The BAST format is grounded in decisions D-BAST-001 through D-BAST-009,
recorded in [the pivot research record](../research/bast-pivot.md#decisions).
The implementation-relevant summary:
| Decision | Summary |
|----------|---------|
| [D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention |
| [D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase |
| [D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape |
| [D-BAST-004](../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) |
| [D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` |
| [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed |
| [D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema |
| [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema |
| [D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths |
## References
- [Pivot research record](../research/bast-pivot.md) — motivation, POC
scope and result, decisions D-BAST-001..009, risks
- [Implementation plan](../plans/bast-implementation.md) — ordered
steps, public-API semver contract, ADR-sync checklist
- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics
(carry forward unchanged; only location moves)
- [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) — Int64/
Uint64 as first-class kinds; JSON precision caveat
- [`schema-layer.md`](schema-layer.md) — the current (v0.1.0) schema
layer; superseded by this document when the pivot lands
- [`validation.md`](validation.md) — the current (v0.1.0) validation
layer; rewritten for the validator split when the pivot lands
+240 -169
View File
@@ -1,29 +1,41 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-15
---
# alktype — Builder API
The builder layer: a fluent Rust API for constructing alktype JSON
Schemas (both `AlkType:*`-bearing binary-layout schemas and plain
JSON-Schema-only operation payload schemas) at runtime, producing
`serde_json::Value`. Decided in [ADR-009](decisions/009-builder-api.md);
resolves [OQ-003](questions/003-builder-api-for-schema-construction.md).
The builder layer: a fluent Rust API for constructing BAST documents
(binary-layout schemas) and standard JSON Schemas (JSON-validation
schemas) at runtime, producing `serde_json::Value`. Decided in
[ADR-009](decisions/009-builder-api.md); resolves
[OQ-003](questions/003-builder-api-for-schema-construction.md). The
two-output-format split is D-BAST-008, recorded in
[ADR-BAST](decisions/bast-bast-format.md).
## What
The `builder` module provides a single `Schema` builder type and a
`Definitions` helper for named `$defs`. The builder's `.build()` method
returns a `serde_json::Value` — the same form alktype already consumes
via `AlkTypeEngine::compile` (for `AlkType:*` schemas) and the same form
`OperationSpec.input_schema` / `output_schema` / `error_schemas` hold
(for plain JSON Schema, no `AlkType:*` kinds).
returns a `serde_json::Value` — one of two forms depending on the
constructor used (D-BAST-008):
- **BAST JSON** (binary layout) — `Schema::struct_().field(...).build()`
produces a BAST TypeDef (`{ "kind": "struct", "fields": [...] }`).
Primitive constructors produce bare TypeRef strings (`"uint32"`).
Feed to [`AlkTypeEngine::compile`](validation.md) (the binary-layout
path) → `validate_bytes`.
- **Standard JSON Schema** (JSON validation) — `Schema::object().field(...)`
produces `{ "type": "object", "properties": {...}, "required": [...] }`.
No BAST `kind`, no custom keywords — a plain JSON Schema. Feed to a
standard `jsonschema::Validator` (or `AlkTypeEngine::compile` with a
JSON Schema for the `validate_json` path, D-BAST-007).
The builder covers:
- All 19 `AlkType:*` kinds (binary-layout schemas) — see
[schema-layer.md](schema-layer.md) for the kinds.
- All 18 BAST kinds (binary-layout schemas) — see
[schema-layer.md](schema-layer.md) for the kinds and
[`bast-format.md`](bast-format.md) for the format.
- All standard JSON Schema keywords needed for operation payload
schemas: `type`, `properties`, `required`, `items`, `enum`, `format`,
`additionalProperties`, `minimum`, `maximum`, `minItems`,
@@ -39,16 +51,20 @@ alktype's first consumer and needs to build schemas at runtime from
Rust code, for two roles:
1. **Binary layout schemas** (channels' 8-byte chunk header, future
binary call frames) — `AlkType:*` schemas, fed to
`AlkTypeEngine::compile` (packed mode, big-endian).
binary call frames) — BAST documents, fed to
`AlkTypeEngine::compile` (packed mode, big-endian) →
`validate_bytes`.
2. **JSON payload schemas** (call's `OperationSpec.input_schema` /
`output_schema` / `error_schemas`) — plain JSON Schema, no
`AlkType:*` kinds, validated via the standard `jsonschema` validator.
`output_schema` / `error_schemas`) — plain JSON Schema, no BAST
`kind`, validated via the standard `jsonschema` validator
(`AlkTypeEngine::compile` with a JSON Schema → `validate_json`).
A single builder serving both roles means alkcall imports one module
for schema construction. See [ADR-009](decisions/009-builder-api.md)
for the decision rationale (why `Value` not a typed `Schema` enum, why
both AlkType and standard JSON Schema in one builder).
both BAST and standard JSON Schema in one builder) and
[ADR-BAST](decisions/bast-bast-format.md) for the two-output-format
decision (D-BAST-008).
## Architecture
@@ -62,35 +78,42 @@ duplicating the JSON form that `AlkTypeEngine::compile`,
### Module placement
`src/builder.rs`, re-exported from the crate root. The builder is a
peer of `schema.rs` (which parses schemas) and `engine.rs` (which
peer of `bast.rs` (which parses BAST documents) and `engine.rs` (which
compiles them). The builder constructs; it does not parse or compile.
```rust
// src/lib.rs (additions)
pub mod builder;
pub use builder::{Schema, Definitions};
pub use builder::{Schema, Definitions, Discriminator};
```
### Field order is load-bearing
### Field order is explicit
`serde_json` with `preserve_order` is already a dependency (ADR-001).
The builder's `Value` output uses `serde_json::Map` (which preserves
insertion order under `preserve_order`), so field declaration order in
the builder is the field order in the binary layout. This is critical
for packed mode (ADR-002) where field order determines offsets.
BAST struct fields are an ordered array (BAST design principle #4 —
see [`bast-format.md`](bast-format.md#design-principles)). The builder's
`struct_()`/`union_()` accumulates fields in call order and emits them
as the `fields` array on `.build()`. Field order in the builder is the
field order in the binary layout. This is critical for packed mode
(ADR-002) where field order determines offsets.
(`serde_json`'s `preserve_order` feature remains a dependency, but
layout correctness no longer depends on it — the `fields` array makes
order explicit. `preserve_order` is still load-bearing for the
`mapping` object's iteration order and for `Definitions`' `$defs`
block, which the parser walks in document order.)
## Public API
### `Schema` builder
`Schema` is the single entry point. Constructors for each AlkType kind
and each standard JSON Schema type; setters for annotations and
`Schema` is the single entry point. Constructors for each BAST kind and
each standard JSON Schema type; setters for annotations and
constraints; `.build()` produces `Value`.
#### AlkType kind constructors
#### BAST kind constructors
One constructor per `AlkTypeKind` variant (see [schema-layer.md](schema-layer.md)
§"The 19 AlkType Kinds"):
§"The 18 BAST Kinds"):
```rust
impl Schema {
@@ -112,43 +135,44 @@ impl Schema {
// Variable-length kinds
pub fn string() -> Self;
pub fn bytes() -> Self;
pub fn timestamp() -> Self;
// Composite kinds
pub fn struct_() -> Self; // fields added via .field()
pub fn union_(disc: Discriminator) -> Self; // variants via .mapping()
pub fn array_of(element: Schema) -> Self;
pub fn array_of(element: Schema) -> Self; // .count() required for valid BAST (D-BAST-004)
pub fn record_of(value: Schema) -> Self;
}
```
Each constructor sets the corresponding `"AlkType:<Kind>": true` key.
For example, `Schema::uint32()` produces `{"AlkType:Uint32": true}`.
Primitive constructors produce the bare BAST TypeRef string on
`.build()`. For example, `Schema::uint32().build()` produces `"uint32"`.
Composite constructors produce the BAST object form.
**`enum_of`** sets both `"AlkType:Enum": true` and the standard
`"enum"` keyword with the provided values (declaration order is the
index order — see [schema-layer.md](schema-layer.md) §"TEnum binary
representation"):
**`enum_of`** produces a BAST enum TypeDef (`{ "kind": "enum", "values":
[...] }`); declaration order is the index order — see
[schema-layer.md](schema-layer.md) §"The 18 BAST Kinds"):
```rust
Schema::enum_of(&["read", "write", "execute"])
// -> { "AlkType:Enum": true, "enum": ["read", "write", "execute"] }
Schema::enum_of(&["read", "write", "execute"]).build()
// -> { "kind": "enum", "values": ["read", "write", "execute"] }
```
**`array_of`** and **`record_of`** take the element/value schema as a
nested `Schema`:
nested `Schema`. `array_of` requires `.count(N)` for valid BAST
(D-BAST-004 — arrays of variable-length elements without a count are
deferred, aligning with OQ-001):
```rust
Schema::array_of(Schema::uint32())
// -> { "AlkType:Array": true, "items": { "AlkType:Uint32": true } }
Schema::array_of(Schema::uint32()).count(3).build()
// -> { "kind": "array", "element": "uint32", "count": 3 }
Schema::record_of(Schema::float32())
// -> { "AlkType:Record": true, "values": { "AlkType:Float32": true } }
Schema::record_of(Schema::float32()).build()
// -> { "kind": "record", "values": "float32" }
```
#### Standard JSON Schema type constructors
For plain JSON Schema (no `AlkType:*` kinds) — call's
`input_schema` / `output_schema` / `error_schemas`:
For plain JSON Schema (no BAST `kind`) — call's `input_schema` /
`output_schema` / `error_schemas`:
```rust
impl Schema {
@@ -163,22 +187,24 @@ impl Schema {
}
```
The `_` suffix disambiguates standard JSON Schema types from AlkType
kinds (`string` is the AlkType kind; `string_` is the standard JSON
Schema type — the AlkType kind constructor sets `"AlkType:String":
true`, the standard constructor sets `"type": "string"`). This is
deliberate: the two are distinct schema forms and the builder makes
the distinction visible at the call site.
The `_` suffix disambiguates standard JSON Schema types from BAST kinds
(`string` is the BAST primitive; `string_` is the standard JSON Schema
type — `string()` would produce `"string"` as a BAST TypeRef,
`string_()` produces `{ "type": "string" }` as a standard JSON Schema).
This is deliberate: the two are distinct schema forms and the builder
makes the distinction visible at the call site.
#### Annotation setters
Annotation setters mirror ADR-003. Each setter is named after the
annotation it produces; calling the setter sets the corresponding JSON
key. Setters return `Self` for chaining.
Annotation setters mirror ADR-003 (semantics unchanged; location moved
to BAST type-level properties under the pivot — see
[ADR-BAST](decisions/bast-bast-format.md)). Each setter is named after
the annotation it produces; calling the setter sets the corresponding
JSON key. Setters return `Self` for chaining.
```rust
impl Schema {
/// Schema-level endianness (ADR-003 §1). Default little.
/// Struct/union-level endianness (ADR-003 §1). Default little.
pub fn endian(mut self, endian: Endian) -> Self;
/// Struct or field alignment (ADR-003 §2). Struct-level sets the
@@ -193,12 +219,19 @@ impl Schema {
/// a variable-length type, reserves this many bytes (strategy 2).
/// In packed mode, validation constraint only.
pub fn max_length(mut self, max: usize) -> Self;
/// Array count (D-BAST-004 — required for valid BAST arrays in v1).
pub fn count(mut self, count: usize) -> Self;
}
```
`Endian` and `VariableEncoding` are re-exported from `schema.rs` (no
new types — the builder uses the existing enums). The setters produce
the exact JSON shapes from ADR-003:
new types — the builder uses the existing enums). When applied to a
struct, `endian`/`align` are struct-level; when the `Schema` is used as
a `.field()` argument, the builder extracts `endian`/`align`/`encoding`/
`maxLength` and places them on the *field* object (BAST field-level
properties). The setters produce the exact BAST JSON shapes from
[`bast-format.md`](bast-format.md):
```rust
Schema::struct_()
@@ -207,31 +240,34 @@ Schema::struct_()
.field("length", Schema::uint32())
.build()
// -> {
// "AlkType:Struct": true,
// "kind": "struct",
// "endian": "big",
// "properties": {
// "channel_id": { "AlkType:Uint32": true },
// "length": { "AlkType:Uint32": true }
// }
// "fields": [
// { "name": "channel_id", "kind": "uint32" },
// { "name": "length", "kind": "uint32" }
// ]
// }
```
#### Composite builders
`struct_()`, `union_()`, `array_of()`, `record_of()` are the
composite constructors. `struct_()` and `union_()` need additional
setters to populate their children:
`struct_()`, `union_()`, `array_of()`, `record_of()` are the composite
constructors. `struct_()` and `union_()` need additional setters to
populate their children:
```rust
impl Schema {
/// Add a field to a struct (or object). Field order is load-bearing
/// for binary layouts (packed mode field order = byte order).
/// Add a field to a struct (or a field-name-discriminator union).
/// Field order is load-bearing for binary layouts (packed mode
/// field order = byte order — the `fields` array is ordered).
/// The field's schema is built from the passed `Schema`.
pub fn field(mut self, name: &str, field: Schema) -> Self;
/// Mark fields as required (standard JSON Schema `required` keyword).
/// Can be called multiple times; required names accumulate.
/// Field names must have been added via `.field()`.
/// Only meaningful for `object()` (standard JSON Schema) — BAST
/// structs require all declared fields present (the validator
/// enforces this). Can be called multiple times; required names
/// accumulate.
pub fn required(mut self, names: &[&str]) -> Self;
/// Set the items schema for a standard `array` type.
@@ -247,9 +283,11 @@ impl Schema {
}
```
**`field`** sets `properties[name] = field.build()`. Repeated calls
append. Field order in the built `Value` is the call order (because
`serde_json::Map` preserves insertion order under `preserve_order`).
**`field`** appends a `{ "name": ..., "kind": <field.build()>, ... }`
entry to the struct/union's `fields` array, extracting field-level
annotations (`endian`, `align`, `encoding`, `maxLength`) from the
passed `Schema`. Repeated calls append in order. Field order in the
built `Value` is the call order.
**`required`** sets the standard JSON Schema `"required"` array. The
builder does not check that the named fields exist (that's a
@@ -260,11 +298,12 @@ accumulates names:
```rust
Schema::object()
.field("path", Schema::string_())
.field("path", Schema::string_().max_length(4096))
.field("offset", Schema::integer().minimum(0))
.field("length", Schema::integer().minimum(0))
.required(["path"])
.required(["offset", "length"])
.build()
// -> {
// "type": "object",
// "properties": { "path": {...}, "offset": {...}, "length": {...} },
@@ -280,35 +319,31 @@ For operation payload schemas (call's `input_schema` etc.):
impl Schema {
/// `minimum` (inclusive lower bound for numbers/integers).
pub fn minimum(mut self, min: f64) -> Self;
/// `maximum` (inclusive upper bound for numbers/integers).
pub fn maximum(mut self, max: f64) -> Self;
/// `minLength` (minimum string length).
pub fn min_length(mut self, min: usize) -> Self;
/// `minItems` (minimum array length).
pub fn min_items(mut self, min: usize) -> Self;
/// `maxItems` (maximum array length).
pub fn max_items(mut self, max: usize) -> Self;
/// `format` (e.g. "date-time", "uri", "email").
pub fn format(mut self, fmt: &str) -> Self;
/// `title` (human-readable description).
pub fn title(mut self, t: &str) -> Self;
/// `description` (human-readable description).
pub fn description(mut self, d: &str) -> Self;
}
```
These set the corresponding standard JSON Schema keywords. They apply
to both AlkType-kind schemas and standard JSON Schema type schemas
(e.g., `Schema::string().max_length(4096)` sets `maxLength`, which
serves as both a validation constraint and, in aligned mode, a
fixed-size reservation — ADR-003 §3).
to standard JSON Schema type schemas (e.g.,
`Schema::string_().max_length(4096)` sets `maxLength`, which on the
`validate_json` path is a JSON-Schema validation constraint). On a BAST
schema, `max_length` also serves as the aligned-mode fixed-size
reservation (ADR-003 §3) and the packed-mode validation constraint
(enforced by the BAST-native validator — see
[validation.md](validation.md)).
#### `.build()`
@@ -343,21 +378,22 @@ to compose them. `from_value` wraps the `Value` so it can be passed to
### `Discriminator` for `union_()`
`union_()` takes a `Discriminator` describing the union's dispatch
mechanism. This mirrors `schema.rs::DiscriminatorKind` but with a
builder-friendly shape (the kind enum is re-exported from `schema.rs`,
not duplicated):
mechanism. This mirrors `bast::BastDiscriminator` (the parser's typed
view) but with a builder-friendly shape:
```rust
pub enum Discriminator {
/// Byte-offset discriminator (ADR-003 §4 Kind A).
/// `offset` is the byte position; `disc_type` is the AlkType kind
/// `offset` is the byte position; `disc_type` is the BAST kind
/// of the discriminator (Uint8/Uint16/Uint32).
Byte {
offset: usize,
disc_type: AlkTypeKind, // restricted to Uint8/Uint16/Uint32
},
/// Field-name discriminator (ADR-003 §4 Kind B).
/// `name` is the field holding the discriminator value.
/// `name` is the field holding the discriminator value. The
/// discriminator field and any shared fields are declared via
/// `.field()` on the union builder.
Field {
name: String,
},
@@ -376,9 +412,13 @@ let packet = Schema::union_(Discriminator::Byte {
.mapping("101", Schema::ref_def("Status"))
.build();
// -> {
// "AlkType:Union": true,
// "discriminator": { "kind": "byte", "offset": 0, "type": "AlkType:Uint8" },
// "mapping": { "5": {"$ref":"#/$defs/Read"}, "6": {...}, "101": {...} }
// "kind": "union",
// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
// "mapping": {
// "5": { "$ref": "#/$defs/Read" },
// "6": { "$ref": "#/$defs/Write" },
// "101": { "$ref": "#/$defs/Status" }
// }
// }
```
@@ -386,22 +426,29 @@ let packet = Schema::union_(Discriminator::Byte {
```rust
let event = Schema::union_(Discriminator::Field { name: "type" })
.field("type", Schema::string())
.mapping("read", Schema::ref_def("Read"))
.mapping("write", Schema::ref_def("Write"))
.build();
// -> {
// "AlkType:Union": true,
// "kind": "union",
// "discriminator": { "kind": "field", "name": "type" },
// "fields": [ { "name": "type", "kind": "string" } ],
// "mapping": { "read": {...}, "write": {...} }
// }
```
(Field-name-discriminator unions require a `fields` array declaring the
discriminator field — D-BAST-005. The builder emits `fields` only when
the discriminator is `Field` and at least one field was added.)
### `Definitions` — named `$defs` for cross-reference
`Definitions` is a helper for building named `$defs` that schemas can
`$ref` by name. This is the ergonomics win for alkcall's
`OperationSpec`, where input/output/error schemas reference shared
definitions (e.g., `FileNotFound`, `RateLimited`).
`$ref` by name, and for assembling a complete BAST document. This is
the ergonomics win for alkcall's `OperationSpec`, where
input/output/error schemas reference shared definitions (e.g.,
`FileNotFound`, `RateLimited`).
```rust
pub struct Definitions { /* ... */ }
@@ -410,50 +457,67 @@ impl Definitions {
pub fn new() -> Self;
/// Define a named schema. Returns a `Schema` that produces
/// `{"$ref": "#/$defs/<name>"}` — the JSON Pointer form that
/// `jsonschema` and `AlkTypeEngine::compile` expect (after
/// `normalize_refs`, which the engine runs at compile time).
/// `{"$ref": "#/$defs/<name>"}` — the JSON Pointer form BAST
/// requires (no `normalize_refs` step; refs are always full
/// pointers).
pub fn define(&mut self, name: &str, schema: Schema) -> Schema;
/// Like `define`, but the schema is an existing `Value` (adopted
/// via `Schema::from_value`).
pub fn define_value(&mut self, name: &str, value: Value) -> Schema;
/// Produce the `{"$defs": { ... }}` object to merge into a
/// top-level schema. Call once at the end.
/// Produce the `{"$defs": { ... }}` object.
pub fn build(self) -> Value;
/// Build a complete BAST document with `root_name` as the root
/// type. The root schema is inserted into `$defs` alongside any
/// previously defined entries. The resulting `Value` is ready for
/// `AlkTypeEngine::compile(&doc, root_name, mode, ...)`.
pub fn build_doc(self, root_name: &str, root: Schema) -> Value;
/// Merge the `$defs` into a top-level schema `Value`. If `top`
/// already has a `$defs` object, the definitions are merged into
/// it; otherwise a `$defs` key is inserted. For BAST documents,
/// prefer `build_doc` — it places the root type inside `$defs`
/// (where BAST requires it).
pub fn merge_into(self, top: &mut Value);
}
```
**Usage:**
**Usage (complete BAST document):**
```rust
let mut defs = Definitions::new();
let file_not_found = defs.define("FileNotFound",
Schema::object()
.field("path", Schema::string_())
.field("errno", Schema::integer())
.required(["path", "errno"])
);
defs.define("Init", Schema::struct_().field("version", Schema::uint32()));
defs.define("Read", Schema::struct_()
.field("handle", Schema::bytes())
.field("offset", Schema::uint64())
.field("len", Schema::uint32()));
let rate_limited = defs.define("RateLimited",
Schema::object()
.field("retry_after_ms", Schema::integer().minimum(0))
.required(["retry_after_ms"])
);
let read_file_error = Schema::object()
.field("code", Schema::string_())
.field("details", Schema::any()) // one of the defined errors
.required(["code"])
.build();
// Merge $defs into the top-level schema that references them
let mut top = Schema::object()
.field("error", read_file_error)
.build();
top.as_object_mut().unwrap().insert("$defs".to_string(), defs.build());
let doc = defs.build_doc("Packet", Schema::struct_()
.field("payload", Schema::union_(Discriminator::Byte {
offset: 0,
disc_type: AlkTypeKind::Uint8,
})
.mapping("1", Schema::ref_def("Init"))
.mapping("5", Schema::ref_def("Read"))));
// -> {
// "$defs": {
// "Init": { "kind": "struct", "fields": [ { "name": "version", "kind": "uint32" } ] },
// "Read": { "kind": "struct", "fields": [ ... ] },
// "Packet": { "kind": "struct", "fields": [
// { "name": "payload", "kind": {
// "kind": "union",
// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
// "mapping": { "1": { "$ref": "#/$defs/Init" }, "5": { "$ref": "#/$defs/Read" } }
// } }
// ] }
// }
// }
//
// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None)
// then validate incoming frames via engine.validate_bytes(&frame).
```
`define` returns a `Schema` (the `$ref` to the definition), so it can
@@ -483,31 +547,36 @@ For cases where the `Definitions::define` return value isn't handy
## Usage Examples
### Example 1: channels' 8-byte chunk header (binary layout)
### Example 1: channels' 8-byte chunk header (binary layout, BAST)
```rust
use alktype::{Schema, Endian};
use alktype::{Schema, Endian, Definitions};
let chunk_header = Schema::struct_()
.endian(Endian::Big)
.field("channel_id", Schema::uint32())
.field("length", Schema::uint32())
.build();
.field("length", Schema::uint32());
// Build a complete BAST document (single-type — one $defs entry).
let doc = Definitions::new().build_doc("ChunkHeader", chunk_header);
// -> {
// "AlkType:Struct": true,
// "endian": "big",
// "properties": {
// "channel_id": { "AlkType:Uint32": true },
// "length": { "AlkType:Uint32": true }
// "$defs": {
// "ChunkHeader": {
// "kind": "struct",
// "endian": "big",
// "fields": [
// { "name": "channel_id", "kind": "uint32" },
// { "name": "length", "kind": "uint32" }
// ]
// }
// }
// }
//
// Feed to AlkTypeEngine::compile(&mut chunk_header, LayoutMode::Packed)
// Feed to AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)
// then validate incoming frames via engine.validate_bytes(&frame).
```
### Example 2: call's `OperationSpec` input schema (JSON payload)
### Example 2: call's `OperationSpec` input schema (JSON payload, standard JSON Schema)
```rust
use alktype::Schema;
@@ -530,18 +599,19 @@ let read_file_input = Schema::object()
// }
//
// Stored in OperationSpec.input_schema; validated via the standard
// jsonschema validator (validate_json for parsed payloads, or via
// serde_json::from_slice then validate_json for wire frames).
// jsonschema validator (AlkTypeEngine::compile with Some(&read_file_input)
// for the validate_json path, or serde_json::from_slice then
// validate_json for wire frames).
```
### Example 3: SFTP `Packet` union (binary layout, byte discriminator)
The SFTP wire shape is `[type:u8][payload-struct]` — a struct with a
union payload field. The engine requires `AlkType:Struct` at the top
level (`OffsetMap::compute` / `SequentialReader::new` both enforce
this; a `Union` is a field type within a struct, not a top-level
schema). The builder constructs the union wrapped in a struct, and
`$defs` are merged into the top-level schema so `$ref`s resolve:
union payload field. The engine requires a struct at the root
(`OffsetMap::compute` / `SequentialReader::new` both enforce this; a
`Union` is a field type within a struct, not a top-level schema). The
builder constructs the union wrapped in a struct, and `$defs` are
placed inside the document via `build_doc` so `$ref`s resolve:
```rust
use alktype::{Definitions, Discriminator, AlkTypeKind, Schema};
@@ -555,27 +625,21 @@ defs.define("Status", Schema::struct_().field("code", Schema::uint32()).field("m
// A "Packet" is a struct with one field — the union. This mirrors
// SFTP's wire shape: [type:u8][payload-struct].
let mut packet = Schema::struct_()
.field(
"payload",
Schema::union_(Discriminator::Byte {
offset: 0,
disc_type: AlkTypeKind::Uint8,
})
.mapping("1", Schema::ref_def("Init"))
.mapping("3", Schema::ref_def("Open"))
.mapping("5", Schema::ref_def("Read"))
.mapping("6", Schema::ref_def("Write"))
.mapping("101", Schema::ref_def("Status")),
)
.build();
// Merge $defs into the top-level schema so $refs resolve at compile time.
defs.merge_into(&mut packet);
// Feed to AlkTypeEngine::compile(&mut packet, LayoutMode::Packed)
let doc = defs.build_doc("Packet", Schema::struct_()
.field("payload", Schema::union_(Discriminator::Byte {
offset: 0,
disc_type: AlkTypeKind::Uint8,
})
.mapping("1", Schema::ref_def("Init"))
.mapping("3", Schema::ref_def("Open"))
.mapping("5", Schema::ref_def("Read"))
.mapping("6", Schema::ref_def("Write"))
.mapping("101", Schema::ref_def("Status"))));
// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None)
// then validate incoming frames via engine.validate_bytes(&frame).
```
### Example 4: OperationSpec error schemas (named `$defs`)
### Example 4: OperationSpec error schemas (named `$defs`, standard JSON Schema)
```rust
use alktype::{Definitions, Schema};
@@ -610,15 +674,19 @@ let op_errors = vec![
http_status: Some(429),
},
];
// $defs is built once and stored alongside the OperationSpec
// `$defs` is built once and stored alongside the OperationSpec.
// (For the validate_json path, compile with Some(&defs.build()) as the
// json_schema argument — but typically OperationSpec schemas are
// validated directly via jsonschema, not via AlkTypeEngine.)
```
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation shapes the builder's setters produce |
| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 |
| BAST format + two output formats | [ADR-BAST](decisions/bast-bast-format.md) | `struct_()` → BAST, `object()` → standard JSON Schema (D-BAST-008) |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation semantics the builder's setters produce (location moved to BAST type-level properties) |
| Load-time validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | The builder does not pre-validate; compile-time is the validation point |
## Open Questions
@@ -634,12 +702,15 @@ and transitively on `Schema::union_`). See
## References
- [ADR-009](decisions/009-builder-api.md) — the decision this spec implements
- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format and the
two-output-format decision (D-BAST-008)
- [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) —
scope boundaries this module extends; "schemas are JSON" principle
- [ADR-003](decisions/003-schema-annotations.md) — the annotation
shapes the builder's setters produce
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds the
builder's constructors produce
semantics the builder's setters produce
- [`bast-format.md`](bast-format.md) — the normative BAST format
specification (the output format for `struct_()`)
- [schema-layer.md](schema-layer.md) — the BAST kinds and parser
- [validation.md](validation.md) — the validation layer that consumes
builder output (via `AlkTypeEngine::compile`)
- `@alkdev/alknet: docs/architecture/crates/call/operation-registry.md`
+24 -19
View File
@@ -100,7 +100,7 @@ impl SequentialReader {
```
`read_field`/`write_field` on `AlkTypeEngine` work for the fixed-size
primitive kinds and the length-prefixed `String`/`Bytes`/`Timestamp`
primitive kinds and the length-prefixed `String`/`Bytes`
fields. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a
`FieldValue` carrying a layout descriptor (byte range, variant start,
or array stride) for the consumer to recurse on — see §"FieldValue" above.
@@ -238,21 +238,23 @@ pub struct UnionDispatch {
}
```
After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)`
to get the variant schema, then reads the variant's fields at
After dispatch, the consumer calls `tunion::resolve_variant(union_node, &dispatch.key)`
to get the variant `BastType`, then reads the variant's fields at
`dispatch.variant_offset` using the normal `data_access` functions (or a
fresh `SequentialReader` scoped to the variant).
fresh `SequentialReader` scoped to the variant). `$ref` variant types are
returned as `BastType::Ref`; the caller resolves them via
`BastDoc::resolve_typeref` when a concrete definition is needed.
### Byte-offset discriminator
```rust
/// Read the discriminator value from a byte-offset TUnion. The discriminator
/// is a fixed-size integer (AlkType:Uint8/Uint16/Uint32) at a known byte
/// offset. Returns the mapping key (stringified integer) and the variant
/// struct offset.
/// is a fixed-size integer (uint8/uint16/uint32) at a known byte offset.
/// Returns the mapping key (stringified integer) and the variant struct
/// offset.
pub fn read_byte_discriminator(
buffer: &[u8],
union_schema: &Value,
union_node: &BastUnion<'_>,
endian: Endian,
) -> Result<UnionDispatch, AlkTypeError>;
```
@@ -268,11 +270,11 @@ starts at `offset + discriminator_size`.
/// Read the discriminator value from a field-name TUnion. The
/// discriminator is a named field within the struct — the consumer
/// provides the field's computed offset (from the OffsetMap or
/// LayoutBuilder). Supports AlkType:String, Uint8, and Enum discriminator
/// LayoutBuilder). Supports string, uint8, and enum discriminator
/// fields.
pub fn read_field_discriminator(
buffer: &[u8],
union_schema: &Value,
union_node: &BastUnion<'_>,
disc_field_offset: usize,
endian: Endian,
) -> Result<UnionDispatch, AlkTypeError>;
@@ -287,16 +289,19 @@ the variant's fields starting at the end of the discriminator field.
### Variant resolution
```rust
/// Look up a variant schema from the union's mapping. Inline schemas
/// are returned directly. $ref pointers of the form "#/$defs/<name>"
/// are resolved against the union schema's own $defs block.
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str)
-> Result<&'a Value, AlkTypeError>;
/// Look up a variant type from the union's mapping. Inline struct/
/// union/enum types are returned directly; `$ref` pointers are
/// returned as `BastType::Ref` — the caller resolves them via
/// `BastDoc::resolve_typeref` when a concrete definition is needed.
pub fn resolve_variant<'a>(
union_node: &'a BastUnion<'a>,
key: &str,
) -> Result<&'a BastType<'a>, AlkTypeError>;
/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a
/// Get the discriminator's byte size (1/2/4 for uint8/16/32) for a
/// byte-offset TUnion. Field-name discriminators have no fixed size
/// and produce a AlkTypeError::Schema.
pub fn discriminator_size(union_schema: &Value) -> Result<usize, AlkTypeError>;
/// and produce an `AlkTypeError::Schema`.
pub fn discriminator_size(union_node: &BastUnion<'_>) -> Result<usize, AlkTypeError>;
```
### TUnion in the layout engines
@@ -324,7 +329,7 @@ dispatch to the primitive `data_access` function for the field's kind.
For aligned-mode access, `AlkTypeEngine::read_field(&buffer, "header.version")`
returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds
the field's `AlkType:*` kind in the schema, and calls the matching
the field's `AlkTypeKind` in the BAST typed tree, and calls the matching
`data_access::read_*` function. `write_field` is the mirror. Composite
kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue`
carrying a layout descriptor; the consumer recurses with a fresh reader
@@ -1,7 +1,19 @@
# ADR-001: alktype — Purpose, Scope, and the jsonschema Engine
## Status
Accepted
**Superseded (format-specific content) by
[ADR-BAST](bast-bast-format.md).** The crate's purpose, scope
boundaries, and the "schema is the format" principle are **retained and
strengthened** — BAST *is* the format. Only the *concrete format*
(custom-keyword JSON Schema → BAST) and the *validation strategy*
(single `jsonschema` custom-keyword validator → two-validator model)
are superseded: the format-specific content by ADR-BAST, the
validation-strategy content by
[ADR-VAL-SPLIT](val-split-two-validator-model.md). This ADR is kept as
the historical record of the v0.1.0 design and the purpose/scope
decision; read it alongside ADR-BAST and ADR-VAL-SPLIT for the current
state.
## Context
@@ -1,7 +1,14 @@
# ADR-002: Two Layout Modes — Packed Sequential vs Aligned Static
## Status
Accepted
Accepted — unchanged under the BAST pivot
([ADR-BAST](bast-bast-format.md)). Layout modes are format-agnostic:
the input format changed from custom-keyword JSON Schema to BAST, but
the two modes, their alignment/packing rules, and the
`LayoutBuilder`/`SequentialReader`/`OffsetMap` API did not. The layout
engines now walk the BAST typed tree ([`BastDoc`](../schema-layer.md))
instead of raw JSON with `get_alktype_kind*`, but the offset
computation algorithm is identical.
## Context
@@ -1,7 +1,18 @@
# ADR-003: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
## Status
Accepted
**Accepted (semantics); amended (location) by
[ADR-BAST](bast-bast-format.md).** The annotation *semantics* decided
here — endianness default, struct/field-level alignment, the three
variable-length encoding strategies, and the two TUnion discriminator
kinds — **carry forward unchanged** under the BAST pivot. Only the
annotation *location* moves: from v0.1.0's custom-keyword objects
(`{"AlkType:String": { "encoding": "..." }}`) to BAST type-level
properties (`{ "name": "handle", "kind": "string", "encoding": "..." }`).
The BAST shapes are normative in
[`bast-format.md`](../bast-format.md#variable-length-encoding); this
ADR is kept as the semantic reference. Read it alongside ADR-BAST.
## Context
@@ -1,7 +1,21 @@
# ADR-004: Error Handling and Validation Strategy
## Status
Accepted
**Accepted (error type); amended (validation strategy) by
[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The `AlkTypeError`
enum, its four variants, the load-time-build / access-time-check split,
and the field-path-carrying errors decided here are **retained
unchanged** under the BAST pivot (D-BAST-009 keeps
`Validation(jsonschema::ValidationError<'static>)`). The "validation
strategy" section — which described v0.1.0's single
`jsonschema`-custom-keyword validator for both paths — is **refined**:
the bytes path now uses the BAST-native validator
(`bast_validation`), the JSON path now uses a standard
`jsonschema::Validator` from a consumer-provided JSON Schema. See
[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the
two-validator model. This ADR is kept as the error-handling reference;
read it alongside ADR-VAL-SPLIT for the current validation strategy.
## Context
+13 -1
View File
@@ -2,7 +2,19 @@
## Status
Accepted
**Accepted (API surface); amended (output format) by
[ADR-BAST](bast-bast-format.md).** The fluent builder API, the
`Schema`/`Definitions`/`Discriminator` types, the constructor and
setter catalog, and the "produces `serde_json::Value`, not a typed
`Schema` enum" decision decided here are **retained unchanged** under
the BAST pivot. Only the `build()` *output format* changes:
`struct_()` now produces BAST JSON (`{ "kind": "struct", "fields": [...] }`)
instead of v0.1.0's custom-keyword JSON (`{ "AlkType:Struct": true,
"properties": {...} }`); `object()` continues to produce standard JSON
Schema. This is D-BAST-008, recorded in ADR-BAST. The builder examples
in [`builder.md`](../builder.md) reflect the current BAST output. This
ADR is kept as the API-surface decision; read it alongside ADR-BAST
for the output format.
## Context
@@ -2,7 +2,23 @@
## Status
Accepted
**Accepted (two-step concept); amended (validation step) by
[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The
`validate_bytes(&[u8])` entry point, the "materialize `Value` from
bytes, then validate" two-step concept, the mode dispatch, the
field-path-carrying errors, and the "not a `Validator` trait / not
framing-aware / not a binary-payload validator for JSON-only schemas"
scope boundaries decided here are **retained unchanged** under the
BAST pivot. Only the validation *step's implementation* changes: the
materialized `Value` is validated by the **BAST-native validator**
(`bast_validation`) instead of v0.1.0's `jsonschema` custom-keyword
validator. The `jsonschema` crate is no longer touched on the bytes
path (it remains for the `validate_json` path and for BAST meta-schema
validation). The error payload type stays
`Validation(jsonschema::ValidationError<'static>)` (D-BAST-009). See
[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the
two-validator model. This ADR is kept as the `validate_bytes` decision;
read it alongside ADR-VAL-SPLIT for the current validation step.
## Context
@@ -0,0 +1,316 @@
# ADR-BAST: BAST (Binary Abstract Syntax Tree) as the Schema Format
## Status
Accepted — supersedes the format-specific content of
[ADR-001](001-alktype-purpose-scope-jsonschema-engine.md). ADR-001's
purpose, scope, and "schema is the format" principle are retained and
strengthened; only the concrete format (custom-keyword JSON Schema →
BAST) is superseded by this ADR.
## Context
alktype v0.1.0 embedded binary layout information inside standard JSON
Schema documents via custom keywords:
```json
{
"AlkType:Struct": true,
"type": "object",
"properties": {
"channel_id": { "AlkType:Uint32": true, "type": "integer" },
"length": { "AlkType:Uint32": true, "type": "integer" }
},
"endian": "big"
}
```
This worked for the Rust engine — it walked the tree, detected
keywords, computed offsets. But it created friction for everything
outside Rust:
1. **Cross-language consumption.** A Python, Go, or TypeScript consumer
that wanted to parse an alktype schema had to re-implement custom
keyword detection. The format was not self-describing — you needed
to know that `AlkType:Uint32` meant "4-byte little/big-endian
unsigned integer" out-of-band.
2. **No meta-schema.** The custom keywords were not part of any JSON
Schema dialect, so `jsonschema` itself could not validate an alktype
schema's *structure*. Editors had no autocomplete; a typo in
`AlkType:Uint32` (e.g., `AlkType:UINT32`) was a runtime engine
error, not a schema-validation error.
3. **Awkward composition.** The keyword-value shape (`true` vs an
annotation object) and the `normalize_refs` step needed to bridge
TypeBox's bare-name `$ref` output and `jsonschema`'s JSON Pointer
requirement were engine internals leaking into the format.
4. **Validator coupling.** The v0.1.0 bytes path validated by
registering 19 `jsonschema::Keyword` factories (~200 lines). The
custom keyword integration was the only way to enforce value-domain
constraints (integer ranges, `maxLength`, enum index bounds) on the
materialized `Value`. The built-in `enum` keyword checked string
membership, but the materializer emitted `Value::Number(index)` —
so out-of-bounds enum indices *silently passed* (a dead constraint).
The engine's core logic (layout computation, data access, union
dispatch, two layout modes) was format-agnostic beneath the accessor
layer. A POC on branch `bast-validator-poc` (commit `f371fe4`,
`src/bast_poc.rs`) proved that a `kind`-based vocabulary with
`$defs`/`$ref`, a BAST-native validator, and lazy variant ref
resolution could replace the custom-keyword machinery end-to-end with
no loss of capability and a net reduction in code. The research record
is [`docs/research/bast-pivot.md`](../../research/bast-pivot.md);
decisions D-BAST-001 through D-BAST-009 are recorded there.
## Decision
**alktype's schema format is BAST (Binary Abstract Syntax Tree): a JSON
document that describes binary data layouts using a `kind`-based
vocabulary with `$defs`/`$ref` for composition.** BAST is itself a
valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser.
The normative format specification is
[`docs/architecture/bast-format.md`](../bast-format.md) (meta-schema,
TypeRef, examples, validation model). This ADR records the decision and
its consequences; the spec records the shape.
### Design principles
1. **BAST is a JSON Schema instance.** A BAST document is valid JSON
that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON
Schema). Any JSON Schema validator can check whether a BAST document
is well-formed; editors with JSON Schema support provide autocomplete
and inline validation for free.
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
top-level `$defs` block. `$ref` handles cross-references and union
variant references — the same pattern as JSON Schema's own `$defs`
and TypeBox's `Type.Module`. No custom reference resolution
mechanism.
3. **`kind`-based vocabulary.** Every type has a `kind` field whose
value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.).
This replaces the `AlkType:*` custom-keyword pattern with a flat,
easily-matched string. The 18 `AlkTypeKind` enum variants are
the BAST kinds; `AlkTypeKind::from_bast_str`/`to_bast_str` map between
the enum and the lowercase BAST strings (D-BAST-002).
4. **Order is explicit.** Struct fields are an ordered array, not an
object with `properties`. Field order is unambiguous — no reliance
on `serde_json`'s `preserve_order` for correctness — and matches the
mental model of binary layouts.
5. **Annotations are type-level properties.** Endianness, alignment,
encoding, and discriminators are properties of the type definition
or field, not custom keywords on a separate schema object. Their
*semantics* carry forward unchanged from
[ADR-003](003-schema-annotations.md); only their *location* moves.
### Document shape
Every BAST document has the same top-level shape:
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
- The `$defs` block is **required** (D-BAST-003). Single-type documents
are a special case with one entry.
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001).
Convention (first entry) is fragile and depends on JSON key order; an
explicit parameter is used instead.
### TypeRef
`TypeRef` is the central mechanism for referencing types. Four forms:
primitive string (`"uint32"`), `$ref` object
(`{ "$ref": "#/$defs/Read" }`), array object
(`{ "kind": "array", "element": "uint32", "count": 3 }`), and record
object (`{ "kind": "record", "values": "string" }`).
The `$ref` form uses standard JSON Pointer syntax **restricted to
`#/$defs/<name>`** — no external references, no fragment-only pointers,
no bare names. The restriction keeps resolution a single hash lookup
and eliminates the `normalize_refs` step the v0.1.0 engine needed for
TypeBox's bare-name refs.
### Meta-schema
The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that
validates the *structure* of BAST documents (is it well-formed?). It
lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded
in the crate as `BAST_META_SCHEMA` (re-exported from the crate root) for
offline use. A *different* validator — the BAST-native validator (see
[ADR-VAL-SPLIT](val-split-two-validator-model.md)) — validates *binary
data* against a BAST document (are the bytes a valid instance?). These
are different validators for different inputs.
### The typed parser
`src/bast.rs` parses a BAST document into a borrowed typed tree
(`BastDoc`/`BastDef`/`BastStruct`/`BastField`/`BastType`/`BastUnion`/
`BastEnum`/`BastArray`/`BastRecord`/`BastRef`). Three consumers (layout
engines, materializer, BAST-native validator) walk the same tree, so a
typed view pays for itself. See
[`schema-layer.md`](../schema-layer.md) for the parser's surface and
[`bast-format.md`](../bast-format.md) for the format.
Variant `$ref`s (union `mapping` entries) are resolved **lazily** by
the materializer and validator via `BastDoc::resolve_typeref` — no
compile-time inlining.
### Untrusted input
Every path that walks a BAST document returns
`Err(AlkTypeError::Schema)` on a malformed document, never
`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream
`alkcall` consumer accepts schemas from arbitrary internet peers in its
hub/spoke topology).
### Bug fix: enum index bounds
The v0.1.0 engine had a dead constraint on the bytes path — the
built-in `enum` keyword checked string membership, but the materializer
emitted `Value::Number(index)`, which never matched. The BAST-native
validator checks the materialized index against `values.len()` bounds,
fixing this. Net improvement, recorded as intended behavior in the
test suite.
## What is removed
Under the BAST pivot, the v0.1.0 custom-keyword machinery is removed:
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
factories) — replaced by the BAST-native validator (~250 lines, a
flat match with no factories, no trait objects, no sub-validator
pre-computation). See [ADR-VAL-SPLIT](val-split-two-validator-model.md).
- `normalize_refs()` / `inline_union_variant_refs()` — BAST refs are
always `#/$defs/<name>`; one hash lookup. Variant refs resolve lazily.
- `get_alktype_kind*` family — superseded by the parser's `kind`-string
dispatch.
- `parse_encoding`/`parse_align`/`parse_max_length`/`parse_endian`/
`parse_discriminator` + `DiscriminatorKind` — replaced by the typed
`BastField`/`BastDiscriminator` views and the parser's internal
BAST-property-form copies.
- `resolve_ref`/`resolve_ref_or_inline` — replaced by
`BastDoc::lookup_def`/`resolve_typeref`.
- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`,
`BYTE_DISCRIMINATOR_TYPES` — replaced by `from_bast_str`/`to_bast_str`
and the parser's typed views.
- The custom-keyword `build_validator` path — `build_validator` is
repurposed to build a *standard* `jsonschema::Validator` from a
consumer-provided JSON Schema (no custom keywords). See
[ADR-VAL-SPLIT](val-split-two-validator-model.md).
The `jsonschema` crate **remains a direct dependency** for
`validate_json` and for validating BAST documents against the BAST
meta-schema. The only thing removed is the custom keyword integration
path.
## Consequences
### Positive
- **Self-describing, cross-language format.** A BAST document carries
its type vocabulary in a meta-schema'd JSON Schema instance. Any
language with a JSON parser and a JSON Schema validator can validate
BAST document structure without knowing alktype's Rust internals.
Editors with `$schema` support provide autocomplete and inline
validation for free.
- **Simpler `$ref` story.** One restricted form (`#/$defs/<name>`), one
hash lookup, no normalization pass. TypeBox interop is a serialization
concern (TypeBox → BAST JSON), not an engine concern.
- **Explicit field order.** The `fields` array makes byte order
unambiguous — no reliance on `serde_json`'s `preserve_order` for
correctness (it remains a dependency for builder output and for
`mapping` iteration order, but layout correctness no longer depends
on it).
- **Architecture simplification.** The BAST-native validator is a flat
recursive match — no factories, no trait objects, no sub-validator
pre-computation, no `with_keyword` registration. ~200 lines of
custom-keyword validators become ~250 lines of straightforward
pattern matching.
- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed.
- **Wasm binary-size win.** The `validate_bytes` path no longer touches
`jsonschema` for validation (it still uses `jsonschema`'s
`ValidationError::custom` type for the error payload, per
D-BAST-009 — but no validator compilation, no keyword registration,
no sub-validators).
### Negative
- **Breaking change to the v0.1.0 public surface.** `compile`'s
signature changes (new `root_name` param, drops `&mut`, takes a BAST
document not a custom-keyword JSON Schema). `validate_json`/
`is_valid_json` change contract (validate against a consumer-provided
JSON Schema, not the alktype schema). `Schema::build`/
`Definitions::build` output format changes. The ~13 `schema::*`
helper re-exports are removed. `build_validator` is repurposed. The
crate is on crates.io at 0.1.0 with zero real consumers, so the bump
is free — but the contract is explicit (see the implementation plan's
Semver Contract table).
- **Two output formats from the builder.** `struct_()` → BAST,
`object()` → standard JSON Schema. The construction API is the same;
only the serialization differs. This is deliberate (D-BAST-008) but
is a thing consumers must learn.
- **`AlkTypeKind::Display` is backed by `to_bast_str`.** The v0.1.0
`as_str`/`Display` rendered the `"AlkType:Uint8"` keyword; the new
`Display` renders the BAST canonical string (`"uint8"`). Error
messages across six modules surface the new name. This is the right
name to surface now, but it is a visible change in error output.
## Scope Boundaries (What This Is Not)
- **Not a replacement for JSON Schema for JSON validation.** A BAST
document cannot validate a JSON payload — it describes binary data
layouts and value-domain constraints for bytes. For JSON validation,
consumers use standard JSON Schema documents (which may be derived
from BAST via future codegen, or authored separately). See
[ADR-VAL-SPLIT](val-split-two-validator-model.md).
- **Not a code generator.** BAST is a data format, not a Rust source
generator. ADR-001's scope boundary stands.
- **Not a schema-evolution / Value system.** TypeBox's `Value.Diff`,
`Value.Migrate`, `Value.Convert` remain out of scope (ADR-001).
- **Not a framing format.** BAST describes one struct/union/enum
instance; it does not strip length prefixes or handle multi-frame
buffers. Framing stays in the consumer (ADR-010).
## Decisions (D-BAST-001..009)
The BAST format is grounded in decisions D-BAST-001 through D-BAST-009,
recorded in
[the pivot research record](../../research/bast-pivot.md#decisions).
Summary:
| Decision | Summary |
|----------|---------|
| [D-BAST-001](../../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention |
| [D-BAST-002](../../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase |
| [D-BAST-003](../../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape |
| [D-BAST-004](../../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) |
| [D-BAST-005](../../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` |
| [D-BAST-006](../../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed |
| [D-BAST-007](../../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema |
| [D-BAST-008](../../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema |
| [D-BAST-009](../../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths |
## References
- [`bast-format.md`](../bast-format.md) — the normative BAST format
specification
- [`schema-layer.md`](../schema-layer.md) — the BAST parser
implementation
- [ADR-VAL-SPLIT](val-split-two-validator-model.md) — the two-validator
model (BAST-native for bytes, standard `jsonschema` for JSON)
- [ADR-001](001-alktype-purpose-scope-jsonschema-engine.md) — purpose,
scope, and the "schema is the format" principle (format-specific
content superseded by this ADR; purpose/scope retained)
- [ADR-003](003-schema-annotations.md) — annotation semantics (carry
forward unchanged; only location moves)
- [ADR-009](009-builder-api.md) — builder API (output format amended
to BAST / standard JSON Schema)
- [ADR-010](010-generalized-validation-validate-bytes.md) —
`validate_bytes` (validation step amended to the BAST-native
validator)
- [BAST pivot research record](../../research/bast-pivot.md) —
motivation, POC scope and result, decisions D-BAST-001..009, risks
- [BAST pivot implementation plan](../../plans/bast-implementation.md)
— ordered steps, semver contract, ADR-sync checklist
@@ -0,0 +1,228 @@
# ADR-VAL-SPLIT: Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON
## Status
Accepted — refines the "validation strategy" section of
[ADR-004](004-error-handling-validation-strategy.md) and the "validation
step" of [ADR-010](010-generalized-validation-validate-bytes.md) for
the BAST pivot. Records decisions D-BAST-006, D-BAST-007, and
D-BAST-009.
## Context
alktype v0.1.0 used a single validation mechanism — the `jsonschema`
crate with 19 custom keyword validators — for both the JSON path
(`validate_json(&Value)`) and the bytes path (`validate_bytes(&[u8])`).
The bytes path materialized a `serde_json::Value` tree from the buffer,
then ran the same `jsonschema::Validator` against it.
Under the BAST pivot ([ADR-BAST](bast-bast-format.md)), the format
changed from custom-keyword JSON Schema to BAST, and the custom-keyword
integration was removed. This forced a re-evaluation of both validation
paths:
1. **The bytes path.** BAST is the complete specification of the binary
format — it describes both the layout (how to read) and the
constraints (what values are valid). An external JSON Schema is not
needed for `validate_bytes`; the BAST document *is* the validation
spec for bytes. The natural validator is a recursive walker over the
BAST type tree that checks the value-domain constraints the
materializer does not (integer ranges, `maxLength`,
enum index bounds, union variant constraints). The POC
(`bast-validator-poc` branch, `src/bast_poc.rs`) proved this out
end-to-end with 20 reference tests.
2. **The JSON path.** BAST describes bytes, not JSON shape. A JSON
`Value` (e.g., an incoming JSON-RPC request) is the wrong input for
a BAST document; the right validator is a standard
`jsonschema::Validator` built from a standard JSON Schema document
the consumer provides. BAST is not involved on this path. This is
the path alkcall uses for its `OperationSpec` JSON validation.
The two paths have different inputs (bytes vs JSON `Value`), different
schema sources (the BAST document vs a consumer-provided JSON Schema),
and different validators (a flat recursive match vs a compiled
`jsonschema::Validator`). But they share the same error variant —
`AlkTypeError::Validation` — so consumers handling both (alkcall uses
`validate_json` for channel 0 JSON-RPC and `validate_bytes` for binary
channels) match one arm.
## Decision
**alktype has two validators for two input types:**
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator (`bast_validation`) | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
### `validate_bytes` — BAST-native validator (D-BAST-006)
`src/bast_validation.rs` is a recursive walker
(`validate_value(doc, &value)`) over the BAST typed tree
([`crate::bast::BastDoc`]/[`BastType`]). The materializer
(`src/materialize.rs`) produces a structurally-correct `Value` tree
from bytes (all declared fields present, types correct, bounds checked,
UTF-8 valid, discriminator in mapping, boolean byte 0 or 1). The
validator enforces only the **value-domain constraints expressed in the
BAST document** — the ones the materializer can't see from the bytes
alone:
| Constraint | Validator arm |
|------------|---------------|
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` |
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` |
| Float finiteness (Float32/64) | `validate_float` |
| String `maxLength` (byte length) | `check_string` |
| Bytes `maxLength` (array length) | `check_bytes` (accepts `Value::String` and `Value::Array`) |
| Enum index bounds | `validate_enum` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `validate_union` reads `__discriminator`, resolves the variant, recurses |
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
| Record values | `validate_record` recurses into each value's `values` type |
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
The validator is a flat `match` — no factories, no trait objects, no
sub-validator pre-computation, no `jsonschema` involvement. ~250 lines
replace ~200 lines of v0.1.0 custom-keyword factories.
No external JSON Schema is required. The BAST document is the complete
specification of the binary format. An optional external JSON Schema
can be layered on top for constraints BAST doesn't express (cross-field
consistency, regex patterns on string content) — additive, not
load-bearing.
### `validate_json` — standard jsonschema (D-BAST-007)
`validate_json(&Value)` / `is_valid_json(&Value)` validate a JSON
`Value` against a standard `jsonschema::Validator` compiled at
`AlkTypeEngine::compile` time from a consumer-provided JSON Schema
(`compile`'s `json_schema: Option<&Value>` parameter). No custom
keywords, no BAST involvement. The JSON Schema is independent of the
BAST document — BAST describes bytes, not JSON shape. It may be
authored separately or derived from BAST via future codegen.
If no JSON Schema was supplied to `compile`, `validate_json` returns
`AlkTypeError::Schema` and `is_valid_json` returns `false`.
`build_validator` (in `src/validation.rs`) is **repurposed**: it builds
a *standard* `jsonschema::Validator` from a plain JSON Schema (no
custom keywords). The engine calls it internally during `compile` when
`json_schema` is `Some`. Consumers that only need a one-off validator
may call `jsonschema::options().build(schema)` directly; `build_validator`
exists so the engine's error mapping (`jsonschema` build error →
`AlkTypeError::Schema`) is reused. The v0.1.0 custom-keyword
`build_validator` is removed.
The `jsonschema` crate remains a direct dependency for this path and
for validating BAST documents against the BAST meta-schema.
### Error payload (D-BAST-009)
`AlkTypeError::Validation(jsonschema::ValidationError<'static>)` is
**retained** as the error variant for both paths. The bytes path no
longer uses `jsonschema` for validation, so its error payload is
constructed via `jsonschema::ValidationError::custom` purely to keep
the variant's type unchanged. The rationale is consumer ergonomics on
the *combined* path: consumers like alkcall use both `validate_json`
and `validate_bytes` and handle `AlkTypeError::Validation` in one
place. A single uniform payload type means one match arm covers both
sources.
The alternative (`Validation(String)`) was rejected — it would force
`validate_json` to flatten its structured errors (instance path, schema
path, keyword) to a `String` via `Display`. The more information-rich
path would lose data to accommodate the less rich one. That is the
wrong direction.
The `no_std`/minimal-build angle (OQ-002) that the alternative was
meant to enable is moot: `validate_json` requires `jsonschema`
regardless, so a bytes-only `no_std` build already has to give up
`validate_json` as a separate, larger decision. The right place to
revisit is when/if OQ-002 is actually pursued.
## Consequences
### Positive
- **Right validator for each input.** Bytes are validated by the BAST
document that describes them; JSON values are validated by a JSON
Schema that describes them. No forced isomorphism between two
different input types.
- **No external JSON Schema needed for `validate_bytes`.** The BAST
document is both the layout spec and the validation spec for bytes.
This is the "schema is the format" principle from ADR-001, now fully
realized.
- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed
— the BAST-native validator checks the materialized index against
`values.len()` directly.
- **Per-variant constraint enforcement (OQ-008) without custom
keywords.** The validator recurses into the selected variant's BAST
definition on `__discriminator` lookup, enforcing every field
constraint the variant declares (e.g., `maxLength` on a `bytes` field
inside a variant struct).
- **Wasm binary-size win.** The `validate_bytes` path no longer touches
`jsonschema` for validation (it still uses `ValidationError::custom`
for the error payload type, per D-BAST-009 — but no validator
compilation, no keyword registration, no sub-validators).
- **Architecture simplification.** ~200 lines of custom-keyword
factories become ~250 lines of straightforward pattern matching. No
`with_keyword` registration; no `inline_union_variant_refs` compile
step.
- **Uniform error payload.** Consumers handle one
`AlkTypeError::Validation` match arm for both paths (D-BAST-009).
### Negative
- **Two validators, not one.** The engine struct carries an
`Option<jsonschema::Validator>` (for `validate_json`) and re-parses
the BAST typed tree on each `validate_bytes` call (the BAST-native
validator is not pre-built — it's a recursive walker over the
on-demand `BastDoc`). This is a small cost; the validators serve
different inputs and don't share structure.
- **`validate_json` requires a consumer-provided JSON Schema.** The
engine no longer builds a validator from the alktype schema; the
consumer must supply a JSON Schema at `compile` time (or accept that
`validate_json` returns `AlkTypeError::Schema`). This is a behavioral
break from v0.1.0, intentional under the pivot.
- **`AlkTypeError::Validation` payload is `jsonschema`'s type even on
the bytes path.** The bytes path constructs it via
`ValidationError::custom`, which is slightly awkward but keeps the
variant uniform. The `no_std` revisit (OQ-002) is the place to
reconsider if a bytes-only minimal build ever materializes.
## Scope Boundaries (What This Is Not)
- **Not a `Validator` trait abstraction.** Two methods on one struct,
not a trait with impls for JSON-only and BAST-binary schemas. The two
impls share little internally (`validate_json` is a single
`jsonschema` call; `validate_bytes` is materialize + BAST-native
walk), so a trait would add a layer without unifying behavior. See
[ADR-010](010-generalized-validation-validate-bytes.md) §"Not a
`Validator` trait abstraction".
- **Not a binary-aware validator that skips the `Value` tree.** The
`Value`-materialization path is the validation path. A future
"validate bytes without materializing" path is a two-way door but
explicitly out of scope for v1 (would re-introduce a hand-rolled
validator, ADR-001).
- **Not framing-aware.** `validate_bytes` validates the bytes of *one*
schema instance. Framing stays in the consumer (ADR-010).
## References
- [`bast-format.md` §Validation Model](../bast-format.md#validation-model)
— the normative validation model
- [ADR-BAST](bast-bast-format.md) — the BAST format decision
- [ADR-004](004-error-handling-validation-strategy.md) — error handling
and validation strategy (load-time build, access-time check,
`AlkTypeError` enum — retained; validation-strategy section refined
by this ADR)
- [ADR-010](010-generalized-validation-validate-bytes.md) —
`validate_bytes` (the two-step concept retained; the validation step
amended to the BAST-native validator by this ADR)
- [`validation.md`](../validation.md) — the validation layer
documentation
- `src/bast_validation.rs` — the BAST-native validator implementation
- `src/validation.rs` — the `build_validator` helper
- [BAST pivot research record](../../research/bast-pivot.md) —
D-BAST-006, D-BAST-007, D-BAST-009
+20 -19
View File
@@ -7,8 +7,8 @@ last_updated: 2026-07-22
The layout engine: offset computation, the two layout modes (packed
sequential vs aligned static), alignment, endianness, and variable-length
field handling. This is the novel code — the recursive walk of the schema
JSON that computes byte positions for each field.
field handling. This is the novel code — the recursive walk of the BAST
typed tree that computes byte positions for each field.
## The Two Layout Modes
@@ -24,8 +24,8 @@ protocols.
**Components:**
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `AlkType:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(bast_doc, root_name)` (requires a `struct` at the root), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `SequentialReader::new(bast_doc, root_name)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
**How it works:**
@@ -69,7 +69,7 @@ and safetensors.
**Component:**
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, AlkTypeError>` (requires `AlkType:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
- **`OffsetMap`** — constructed via `OffsetMap::compute(&doc) -> Result<Self, AlkTypeError>` (requires a `struct` at the root). Walks the BAST typed tree once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
**How it works:**
@@ -118,14 +118,15 @@ fields must use `maxLength` (fixed-size reservation) or
## Offset Computation Algorithm
The offset computation is a recursive walk of the schema JSON. The
The offset computation is a recursive walk of the BAST typed tree
([`BastDoc`](schema-layer.md#the-bast-parser-bast-module)). The
algorithm is the same for both modes; the difference is whether alignment
padding is inserted between fields.
### Fixed-size types
For each fixed-size type, the algorithm:
1. Determines the type's byte size from the `AlkType:*` kind.
1. Determines the type's byte size from the `AlkTypeKind`.
2. In aligned mode: inserts padding to satisfy the type's alignment
(or the field's `align` annotation, or the struct's `align` default).
3. Records the field's `(start, end)` range.
@@ -133,14 +134,14 @@ For each fixed-size type, the algorithm:
### Composite types
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
**`struct`:** Recurse into the struct's `fields` array. The inner fields
are computed relative to the struct's start offset. The struct's total
size is the sum of its fields' sizes (plus alignment padding in aligned
mode). The struct itself may have an `align` annotation that rounds up
its total size.
**`TUnion`:** TUnion is supported in packed sequential mode only. In
aligned static mode, `OffsetMap::compute` rejects `TUnion` fields with
**`union`:** TUnion is supported in packed sequential mode only. In
aligned static mode, `OffsetMap::compute` rejects `union` fields with
`AlkTypeError::Offset` — see
[ADR-008](decisions/008-reject-tunion-in-aligned-mode.md). Unions
are the protocol dispatch pattern (SFTP type bytes, call protocol event
@@ -161,12 +162,12 @@ total size. The `SequentialReader` reads the discriminator first, looks
up the variant schema, then reads the variant struct sequentially — it
doesn't need to know the union's total size upfront.
**`TArray` of fixed-size elements:** Element stride = element size (plus
**`array` of fixed-size elements:** Element stride = element size (plus
alignment padding in aligned mode). Element `i` starts at
`array_offset + i × stride`. The array's total size is `count × stride`.
**`TArray` of variable-length-element structs:** Deferred for v1
(OQ-001).
**`array` of variable-length-element structs:** Deferred for v1
(OQ-001, D-BAST-004 — BAST arrays require `count` in v1).
### Variable-length types
@@ -289,7 +290,7 @@ pub struct FieldPosition {
A field's computed position in a packed layout, produced by
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
length prefix); for fixed-size fields, `size` is the type's byte size.
`kind` records the field's `AlkType:*` kind so the consumer can dispatch
`kind` records the field's `AlkTypeKind` so the consumer can dispatch
to the correct `data_access` read/write function.
### `PackedLayout` (packed mode)
@@ -317,16 +318,16 @@ A flat table of `(field_path, byte_range)` pairs computed from a schema.
```rust
impl OffsetMap {
pub fn compute(schema: &Value) -> Result<Self, AlkTypeError>;
pub fn compute<'a>(doc: &'a BastDoc<'a>) -> Result<Self, AlkTypeError>;
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
}
```
`compute` requires a `AlkType:Struct` at the top level. `total_size`
includes trailing alignment padding. `iter` yields fields in insertion
order (schema `properties` order, nested struct fields appearing inline).
`compute` requires a `struct` at the root. `total_size`
includes trailing alignment padding. `iter` yields fields in the BAST
`fields` array order (nested struct fields appearing inline).
## Design Decisions
@@ -355,7 +356,7 @@ See [open-questions.md](open-questions.md) for full details.
the two layout modes decision
- [ADR-003](decisions/003-schema-annotations.md) — schema
annotations
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds and their
- [schema-layer.md](schema-layer.md) — the 18 AlkType kinds and their
byte sizes
- [data-access.md](data-access.md) — read/write functions that use the
computed offsets
+19 -15
View File
@@ -106,31 +106,35 @@ architect's desk" is answerable at a glance.
### OQ-006: Builder spec Example 3 — wrap `Union` in a `Struct` — RESOLVED
- **Status**: resolved. [builder.md](builder.md) Example 3 now wraps
the `Union` in a `Schema::struct_().field("payload", ...)` and merges
`$defs` via `Definitions::merge_into`. Matches the realistic SFTP
wire shape and the engine's `AlkType:Struct`-at-root constraint.
the `Union` in a `Schema::struct_().field("payload", ...)` and builds
a complete BAST document via `Definitions::build_doc`. Matches the
realistic SFTP wire shape and the engine's struct-at-root constraint
(the root `$defs` entry must be a `struct`).
- **Full file**: [OQ-006](questions/006-builder-spec-example-3-wrap-union.md)
### OQ-007: `Bytes` materialization — lossy UTF-8 conversion — RESOLVED
- **Status**: resolved. Array of u8: the materializer produces
`Value::Array` of `Value::Number` (one entry per byte, 0..=255) for
`AlkType:Bytes` fields. The `BytesValidator` accepts both
`Value::String` (for `validate_json`) and `Value::Array` (for
`validate_bytes`). `maxLength` = max byte count. Implemented in
`src/materialize.rs` and `src/validation.rs`.
`bytes` fields. The BAST-native validator (`bast_validation::check_bytes`)
accepts both `Value::String` (for `validate_json`-style inputs) and
`Value::Array` (for `validate_bytes`). `maxLength` = max byte count.
Implemented in `src/materialize.rs` and `src/bast_validation.rs`.
- **Full file**: [OQ-007](questions/007-bytes-materialization-lossy-utf8.md)
### OQ-008: `UnionValidator` variant dispatch — RESOLVED
- **Status**: resolved. `UnionValidator` now builds a sub-validator for
each variant at factory time and dispatches on `__discriminator` at
validation time. `AlkTypeEngine::compile` calls
`schema::inline_union_variant_refs` before `build_validator` to inline
`$ref`s in union `mapping` entries (necessary because the
`union_factory` receives the union node, but `$defs` live at the
schema root). Implemented in `src/validation.rs`, `src/schema.rs`,
and `src/engine.rs`.
- **Status**: resolved. Under the BAST pivot, the BAST-native validator
(`bast_validation::validate_union`) reads `__discriminator`, looks up
the variant `BastType` in the union's `mapping`, and recurses into the
variant's BAST definition via `validate_typeref` — enforcing every
field constraint the variant declares (e.g. `maxLength` on a `bytes`
field inside a variant struct). Variant `$ref`s resolve lazily via
`BastDoc::resolve_typeref` — no `inline_union_variant_refs` compile
step (removed under BAST). No custom keywords, no `jsonschema`
involvement on the bytes path. Implemented in `src/bast_validation.rs`
and `src/bast.rs`. See
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
- **Full file**: [OQ-008](questions/008-unionvalidator-variant-dispatch.md)
## Deferred / Blocked
+114 -77
View File
@@ -1,12 +1,12 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-15
---
# alktype — Overview
The binary struct engine: a small Rust crate that takes a JSON Schema
with `AlkType:*` custom keywords and produces an offset map, read/write
The binary struct engine: a small Rust crate that takes a BAST (Binary
Abstract Syntax Tree) document and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
@@ -16,29 +16,43 @@ Component details are in the sibling documents.
## What
`alktype` is a library crate that consumes JSON Schemas annotated
with `AlkType:*` custom keywords (the same kinds defined in TypeBox's
`typedef.ts`, plus `AlkType:Bytes`, `AlkType:Int64`, and `AlkType:Uint64`
as alktype additions) and produces three capabilities:
`alktype` is a library crate that consumes BAST documents and produces
three capabilities:
1. **An offset map** — walks the schema, computes byte offsets for each
field based on type sizes, field order, and alignment.
1. **An offset map** — walks the BAST typed tree, computes byte offsets
for each field based on type sizes, field order, and alignment.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that a
buffer's bytes match the schema's type constraints.
3. **Validation** — two validators for two input types:
- `validate_bytes(&[u8])` uses the BAST-native validator (a recursive
walker over the BAST type tree) to check the value-domain
constraints the materializer doesn't (integer ranges, `maxLength`,
enum index bounds, union variant constraints).
- `validate_json(&Value)` uses a standard `jsonschema::Validator`
compiled from a consumer-provided JSON Schema (BAST is not involved
— BAST describes bytes, not JSON shape).
The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing). The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are small (a few lines
each, generated from shared macros — see [validation.md](validation.md)).
BAST is a JSON document that describes binary data layouts using a
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
itself a valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser. See [ADR-BAST](decisions/bast-bast-format.md)
and [`bast-format.md`](bast-format.md).
The heavy lifting is done by the `jsonschema` crate (the
`validate_json` path and BAST document meta-schema validation) and
`serde_json` (BAST document parsing). The novel code is the offset
computation — a recursive walk of the BAST typed tree that computes
byte positions for each field — and the BAST-native validator — a flat
recursive match over the same tree. See
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md)
(purpose/scope) and [ADR-BAST](decisions/bast-bast-format.md) (format).
The crate replaces two prior attempts that built their own jsonschema
engines — typebox-rs (~8,400 lines) and the @alkdev/alktype prototype
(~5,600 lines) — with `jsonschema` + an offset map + small custom keyword
implementations. See
(~5,600 lines) — with a BAST parser + an offset map + a BAST-native
validator + the `jsonschema` crate for the JSON-validation path. See
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
## Why
@@ -48,21 +62,20 @@ read or write binary data at computed offsets. Instead of per-protocol
serde structs (russh-sftp's 29 packet types), per-handler wire format
code (TTY's 5-byte format parser), or per-format offset computation
(metatensor's tensor access), all of these become instances of the same
engine with different schemas.
engine with different BAST documents.
The guiding insight:
> **The schema is the format.** A JSON Schema with `AlkType:Float32`,
> `AlkType:Struct`, `AlkType:Union` etc. is both the validation spec and
> the layout spec. No separate format definition, no separate parser, no
> separate validator. One schema, three uses: validate, compute offsets,
> access data.
> **The schema is the format.** A BAST document is both the layout spec
> and the validation spec for bytes. No separate format definition, no
> separate parser, no separate validator. One schema, three uses:
> validate, compute offsets, access data.
This is the convergence of three threads identified in the
call-channels-unification research: the `typedef.ts` schema kinds from
TypeBox, the russh-sftp protocol packets, and the metatensor format. The
common pattern: a JSON Schema describes the shape of binary data, and
the binary data is the struct's bytes at computed offsets.
common pattern: a schema describes the shape of binary data, and the
binary data is the struct's bytes at computed offsets.
The crate was bumped up in the timeline when the call-channels-unification
research surfaced that channels, TTY, and the binary call protocol are
@@ -75,41 +88,45 @@ read/write the binary payload."
## The "Schema Is the Format" Principle
A JSON Schema with `AlkType:*` custom keywords serves three roles
simultaneously:
A BAST document serves three roles simultaneously:
| Role | Mechanism | When |
|------|-----------|------|
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
| **Validation spec (bytes)** | BAST-native validator (recursive walker over the BAST type tree) | Load time (parse typed tree), access time (`validate_bytes`) |
| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) |
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) |
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
No separate format definition, no separate parser, no separate validator.
The schema is the single source of truth for the binary format. Adding a
new field to a protocol is adding a property to the schema JSON — the
engine computes the new offsets automatically.
No separate format definition, no separate parser, no separate
validator. The BAST document is the single source of truth for the
binary format. Adding a new field to a protocol is adding an entry to
the BAST `fields` array — the engine computes the new offsets
automatically.
This is the same principle as `#[repr(C)]` struct field access, but at
runtime from a portable JSON Schema instead of at compile-time from
language-specific annotations. The schema is the ABI contract.
runtime from a portable JSON document instead of at compile-time from
language-specific annotations. The BAST document is the ABI contract.
## Dependencies
```
alktype
├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support
├── serde_json (with preserve_order) — schema parsing; field order is load-bearing
└── (no tokio, no platform deps) — WASM-clean by construction
├── jsonschema (v0.46, Draft 2020-12, default-features=false) — validate_json path + BAST meta-schema validation
├── serde_json (with preserve_order) — BAST document parsing; mapping iteration order is load-bearing
└── (no tokio, no platform deps) — WASM-clean by construction
```
`alktype` is dependency-light: `jsonschema` + `serde_json` only.
No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for
browser use. The `jsonschema` crate is already in the workspace at
`@alkdev/alknet: jsonschema/` — alktype is its first consumer.
browser use. The `validate_bytes` path does not touch `jsonschema` for
validation (it uses `jsonschema::ValidationError::custom` only for the
error payload type, D-BAST-009) — a small wasm binary-size win.
`serde_json` requires the `preserve_order` feature because field order
is load-bearing for binary layouts. The order of properties in the
schema JSON determines the order of fields in the binary struct.
`serde_json`'s `preserve_order` feature remains a dependency. Under
BAST, struct field order is explicit (the `fields` array), so layout
correctness no longer depends on it; but `mapping` iteration order and
the `Definitions` `$defs` block order are still load-bearing for the
parser's lazy resolution and the builder's output.
## Consumers
@@ -133,14 +150,17 @@ schema roles (binary layout + JSON payloads) from one library. See
The russh-sftp case is the most instructive and the highest-value POC
target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written
dispatch on a type byte followed by serde deserialization. Under alktype,
the dispatch is `TUnion` with a byte-offset discriminator — the schema
says "byte 0 is the discriminator, bytes 1..N are the variant struct."
The engine reads the discriminator, looks up the variant schema, computes
offsets, reads fields. Same result, no per-packet-type code.
the dispatch is a BAST `union` with a byte-offset discriminator — the
schema says "byte 0 is the discriminator, bytes 1..N are the variant
struct." The engine reads the discriminator, looks up the variant
schema, computes offsets, reads fields. Same result, no per-packet-type
code.
## Scope Boundaries (What This Is Not)
These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
These boundaries are decided in
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md)
and [ADR-BAST](decisions/bast-bast-format.md).
- **Not metatensor.** alktype is the binary struct *engine*. Metatensor
is a *format* (8-byte header + JSON header + binary data) that uses the
@@ -150,60 +170,73 @@ These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-js
should not do anything that explicitly blocks adding a Value system
later.
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
concern. The alktype engine consumes schemas; it does not generate them.
concern. The alktype engine consumes BAST documents; it does not
generate them.
- **Schema builder is in scope as of v0.1.0.** A fluent Rust API for
constructing schemas at runtime, producing `serde_json::Value`, is
shipped in v0.1.0 ([ADR-009](decisions/009-builder-api.md), resolves
OQ-003). The builder covers AlkType kinds and standard JSON Schema;
see [builder.md](builder.md). Schemas may still be authored in
TypeBox, generated by ujsx components, or hand-written — the builder
is an additional construction path, not a replacement.
constructing BAST documents and standard JSON Schemas at runtime,
producing `serde_json::Value`, is shipped in v0.1.0
([ADR-009](decisions/009-builder-api.md), resolves OQ-003). The
builder covers BAST kinds (`struct_()`) and standard JSON Schema
(`object()`); see [builder.md](builder.md). BAST documents may still
be authored in TypeBox, generated by ujsx components, or hand-written
— the builder is an additional construction path, not a replacement.
- **Not a serialization framework.** The alktype engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use alktype.
computed offsets — no intermediate `Value` tree (except for the
`validate_bytes` materialization step), no reflection, no dynamic
dispatch per field. For JSON data, use serde. For binary data with a
known schema, use alktype.
- **Not a JSON-payload validator.** BAST describes bytes, not JSON
shape. `validate_json` validates a JSON `Value` against a
consumer-provided standard JSON Schema, not against the BAST document.
See [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
## Architecture (component pointers)
- **[schema-layer.md](schema-layer.md)** — the 19 `AlkType:*` kinds,
jsonschema custom keyword integration, TypeBox interop, schema
annotations (endianness, alignment, encoding, TUnion discriminators).
- **[schema-layer.md](schema-layer.md)** — the BAST parser (the typed
surface every engine module walks), the 18 BAST kinds, the
`AlkTypeKind` enum, and the foundational annotation types.
- **[`bast-format.md`](bast-format.md)** — the normative BAST format
specification (meta-schema, TypeRef, examples, validation model).
- **[layout-engine.md](layout-engine.md)** — offset computation, the two
layout modes (packed sequential vs aligned static), alignment,
endianness, variable-length field handling.
- **[data-access.md](data-access.md)** — read/write functions, TUnion
dispatch, field paths, zero-copy access for fixed-size types,
length-prefix reading for variable-length types.
- **[validation.md](validation.md)** — custom keyword validators for all
19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time
validation, `AlkTypeEngine` as the compiled form of a schema.
`validate_json` for JSON values; `validate_bytes` for binary buffers
(ADR-010).
- **[builder.md](builder.md)** — fluent Rust API for constructing
alktype JSON Schemas at runtime, producing `serde_json::Value`.
Covers AlkType kinds and standard JSON Schema (ADR-009).
- **[validation.md](validation.md)** — the two-validator model
(BAST-native for `validate_bytes`, standard `jsonschema` for
`validate_json`), `AlkTypeError`, load-time vs access-time validation,
`AlkTypeEngine` as the compiled form of a BAST document.
- **[builder.md](builder.md)** — fluent Rust API for constructing BAST
documents and standard JSON Schemas at runtime, producing
`serde_json::Value`. Covers BAST kinds and standard JSON Schema
(ADR-009, D-BAST-008).
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries (format-specific content superseded by ADR-BAST) |
| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | BAST as the schema format; meta-schema, `$defs`/`$ref`/`kind` vocabulary; supersedes ADR-001's format-specific content |
| Two layout modes | [ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) |
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) — semantics carry forward; location moved to BAST type-level properties |
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping (validation-strategy section refined by ADR-VAL-SPLIT) |
| Two-validator model | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | BAST-native validator for `validate_bytes`; standard `jsonschema::Validator` for `validate_json`; D-BAST-006/007/009 |
| Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (SFTP offsets, metatensor data_offsets) |
| Non-final inline variable fields | [ADR-006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields) |
| Packed-mode read factory | [ADR-007](decisions/007-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader |
| TUnion in aligned mode | [ADR-008](decisions/008-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) |
| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate |
| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 (output format amended to BAST / standard JSON Schema by ADR-BAST) |
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate (validation step amended to the BAST-native validator by ADR-VAL-SPLIT) |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-001** (deferred(scope)): Arrays of variable-length-element structs.
- **OQ-001** (deferred(scope)): Arrays of variable-length-element
structs — BAST arrays require `count` in v1 (D-BAST-004), aligning
with this deferral.
- **OQ-002** (deferred(scope)): `no_std` + `alloc` support.
- **OQ-003** (resolved by [ADR-009](decisions/009-builder-api.md)):
Builder API for schema construction. Shipped in v0.1.0; see
@@ -223,8 +256,12 @@ See [open-questions.md](open-questions.md) for full details.
- `@alkdev/alknet: alknet-typedef-poc/` — the POC code (disposable)
- `@alkdev/alknet: typebox-rs/` — prior attempt, replaced by alktype
- `@alkdev/alknet: alktype-prototype/` — prior attempt (the @alkdev/alktype prototype; not to be confused with this crate, which reuses the name but is backed by the `jsonschema` crate)
- [BAST pivot research record](../research/bast-pivot.md) — motivation,
POC scope and result, decisions D-BAST-001..009, risks
- [BAST pivot implementation plan](../plans/bast-implementation.md) —
ordered implementation steps, semver contract, ADR-sync checklist
> **Note**: The research findings, POC code, and prior-attempt paths above
> refer to the parent `@alkdev/alknet` workspace where this crate originated.
> They are preserved here as historical context for the architectural
> decisions; the artifacts themselves are not part of this standalone repo.
> decisions; the artifacts themselves are not part of this standalone repo.
@@ -1,5 +1,19 @@
# OQ-008: `UnionValidator` variant dispatch — validate variant fields against variant schema
> **Note (post-BAST-pivot):** The v0.1.0 resolution below —
> `UnionValidator` + `inline_union_variant_refs` + custom-keyword
> `jsonschema` sub-validators — was superseded by the BAST pivot. The
> current implementation is the BAST-native validator
> (`bast_validation::validate_union`), which reads `__discriminator`,
> looks up the variant `BastType` in the union's `mapping`, and
> recurses via `validate_typeref` into the variant's BAST definition.
> Variant `$ref`s resolve lazily via `BastDoc::resolve_typeref` — no
> `inline_union_variant_refs` compile step (removed under BAST). No
> custom keywords, no `jsonschema` involvement on the bytes path. See
> [ADR-VAL-SPLIT](../decisions/val-split-two-validator-model.md). The
> v0.1.0 resolution text is preserved below as the historical record
> of how the question was originally resolved.
- **Origin**: Raised during the v0.1.0 POC round 2 (SFTP Packet
`validate_bytes` POC). Surfaced when the over-`maxLength` `Bytes`
test failed: the materializer read the bytes correctly, but the
+211 -425
View File
@@ -1,58 +1,60 @@
---
status: draft
last_updated: 2026-07-22
status: accepted
last_updated: 2026-08-15
---
# alktype — Schema Layer
The schema layer: the 19 `AlkType:*` custom type kinds, their mapping to
Rust types and byte sizes, the `jsonschema` custom keyword integration,
TypeBox interop, and the concrete JSON shapes for schema-level
annotations.
The schema layer: the BAST (Binary Abstract Syntax Tree) format and the
typed parser that the layout engines, materializer, and BAST-native
validator walk. BAST replaces the v0.1.0 `AlkType:*` custom-keyword JSON
Schema format decided in ADR-001; the pivot is recorded in
[ADR-BAST](decisions/bast-bast-format.md) and grounded in
[D-BAST-001..009](../research/bast-pivot.md#decisions).
## The 19 AlkType Kinds
The **normative format specification** is
[`bast-format.md`](bast-format.md) (meta-schema, TypeRef, examples,
validation model). This document describes the *implementation* — the
typed parser in `src/bast.rs` and the foundational `AlkTypeKind` enum
in `src/schema.rs` — and points at the format spec for shape details.
These are the custom schema kinds defined in TypeBox's `typedef.ts`
(`@alkdev/alknet: typebox/example/typedef/typedef.ts`, 619 lines) and
ported to Rust via `jsonschema` custom keywords. Each kind carries binary
layout semantics — a known byte size (for fixed-size types) or a known
encoding strategy (for variable-length types).
## The 18 BAST Kinds
| Kind | TypeBox key | Rust type | Size | Category |
|------|-------------|-----------|------|----------|
| `TFloat32` | `AlkType:Float32` | `f32` | 4 | fixed |
| `TFloat64` | `AlkType:Float64` | `f64` | 8 | fixed |
| `TInt8` | `AlkType:Int8` | `i8` | 1 | fixed |
| `TInt16` | `AlkType:Int16` | `i16` | 2 | fixed |
| `TInt32` | `AlkType:Int32` | `i32` | 4 | fixed |
| `TInt64` | `AlkType:Int64` | `i64` | 8 | fixed |
| `TUint8` | `AlkType:Uint8` | `u8` | 1 | fixed |
| `TUint16` | `AlkType:Uint16` | `u16` | 2 | fixed |
| `TUint32` | `AlkType:Uint32` | `u32` | 4 | fixed |
| `TUint64` | `AlkType:Uint64` | `u64` | 8 | fixed |
| `TBoolean` | `AlkType:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed |
| `TString` | `AlkType:String` | length-prefixed UTF-8 | variable | variable |
| `TBytes` | `AlkType:Bytes` | length-prefixed raw bytes | variable | variable |
| `TStruct` | `AlkType:Struct` | record of fields | sum of field sizes | composite |
| `TUnion` | `AlkType:Union` | tagged union | discriminator + variant | composite |
| `TArray` | `AlkType:Array` | repeated element | count × element size | composite |
| `TEnum` | `AlkType:Enum` | u32 index into enum values | 4 (fixed) | fixed |
| `TRecord` | `AlkType:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
| `TTimestamp` | `AlkType:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
BAST uses lowercase `kind` strings (`"uint32"`, `"struct"`, `"union"`,
etc.). The engine represents them as the `AlkTypeKind` Rust enum — one
variant per kind — providing compile-time exhaustiveness checking and
integer-discriminant dispatch (a jump table) instead of string
comparison at every field access.
`AlkType:Int64` and `AlkType:Uint64` are alktype additions —
TypeBox's `typedef.ts` tops out at 32-bit integers. They are required by
the primary POC targets: SFTP `Read`/`Write` packets have `offset: u64`,
and metatensor `data_offsets` are `u64`. See
| BAST kind | `AlkTypeKind` | Rust type | Size | Category |
|-----------|---------------|-----------|------|----------|
| `int8` | `Int8` | `i8` | 1 | fixed |
| `int16` | `Int16` | `i16` | 2 | fixed |
| `int32` | `Int32` | `i32` | 4 | fixed |
| `int64` | `Int64` | `i64` | 8 | fixed |
| `uint8` | `Uint8` | `u8` | 1 | fixed |
| `uint16` | `Uint16` | `u16` | 2 | fixed |
| `uint32` | `Uint32` | `u32` | 4 | fixed |
| `uint64` | `Uint64` | `u64` | 8 | fixed |
| `float32` | `Float32` | `f32` | 4 | fixed |
| `float64` | `Float64` | `f64` | 8 | fixed |
| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed |
| `string` | `String` | length-prefixed UTF-8 | variable | variable |
| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable |
| `struct` | `Struct` | record of fields | sum of field sizes | composite |
| `union` | `Union` | tagged union | discriminator + variant | composite |
| `array` | `Array` | repeated element | count × element size | composite |
| `enum` | `Enum` | u32 index into enum values | 4 (fixed) | fixed |
| `record` | `Record` | count-prefixed (key, value) pairs | variable | variable |
`int64`/`uint64` are alktype additions — TypeBox's `typedef.ts` tops
out at 32-bit integers. Required by SFTP `Read`/`Write` `offset: u64`
and metatensor `data_offsets`. See
[ADR-005](decisions/005-int64-uint64-first-class-kinds.md).
### The `AlkTypeKind` enum
The engine represents the 19 kinds as a Rust enum — `AlkTypeKind` — with
one variant per kind (`AlkTypeKind::Float32`, `AlkTypeKind::Struct`, etc.).
The enum provides compile-time exhaustiveness checking and integer
discriminant dispatch (a jump table) instead of string comparison at
every field access. It is `pub` and re-exported from the crate root.
`src/schema.rs` defines the enum:
```rust
pub enum AlkTypeKind {
@@ -60,7 +62,7 @@ pub enum AlkTypeKind {
Uint8, Uint16, Uint32, Uint64,
Float32, Float64,
Boolean, Enum,
String, Bytes, Timestamp,
String, Bytes,
Struct, Union, Array, Record,
}
```
@@ -69,439 +71,223 @@ The enum carries the kind's binary-layout metadata as inherent methods:
| Method | Returns | Notes |
|--------|---------|-------|
| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"AlkType:Uint8"` |
| `to_bast_str(self)` | `&'static str` | The lowercase BAST kind string (`"uint32"`) |
| `from_bast_str(s)` | `Result<AlkTypeKind, AlkTypeError>` | Parses a lowercase BAST kind string; `AlkTypeError::Schema` for unknowns |
| `type_size(self)` | `Option<usize>` | `Some(N)` for fixed-size kinds; `None` for variable/composite |
| `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array |
| `is_fixed_size(self)` | `bool` | True for the 12 fixed-size primitive kinds |
| `is_composite(self)` | `bool` | True for Struct, Union, Array, Record |
| `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record |
| `is_variable_length(self)` | `bool` | True for String, Bytes, Record |
| `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter |
`AlkTypeKind` implements `Display` (renders the keyword string) and
`FromStr` (parses the keyword string back into the variant, returning
`AlkTypeError::Schema` for unknown kinds). The layout engines and the
validator dispatch on the enum, not on strings.
`AlkTypeKind` implements `Display`, backed by `to_bast_str` so the
layout engines, materializer, validator, and parser surface the
canonical BAST name in error messages. `from_bast_str` is the inverse
and is the dispatch point the BAST parser uses to map a `kind` string
to the enum variant (D-BAST-002).
### Fixed-size types
### Foundational annotation types
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset
computation uses these sizes directly. Read/write is zero-copy pointer
cast for these types.
**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other
values are invalid and produce a `AlkTypeError::Access` on read.
**`TEnum` binary representation:** A `u32` index into the enum's declared
values, in declaration order. The first declared value is index 0, the
second is index 1, etc. The enum's values are declared via the standard
JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`).
The `AlkType:Enum` custom keyword signals that the type is an enum for
layout purposes; the built-in `enum` keyword provides the value list.
**Design note:** TypeBox's `TEnum` is a string enum (variable-length). The
alktype engine uses a `u32` index instead — a deliberate deviation from
TypeBox fidelity in favor of binary efficiency. Most enums have a small
number of variants (e.g., the call protocol's 5 event types); a `u32`
index is compact, fixed-size, and sufficient for any realistic enum. The
JSON representation (for validation) remains a string; the binary
representation is the `u32` index.
The `u32` index follows the schema's endianness annotation (ADR-003), like
all other fixed-size types. In little-endian mode the index is
`u32::from_le_bytes`; in big-endian mode it is `u32::from_be_bytes`.
### Variable-length types
`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes.
The alktype engine supports three strategies for handling variable-length
types in binary layouts, selected by the `encoding` annotation and the
standard JSON Schema `maxLength` keyword:
| Strategy | Encoding annotation | Layout behavior | Use case |
|----------|-------------------|-----------------|----------|
| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) |
| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) |
| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) |
**Strategy 1: Inline length-prefixing (default).** The field's fixed
portion is a 4-byte length prefix at a computed offset. The variable data
follows immediately after. In packed sequential mode, the length prefix
determines the position of subsequent fields. In aligned static mode, the
length prefix is at a known offset; the variable data is not included in
the static layout. This is the universal pattern used by channels, SFTP,
TTY, and most binary protocols.
**Strategy 2: Fixed-size reservation.** When a variable-length field
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
than `maxLength` is zero-padded; data longer than `maxLength` is a
validation error. This makes the field fixed-size from the layout
perspective — subsequent fields have known, unchanging offsets. This is
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
pattern for fields with known maximum sizes.
In packed sequential mode, `maxLength` is a validation constraint only —
the engine still uses inline length-prefixing (strategy 1) because
protocols don't benefit from fixed-size reservation.
**Strategy 3: Offset indirection.** The field is a struct
`{offset: u32, length: u32}` at a known position. The consumer provides
the data region separately; the engine reads the offset and length, then
slices the data region. This is the metatensor blob tensor pattern — the
index struct lives in one region, the blob data lives in another. Enables
mmap-friendly random access to variable-length data without parsing
length prefixes and without reserving worst-case space.
**Default strategy selection:**
- In packed sequential mode: always strategy 1 (inline length-prefixing).
`maxLength` is a validation constraint only.
- In aligned static mode: strategy 2 (fixed-size reservation) if
`maxLength` is declared; strategy 3 (offset indirection) if
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
length-prefixing) otherwise.
**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3)
respects the schema's `"endian"` annotation (ADR-003). In little-endian
mode, the length is `u32::from_le_bytes`. In big-endian mode, the length
is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have
consistent byte order for both field values and length prefixes.
**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`.
Otherwise identical to `TString` in layout (same three strategies).
**Design note:** `AlkType:Bytes` is an alktype addition — it does
not exist in TypeBox's `typedef.ts` (which defines 16 kinds). It is
included because raw byte arrays are a common binary protocol primitive
(SFTP data payloads, channels payloads, tensor data) and are semantically
distinct from UTF-8 strings. In the binary representation, TBytes is raw
bytes with no encoding (not base64, not hex). In the JSON representation
(for validation), TBytes is a string (JSON has no native byte type).
**`TRecord`:** A string-keyed map. The value type is declared via the
schema's `"values"` property (e.g., `"values": { "AlkType:Float32": true }`).
Binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
The count is the number of entries. Each key is a length-prefixed UTF-8
string. Each value is encoded according to its declared `AlkType:*` kind
— a `Record<Uint32>` value is 4 raw bytes; a `Record<String>` value is
itself a length-prefixed string; a `Record<Struct>` value is the struct's
fields laid out inline. There is **no separate `value_len` prefix** —
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix). The
count and key-length prefixes respect the schema's endianness. In
aligned static mode with `maxLength`, the entire record is reserved at
`maxLength` bytes (zero-padded).
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
fixed-size reservation (strategy 2 with `maxLength`). The data-access
layer treats timestamps as opaque length-prefixed strings — it does not
parse or validate the timestamp format. The jsonschema custom keyword
validator checks RFC 3339 conformance at the JSON level (see
[validation.md](validation.md)).
`TArray` is variable-length when the element type is variable-length or
when the count is not known at schema time. For fixed-size element arrays
with a known count, the size is `element_size × count`.
**`TArray` count declaration:** The array count is declared via the
standard JSON Schema `"minItems"` and `"maxItems"` keywords. When
`minItems == maxItems`, the array has a fixed count known at schema time.
When they differ or are absent, the count is variable and the array uses
a length-prefixed encoding: `[count: u32][element_0]...[element_N]`.
The count prefix respects the schema's endianness.
### Composite types
`TStruct` and `TUnion` are composite — their size is the sum of their
fields' sizes (plus alignment padding in aligned static mode). The offset
computation recurses into their properties.
## Schema-Layer Public API
The `schema` module exposes the foundational types and functions every
other module depends on. These are re-exported from the crate root.
### `get_alktype_kind` vs `get_alktype_kind_loose`
The engine recognizes a `AlkType:*` kind on a schema node two ways,
because the keyword value may be either a boolean (`true`) or an
annotation object (`{ "encoding": "..." }`):
| Function | Recognizes | Returns |
|----------|------------|---------|
| `get_alktype_kind(node) -> Option<&str>` | Boolean form only (`{ "AlkType:String": true }`) | The keyword string, e.g. `"AlkType:String"` |
| `get_alktype_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string |
| `get_alktype_kind_enum(node) -> Option<AlkTypeKind>` | Boolean form only | The parsed enum variant |
| `get_alktype_kind_loose_enum(node) -> Option<AlkTypeKind>` | Boolean form **and** object form | The parsed enum variant |
The boolean-form-only functions are used by the validator factories
(which reject the object form as a schema error) and the top-level
kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new`
(which require `AlkType:Struct` at the root). The "loose" variants are
used by the layout engines during field traversal, so that a variable-
length field with an `encoding` annotation (`{ "AlkType:String":
{ "encoding": "offset-indirect" } }`) is still recognized as a `String`.
### Annotation parsers
Each schema-level annotation has a dedicated parser that reads it from a
`serde_json::Value` node and returns a sensible default when absent:
| Function | Annotation | Default |
|----------|------------|---------|
| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` |
| `parse_align(node) -> Option<usize>` | `"align"` | `None` |
| `parse_max_length(node) -> Option<usize>` | `"maxLength"` | `None` |
| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` |
| `parse_discriminator(node) -> Result<DiscriminatorKind, AlkTypeError>` | `"discriminator"` | (required — returns `AlkTypeError::Schema` if absent) |
### Public enums
`src/schema.rs` also defines the two annotation enums (semantics
unchanged from ADR-003; only their *location* in the document moved —
see [ADR-BAST](decisions/bast-bast-format.md) and
[bast-format.md §Variable-Length Encoding](bast-format.md#variable-length-encoding)):
```rust
pub enum Endian { Little, Big }
pub enum VariableEncoding { LengthPrefixed, OffsetIndirect }
pub enum DiscriminatorKind {
Byte { offset: usize, disc_type: AlkTypeKind },
Field { name: String },
}
```
`DiscriminatorKind::Byte` carries the byte position (`offset`) and the
discriminator's `AlkType:*` kind (`disc_type`, restricted to `Uint8`/
`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator
field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for
how these drive dispatch.
The `Discriminator` builder enum lives in
[`src/builder.rs`](../../src/builder.rs) (the builder's domain); the BAST
parser's typed discriminator view is
[`BastDiscriminator`](#the-bast-parser-bast-module).
### `$ref` resolution and normalization
## The BAST Parser (`bast` module)
| Function | Purpose |
|----------|---------|
| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/<name>"`. Idempotent. Runs once at `AlkTypeEngine::compile` time. |
| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. |
| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). |
`src/bast.rs` is the typed surface over a BAST document. Three
consumers walk the same tree — the layout engines
([`offset_map`](layout-engine.md), [`layout_builder`](layout-engine.md),
[`sequential_reader`](layout-engine.md)), the
[`materialize`](data-access.md) layer, and the
[`bast_validation`](validation.md) validator — so a typed view pays for
itself: each walks matched arms over `BastType` instead of re-parsing
raw JSON at every node. Borrowing (not cloning) the source
`serde_json::Value` keeps the parse allocation-free beyond the small
typed nodes themselves.
`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s
JSON Pointer requirement. The layout engines call `resolve_ref_or_inline`
on every `$ref`-bearing node they encounter during traversal.
### Document shape
## jsonschema Custom Keyword Integration
Every BAST document has the same top-level shape:
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
via the `with_keyword` API. Each `AlkType:*` kind is registered as a
custom keyword:
```rust
let validator = jsonschema::options()
.with_keyword("AlkType:Float32", factory)
.with_keyword("AlkType:Int32", factory)
.with_keyword("AlkType:Struct", factory)
// ... all 19 kinds
.build(&schema)?;
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path — enabling cross-keyword awareness. The
`AlkType:Struct` validator, for example, inspects the parent's
`properties` to validate each field against its declared `AlkType:*` kind.
- The `$defs` block is **required** (D-BAST-003). Single-type documents
are a special case with one entry.
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001).
It selects which `$defs` entry is the top-level type; convention
(first entry) is fragile and depends on JSON key order, so an
explicit parameter is used instead.
Each custom keyword implementation is ~10 lines. The `jsonschema` crate
handles all structural validation (object properties, required fields,
array items, enum values) — the custom keywords only need to validate
the leaf type constraints. See [validation.md](validation.md) for the
validator implementations.
See [`bast-format.md`](bast-format.md) for the normative TypeDef shapes
(Struct, Union, Enum, FieldDef, TypeRef) and the meta-schema.
This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side.
Same semantics, different language, same JSON Schema wire format. A
TypeBox schema serialized to JSON feeds into the alktype engine after a
single pre-processing step: normalizing `$ref` values (see below).
### Typed tree
## TypeBox Interop
The parser produces a borrowed typed tree:
TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox
schema like:
| Type | Role |
|------|------|
| `BastDoc<'a>` | The parsed document: the root `Value`, the chosen root name, and the parsed root `BastDef`. Entry point via `BastDoc::new(root, root_name)`. |
| `BastDef<'a>` | A named `$defs` entry — `{ name, kind: BastDefKind, source }`. Only `struct`/`union`/`enum` can live at the top level. |
| `BastDefKind<'a>` | `Struct(BastStruct)` / `Union(BastUnion)` / `Enum(BastEnum)`. |
| `BastStruct<'a>` | `{ endian, align, fields: Vec<BastField>, source }`. Field order is the `fields` array order (BAST design principle #4 — no reliance on `serde_json`'s `preserve_order`). |
| `BastField<'a>` | `{ name, ty: BastType, endian, align, encoding, max_length, source }`. Annotations are field-level properties (ADR-003 semantics, BAST location). |
| `BastUnion<'a>` | `{ endian, discriminator, fields, mapping: Vec<(key, BastType)>, source }`. Variant `$ref`s resolve **lazily** — no compile-time inlining. |
| `BastDiscriminator<'a>` | `Byte { offset, disc_type }` / `Field { name }`. The typed view of the `discriminator` object. |
| `BastEnum<'a>` | `{ values: Vec<&'a str>, source }`. Non-empty (enforced). |
| `BastType<'a>` | A TypeRef — `Primitive(AlkTypeKind)` / `Ref(BastRef)` / `Array(BastArray)` / `Record(BastRecord)` / `Struct(...)` / `Union(...)` / `Enum(...)`. The central mechanism for typing fields, array elements, record values, and union variants. |
| `BastRef<'a>` | A `$ref` restricted to `#/$defs/<name>`. Carries just the name. |
| `BastArray<'a>` | `{ element: Box<BastType>, count, source }`. `count` is required in v1 (D-BAST-004). |
| `BastRecord<'a>` | `{ values: Box<BastType>, source }`. |
```typescript
const TensorRef = Type.Object({
dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]),
shape: Type.Array(Type.Number()),
data_offsets: Type.Tuple([Type.Number(), Type.Number()])
});
```
All of these are re-exported from the crate root (`pub use bast::{...}`
in `src/lib.rs`).
serialized to JSON is a standard JSON Schema with `type: "object"`,
`properties`, and `required`. That JSON feeds into the alktype engine
after `$ref` normalization. The `AlkType:*` custom keywords are added by
TypeBox's `TypeRegistry.Set` — they appear in the serialized JSON as
additional properties on the schema object.
### `$ref` resolution
### `$ref` normalization
BAST `$ref`s are always full JSON Pointers restricted to
`#/$defs/<name>` — no external references, no fragment-only pointers,
no bare names (rejected by the parser). The restriction keeps
resolution a single hash lookup and eliminates the v0.1.0
`normalize_refs` pass that rewrote TypeBox's bare-name refs.
TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`),
referencing sibling definitions within the same `$defs` block. The
`jsonschema` crate requires full JSON Pointer paths (e.g.,
`"$ref": "#/$defs/Read"`). The alktype engine normalizes TypeBox-style
refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization)
— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to
`"#/$defs/<name>"`. The normalization is idempotent — full JSON Pointer
refs pass through unchanged. It runs once at `AlkTypeEngine::compile`
time, before the schema is passed to `jsonschema` or the offset
computation.
`BastDoc` exposes three resolution helpers:
**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs
with `Resource 'Read' is not present in a registry`. Full JSON Pointer
refs (`#/$defs/Read`) resolve correctly. The normalization step bridges
the gap between TypeBox's output and jsonschema's input.
| Method | Purpose |
|--------|---------|
| `lookup_def(name) -> Result<&'a Value, AlkTypeError>` | The single hash lookup into `$defs`. |
| `resolve_ref(r: &BastRef) -> Result<BastDef, AlkTypeError>` | Resolve a `BastRef` to its `BastDef`. |
| `resolve_typeref(ty: &BastType) -> Result<BastType, AlkTypeError>` | Deref one `$ref` level, or return the inline type unchanged. The composite-walkers call this. |
| `resolve_typeref_as_def(ty, path) -> Result<BastDef, AlkTypeError>` | Resolve a `BastType` to a `BastDef`, wrapping inline composites in a synthetic def. Convenient for the validator/materializer. |
The alktype engine does not depend on TypeBox or any JS toolchain. It
consumes JSON — whether that JSON was authored in TypeBox, generated by
a ujsx component, or hand-written. The schema is the interface.
Variant `$ref`s (union `mapping` entries) are resolved **lazily** by
the materializer and validator via these helpers — no
`inline_union_variant_refs` compile step (removed under BAST). The
parser only records the `BastRef` target name.
### Untrusted input
Every path that walks a BAST document returns
`Err(AlkTypeError::Schema)` on a malformed document, never
`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream
`alkcall` consumer accepts schemas from arbitrary internet peers in its
hub/spoke topology). Overflow-safe arithmetic (`checked_add`,
`usize::try_from`) is used for any offset/count cast (AGENTS.md §4).
A malformed document (missing `$defs`, missing `kind`, unknown kind
string, dangling `$ref`, empty `mapping`, non-struct/union/enum at the
top level, a field-name union without a `fields` array, etc.) surfaces
as `AlkTypeError::Schema` with a path-annotated message.
### What the parser does *not* do
- **No meta-schema validation.** `BastDoc::new` parses structurally
(every reachable def parses to a typed `BastDef`) but does not run
the BAST meta-schema. Consumers that want full structural validation
can run the meta-schema via `jsonschema` directly
([`BAST_META_SCHEMA`](bast-format.md#the-meta-schema) is re-exported
from the crate root). The parser's structural checks catch the cases
that matter for layout/materialize/validate; the meta-schema is the
authoritative well-formedness check.
- **No eager full-document parse.** Only the root definition and the
definitions it (transitively) references are parsed eagerly; orphan
`$defs` entries are not checked. Lazy `$ref` resolution reaches the
rest at access time.
- **No annotation interpretation.** The parser *records* `endian`/
`align`/`encoding`/`maxLength` on `BastField`/`BastStruct`; the
layout engines and validator *interpret* them (ADR-003 semantics).
## Schema Annotations
Schema-level annotations control binary layout behavior. These are
decided in [ADR-003](decisions/003-schema-annotations.md).
Annotation *semantics* carry forward unchanged from ADR-003; only their
*location* moved from v0.1.0's custom-keyword objects to BAST
type-level properties. The concrete BAST shapes are in
[`bast-format.md`](bast-format.md):
### Endianness
- [Endianness](bast-format.md#endianness) — struct/union-level `endian`
with field-level override.
- [Alignment](bast-format.md#alignment) — struct/field-level `align`
(aligned mode only).
- [Variable-length encoding](bast-format.md#variable-length-encoding) —
field-level `encoding` and `maxLength`.
- [Union discriminators](bast-format.md#union) — `discriminator` object
on the union def (`byte` or `field`).
Schema-level annotation with a default of little-endian:
The `maxLength` keyword is *not* a BAST invention — it is the standard
JSON Schema `maxLength`, repurposed as a byte-length cap. In aligned
mode it reserves a fixed-size slot; in packed mode it is a validation
constraint only. See [bast-format.md §Variable-Length
Encoding](bast-format.md#variable-length-encoding) and
[ADR-003](decisions/003-schema-annotations.md).
```json
{ "AlkType:Struct": true, "endian": "big", "properties": { ... } }
```
## What Was Removed
- `"endian": "little"` (default) — read/write in little-endian byte order.
- `"endian": "big"` — read/write in big-endian byte order.
- Applies to the entire schema and all nested types.
The v0.1.0 custom-keyword accessor layer was removed in step 8 of the
BAST pivot. The `schema` module retains only the foundational types
(`AlkTypeKind`, `Endian`, `VariableEncoding`, shared constants); the
BAST parser is the typed surface every engine module walks. Removed:
### Alignment
- `get_alktype_kind` / `get_alktype_kind_enum` / `get_alktype_kind_loose`
/ `get_alktype_kind_loose_enum` — superseded by the parser's
`kind`-string dispatch.
- `normalize_refs` / `inline_union_variant_refs` (+ recursive helpers)
— BAST refs are always `#/$defs/<name>`; one hash lookup, variant
refs resolve lazily.
- `parse_encoding` / `parse_align` / `parse_max_length` / `parse_endian`
— `bast.rs` has its own BAST-property-form copies (internal to the
parser).
- `parse_discriminator` + `DiscriminatorKind` — replaced by
`bast::BastDiscriminator`; the builder has its own `Discriminator`
enum.
- `resolve_ref` / `resolve_ref_or_inline` — replaced by
`BastDoc::lookup_def` / `resolve_typeref`.
- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`,
`BYTE_DISCRIMINATOR_TYPES`, and their unit tests.
Both struct-level and field-level, with field-level overriding:
```json
{
"AlkType:Struct": true,
"align": 256,
"properties": {
"weight": { "AlkType:Float32": true, "align": 16 }
}
}
```
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/
enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix),
1 for struct/union/array.
- Only meaningful in aligned static mode (ADR-002). Ignored in packed
sequential mode.
### Variable-length encoding
The alktype engine supports three strategies for variable-length types
(see §Variable-length types above for full details). The strategy is
selected by the `encoding` annotation and the standard JSON Schema
`maxLength` keyword:
```json
// Strategy 1: Inline length-prefixing (default, shorthand)
{ "AlkType:String": true }
// Strategy 1: Explicit inline length-prefixing
{ "AlkType:String": { "encoding": "length-prefixed" } }
// Strategy 2: Fixed-size reservation (uses standard maxLength)
{ "AlkType:String": true, "maxLength": 256 }
// Strategy 3: Offset indirection (opt-in)
{ "AlkType:String": { "encoding": "offset-indirect" } }
```
- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at
computed offset, variable data follows immediately. Used by protocol
wire formats.
- `maxLength` (standard JSON Schema keyword) — in aligned static mode,
reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the
field fixed-size from the layout perspective. In packed sequential
mode, `maxLength` is a validation constraint only.
- `"encoding": "offset-indirect"` — field is a struct
`{offset: u32, length: u32}` pointing into a separate data region.
The consumer provides the data region separately. Used by metatensor
blob tensors.
- Applies to all variable-length types: `AlkType:String`, `AlkType:Bytes`,
`AlkType:Array`, `AlkType:Record`, `AlkType:Timestamp`.
### TUnion discriminators
Two discriminator kinds: byte-offset (protocol dispatch) and field-name
(typedef.ts pattern).
**Byte-offset discriminator** (SFTP type bytes, call protocol event types):
```json
{
"AlkType:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "AlkType:Uint8"
},
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" },
"101": { "$ref": "#/$defs/Status" }
}
}
```
- `"offset"` — byte position of the discriminator.
- `"type"` — the `AlkType:*` kind of the discriminator (typically
`AlkType:Uint8`).
- Mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`.
**Field-name discriminator** (typedef.ts pattern):
```json
{
"AlkType:Union": true,
"discriminator": {
"kind": "field",
"name": "type"
},
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
}
```
- `"name"` — the field name holding the discriminator value.
- Mapping keys are string values matching the discriminator field's value.
- The discriminator field is just another field in the struct.
Mapping values may be either inline schemas or `$ref` pointers. Both work.
See [ADR-BAST](decisions/bast-bast-format.md) §"What is removed" and
[`bast-format.md` §What is removed](bast-format.md#what-is-removed).
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators |
| BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary | [ADR-BAST](decisions/bast-bast-format.md) | Supersedes ADR-001's format-specific content; records D-BAST-001..009 |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Annotation semantics (carry forward unchanged; only location moves) |
| Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (required by SFTP offsets and metatensor data_offsets) |
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle |
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle (format-specific content superseded by ADR-BAST) |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-003** (deferred(scope)): Builder API for schema construction.
- **OQ-001** (deferred(scope)): Arrays of variable-length-element
structs — BAST arrays require `count` in v1 (D-BAST-004), aligning
with this deferral.
## References
- `@alkdev/alknet: typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `@alkdev/alknet: jsonschema/` — the jsonschema crate (v0.46.5, Draft
2020-12)
- [ADR-003](decisions/003-schema-annotations.md) — schema
annotation shapes
- [validation.md](validation.md) — custom keyword validator implementations
- [`bast-format.md`](bast-format.md) — the normative BAST format
specification (meta-schema, TypeRef, examples, validation model)
- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format decision
- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics
- [BAST pivot research record](../research/bast-pivot.md) — motivation,
POC scope and result, decisions D-BAST-001..009
- [validation.md](validation.md) — the BAST-native validator and the
`validate_json` JSON-Schema path
- [`src/bast.rs`](../../src/bast.rs) — the parser implementation
- [`src/schema.rs`](../../src/schema.rs) — the `AlkTypeKind` enum and
foundational annotation types
+270 -287
View File
@@ -1,63 +1,146 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-15
---
# alktype — Validation
The validation layer: custom keyword validators for all 19 `AlkType:*`
kinds, the `AlkTypeError` enum, load-time vs access-time validation
strategy, and the `AlkTypeEngine` as the compiled form of a schema.
The validation layer: two validators for two input types, the
`AlkTypeError` enum, the load-time-build / access-time-check strategy,
and the `AlkTypeEngine` as the compiled form of a BAST document.
## The Validator Split
BAST separates two concerns that the v0.1.0 format conflated, and in
doing so reveals that the engine has **two distinct validation paths**
with different inputs and guarantees. This is the validator split,
decided in [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model),
[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model),
and
[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape),
and recorded in [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator (`bast_validation`) | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
### `validate_bytes` — bytes in, BAST is the validator
The materializer produces a `serde_json::Value` tree from bytes
(walking the layout engine). By construction, this `Value` is
*structurally correct*: all declared fields are present (the
materializer iterates the field list), types are correct (`read_u32`
produces `Value::Number`), bounds are checked (via
`data_access::check_bounds`), UTF-8 is valid (via `from_utf8`), the
discriminator is in the mapping, and the boolean byte is 0 or 1.
What the materializer does NOT check — and what the BAST-native
validator checks afterward — are **value-domain constraints expressed
in the BAST document**. The BAST-native validator
(`src/bast_validation.rs`) is a recursive walker over the BAST typed
tree ([`crate::bast::BastDoc`]/[`BastType`]) that checks exactly these:
| Constraint | Validator arm |
|------------|---------------|
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` with `as_i64`/`as_u64` + range check |
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `validate_float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `check_string` reads the field-level `maxLength` |
| Bytes `maxLength` (array length) | `check_bytes` accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `validate_enum` checks `idx < values.len()` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `validate_union` reads `__discriminator`, looks up the variant, recurses via `validate_typeref` |
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
| Record values | `validate_record` recurses into each value's `values` type |
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
No external JSON Schema is required for `validate_bytes`. The BAST
document is the complete specification of the binary format — it
describes both the layout (how to read) and the constraints (what
values are valid). This is the "schema is the format" principle from
ADR-001, now fully realized.
An optional external JSON Schema can be layered on top for constraints
BAST doesn't express (cross-field consistency, regex patterns on string
content). This is additive, not load-bearing.
### `validate_json` — JSON in, JSON Schema is the validator
The consumer provides a JSON `Value` (e.g., an incoming JSON-RPC
request). The BAST document is irrelevant — BAST describes bytes, not
JSON shape. The right validator for a JSON value is a standard
`jsonschema::Validator` built from a standard JSON Schema document the
consumer provides at `AlkTypeEngine::compile` time. This is the path
alkcall uses for its `OperationSpec` JSON validation. No custom
keywords; BAST is not involved.
If no JSON Schema was supplied to `compile`, the JSON-validation
methods return `AlkTypeError::Schema` (`validate_json`) or `false`
(`is_valid_json`).
### `AlkTypeError::Validation` payload shape (D-BAST-009)
The `validate_bytes` path no longer uses `jsonschema`, so its error
payload is constructed via `jsonschema::ValidationError::custom` purely
to keep the `Validation` variant's type unchanged. The rationale is
consumer ergonomics on the *combined* path: consumers like alkcall use
both `validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary
channels) and handle `AlkTypeError::Validation` in one place. A single
uniform payload type means one match arm covers both sources.
The alternative (`Validation(String)`) would force `validate_json` to
flatten its structured errors (instance path, schema path, keyword) to
a `String` via `Display` — the more information-rich path loses data to
accommodate the less rich one. That is the wrong direction.
### What is removed
Under the BAST pivot, the v0.1.0 validation machinery is removed from
the `validate_bytes` path:
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
factories) — replaced by the BAST-native validator (~250 lines, a
flat match with no factories, no trait objects, no sub-validator
pre-computation).
- `inline_union_variant_refs()` — union variant refs are resolved lazily
by the validator and materializer.
- The custom-keyword `build_validator` path — `build_validator` is
repurposed to build a *standard* `jsonschema::Validator` from a
consumer-provided JSON Schema (no custom keywords). See
[`build_validator`](#build_validator).
The `jsonschema` crate **remains a direct dependency** for
`validate_json` and for validating BAST documents against the BAST
meta-schema. The only thing removed is the custom keyword integration
path. The `validate_bytes` path no longer touches `jsonschema` — a
small wasm binary-size win in addition to the architecture
simplification.
## Validation Strategy
Validation is delegated to the `jsonschema` crate (v0.46.5, Draft
2020-12). The alktype engine does not implement its own validation —
it registers custom keyword validators for each `AlkType:*` kind and
lets `jsonschema` handle the structural validation (object properties,
required fields, array items, enum values).
The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md)
and refined by [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md):
The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md):
1. **Load time:** Parse the schema JSON, build the layout engine, build the
jsonschema validator. This is the `AlkTypeEngine::compile(schema)` constructor.
1. **Load time:** Parse the BAST document into the typed tree, compute
the layout engine, and (optionally) build the standard
`jsonschema::Validator` for the JSON-validation path. This is the
`AlkTypeEngine::compile` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation.
### What validation validates
The jsonschema validator operates on `serde_json::Value` instances — it
validates JSON representations of data, not raw byte buffers. This is
the correct separation of concerns:
- **JSON validation** (jsonschema): validates that a JSON document
conforms to the schema. Used for validating hand-written schemas,
TypeBox output, JSON payloads, or the JSON representation of a binary
struct after deserialization.
- **Binary access validation** (data access layer): the read/write
functions perform type-level validation at access time — range checks
for integers, UTF-8 validity for strings, buffer bounds checking.
These return `AlkTypeError::Access` with field paths.
The "schema is the format" principle means the same schema describes
both the JSON shape and the binary layout. The jsonschema validator
checks the JSON shape; the data access layer checks the binary layout.
A consumer that wants to validate a binary buffer end-to-end reads the
buffer into a `Value` tree via the data access layer, then validates
that `Value` against the jsonschema validator. This is a two-step
process, not a single `validate(buffer)` call.
### The `AlkTypeEngine` struct
The `AlkTypeEngine` is the compiled form of a schema. It supports both
layout modes (ADR-002) via an internal `Layout` enum:
The `AlkTypeEngine` is the compiled form of a BAST document. It
supports both layout modes (ADR-002) via an internal `Layout` enum:
```rust
pub struct AlkTypeEngine {
layout: Layout, // packed or aligned (private enum)
validator: jsonschema::Validator, // compiled once at load time
endian: Endian, // parsed from the schema's "endian" annotation
schema: Value, // the normalized schema (refs resolved)
layout: Layout, // packed or aligned (private enum)
json_validator: Option<jsonschema::Validator>, // None when no JSON Schema supplied
endian: Endian, // parsed from the root struct's "endian"
bast_doc: Value, // retained for sequential_reader/read_field
root_name: String, // the selected $defs entry
}
// Private — the consumer selects via LayoutMode at compile time.
@@ -68,205 +151,105 @@ enum Layout {
```
The consumer selects the mode at construction time via `LayoutMode`
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout`
enum is private — the engine exposes mode-appropriate accessors instead:
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The
`Layout` enum is private — the engine exposes mode-appropriate
accessors instead:
```rust
impl AlkTypeEngine {
pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result<Self, AlkTypeError>;
pub fn compile(
bast_doc: &Value,
root_name: &str,
mode: LayoutMode,
json_schema: Option<&Value>,
) -> Result<Self, AlkTypeError>;
pub fn mode(&self) -> LayoutMode;
pub fn endian(&self) -> Endian;
pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode
pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode
pub fn sequential_reader(&self) -> Option<SequentialReader>; // owned fresh reader (ADR-007)
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // ADR-004
pub fn is_valid_json(&self, instance: &Value) -> bool; // ADR-004
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // ADR-010
pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str)
-> Result<FieldValue<'a>, AlkTypeError>; // aligned mode
pub fn write_field(&self, buffer: &mut [u8], field_path: &str,
value: &FieldValue<'_>) -> Result<(), AlkTypeError>; // aligned mode
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // D-BAST-007
pub fn is_valid_json(&self, instance: &Value) -> bool; // D-BAST-007
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // D-BAST-006
}
```
`compile` takes `&mut Value` because it normalizes `$ref` values in place
(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization))
before computing the layout and building the validator. The `schema`
field retains the normalized schema for `read_field`'s kind lookup and
for `sequential_reader()`'s factory construction. The validator is
mode-agnostic (it operates on `Value`, not raw bytes).
`compile` takes `&Value` (not `&mut Value`) — BAST needs no in-place
`normalize_refs`. `root_name` selects which `$defs` entry is the
top-level type (D-BAST-001). `json_schema` is the optional
consumer-provided standard JSON Schema for the `validate_json` path
(D-BAST-007); pass `None` when JSON validation is not needed. The
engine retains a clone of the BAST `Value` so `sequential_reader` and
`read_field` can re-parse the typed tree on demand without lifetime
entanglement with the caller's `Value`.
The `Layout::Packed` variant stores only the `LayoutBuilder` (write-side).
The `SequentialReader` (read-side) is not stored — it has mutable cursor
state that the consumer owns, so `sequential_reader()` constructs a fresh
reader on each call (ADR-007).
The `Layout::Packed` variant stores only the `LayoutBuilder`
(write-side). The `SequentialReader` (read-side) is not stored — it has
mutable cursor state that the consumer owns, so `sequential_reader()`
constructs a fresh reader on each call (ADR-007).
The `read_field`/`write_field` methods on `AlkTypeEngine` are the
aligned-mode data-access API — see [data-access.md](data-access.md)
§"Higher-level read/write".
## Custom Keyword Validators
## `build_validator`
Each `AlkType:*` kind gets a `Keyword` implementation registered via
`jsonschema::options().with_keyword(...)`. The validators check leaf
type constraints; `jsonschema` handles all structural validation.
### Numeric type validators
**`AlkType:Float32` / `AlkType:Float64`:**
- Value must be a finite number.
- For `Float32`: value must be representable as `f32` (no precision loss
beyond `f32`'s mantissa).
**`AlkType:Int8` / `AlkType:Int16` / `AlkType:Int32`:**
- Value must be an integer within the type's range.
- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
**`AlkType:Uint8` / `AlkType:Uint16` / `AlkType:Uint32`:**
- Value must be a non-negative integer within the type's range.
- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295.
### String and binary validators
**`AlkType:String`:**
- Value must be a valid UTF-8 string.
- If `maxLength` is specified in the schema, the string's byte length
must not exceed it.
**`AlkType:Bytes`:**
- Value must be a string (the JSON form for `validate_json` consumers)
or an array of integers 0..=255 (the materialized form for
`validate_bytes`). JSON has no native byte type; the string form is
the JSON convention, the array form is the round-trippable form for
non-UTF-8 bytes (see [OQ-007](questions/007-bytes-materialization-lossy-utf8.md)).
- If `maxLength` is specified, the byte length must not exceed it. For
the string form, this is the string's byte length; for the array
form, this is the array length (one entry per byte).
- **Binary representation:** In the binary layout, `TBytes` is raw bytes
with no encoding (not base64, not hex). The JSON representation (for
validation) uses a string or array; the binary representation (for
data access) uses `&[u8]` directly.
**`AlkType:Enum`:**
- The `AlkType:Enum` custom keyword signals that the type is an enum for
*layout* purposes (the engine needs to know it's a fixed-size u32 index,
not a variable-length string). The built-in `enum` keyword provides the
value list and handles value-membership validation. The custom keyword
validator is a no-op beyond the built-in check — it exists solely for
the layout engine to recognize the type.
**`AlkType:Timestamp`:**
- Value must be a valid RFC 3339 timestamp string (the internet profile
of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
### Composite type validators
**`AlkType:Struct`:**
- Value must be an object.
- Each property must match its declared `AlkType:*` kind.
- Required fields must be present.
- The `jsonschema` crate's built-in `properties` and `required` keywords
handle the structural checks — the custom keyword only needs to
validate that each field's value matches its `AlkType:*` kind.
**`AlkType:Union`:**
- The instance must be an object with a `__discriminator` field
carrying the mapping key (stringified discriminator value for
byte-offset discriminators, string value for field-name
discriminators). This is the shape the materializer produces for
`validate_bytes`; `validate_json` consumers produce the same shape
when validating a union instance.
- The `UnionValidator` builds a sub-validator for each variant at
factory time (when the parent validator tree is constructed) and
dispatches on `__discriminator` at validation time, validating the
full instance (including the variant fields) against the selected
variant's schema. This closes the OQ-008 gap: variant field
constraints (e.g. `maxLength` on a `Bytes` field inside a variant)
are checked.
- `$ref`s in the union's `mapping` are inlined by
`schema::inline_union_variant_refs` during `AlkTypeEngine::compile`
(before `build_validator`), so the `union_factory` sees full inline
variant schemas. See [OQ-008](questions/008-unionvalidator-variant-dispatch.md).
**`AlkType:Array`:**
- Value must be an array.
- Each element must match the array's declared element type.
- If `minItems`/`maxItems` is specified, the array length must be within
bounds.
### Other validators
**`AlkType:Boolean`:**
- Value must be `true` or `false`.
**`AlkType:Record`:**
- Value must be an object.
- All values must match the record's declared value type (specified via
the `"values"` property in the schema, e.g.,
`"values": { "AlkType:Float32": true }`).
### Validator implementation pattern
Each custom keyword implementation is ~10 lines. Example for
`AlkType:Float32`:
`src/validation.rs` exposes one function:
```rust
struct Float32Validator;
impl Keyword for Float32Validator {
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
match instance {
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
}
}
fn is_valid(&self, instance: &Value) -> bool {
instance.as_f64().map_or(false, |f| f.is_finite())
}
}
pub fn build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError>;
```
Registration:
Under the pivot this is **repurposed** (D-BAST-007): it builds a
*standard* `jsonschema::Validator` from a consumer-provided plain JSON
Schema — no custom keywords, no BAST involvement. The engine calls it
internally during `compile` when `json_schema` is `Some`. Consumers
that only need a one-off validator may call `jsonschema::options().build(schema)`
directly; `build_validator` exists so the engine's error mapping
(`jsonschema` build error → `AlkTypeError::Schema`) is reused.
```rust
let validator = jsonschema::options()
.with_keyword("AlkType:Float32", |parent, value, path| {
Ok(Box::new(Float32Validator))
})
.build(&schema)?;
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path. This enables cross-keyword awareness — for
example, a `AlkType:Struct` validator can inspect the parent's
`properties` to validate each field against its declared `AlkType:*` kind.
The v0.1.0 custom-keyword `build_validator` (registered 19
`with_keyword(...)` factories) is removed.
## AlkTypeError
A single `AlkTypeError` enum covers all error conditions across the
engine's three phases (schema parsing, offset computation, read/write)
plus validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md).
engine's phases (schema parsing, offset computation, read/write) plus
validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md);
the variant shapes are unchanged under the pivot (D-BAST-009).
```rust
pub enum AlkTypeError {
/// Schema parsing errors (invalid JSON, missing keywords, unknown AlkType kinds).
/// Schema parsing errors (malformed BAST, dangling $ref, unknown kind).
Schema(String),
/// Offset computation errors (field not found, unsupported type).
Offset { field_path: String, reason: String },
/// Read/write errors (buffer too short, invalid UTF-8, value out of range).
Access { field_path: String, reason: String },
/// Validation errors (delegated to jsonschema).
Validation(ValidationError<'static>),
/// Validation errors (both paths — D-BAST-009 uniform payload).
Validation(jsonschema::ValidationError<'static>),
}
```
- **`Schema`** — for errors during `AlkTypeEngine::compile()`. Invalid
JSON, missing required keywords, unknown `AlkType:*` kinds.
- **`Schema`** — for errors during `AlkTypeEngine::compile()` or any
BAST-walking path. Malformed BAST, missing `$defs`, unknown `kind`
string, dangling `$ref`, empty `mapping`, etc.
- **`Offset`** — for errors during offset computation. Field not found
in the schema, type not supported for offset computation, recursive
depth exceeded. Carries the field path.
- **`Access`** — for errors during read/write. Buffer too short, invalid
UTF-8 in a string field, value out of range for the target type.
Carries the field path.
- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The
`'static` lifetime is correct — the validator owns its schema reference
and lives for the lifetime of the `AlkTypeEngine`.
in the BAST tree, type not supported for offset computation. Carries
the field path.
- **`Access`** — for errors during read/write. Buffer too short,
invalid UTF-8 in a string field, value out of range for the target
type. Carries the field path.
- **`Validation`** — wraps a `jsonschema::ValidationError<'static>`.
On the `validate_json` path, this is the `jsonschema` crate's own
structured error. On the `validate_bytes` path, it is constructed via
`jsonschema::ValidationError::custom` from the BAST-native
validator's path + reason string. The `'static` lifetime is correct —
the payload owns its data.
### Field-path-carrying errors
@@ -287,20 +270,54 @@ you exactly which field failed and why.
### Load time: `AlkTypeEngine::compile()`
The expensive work happens once at schema load time:
1. Normalize `$ref` values in the schema (`normalize_refs`).
2. Parse the schema's `"endian"` annotation.
3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned).
4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
The result is a `AlkTypeEngine` that can be used for repeated operations.
1. Parse the BAST document into the typed tree (`BastDoc::new`).
2. Parse the root struct's `"endian"` annotation.
3. Compute the layout (`LayoutBuilder` for packed, `OffsetMap` for
aligned).
4. If `json_schema` is `Some`, build the standard
`jsonschema::Validator` via `validation::build_validator`.
The result is an `AlkTypeEngine` that can be used for repeated
operations. The BAST-native validator is not pre-built — it is a
recursive walker that runs on the materialized `Value` at access time,
re-using the `BastDoc` (re-parsed on demand from the retained
`bast_doc`).
### Access time: `engine.validate_bytes(&[u8])`
For binary-layout schemas, `validate_bytes` runs the two phases in
sequence (D-BAST-006):
1. **Materialize `Value` from bytes.** `materialize::materialize_packed`
or `materialize::materialize_aligned` walks the buffer against the
BAST typed tree and the engine's `Endian`, producing a
`serde_json::Value` tree. Composites are recursed into (`Struct` →
object of field values; `Array` → array of element values; `Union`
→ dispatch then recurse; `Record` → object of key/value entries).
The read phase reuses the existing data-access functions and returns
`AlkTypeError::Access` (with field paths) on read failures.
2. **Validate the `Value`.** The materialized `Value` is passed to
`bast_validation::validate_value(&doc, &value)`, producing
`AlkTypeError::Validation` on the first violated value-domain
constraint.
Mode dispatch:
- **Packed mode** — materializes fields in declaration order.
- **Aligned mode** — uses the `OffsetMap` to read fields at their
computed offsets.
Both modes produce the same `Value` form; the BAST-native validator is
mode-agnostic.
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
Validation is opt-in per operation. The consumer calls
`engine.validate_json(instance)` when validation is desired, or
`engine.is_valid_json(instance)` for a boolean check. The jsonschema
validator is already compiled — these are fast checks against the
compiled validator.
`engine.is_valid_json(instance)` for a boolean check. The
`jsonschema::Validator` is already compiled — these are fast checks
against the compiled validator.
```rust
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>;
@@ -308,79 +325,37 @@ pub fn is_valid_json(&self, instance: &Value) -> bool;
```
The argument is a `serde_json::Value` (the JSON representation of the
data), not a raw byte buffer — see §"What validation validates" above.
To validate a binary buffer end-to-end, the consumer reads it into a
`Value` tree via the data access layer, then validates that `Value`.
data), not a raw byte buffer. `validate_json` validates against the
consumer-provided JSON Schema supplied at `compile` time (D-BAST-007);
the BAST document is not involved. If no JSON Schema was supplied,
`validate_json` returns `AlkTypeError::Schema` and `is_valid_json`
returns `false`.
High-throughput paths can skip validation. Security-sensitive paths
(parsing incoming frames from untrusted peers) can validate every frame.
The choice is the consumer's.
(parsing incoming frames from untrusted peers) can validate every
frame. The choice is the consumer's.
### Access time: `engine.validate_bytes(&[u8])` — binary buffer validation
For binary-layout schemas (schemas declaring `AlkType:*` kinds), the
engine offers a single-call form of the two-step dance: walk the bytes
against the layout to materialize a `Value` tree, then validate that
`Value` against the compiled jsonschema validator. Decided in
[ADR-010](decisions/010-generalized-validation-validate-bytes.md).
```rust
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>;
```
`validate_bytes` runs the existing machinery in sequence:
1. **Materialize `Value` from bytes.** A new internal helper
(`materialize_value`, alongside `SequentialReader::read_field_value`
in `src/sequential_reader.rs`) walks the buffer against the schema
and the engine's `Endian`, producing a `serde_json::Value` tree.
Composites are recursed into (`Struct` → object of field values;
`Array` → array of element values; `Union` → dispatch then recurse;
`Record` → object of key/value entries). The read phase reuses the
existing data-access functions and returns `AlkTypeError::Access`
(with field paths) on read failures.
2. **Validate the `Value`.** The materialized `Value` is passed to the
existing `self.validator.validate(&value)`, producing
`AlkTypeError::Validation` on failure.
Mode dispatch:
- **Packed mode** — walks with a fresh `SequentialReader` (the engine
is already a reader factory per ADR-007), materializing fields in
declaration order.
- **Aligned mode** — uses the `OffsetMap` to read fields at their
computed offsets, then materializes composites by recursing into the
offset map's nested entries.
Both modes produce the same `Value` form; the validator is
mode-agnostic (it operates on `Value`, not bytes — ADR-004).
#### When to use which entry point
### When to use which entry point
| Entry point | Schema form | Input form | When |
|-------------|--------------|------------|------|
| `validate_json(&Value)` | Any (AlkType or plain JSON Schema) | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); TypeBox output; anything off `serde_json::from_slice` / `from_str` |
| `validate_bytes(&[u8])` | AlkType binary-layout schema | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs |
| `validate_json(&Value)` | Consumer-provided standard JSON Schema | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); anything off `serde_json::from_slice` / `from_str` |
| `validate_bytes(&[u8])` | BAST document (binary layout) | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs |
`validate_bytes` requires the engine's schema to declare `AlkType:*`
kinds — it materializes `Value` via the layout engine, which needs
binary-layout semantics. A pure JSON Schema (call's `input_schema`,
no AlkType kinds) compiled via `AlkTypeEngine::compile` would fail at
the materialize step (no `AlkType:Struct` at the root). For pure JSON
payloads, the consumer uses `serde_json::from_slice` then
`validate_json`. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md)
§"Not a binary-payload validator for JSON-only schemas".
`validate_bytes` requires the engine's root type to be a struct (the
layout engine enforces this) — it materializes `Value` via the layout
engine, which needs binary-layout semantics. For pure JSON payloads,
the consumer uses `serde_json::from_slice` then `validate_json`. See
[ADR-010](decisions/010-generalized-validation-validate-bytes.md) and
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
#### What `validate_bytes` is not
- **Not a new validation engine.** It runs the existing `jsonschema`
validator against the existing materialized `Value`. No new
validator code, no parallel validation path (ADR-001).
- **Not framing-aware.** It validates the bytes of *one* schema
instance. It does not strip length prefixes, parse
`[length: u32][payload]` framing, or handle multiple frames in a
buffer. That's the consumer's job. alktype validates what one
schema describes; it does not parse the wire envelope around it.
buffer. That's the consumer's job. alktype validates what one schema
describes; it does not parse the wire envelope around it.
- **Not a `Validator` trait.** Two methods on one struct, not a trait
abstraction. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md)
§"Not a `Validator` trait abstraction".
@@ -390,24 +365,24 @@ payloads, the consumer uses `serde_json::from_slice` then
Validation and data access are independent operations on the same data.
The consumer can:
1. Validate the JSON representation of a buffer to ensure it conforms to
the schema.
1. Validate the bytes of a buffer to ensure it conforms to the BAST
document's value constraints.
2. Read fields from the binary buffer at computed offsets.
3. Both — validate the JSON representation first, then read the binary
buffer (defense in depth).
3. Both — validate first, then read (defense in depth).
The engine does not couple validation and access. A consumer that trusts
its data source can skip validation and go straight to read/write. A
consumer that parses untrusted input can validate the JSON
representation first, then access the binary buffer.
The engine does not couple validation and access. A consumer that
trusts its data source can skip validation and go straight to
read/write. A consumer that parses untrusted input can validate first,
then access the binary buffer.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Two-validator model (BAST-native + standard jsonschema) | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | `validate_bytes` uses the BAST-native validator; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema; D-BAST-006/007/009 |
| Error handling and validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation (materialize `Value` from bytes, then validate); two methods on one struct, not a trait |
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine |
| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | The BAST document is the complete binary-format spec (layout + value constraints) |
## Open Questions
@@ -418,11 +393,19 @@ see [builder.md](builder.md).
## References
- `@alkdev/alknet: docs/research/alknet-typedef/findings.md`
§"Validation" — the POC's custom keyword validators for all 17 kinds
- [`bast-format.md` §Validation Model](bast-format.md#validation-model)
— the normative validation model
- [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) — the
two-validator decision
- [ADR-004](decisions/004-error-handling-validation-strategy.md) —
error handling and validation strategy
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds that the
validators check
- [data-access.md](data-access.md) — read/write functions that operate
on the same buffers
- [ADR-010](decisions/010-generalized-validation-validate-bytes.md) —
`validate_bytes` (the collapsed two-step dance)
- [schema-layer.md](schema-layer.md) — the BAST parser that the
BAST-native validator walks
- [data-access.md](data-access.md) — read/write functions and the
materializer that produce the `Value` the validator checks
- [`src/bast_validation.rs`](../../src/bast_validation.rs) — the
BAST-native validator implementation
- [`src/validation.rs`](../../src/validation.rs) — the `build_validator`
helper
+549
View File
@@ -0,0 +1,549 @@
---
status: complete
created: 2026-08-15
last_updated: 2026-08-15
---
# BAST Pivot — Implementation Plan
**Status: complete.** All 10 steps are implemented and pushed to
`origin/main` (steps 1–8 in commits `66ab9d7` → `54fd112`; step 9 was
a no-op — steps 4–8 converted the tests as they went, leaving only the
intentional `from_bast_str` rejection test referencing the
`"AlkType:Uint32"` string; step 10 synced the architecture docs and
ADRs in this commit). The two new ADRs
([ADR-BAST](../architecture/decisions/bast-bast-format.md),
[ADR-VAL-SPLIT](../architecture/decisions/val-split-two-validator-model.md))
record the decisions; the amended ADRs (001, 002, 003, 004, 009, 010)
carry supersession/amendment notes. The research record
([`bast-pivot.md`](../research/bast-pivot.md)) is flipped to
`implemented`. What follows is the original plan, preserved as the
historical execution record.
---
This is the execution plan for the BAST pivot: replacing alktype's
v0.1.0 `AlkType:*` custom-keyword JSON Schema format with the BAST
(Binary Abstract Syntax Tree) format. It is the **entry point** an
implementing agent reads first.
Companion documents:
- [`docs/architecture/bast-format.md`](../architecture/bast-format.md) —
the normative BAST format spec (meta-schema, TypeRef, examples,
validation model). Read this for *what* the format is.
- [`docs/research/bast-pivot.md`](../research/bast-pivot.md) — the
research record: motivation, POC scope and result, decisions
D-BAST-001..009, risks. Read this for *why* and *what was proved*.
The POC lives on branch `bast-validator-poc` (commit `f371fe4`) as
`src/bast_poc.rs` — reference scaffolding, deliberately not merged.
**Working order:** read this plan top-to-bottom. The Semver Contract
section is the scope-creep guardrail — consult it before each step.
Each step links to the specific spec section it implements and the
relevant D-BAST-* decision anchor. Implement steps in order; each step
lists its verification gate.
## Semver Contract
The crate is on crates.io at 0.1.0 with zero real consumers, so a
breaking bump is free — but the contract is explicit so the
implementation doesn't drift. Per AGENTS.md, the 0.1.0 public surface
is the items re-exported from `src/lib.rs`. This table is the
authoritative scope-creep guardrail for the pivot.
| Public item (from `lib.rs` re-exports) | Class | Change |
|---|---|---|
| `AlkTypeKind` (enum + variants + methods) | **Additive** | Unchanged. 19 variants, same methods. New `from_str()`/`to_str()` mapping for lowercase BAST kind strings (`"uint32"` ↔ `AlkTypeKind::Uint32`) — additive methods. |
| `Endian`, `VariableEncoding`, `DiscriminatorKind` | **Unchanged** | — |
| `AlkTypeEngine::compile` | **Breaking** | Signature: `compile(schema: &mut Value, mode)` → `compile(bast_doc: &Value, root_name: &str, mode)`. Adds required `root_name` param (D-BAST-001); drops `&mut` (BAST needs no in-place `normalize_refs`); input is a BAST document, not a custom-keyword JSON Schema. |
| `AlkTypeEngine::validate_json` | **Breaking (behavioral)** | Signature unchanged `(instance: &Value) -> Result<...>`, but the validator it runs is now a standard `jsonschema::Validator` from a consumer-provided JSON Schema, not a custom-keyword validator built from the alktype schema. The *contract* of what schema validates the instance changes. |
| `AlkTypeEngine::validate_bytes` | **Unchanged (contract)** | Same signature. Internally the validation step switches from `jsonschema::Validator` to the BAST-native validator. Error type unchanged (D-BAST-009). |
| `AlkTypeEngine::is_valid_json` | **Breaking (behavioral)** | Same caveat as `validate_json` — validates against the consumer JSON Schema, not the alktype schema. |
| `AlkTypeEngine` accessors (`endian`, `mode`, `offset_map`, `layout_builder`, `sequential_reader`, `read_field`, `write_field`, etc.) | **Unchanged** | Layout-layer accessors are format-agnostic. |
| `LayoutMode`, `OffsetMap`, `ByteRange` | **Unchanged** | — |
| `LayoutBuilder`, `PackedLayout`, `FieldPosition` | **Unchanged** | — |
| `SequentialReader`, `FieldValue` | **Unchanged** | — |
| `UnionDispatch` | **Unchanged** | — |
| `data_access::*` functions | **Unchanged** | — |
| `AlkTypeError` (all 4 variants) | **Unchanged** | D-BAST-009 keeps `Validation(jsonschema::ValidationError<'static>)`. |
| `Schema` builder (`struct_`, `object`, `field`, `build`, all setters) | **Breaking (output format)** | Public method signatures unchanged. `build()` output changes from custom-keyword JSON to BAST JSON (for `struct_`) / standard JSON Schema (for `object`). Callers that introspect the built `Value` break; callers that pass it straight to `compile` are source-compatible once `compile` takes BAST. |
| `Definitions` builder (`new`, `define`, `define_value`, `build`, `merge_into`) | **Breaking (output format)** | Same as `Schema` — signatures unchanged, `build()`/`merge_into()` output shape changes to BAST `$defs`. |
| `Discriminator` builder enum | **Unchanged** | — |
| `build_validator` (from `validation`) | **Breaking (signature or removal)** | Currently `build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError>` builds a custom-keyword validator. Under the pivot it either (a) is removed (consumers call `jsonschema` directly for standard JSON Schema) or (b) is repurposed to build a standard `jsonschema::Validator` from a consumer-provided standard JSON Schema (no custom keywords). Decision belongs to step 6. Either way the current signature's contract breaks. |
| `get_alktype_kind`, `get_alktype_kind_enum`, `get_alktype_kind_loose`, `get_alktype_kind_loose_enum`, `normalize_refs`, `inline_union_variant_refs`, `resolve_ref`, `resolve_ref_or_inline`, `parse_align`, `parse_discriminator`, `parse_encoding`, `parse_endian`, `parse_max_length` | **Breaking (removal or rework)** | All currently re-exported from `lib.rs`. `normalize_refs` and `inline_union_variant_refs` are removed (BAST needs neither). The `get_alktype_kind*` family is removed (replaced by direct `kind` parsing). The `parse_*` and `resolve_*` functions are reworked to read BAST properties instead of keyword-value objects, or removed if subsumed by the BAST parser. **Open: which of these stay public vs become internal.** Current leaning — drop all from `lib.rs` re-exports (they're engine-internal accessors, not consumer API); the BAST parser exposes a new typed surface instead. Confirmed during step 3. |
**Net breaking surface:** `compile`, `validate_json`/`is_valid_json`
(contract), `Schema::build`/`Definitions::build` (output format),
`build_validator` (signature/removal), and the ~13 `schema::*` helper
re-exports. **Net additive:** BAST parser, BAST-native validator,
`AlkTypeKind::from_str`/`to_str`. **Net unchanged:** the entire layout
+ data-access + materialize + tunion layer, `AlkTypeError`, the
`Discriminator` builder, `AlkTypeKind` variants.
### Decisions deferred to their implementation steps
These are small enough to decide when the step is reached, but are
flagged here so they don't become drive-by semver changes:
1. **`validate_json` JSON Schema source** (step 6): does the consumer
pass the JSON Schema to `compile` (engine carries a second
validator) or to `validate_json` at call time? The former preserves
the current single-call ergonomics; the latter is more flexible. Not
semver-relevant either way if `validate_json`'s signature can absorb
a new param or stay as-is — needs the call-site analysis.
2. **`build_validator` fate** (step 6): removed vs repurposed. If
repurposed, its signature stays but its contract (no custom
keywords) changes — a behavioral break, not a type break.
3. **`schema::*` helper re-exports** (step 3): drop from `lib.rs`
(engine-internal) vs keep public for consumers that walk schemas.
Leaning: drop — they're accessors for the old format, and the BAST
parser exposes a cleaner typed surface. Confirmed during step 3.
## Steps
### Step 1 — Add `AlkTypeKind::from_str`/`to_str` for BAST kind strings
**Goal:** Add the lowercase-string mapping (`"uint32"` ↔
`AlkTypeKind::Uint32`) that the BAST parser and validator dispatch on.
This is the additive-only, zero-risk foundation — no existing code
changes.
**Spec reference:** [bast-format.md §Primitives](../architecture/bast-format.md#primitives),
[D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set).
**Files:** `src/schema.rs` (the `AlkTypeKind` impl block). No `lib.rs`
change needed — the methods are inherent on the already-re-exported
enum.
**Implementation notes:**
- `to_str(self) -> &'static str` returns the lowercase BAST string.
- `from_str(s: &str) -> Result<AlkTypeKind, AlkTypeError>` returns
`AlkTypeError::Schema` for unknown strings. This is a new inherent
method, distinct from the existing `FromStr` impl that parses the
v0.1.0 `"AlkType:Uint32"` keyword form. Do not remove the existing
`FromStr` yet — step 8 removes the v0.1.0 accessors.
- Cover all 14 primitive kinds plus `struct`, `union`, `array`,
`record`, `enum` (19 total, matching the enum variants). The
lowercase strings are in the [primitives table](../architecture/bast-format.md#primitives);
composite kinds are `"struct"`, `"union"`, `"array"`, `"record"`,
`"enum"`.
**Verification:** `cargo test --release` (new unit tests for the
mapping, both directions; existing tests unaffected). `cargo clippy
--all-targets -- -D warnings`.
---
### Step 2 — Embed the BAST meta-schema
**Goal:** Embed the BAST meta-schema as a `serde_json::Value` constant
in the crate, available for validating BAST documents at compile time
and for publishing at `https://alk.dev/bast/v1/schema`.
**Spec reference:** [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
**Files:** New `src/bast_meta.rs` (or a `const` in `src/schema.rs` —
match existing module conventions). Re-export the meta-schema `Value`
from `lib.rs` if consumers should be able to validate BAST documents
themselves (likely yes — additive, not semver-relevant).
**Implementation notes:**
- The meta-schema JSON is in [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
Copy it verbatim into a `serde_json::json! {...}` macro invocation or
parse it from an embedded string via `serde_json::from_str`.
- No feature flags (AGENTS.md §6). The meta-schema is a compile-time
constant, no I/O.
- WASM-clean: no `include_str!` of an external file is needed if the
`json!` macro is used; either way is wasm-safe.
**Verification:** `cargo test --release`. `cargo build --target
wasm32-unknown-unknown --release` (meta-schema is a `Value` constant —
wasm-relevant). `cargo clippy --all-targets -- -D warnings`.
---
### Step 3 — BAST document parser
**Goal:** Implement the BAST document parser that the layout engines
and materializer use instead of the `get_alktype_kind*` custom-keyword
accessors. This is the natural entry point for the pivot — the largest
step, and the one the rest of the steps build on.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(the whole document — the parser implements the format spec).
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection),
[D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement),
[D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions).
**Files:** New `src/bast.rs` (the parser). The existing `src/schema.rs`
stays for now — steps 4–8 migrate callers off it. Update `src/lib.rs`
to add `pub mod bast;` and re-export the parser's public surface.
**Implementation notes:**
- The parser reads `kind`/`fields`/annotation properties from BAST
nodes. It produces a typed surface (a small `BastNode` enum or
equivalent) that the layout engines, materializer, and validator can
walk without re-parsing the raw JSON at every node. The POC parsed
lazily from raw JSON in both passes to keep the model honest; a typed
tree is a straightforward follow-on optimization (POC observation 5).
Either is acceptable for the production version; the typed tree is
recommended since three consumers (layout, materialize, validate)
walk the same tree.
- `$ref` resolution: `#/$defs/<name>` only — a single hash lookup. No
`normalize_refs` (BAST refs are always full JSON Pointers), no
`inline_union_variant_refs` (union variant refs resolved lazily by
the validator and materializer). See [bast-format.md §TypeRef](../architecture/bast-format.md#typeref).
- Untrusted input: every path that walks a BAST document must return
`Err(AlkTypeError::Schema)` on a malformed document, never `panic!`/
`unreachable!` (AGENTS.md §3). The POC's
`malformed_document_produces_schema_error_not_panic` test is the
template.
- **Decide deferred decision #3 here:** drop the `schema::*` helper
re-exports from `lib.rs`, or keep them public. Leaning: drop. The
BAST parser exposes a cleaner typed surface; the v0.1.0 accessors
are engine-internal and not consumer API.
**Verification:** `cargo test --release` (port the POC's parser tests
— the malformed-document test, the type-ref resolution tests).
`cargo clippy --all-targets -- -D warnings`. The layout engines don't
use the parser yet (step 4 wires it in), so the existing suite still
passes on the old path.
---
### Step 4 — Wire `compile()` to accept a BAST document + root name
**Goal:** Change `AlkTypeEngine::compile` to the new signature and
have it use the BAST parser instead of the custom-keyword accessors.
The layout engines (`offset_map`, `layout_builder`,
`sequential_reader`) consume the BAST parser's typed output instead of
walking raw JSON with `get_alktype_kind*`.
**Spec reference:** [bast-format.md §Document Shape](../architecture/bast-format.md#document-shape),
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection).
Semver contract: `compile` is **Breaking**.
**Files:** `src/engine.rs` (the `compile` signature and body). The
layout modules (`src/offset_map.rs`, `src/layout_builder.rs`,
`src/sequential_reader.rs`) — their schema-walking code changes from
`get_alktype_kind*` calls to BAST parser calls. `src/lib.rs` if the
parser's public surface needs re-exporting (step 3 may have done this).
**Implementation notes:**
- New signature: `pub fn compile(bast_doc: &Value, root_name: &str,
mode: LayoutMode) -> Result<Self, AlkTypeError>`. Note `&Value` (not
`&mut Value`) — BAST needs no in-place `normalize_refs`.
- The engine stores the BAST document (or the parsed typed tree) for
`sequential_reader()`'s factory construction and `read_field`'s kind
lookup. The `Layout` enum and mode dispatch are unchanged.
- `parse_endian`, `parse_align`, `parse_encoding`, `parse_discriminator`
are reworked to read BAST properties (struct/field-level) instead of
keyword-value objects. Their *semantics* are unchanged (ADR-003);
only their *input location* moves. Whether they stay as free
functions or become methods on the typed `BastNode` is an
implementation choice — the POC read properties inline.
- The layout engines are format-agnostic beneath the accessors
(checked offset arithmetic, the two modes, union dispatch). This
step is an accessor swap, not a layout-engine rewrite.
**Verification:** `cargo test --release` (test inputs must be converted
to BAST format — see step 9 for the full test conversion; this step
converts the layout tests as a sanity check). `cargo clippy
--all-targets -- -D warnings`. `cargo build --target
wasm32-unknown-unknown --release` (layout/wasm-relevant).
---
### Step 5 — BAST-native validator (production version)
**Goal:** Port the POC's BAST-native validator into a production module
and wire it into `validate_bytes` as the validation step, replacing the
`jsonschema::Validator` call on the bytes path.
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
[D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model),
[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape).
POC reference: `src/bast_poc.rs` on branch `bast-validator-poc`.
**Files:** New `src/bast_validation.rs`. `src/engine.rs`
(`validate_bytes` body — swap the `self.validator.validate(&value)` call
for the BAST-native validator). `src/lib.rs` — add `pub mod
bast_validation;` (the validator is engine-internal; whether it's
re-exported is an implementation choice, leaning no).
**Implementation notes:**
- The POC is the reference. The validator is a single recursive
function (`validate_typeref`) that dispatches on the BAST `kind`. The
constraint table is in [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model).
- Construct `AlkTypeError::Validation` via
`jsonschema::ValidationError::custom` — the variant's payload type is
unchanged (D-BAST-009). The bytes path no longer touches `jsonschema`
for validation, but the error type retains the `jsonschema` type for
uniformity with the `validate_json` path.
- The validator and materializer share the BAST-walking code structure.
If step 3 produced a typed `BastNode` tree, both consume it. If step
3 parses lazily, the validator parses lazily too (POC approach).
- Enum index bounds: check the materialized index against
`values.len()` — this **fixes the v0.1.0 dead constraint** (the
built-in `enum` keyword checked string membership, but the
materializer emits `Value::Number(index)`, which never matched). Net
improvement.
- Union variant dispatch: read `__discriminator`, look up the variant's
BAST definition, recurse. Recovers OQ-008 per-variant constraint
enforcement without custom keywords.
**Verification:** `cargo test --release` — the existing `validate_bytes`
tests are the regression target (test *inputs* change to BAST format
in step 9; expected validation outcomes must be identical). The POC's
20 tests are the reference. `cargo clippy --all-targets -- -D warnings`.
`cargo build --target wasm32-unknown-unknown --release`.
---
### Step 6 — `validate_json` against a consumer-provided JSON Schema
**Goal:** Update `validate_json`/`is_valid_json` to validate against a
standard `jsonschema::Validator` compiled from a consumer-provided JSON
Schema, not a custom-keyword validator built from the alktype schema.
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model).
Semver contract: `validate_json`/`is_valid_json` are **Breaking
(behavioral)**; `build_validator` is **Breaking (signature or
removal)**.
**Files:** `src/engine.rs` (`validate_json`/`is_valid_json` bodies, and
the engine's stored validator field if the JSON Schema is supplied at
compile time). `src/validation.rs` (`build_validator` — repurposed or
removed). `src/lib.rs` (the `build_validator` re-export if removed).
**Implementation notes:**
- **Decide deferred decision #1 here:** does the consumer pass the JSON
Schema to `compile` (engine carries a second validator) or to
`validate_json` at call time? The former preserves single-call
ergonomics; the latter is more flexible. Needs the alkcall call-site
analysis. Not semver-relevant either way if the signature can absorb
the change.
- **Decide deferred decision #2 here:** `build_validator` removed vs
repurposed. If repurposed, its signature stays but its contract
changes (no custom keywords) — a behavioral break. If removed, drop
the `lib.rs` re-export.
- The `jsonschema` crate remains a direct dependency (for `validate_json`
and for validating BAST documents against the meta-schema). Only the
custom keyword integration is removed.
- The engine may carry two validators: the BAST-native validator (for
`validate_bytes`, from step 5) and the standard `jsonschema::Validator`
(for `validate_json`, from this step). Or `validate_json` takes the
JSON Schema at call time and builds a transient validator. The
decision shapes the engine struct's fields.
**Verification:** `cargo test --release` (new tests for the
consumer-provided JSON Schema path; existing `validate_json` tests
converted — their schemas were custom-keyword, now standard). `cargo
clippy --all-targets -- -D warnings`.
---
### Step 7 — Builder API produces BAST JSON
**Goal:** Update the builder's `build()` methods to produce BAST JSON
(for `struct_()`) and standard JSON Schema (for `object()`). Public
method signatures are unchanged; only the output `Value` shape changes.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(the output format), [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats).
Semver contract: `Schema::build`/`Definitions::build` are **Breaking
(output format)**.
**Files:** `src/builder.rs`. `src/lib.rs` if the builder's public
surface changes (it shouldn't — method signatures are unchanged).
**Implementation notes:**
- `Schema::struct_().field(...).build()` → BAST JSON (a `$defs` entry
with `kind: "struct"`, ordered `fields` array, type-level
annotations).
- `Schema::object().field(...).build()` → standard JSON Schema (no
`AlkType:*` keywords, no BAST `kind` — just `type`/`properties`/
`required`).
- `Definitions::build()`/`merge_into()` → a BAST `$defs` block.
- The builder already distinguishes AlkType kinds from JSON Schema types
via naming conventions (`string()` vs `string_()`). The construction
API is the same; only the serialization differs.
- The `Discriminator` builder is unchanged (semver contract:
**Unchanged**).
**Verification:** `cargo test --release` (builder tests assert on the
output `Value` — update the expected shapes). `cargo clippy
--all-targets -- -D warnings`.
---
### Step 8 — Remove v0.1.0 custom-keyword machinery
**Goal:** Remove the dead code now that all callers use the BAST parser
and BAST-native validator.
**Spec reference:** [bast-format.md §What is removed](../architecture/bast-format.md#what-is-removed).
Semver contract: the ~13 `schema::*` helper re-exports are **Breaking
(removal or rework)** (decision #3, confirmed in step 3).
**Files:** `src/schema.rs` (remove `get_alktype_kind*`,
`normalize_refs`, `inline_union_variant_refs`; rework or remove
`parse_*`/`resolve_*`). `src/validation.rs` (remove the 19
`jsonschema::Keyword` implementations if not already removed in step 5/6).
`src/lib.rs` (drop the removed items from the `pub use` block).
**Implementation notes:**
- Remove: all 19 `jsonschema::Keyword` implementations (~200 lines),
`normalize_refs()`, `inline_union_variant_refs()`, the
`get_alktype_kind*` family.
- Rework or remove: `parse_align`, `parse_discriminator`,
`parse_encoding`, `parse_endian`, `parse_max_length`, `resolve_ref`,
`resolve_ref_or_inline`. If the BAST parser subsumes them (likely),
remove them. If any remain useful as free functions over the typed
`BastNode`, keep them internal (not re-exported from `lib.rs`).
- The `jsonschema` crate's `with_keyword(...)` registration calls are
removed from `compile`/`build_validator`. The crate itself stays.
- Drop the removed items from `lib.rs`'s `pub use schema::{ ... }`
block. The BAST parser's public surface replaces them.
**Verification:** `cargo test --release`. `cargo clippy --all-targets
-- -D warnings`. `cargo doc --no-deps` (the public API surface
changed — doc comments must build). `cargo build --target
wasm32-unknown-unknown --release` (removing code shouldn't add
platform deps).
---
### Step 9 — Convert all tests to BAST format
**Goal:** Update the full test suite to use BAST format for inputs.
Test assertions (expected validation outcomes, expected offsets,
expected materialized values) must be identical — only the input
schema shape changes.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(input format).
**Files:** `tests/*.rs` (integration tests), `src/*.rs` inline `#[cfg(test)]`
modules (unit tests).
**Implementation notes:**
- This may be partially done by steps 4–8 (each step converts the tests
it touches as a sanity check). This step is the sweep: every test
using `AlkType:*` keywords converts to BAST `kind`/`fields`.
- The POC's 20 tests are the reference for BAST-shaped test inputs.
- Expected validation outcomes are the regression target. The
enum-index-bounds test is new behavior (the v0.1.0 dead constraint
is now enforced) — that test's expectation *changes* (was: silently
passed; now: `AlkTypeError::Validation`). This is the intended fix,
not a regression.
- Coverage: 310 crate + 86 integration tests (~396 total). All must
pass.
**Verification:** `cargo test --release` (the full suite — this is the
gate). `cargo clippy --all-targets -- -D warnings`.
---
### Step 10 — Sync architecture docs and ADRs
**Goal:** Sync the descriptive docs and ADRs to the shipped code. This
is the final step — per AGENTS.md, ADRs are written post-implementation,
grounded in shipped code.
**Spec reference:** [Semver Contract §ADR impact](#adr-impact-checklist)
below.
**Files:** `docs/architecture/README.md`, `docs/architecture/overview.md`,
`docs/architecture/schema-layer.md` (rewrite for the BAST parser),
`docs/architecture/validation.md` (rewrite for the validator split),
`docs/architecture/builder.md` (update `build()` output examples),
`src/lib.rs` (module doc comment). New ADRs: ADR-BAST, ADR-VAL-SPLIT.
Amended ADRs: 001 (superseded), 003, 004, 009, 010.
**Implementation notes:**
- Rewrite `schema-layer.md` to describe the BAST parser (replaces the
custom-keyword accessor walk-through). The current `schema-layer.md`
content is the v0.1.0 reference; `bast-format.md` already contains
the target spec. Either fold `bast-format.md` into `schema-layer.md`
or keep both with `schema-layer.md` pointing at `bast-format.md` for
the format and describing the parser module.
- Rewrite `validation.md` for the validator split (the [bast-format.md
§Validation Model](../architecture/bast-format.md#validation-model)
content moves here, expanded with the production validator's
details).
- Update `builder.md` output examples to BAST JSON.
- Update `src/lib.rs` module doc comment: "Takes a JSON Schema with
`AlkType:*` custom keywords" → "Takes a BAST document".
- Update `docs/architecture/README.md` index — the document table, the
ADR table (new ADRs, superseded ADR-001), the key design principles
(#1, #2, #7, #10 change wording).
- Remove stale TODOs referencing custom-keyword normalization,
`inline_union_variant_refs`, or the rejected bare-name-ref design
(AGENTS.md §"Architecture Context").
- `docs/research/bast-pivot.md` is the research record — its status
flips from `draft` to `accepted`/`implemented` and it gains a pointer
to the ADRs that superseded its decisions.
**Verification:** `cargo doc --no-deps` (doc comments build).
Cross-reference check: every link in this plan, `bast-format.md`, and
the new/updated ADRs resolves. `cargo test --release` (no code change,
but the doc sweep shouldn't break anything).
## ADR Impact Checklist
Sync these ADRs when step 10 lands. Per AGENTS.md, ADRs are written
post-implementation, grounded in shipped code.
| ADR | Action | Reason |
|---|---|---|
| [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) (purpose, scope, "schema is the format") | **Supersede** | The "schema is the format" principle is retained and strengthened (BAST *is* the format), but the concrete format changes from custom-keyword JSON Schema to BAST. A new ADR (ADR-BAST) records the BAST format as the realization of the principle. ADR-001 Status → Superseded by ADR-BAST. |
| [ADR-002](../architecture/decisions/002-two-layout-modes-packed-vs-aligned.md) (two layout modes) | **Unchanged** | Layout modes are format-agnostic. One-line note that the input format changed but the modes didn't. |
| [ADR-003](../architecture/decisions/003-schema-annotations.md) (annotations) | **Amend** | Annotation *semantics* carry forward unchanged; annotation *location* moves from custom-keyword objects to BAST type-level properties. Amend the "where annotations live" sections, keep the semantics. |
| [ADR-004](../architecture/decisions/004-error-handling-validation-strategy.md) (error handling, validation strategy) | **Amend** | Error enum shape unchanged (D-BAST-009). The "validation strategy" section updates: bytes path uses BAST-native validator, JSON path uses standard `jsonschema`. The load-time/access-time split is retained. |
| [ADR-005](../architecture/decisions/005-int64-uint64-first-class-kinds.md) (Int64/Uint64) | **Unchanged** | Kinds carry forward; JSON precision caveat unchanged. |
| [ADR-006](../architecture/decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) (reject non-final inline in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
| [ADR-007](../architecture/decisions/007-packed-mode-read-factory.md) (packed-mode read factory) | **Unchanged** | Reader factory semantics are format-agnostic. |
| [ADR-008](../architecture/decisions/008-reject-tunion-in-aligned-mode.md) (reject TUnion in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
| [ADR-009](../architecture/decisions/009-builder-api.md) (builder API) | **Amend** | Public method surface unchanged; `build()` output format changes (BAST for `struct_`, standard JSON Schema for `object`). Amend the "output format" section; keep the method catalog. |
| [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md) (`validate_bytes`) | **Amend** | The two-step concept (materialize → validate) is retained. The validation step's *implementation* changes from `jsonschema` custom keywords to the BAST-native validator. Amend the "validation step" section; add a pointer to D-BAST-006/D-BAST-009 and ADR-VAL-SPLIT. |
**New ADRs to write (post-implementation, grounded in shipped code):**
- **ADR-BAST** — the BAST format, meta-schema, and `$defs`/`$ref`/
`kind` vocabulary. Supersedes ADR-001's format-specific content.
- **ADR-VAL-SPLIT** (or fold into ADR-004's amend) — the two-validator
model: BAST-native for `validate_bytes`, standard `jsonschema` for
`validate_json`. Records D-BAST-006, D-BAST-007, D-BAST-009.
**Descriptive docs to sync (post-implementation):**
- `docs/architecture/schema-layer.md` — rewrite for the BAST parser
(replaces the custom-keyword accessor walk-through).
- `docs/architecture/validation.md` — rewrite for the validator split.
- `docs/architecture/builder.md` — update the `build()` output examples
to BAST JSON.
- `src/lib.rs` module doc comment — update the "Takes a JSON Schema
with `AlkType:*` custom keywords" preamble to BAST.
- `docs/architecture/README.md` — update the document table, ADR table,
and key design principles for the pivot.
- `docs/architecture/overview.md` — update the "what" and "why" for
BAST (the crate now takes a BAST document, not a custom-keyword JSON
Schema).
**Stale TODOs to remove:** any TODO referencing custom-keyword
normalization, `inline_union_variant_refs`, or the rejected
bare-name-ref design — align with the ADRs as AGENTS.md §"Architecture
Context" requires.
## Verification Commands
Run these before committing each step. All must pass. Per AGENTS.md:
```bash
cargo test --release # full suite (~396 tests: 310 crate + 86 integration)
cargo clippy --all-targets -- -D warnings
cargo doc --no-deps # if docs changed (step 8, step 10)
cargo build --target wasm32-unknown-unknown --release # if layout/wasm-relevant code changed (step 2, 4, 5, 8)
cargo publish --dry-run --allow-dirty # before a release (post-step 10)
```
+551
View File
@@ -0,0 +1,551 @@
---
status: implemented
created: 2026-08-14
last_updated: 2026-08-15
---
# BAST Pivot — Research Record
**Status: implemented.** The BAST pivot landed in steps 1–10 of the
[implementation plan](../plans/bast-implementation.md) (commits
`66ab9d7` → `54fd112` on `origin/main`). The decisions D-BAST-001..009
are recorded in two new ADRs —
[ADR-BAST](../architecture/decisions/bast-bast-format.md) (the format)
and [ADR-VAL-SPLIT](../architecture/decisions/val-split-two-validator-model.md)
(the two-validator model) — which supersede the format-specific content
of [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md)
and refine the validation strategy of
[ADR-004](../architecture/decisions/004-error-handling-validation-strategy.md)
and [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md).
The normative format spec is
[`docs/architecture/bast-format.md`](../architecture/bast-format.md);
the parser is documented in
[`docs/architecture/schema-layer.md`](../architecture/schema-layer.md).
What follows is the original research record — the *why* and *what was
proved*, preserved as the historical grounding for the decisions.
---
Replace alktype's custom JSON Schema keywords (`AlkType:Uint32`,
`AlkType:Struct`, etc.) with a standalone JSON format — BAST (Binary
Abstract Syntax Tree) — that describes binary data layouts using a
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
itself a valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser.
The engine's core logic (layout computation, data access, union
dispatch, two layout modes) is unchanged. Only the schema-walking
accessor layer changes: instead of detecting `AlkType:*` keywords
scattered through a JSON Schema tree, the walkers read `kind`/`fields`/
annotation properties from a purpose-built format.
The builder API's public surface stays the same; only the JSON output
format changes internally.
> **Document role.** This is the research record: motivation, POC
> scope and result, decisions, risks, references. The normative format
> specification lives in [`docs/architecture/bast-format.md`](../architecture/bast-format.md).
> The execution plan — ordered implementation steps, the public-API
> semver contract, and the ADR-sync checklist — lives in
> [`docs/plans/bast-implementation.md`](../plans/bast-implementation.md).
> Those documents supersede the format-spec, what-changes, and
> migration-path sections that previously lived here; this record
> keeps the *why* and *what was proved*, not the *how to implement*.
## Motivation
### Current state
alktype v0.1.0 embeds binary layout information inside standard JSON
Schema documents via custom keywords:
```json
{
"AlkType:Struct": true,
"type": "object",
"properties": {
"channel_id": { "AlkType:Uint32": true, "type": "integer" },
"length": { "AlkType:Uint32": true, "type": "integer" }
},
"endian": "big"
}
```
This works for the Rust engine — it walks the tree, detects keywords,
computes offsets. But it creates friction for everything outside Rust:
1. **Cross-language consumption.** A Python, Go, or TypeScript consumer
that wants to parse an alktype schema must re-implement custom keyword
detection. The format is not self-describing — you need to know that
`AlkType:Uint32` means "4-byte unsigned integer" and that it can
appear as either `true` or `{ "encoding": "..." }`.
2. **Code generation.** Generating Rust/TypeScript/Python readers and
writers from a schema requires walking an arbitrary JSON Schema tree
looking for custom keywords. A `kind`-based format with known keys
makes this a straightforward structural walk.
3. **Tooling.** Editors, linters, and schema validators don't understand
`AlkType:*` keywords. A BAST document with a published meta-schema
gets autocomplete, validation, and documentation in any JSON Schema-
aware editor for free.
4. **Two concerns in one document.** The current format conflates binary
layout (what the engine needs) with JSON validation (what jsonschema
needs). A `type: "object"` with `properties` and `required` is a JSON
validation concern; `AlkType:Uint32` is a binary layout concern. They
live in the same JSON object but serve different masters.
### The downstream pain is real
The alkcall agent's review identified that the channels 8-byte chunk
header is hand-rolled with manual bit shifts — alktype's binary layout
capability is unused because the custom-keyword format is awkward to
integrate for a simple 2-field struct. alktty plans to hand-roll its
5-byte TTY chunk format for the same reason. SFTP's 29 packet types
were proven byte-identical with alktype in the POC, but the production
path requires defining 29 schemas in the custom-keyword format.
All three cases are the same pattern: a small binary struct that needs
a schema-driven reader/writer. BAST makes this trivial — a 10-line JSON
file replaces hand-rolled bit shifts.
### Timing
v0.1.0 was published but has zero real consumers (only bots/scanners
have downloaded it). A breaking change now is free. Waiting until
adoption creates migration cost.
## POC Scope
The layout swap needs no POC — it is a backend swap (custom keywords →
`kind`-based format) on top of a proven layout engine. The layout
engine's byte-identity is already proven (alknet-typedef-poc,
alktype-builder-poc) and the layout code is unchanged, so there is
nothing empirical to de-risk there.
One targeted POC **was** needed to de-risk the validation model. The
risk was specific and falsifiable: can a BAST-native validator — a
recursive walker over the BAST type tree — fully replace the 19 custom
keyword validators on the `validate_bytes` path, including the OQ-008
union variant dispatch, without regression?
**The POC has been run and succeeded.** The outcome is recorded in
[POC Result](#poc-result--bast-native-validator) below. The POC code is
on branch `bast-validator-poc` (commit `f371fe4`), not merged to main —
it is reference scaffolding superseded by the production module in
[implementation step 5](../plans/bast-implementation.md#step-5--bast-native-validator-production-version).
### POC: BAST-native validator for `validate_bytes`
**Hypothesis:** A recursive walker over the BAST type tree can enforce
all value-domain constraints that the 19 custom keyword validators
currently enforce, recovering the OQ-008 union variant dispatch
behavior, and fixing the enum-membership dead constraint on the bytes
path — all without `jsonschema` custom keywords and without requiring
the consumer to provide an external JSON Schema.
**Scope:**
1. Implement the BAST-native validator as a new module
(`src/bast_validation.rs` or similar)
2. The validator walks a materialized `Value` tree against the BAST
type definitions, checking:
- Integer ranges (Int8..Uint64)
- Float finiteness (Float32/64)
- String `maxLength` (UTF-8 byte length)
- Bytes `maxLength` (byte length)
- RFC 3339 timestamp shape (non-strict, matching current behavior)
- Enum index bounds (0..values.len()-1 — **fixes the dead constraint**)
- Union variant dispatch (read `__discriminator`, look up variant
BAST definition, recurse)
- Boolean validity (materializer already checks, but the validator
should confirm)
3. Wire it into `validate_bytes()` as the validation step (replacing
the `jsonschema::Validator` call)
4. Run the **existing test suite** — the tests encode all current
expected validation behavior. If they pass, the POC succeeds.
**Success criteria:**
- All existing `validate_bytes` tests pass without modification to their
assertions (test *inputs* will change to BAST format, but the
expected validation outcomes must be identical)
- The union variant dispatch tests (OQ-008) pass — `maxLength` on a
`Bytes` field inside a union variant is enforced
- The enum index-bounds validation works (new behavior — currently
broken, so this is a fix, not a regression)
**Failure path:** If the POC reveals that the BAST-native validator
cannot cleanly express some constraint (e.g., a constraint that relies
on JSON Schema's structural keywords in a way that's hard to
reimplement), the fallback is the "structural-only + external JSON
Schema" model from the original Gap 1 — but this is unlikely given that
the materializer already guarantees structure, leaving only value-domain
checks.
**Out of scope for this POC:**
- `validate_json` — this path uses a standard `jsonschema::Validator`
from a consumer-provided JSON Schema, not the BAST-native validator.
No POC needed; it's a standard `jsonschema` usage.
- BAST document parsing / meta-schema validation — the parser is
straightforward JSON walking; no empirical risk.
- Layout computation — unchanged, already proven.
## POC Result — BAST-native validator
**Status: succeeded.** The POC is on branch `bast-validator-poc` in
`src/bast_poc.rs` (20 tests, all passing; full crate suite — 416 tests —
green; `cargo clippy --all-targets -- -D warnings` clean;
`cargo build --target wasm32-unknown-unknown --release` clean).
The POC implements the `validate_bytes` validation model from
[D-BAST-006](#d-bast-006-validate_bytes-validation-model) as a
self-contained module that does **not** touch the production schema /
materializer / validator paths. It reuses only `data_access` (read
primitives), `AlkTypeError` (error type), and `Endian`. The BAST
document parser, a packed-mode materializer, and the BAST-native
validator are all implemented from scratch — that is the point: prove
the model works end-to-end before refactoring the production code.
### What the POC proves
The hypothesis from the [POC section](#poc-bast-native-validator-for-validate_bytes)
is confirmed: a recursive walker over the BAST type tree fully
replaces the 19 custom keyword validators on the `validate_bytes` path,
including the OQ-008 union variant dispatch, and fixes the
enum-membership dead constraint — all without `jsonschema` custom
keywords and without an external JSON Schema.
The hard cases that were the actual de-risking targets all pass:
- **Union byte-offset discriminator + `maxLength` inside a variant
(OQ-008).** `union_byte_disc_max_length_inside_variant_enforced`
materializes a union with two `$ref` variants, dispatches on a
byte-offset `uint8` discriminator, and enforces `maxLength` on a
`bytes` field inside the selected variant. The validator reads
`__discriminator`, looks up the variant's BAST definition, and
recurses — same behavior as the current `UnionValidator`'s
per-variant sub-validators, but with no `jsonschema` involvement.
- **Union field-name discriminator + `maxLength` inside a variant.**
`union_field_disc_max_length_inside_variant_enforced` covers the
typedef.ts-style discriminator (a length-prefixed string field
selects the variant). Same recursion model.
- **Enum index-bounds fix.** `enum_index_out_of_bounds_rejected`
exercises the constraint that is **broken in the current engine**
(the built-in `enum` keyword checks string membership; the
materializer emits `Value::Number(index)`, which never matches — a
dead constraint). The BAST-native validator checks the materialized
index against the `values` array bounds (0..len-1), which is the
correct validation for a binary enum encoded as an index. Net
improvement, not a regression.
- **Nested struct wrapping a union wrapping a struct.**
`nested_struct_with_union_variant` confirms the recursion composes
through multiple type layers.
- **Arrays of fixed-size structs with `count`.**
`array_of_structs_with_count` covers the `Vector3`-style array
(D-BAST-004).
- **Records (count-prefixed string-keyed maps).**
`record_of_uint32` covers the `TRecord` shape.
- **Untrusted schema input.** `malformed_document_produces_schema_error_not_panic`
confirms a malformed BAST document surfaces as `AlkTypeError::Schema`,
not a panic (AGENTS.md §3).
- **Basic cases** (chunk header, int8/uint32 ranges, string/bytes
`maxLength`, timestamp, bool, short buffer) all pass — if a couple of
basic examples work, all of them do, since the validator is a flat
per-kind dispatch with no per-kind special-casing beyond the range
bounds.
### How the validator works
The validator is a single recursive function (`validate_typeref`) that
dispatches on the BAST `kind`. Each arm checks the value-domain
constraint for that kind and, for composites, recurses into the child
type definitions. The materializer (also implemented in the POC)
guarantees structural correctness — bounds, UTF-8, bool byte,
discriminator lookup, all fields present — so the validator only
enforces what the materializer cannot. The full constraint table is in
[`bast-format.md` §Validation Model](../architecture/bast-format.md#validation-model).
### Observations for the production implementation
1. **No `jsonschema` dependency for `validate_bytes`.** The validator
only needs `serde_json` (for `Value`) and the BAST document. The
`jsonschema` crate is still a direct dependency for `validate_json`
and for validating BAST documents against the BAST meta-schema, but
the `validate_bytes` path no longer touches it. This is a small wasm
binary-size win in addition to the architecture simplification.
2. **`$ref` resolution is a single hash lookup.** The POC's
`resolve_ref_or_inline` handles only `#/$defs/Name` pointers — the
only form BAST allows. The current engine's `normalize_refs` /
`inline_union_variant_refs` / `resolve_ref_or_inline` machinery for
bare-name refs and inlined union variants is no longer needed: BAST
`$ref`s are always full JSON Pointers, and union variant refs are
resolved lazily by the validator (the materializer already does this
for the read path). The `inline_union_variant_refs` compile step can
be removed entirely.
3. **The validator is ~250 lines.** The 19 custom keyword validators
(`src/validation.rs`) plus the macro definitions are ~500 lines and
require the `jsonschema::Keyword` trait plumbing (factory closures,
`Box<dyn Keyword>`, sub-validator construction at factory time). The
BAST-native validator is a flat match — no factories, no trait
objects, no sub-validator pre-computation. The recursion is direct.
4. **The `AlkTypeError::Validation` variant still wraps
`jsonschema::ValidationError<'static>`.** The POC uses
`jsonschema::ValidationError::custom` to construct these so the
error type is unchanged. This is now the decided shape for the
production refactor — see [D-BAST-009](#d-bast-009-alktypeerrorvalidation-payload-shape).
The rationale is consumer ergonomics: a single uniform payload type
means one match arm covers both `validate_json` and `validate_bytes`
errors downstream, and `validate_json`'s structured errors are worth
preserving rather than flattening to a `String`.
5. **The materializer and validator share the BAST-walking code
structure.** Both walk the same `kind`/`fields`/`mapping` tree. The
production refactor could share a typed BAST tree (a small
`BastNode` enum) between them so the walk is parsed once. The POC
parses lazily from the raw JSON in both passes to keep the model
honest; a typed tree is a straightforward follow-on optimization, not
a risk.
### Verdict
The "how do we reproduce the same behavior?" question is answered:
walk the BAST tree the same way the materializer does, checking the
same value-domain constraints the custom keyword validators check
today. The model is a strict simplification — fewer moving parts, no
`jsonschema` integration on the bytes path, no compile-time
`inline_union_variant_refs` step, no factory closures or trait
objects, and the enum dead-constraint is fixed as a side effect.
The POC does not wire into `AlkTypeEngine::validate_bytes` — that is
the production refactor ([implementation step 5](../plans/bast-implementation.md#step-5--bast-native-validator-production-version)),
which replaces `validation::build_validator` usage on the bytes path
with the BAST-native validator. The POC's job was to de-risk the model
before that refactor; that job is done.
## Decisions
The following were open questions in earlier drafts. Each is now
resolved. They are recorded here as decisions, not re-litigated. The
normative format specification that realizes these decisions is in
[`bast-format.md`](../architecture/bast-format.md); the semver
classification of each is in the
[implementation plan's Semver Contract](../plans/bast-implementation.md#semver-contract).
### D-BAST-001: Root type selection
**Decision:** Explicit. The root type name is a required parameter to
`compile()`: `AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed)`.
This is unambiguous and matches how consumers think about it ("compile
the ChunkHeader schema"). Convention (first entry in `$defs`) is fragile
and depends on JSON key order; a `$root` marker is redundant with an
explicit parameter.
### D-BAST-002: Primitive type string set
**Decision:** Lowercase (`"uint32"`, `"int8"`, `"float64"`, `"bool"`,
`"string"`, `"bytes"`, `"timestamp"`). Matches JSON Schema's own
convention (`"string"`, `"integer"`, `"boolean"`), is easier to type,
and is the convention in the TypeBox research examples. The
`AlkTypeKind` enum variants remain PascalCase in Rust — the mapping is
a simple `from_str()` impl.
### D-BAST-003: Top-level `$defs` requirement
**Decision:** Always `$defs`. Every BAST document has the same
top-level shape: `{ "$defs": { ... } }`. Single-type documents are a
special case with one entry. The `$defs` block is the namespace; the
root type name (D-BAST-001) selects the entry point. A bare struct at
the top level would be a special case with different parsing logic and
no home for additional definitions.
### D-BAST-004: Arrays of variable-length elements (deferred)
**Decision:** Arrays of variable-length elements are **not supported
in v1**. The meta-schema requires `count` on all array types, making
arrays fixed-size only. This matches the engine's current behavior (it
rejects arrays of variable-length elements) and aligns with OQ-001
(deferred, blocked on a concrete consumer needing interleaved
variable-stride arrays).
Variable-length collections are still available via `record` (a
count-prefixed string-keyed map), which the engine supports. If a
consumer needs a variable-length array of fixed-size elements, they
can use a record with integer-stringified keys as a workaround, or
wait for OQ-001 to be addressed.
### D-BAST-005: Field-name discriminator unions
**Decision:** Supported. The meta-schema includes an optional `fields`
array on `UnionDef`. When `discriminator.kind == "field"`, the
`fields` array provides the union's field definitions (including the
discriminator field). When `discriminator.kind == "byte"`, `fields` is
absent — the union's layout is purely the variant layout. This
preserves a feature the engine already supports. The meta-schema is in
[`bast-format.md` §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
### D-BAST-006: `validate_bytes` validation model
**Decision:** BAST-native validator. The `validate_bytes` path uses a
recursive walker over the BAST type tree to check value-domain
constraints on the materialized `Value` — no external JSON Schema
needed. This recovers the OQ-008 union variant dispatch behavior (the
validator recurses into the variant's BAST definition) and fixes the
enum-membership dead constraint (the validator checks the materialized
index against the `values` array bounds). See [`bast-format.md` §Validation
Model](../architecture/bast-format.md#validation-model) and the [POC](#poc-bast-native-validator-for-validate_bytes).
An optional external JSON Schema can be layered on top for constraints
BAST doesn't express (cross-field consistency, regex patterns on string
content). This is additive, not load-bearing.
### D-BAST-007: `validate_json` validation model
**Decision:** Standard JSON Schema validator. `validate_json` on
`AlkTypeEngine` validates a consumer-provided JSON `Value` against a
`jsonschema::Validator` compiled from a standard JSON Schema document
the consumer provides at compile time. No custom keywords. The BAST
document is not involved in this path — BAST describes bytes, not JSON
shape. This preserves the `validate_json` / `validate_bytes` symmetry
from ADR-010, but the two paths now use different validators (standard
`jsonschema` for JSON, BAST-native for bytes), reflecting their
different inputs and guarantees.
The JSON Schema may be authored separately or derived from BAST via
future codegen. For alkcall's channel 0 (JSON-RPC), the JSON Schema is
the `OperationSpec` schema, authored independently of any BAST
document.
### D-BAST-008: Builder API — two output formats
**Decision:** One builder, two build methods. The construction API is
the same (field names, types, annotations); only the output format
differs. `Schema::struct_().field(...).build()` → BAST JSON (binary
layout). `Schema::object().field(...).build()` → standard JSON Schema
(JSON validation). The builder already distinguishes AlkType kinds from
JSON Schema types via naming conventions (`string()` vs `string_()`).
Both output formats live in the same crate. This is the point of
alktype: one small wasm-compatible codebase that handles both binary
layout and JSON validation for protocol crates. alkcall uses both —
channel 0 is JSON (standard JSON Schema), binary channels use BAST.
Future crates (alktty, tunnels, sftp, git) will use BAST for their
binary formats. The codegen feature (future) will generate
readers/writers from BAST documents for these crates.
### D-BAST-009: `AlkTypeError::Validation` payload shape
**Status: decided.** Keep `Validation(jsonschema::ValidationError<'static>)`.
`AlkTypeError::Validation` currently wraps
`jsonschema::ValidationError<'static>`. Under the BAST pivot the
`validate_bytes` path no longer uses `jsonschema` at all (confirmed by
the [POC](#poc-result--bast-native-validator) — observation 1), so the
error payload on that path is constructed via
`jsonschema::ValidationError::custom` purely to keep the variant's type
unchanged. The two options were:
1. **Keep `Validation(jsonschema::ValidationError<'static>)`.** Simplest —
`ValidationError::custom` is public and `'static`, so the bytes path
can construct it without a real `jsonschema` validator. Cost: the
error type retains its `jsonschema` dependency even though the bytes
path no longer drives it. `validate_json` still uses `jsonschema`, so
the dependency isn't removable either way — but the error type
carries `jsonschema` only for one of its two callers.
2. **Introduce `Validation(String)` (or a small structured payload).**
Drops the `jsonschema` type from the public error enum. This is a
**semver-relevant public-API change** (the `Validation` variant's
payload type changes), so per AGENTS.md it requires an explicit
decision, not a drive-by. Benefit: the error type is
`jsonschema`-free, which matters if a future `no_std`/minimal build
wants to drop `jsonschema` from the bytes-only path (relates to
OQ-002).
**Rationale for option 1:** The deciding factor is consumer ergonomics
on the *combined* path. Consumers like alkcall use both `validate_json`
(channel 0, JSON-RPC) and `validate_bytes` (binary channels) and handle
`AlkTypeError::Validation` in one place. A single uniform payload type
means one match arm covers both sources — no `Validation(jsonschema_err)
vs Validation(string)` branching downstream. Option 2 would force
`validate_json` to flatten its structured errors (instance path, schema
path, keyword) to a `String` via `Display` just to match a bytes-path
shape — the more information-rich path loses data to accommodate the
less rich one. That is the wrong direction.
The `no_std`/minimal-build angle (OQ-002) that option 2 was meant to
enable is moot in practice: `validate_json` requires `jsonschema`
regardless, so a bytes-only `no_std` build already has to give up
`validate_json` as a separate, larger decision. Dropping the type from
one error variant does not unlock that build — the dependency is load-
bearing on the other validation path. The right place to revisit this is
when/if OQ-002 is actually pursued, not preemptively.
The POC already used option 1 (via `ValidationError::custom`); the
production refactor ([implementation step 5](../plans/bast-implementation.md#step-5--bast-native-validator-production-version))
follows the same construction pattern. No semver-relevant change to the
`Validation` variant.
## Risks and Mitigations
| Risk | Mitigation |
|------|-----------|
| BAST format doesn't cover all 19 type kinds | The format is designed to cover all 19. The meta-schema is the spec — if a kind can't be expressed, the meta-schema is wrong. |
| `$ref` resolution complexity moves from engine to schema authoring | BAST `$ref` values are always full JSON Pointers (`#/$defs/Name`). No normalization, no bare names. Resolution is a single hash lookup. |
| BAST-native validator misses a constraint the custom keywords enforced | The [POC](#poc-bast-native-validator-for-validate_bytes) runs the existing test suite, which encodes all current expected validation behavior. If a constraint is missed, a test fails before the pivot lands. |
| Losing OQ-008 union variant dispatch | The BAST-native validator recurses into the variant's BAST definition on `__discriminator` lookup — same behavior, no custom keywords. Covered by the POC. |
| Enum membership broken on bytes path | Already broken today (dead constraint). The BAST-native validator fixes it by checking the materialized index against the `values` array bounds. Net improvement. |
| `validate_json` loses custom keyword validation | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema. Consumers that relied on custom keywords for JSON validation need to provide equivalent standard JSON Schema keywords. No real consumers exist yet. |
| Builder API output format change breaks consumers | No real consumers exist yet (v0.1.0 has zero adoption). The builder's public methods are unchanged; only the JSON output format changes. |
| Meta-schema maintenance burden | The meta-schema is small and changes rarely. It's embedded in the crate and published at a stable URL. |
## Future Directions
These are enabled by BAST but out of scope for the pivot itself. They
are mentioned to show that BAST makes them possible, not to commit to a
specific implementation timeline.
### Codegen
BAST enables code generation that the custom-keyword format makes
awkward. A codegen module (feature-gated behind `codegen`) would walk
`$defs` entries, inspect `kind` values, and emit Rust/TypeScript/Python
readers and writers via Handlebars templates. The typebox-rs `codegen/`
module (`/workspace/@alkimiadev/typebox-rs`) is the reference
architecture: `SchemaRegistry` for named types, `RustGenerator`/
`TypeScriptGenerator` wrapping `Handlebars`, `schema_to_rust_type()`/
`schema_to_ts_type()` mapping functions. alktype's codegen would follow
the same pattern but walk BAST `kind` values. The handlebars-rs
dependency is WASM-compatible. The pivot changes the schema format;
codegen builds on top of the new format.
### ABI Adapter
A BAST document describes the binary interface of a protocol — it is
essentially an ABI specification in JSON. This enables version
negotiation (two peers exchange BAST documents to agree on a protocol
version; the engine detects mismatches and either rejects or adapts),
schema migration (a consumer with schema v1 can read data written by
schema v2 if the changes are compatible), and WASM interop (a WASM
component can export its BAST schema as part of its WIT interface).
## References
- [`docs/architecture/bast-format.md`](../architecture/bast-format.md) —
the normative BAST format spec (meta-schema, TypeRef, examples,
validation model)
- [`docs/plans/bast-implementation.md`](../plans/bast-implementation.md) —
the execution plan (ordered steps, semver contract, ADR-sync checklist)
- [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) — current "schema is the format" principle (to be superseded by ADR-BAST post-implementation)
- [ADR-003](../architecture/decisions/003-schema-annotations.md) — annotation semantics (carry forward to BAST unchanged)
- [ADR-009](../architecture/decisions/009-builder-api.md) — builder API (public surface unchanged, output format changes)
- [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md) — `validate_bytes` (unchanged in concept)
- `/workspace/research/typebox_research/ujsx/jpath.gen.ts` — TypeBox `Type.Module` pattern (the `$defs`/`$ref` model BAST follows)
- `/workspace/research/typebox_research/ujsx/mdast.gen.ts` — TypeBox cross-module references and composite types
- `/workspace/research/typebox_research/codegen/ts-to-module.ts` — TypeScript-to-TypeBox codegen (reference for future BAST codegen)
- `/workspace/@alkimiadev/typebox-rs/src/codegen/` — Rust/TypeScript codegen from schemas (reference architecture)
- `/workspace/alknet-typedef-poc/tests/sftp_roundtrip_test.rs` — SFTP POC proving byte-identical output (to be replicated with BAST)
- `/workspace/@alkdev/alkcall/src/channels/wire.rs` — hand-rolled chunk header (target for BAST replacement)
+419
View File
@@ -0,0 +1,419 @@
---
status: open
last_updated: 2026-08-15
reviewed_artifacts:
- src/lib.rs
- src/bast.rs
- src/bast_meta.rs
- src/bast_validation.rs
- src/builder.rs
- src/data_access.rs
- src/engine.rs
- src/error.rs
- src/layout_builder.rs
- src/materialize.rs
- src/offset_map.rs
- src/schema.rs
- src/sequential_reader.rs
- src/tunion.rs
- src/validation.rs
- src/macros.rs
- tests/{engine_integration,error_paths,poc_roundtrip,tunion_dispatch}.rs
- Cargo.toml
tool: manual source read + cargo test/clippy + cargo-llvm-cov
reviewer: post-BAST-pivot code review
---
# Code Review #003 — Post-BAST-Pivot Review
## Purpose
First logic/correctness review after the BAST pivot (the v0.1.0
`AlkType:*` custom-keyword JSON Schema format was replaced with the BAST
format; see `docs/plans/bast-implementation.md`). The pivot touched
every schema-walking path, so this pass re-reads the whole crate for
correctness, code smell, panic safety, and coverage — the same scope as
review #002, but against the new BAST surface.
Two things motivated this review beyond the routine sweep:
1. The pivot was a large, multi-step change (10 steps, 8 commits). A
couple of pre-existing bugs were fixed *during* the pivot (the enum
index-bounds dead constraint, the `write_bytes` u32 truncation), so
the same class of bug could be lurking in the newly-rewritten paths.
2. The publisher asked specifically for a coverage pass
(`cargo-llvm-cov`) with an eye toward *important* things being
covered rather than raw numbers.
## Methodology
- Full read of all 16 `src/*.rs` files (production + test modules) and
all 4 integration test files.
- `cargo test --release`, `cargo clippy --all-targets -- -D warnings`.
- `cargo llvm-cov --release` (summary + per-file + uncovered-lines) to
attribute coverage gaps to specific code paths.
- Targeted reproduction of the suspicious paths (field-level endian
override, aligned-mode variable-length/array materialization) via
throwaway integration tests.
- Cross-reference every error path against its caller to confirm errors
propagate (not swallowed) and carry useful attribution.
- Read `docs/reviews/002-code-review.md` for prior context and
resolved/unresolved items.
## Verification Baseline
All verification run on the reviewed tree (commit `562284f`):
- `cargo test --release`: **389 tests pass** (312 crate unit tests +
77 integration tests across 4 files). Zero failures.
- `cargo clippy --all-targets -- -D warnings`: **clean**.
- `cargo llvm-cov --release`: **90.14% line coverage** (7903/8682),
**86.68% function coverage** (743/842). Per-file breakdown below.
- No `unsafe` anywhere in the crate.
- No `TODO`/`FIXME`/`HACK`/`XXX` markers in source.
- All `unwrap`/`expect`/`panic!`/`unreachable!` are confined to
`#[cfg(test)]` modules, verified by line-context cross-reference.
### Coverage breakdown
| Module | Lines | Functions |
|---|---:|---:|
| bast.rs | 88.3% | 90.3% |
| bast_validation.rs | 93.0% | 90.5% |
| builder.rs | 91.3% | 87.6% |
| data_access.rs | 83.5% | 76.2% |
| engine.rs | 96.9% | 98.3% |
| layout_builder.rs | 91.6% | 81.8% |
| materialize.rs | **81.8%** | **76.1%** |
| offset_map.rs | 89.5% | 79.2% |
| sequential_reader.rs | **84.6%** | **75.0%** |
| tunion.rs | 92.7% | 91.7% |
| **TOTAL** | **90.1%** | **86.7%** |
The low-function-count modules are not test-helper noise — they are
exactly where the correctness bugs below live. The uncovered lines in
`materialize.rs` and `sequential_reader.rs` are the aligned-mode
variable-length/array paths and the field-level-endian paths, which are
**untested and broken** (see M1, M2). The `data_access.rs` 76% function
coverage is mostly the `read_*_indirect` family, which has no production
caller (see L1).
## Summary Statistics
| Severity | Count |
|----------|------:|
| Critical | 0 |
| Medium | 3 (M1, M2, M3) |
| Low | 2 (L1, L2) |
| Nit | 4 (N1, N2, N3, N4) |
No critical findings. The crate is in good shape, but the three Medium
findings are **silent data-corruption / silent-misinterpretation bugs**
in the newly-rewritten paths — they must be fixed before the next
release. The Low findings are dead code and a robustness gap; the Nits
are hygiene.
---
## Findings
### M1. Field-level `endian` override is ignored by the reader and aligned materializer
**Files**: `src/sequential_reader.rs:290`, `src/engine.rs:338,455`,
`src/materialize.rs:461`
**Problem**: The spec documents per-field endian override
(`docs/architecture/bast-format.md` §Endianness — "Field-level `endian`
overrides the struct/union default"), and the packed materializer honors
it (`materialize.rs:120` uses `field.effective_endian(endian)`). But
three paths use only the struct-level endian:
- `sequential_reader.rs:290` `read_field_value` — uses `self.endian`,
never `field.effective_endian`.
- `engine.rs:338` `read_field` and `engine.rs:455` `write_field` —
`let endian = self.endian;`.
- `materialize.rs:461` `materialize_struct_aligned` — passes the struct
`endian` to `materialize_leaf_at`, never `field.effective_endian`.
**Reproduction** (throwaway integration test, confirmed): a big-endian
struct with a `"crc": { "kind": "uint32", "endian": "little" }` field
reads `0x01020304` as `67305985` (big-endian interpretation) in both
`sequential_reader` and `read_field`. The bytes are correct; the
interpretation is wrong — silent data corruption.
**Fix**: thread `field.effective_endian(endian)` through all three paths.
`read_field_value` already receives the `BastField`; `engine::read_field`
/ `write_field` need to look up the field's effective endian (they
already walk the BAST tree via `lookup_field_kind`); `materialize_struct_aligned`
needs to pass `field.effective_endian(endian)` to `materialize_leaf_at`
instead of the struct default.
**Lift**: closes a silent-corruption path on a documented feature. Small
effort (~10 lines + regression tests).
---
### M2. Aligned-mode `validate_bytes` is broken for arrays, `maxLength` fields, and `offset-indirect` fields
**File**: `src/materialize.rs:461-512` (`materialize_struct_aligned`)
**Problem**: `materialize_struct_aligned` routes fixed-size and
variable-length leaves through `materialize_leaf_at`, which calls
`materialize_typeref_packed` — i.e. it always reads a **length-prefixed**
value. But the aligned `OffsetMap` stores three different shapes:
- `maxLength` fields are a raw reservation (no length prefix) — the
materializer reads the first 4 bytes of the *data* as a length prefix.
Reproduced: `"hello"` in an 8-byte reservation read a length of
`1819043180` and failed with a bounds error.
- `offset-indirect` fields are an 8-byte `{offset, length}` pair — read
as a length prefix, garbage.
- arrays are recorded as `vals[0]`/`vals[1]` entries only, so
`offset_map.get("vals")` returns `None` → `Offset` error. Reproduced.
Only the default inline length-prefixed variable field (and only as the
final field, per ADR-006) works in aligned mode. There are **no tests**
covering aligned `validate_bytes` with arrays or non-default variable
encodings — that is why this slipped through the pivot.
**Fix**: two options, decide with the publisher:
1. **Implement** aligned materialization for the three shapes: read
`maxLength` fields as a fixed-size slice, `offset-indirect` fields via
`read_*_indirect` (with a data region), and arrays by iterating the
`vals[i]` offset-map entries.
2. **Reject** these combinations at compile time (return
`AlkTypeError::Schema`/`Offset` from `OffsetMap::compute` or
`compile`) if they are out of scope for v1, so the failure is loud
and at load time rather than a silent misread at access time.
Option 2 is the smaller, safer fix and matches the existing ADR-006/
ADR-008 pattern of rejecting unsupported aligned-mode combinations. The
`offset-indirect` encoding is already dead code on the read path (see
L1), which argues for rejecting it in aligned mode until it is actually
implemented.
**Lift**: closes a silent-misread path. Medium effort either way.
---
### M3. The BAST meta-schema is never applied at compile time; annotation parsers silently tolerate malformed values
**Files**: `src/engine.rs:123` (`compile`), `src/bast.rs:899-926`
(`parse_endian_opt`, `parse_align`, `parse_encoding`, `parse_max_length`)
**Problem**: `BAST_META_SCHEMA` is exported and self-tested, but
`AlkTypeEngine::compile` never validates the document against it. The
parser (`bast.rs`) is the only gate, and it silently tolerates malformed
annotations:
- `parse_endian_opt` (`bast.rs:899`) — `"endian": "middle"` → `None` →
silently defaults to little.
- `parse_encoding` (`bast.rs:921`) — unknown encoding → silently
`LengthPrefixed`.
- `parse_align` / `parse_max_length` — non-integer / negative → silently
`None`.
These are exactly the cases the meta-schema's `enum` / `minimum`
constraints exist to reject. Per AGENTS.md §3, schemas are untrusted
input (the `alkcall` consumer accepts them from arbitrary internet
peers). A malicious peer can send `"endian": "bogus"` and get a
silently-misinterpreted layout instead of a `Schema` error.
**Fix**: validate the document against `BAST_META_SCHEMA` in `compile`
(one-time, cheap — the meta-schema is a `LazyLock<Value>`), *or* make
the annotation parsers return `Err(AlkTypeError::Schema)` on unknown
values. The meta-schema route is preferred: it is the single source of
truth and catches the whole class of malformed-annotation bugs at once.
**Lift**: closes a silent-misinterpretation path on untrusted input.
Small effort (~5 lines + tests).
---
### L1. `offset-indirect` is dead code on the read path
**Files**: `src/data_access.rs:307-355`, `src/engine.rs:388-399`
**Problem**: `data_access::read_string_indirect` / `read_bytes_indirect`
have no production caller (only their own unit tests). `engine.read_field`
always calls `read_string` / `read_bytes` (length-prefixed) regardless of
the field's `encoding`. So a schema declaring
`"encoding": "offset-indirect"` compiles and lays out correctly in the
offset map, but can never be read back.
This is the same root cause as M2's `offset-indirect` arm. Decide
together with M2: either wire `read_*_indirect` into the read path (and
the aligned materializer), or drop the `offset-indirect` encoding
entirely until a consumer needs it. Leaving it half-wired is the worst
state — it looks supported but silently misreads.
**Lift**: removes dead code or completes a feature. Small effort.
---
### L2. `materialize_packed` / `materialize_aligned` take a dead `endian` parameter
**File**: `src/materialize.rs:44-88`
**Problem**: both functions take `endian: Endian` and immediately
`let _ = endian;`, using `struct_node.endian()` instead. The caller's
`self.endian` (from `engine.rs:285,287`) is ignored. The signature is
misleading — a reader assumes the passed endian is honored.
**Fix**: drop the parameter and read the endian from the root struct
inside the function (it already does). ~4 lines. Purely a clarity fix;
no behavior change.
---
### N1. `number_from_f64` maps NaN/Inf to `Value::Null`
**File**: `src/materialize.rs:451-455`
**Problem**: `serde_json::Number::from_f64` returns `None` for NaN/Inf,
so `number_from_f64` substitutes `Value::Null`. A NaN float in the buffer
then surfaces as "expected a number" from `validate_float`
(`bast_validation.rs:176`) rather than "expected a finite number". The
error is misleading, though the outcome (rejection) is correct.
**Fix** (optional): have the materializer propagate a non-finite float
as an `AlkTypeError::Access` at read time, or leave as-is and accept the
slightly-off error message. Not a correctness bug.
---
### N2. `BastType::alk_kind()` returns `Struct` for any `$ref`
**File**: `src/bast.rs:715-725`
**Problem**: `BastType::Ref(_) => AlkTypeKind::Struct` is documented but
a footgun — a `$ref` to a union/enum misreports its kind unless the
caller resolves first. Most callers do resolve first, but the invariant
is fragile and easy to break in a future edit.
**Fix** (optional): leave as-is (documented) or make `alk_kind` return
`Option<AlkTypeKind>` / require resolution. Defer unless it bites.
---
### N3. `check_bytes` accepts both `String` and `Array` forms
**File**: `src/bast_validation.rs:213-256`
**Problem**: `check_bytes` handles `Value::String` and `Value::Array`,
but the materializer only ever emits `Array` for bytes
(`materialize.rs:212`). The `String` arm is dead/legacy. Harmless, but
it widens the accepted surface for no reason.
**Fix** (optional): drop the `String` arm, or keep it if a future
materializer emits bytes as a string. Defer.
---
### N4. Stale ADR references in doc comments
**Files**: `src/error.rs:3` ("ADR-098"), `src/tunion.rs:1` ("ADR-097"),
`src/engine.rs:51,196` ("ADR-101"), `src/offset_map.rs:1`,
`src/layout_builder.rs:1`, `src/sequential_reader.rs:1` ("ADR-096")
**Problem**: none of these ADR numbers exist in
`docs/architecture/decisions/` (which has 001–010 + `bast-bast-format` +
`val-split-two-validator-model`). The pivot renumbered/renamed ADRs but
the code comments were not synced. A reader following the reference hits
a dead end.
**Fix**: map each stale reference to the correct ADR (e.g. "ADR-096" →
ADR-002 for the two layout modes, "ADR-101" → ADR-007 for the packed
read factory, "ADR-098" → ADR-004 for error handling, "ADR-097" →
ADR-003 for annotations) and update the comments. ~6 lines.
---
## The `Timestamp` kind
`AlkTypeKind::Timestamp` is a first-class kind that is byte-identical to
`String` everywhere (length-prefixed UTF-8), and its only distinguishing
behavior is `is_rfc3339_timestamp` (`bast_validation.rs:406`) — a
hand-rolled, non-strict check that the doc itself admits "Feb 31 passes;
seconds range isn't checked; leap seconds aren't handled." The parsing
is fragile (the `rfind('-')` timezone-offset heuristic, no
fractional-second handling).
It is documented as matching v0.1.0, so it is not a regression, but it
is the weakest part of the validator and adds a 19th kind plus a
`needs_endian` / `is_variable_length` / `natural_alignment` arm, all to
validate a string that a consumer could validate with a standard JSON
Schema `format: "date-time"` on the `validate_json` path.
**Decision (publisher)**: remove it. It is a residual from an early
research reference that included a timestamp; it is largely irrelevant
at the BAST level, and JSON-level timestamp validation is `jsonschema`'s
job, not alktype's. Tracked as a follow-up task, not part of this
review's findings.
---
## What's Good
The crate is in notably good shape after the pivot. Highlights:
- **The BAST parser is clean and defensive.** `bast.rs` returns
`AlkTypeError::Schema` on every malformed-document path, never panics,
and uses `checked_add` / `usize::try_from` for all count/offset casts.
The typed tree (`BastDoc`/`BastDef`/`BastType`) is a real improvement
over the v0.1.0 raw-JSON accessors.
- **The enum index-bounds fix is correct.** `validate_enum`
(`bast_validation.rs:274`) checks the materialized index against
`values.len()`, closing the v0.1.0 dead constraint. Well-tested.
- **Overflow safety is thorough.** `checked_add` everywhere in the hot
paths; the `write_bytes` u32 truncation from review #002 (M2) is
fixed and the guard pattern is now the norm.
- **Error attribution is excellent.** Every `Access`/`Offset` error
carries a `field_path`; the BAST parser errors carry a dotted path
into the document (`"bast: struct at .fields[2] ..."`).
- **The two-validator split is clean.** `bast_validation` (bytes) and
`validation` (JSON) are clearly separated, and the
`AlkTypeError::Validation` payload stays uniform across both
(D-BAST-009).
- **Tests are strong where they exist.** 389 tests, good coverage of
error paths, both endiannesses, short buffers, unknown discriminators,
invalid UTF-8. The gaps are precisely the paths M1/M2 identify.
- **No `unsafe`, no `TODO`/`FIXME`** — clean codebase hygiene.
---
## Recommended Order
1. **M1** (field-level endian override) — ~10 lines + tests, closes a
silent-corruption path on a documented feature. Smallest and
highest-value.
2. **M2** (aligned-mode materialization) — decide implement-vs-reject
with the publisher; the reject option is small and matches the
ADR-006/ADR-008 pattern.
3. **M3** (meta-schema at compile time) — ~5 lines + tests, closes a
silent-misinterpretation path on untrusted input.
4. **L1** (offset-indirect dead code) — decide together with M2.
5. **L2** (dead `endian` parameter) — ~4 lines, clarity only.
6. **N1–N4** — hygiene; N4 (stale ADR refs) is worth doing in the same
pass as the `Timestamp` removal since both touch doc comments.
The `Timestamp` removal is a separate, self-contained task the publisher
has already decided on; it can be done independently of the above.
---
## Notes
- All line numbers refer to the tree at commit `562284f` (the last
commit on `main` at review time).
- The coverage numbers are from `cargo llvm-cov --release` on the same
tree. The `--summary-only` and `--show-missing-lines` outputs were
used to attribute gaps; the full HTML report is at
`target/llvm-cov/html`.
- This review does not cover documentation quality (README, inline docs,
docs.rs rendering) beyond the stale-ADR-reference nit (N4). Per the
publisher's workflow, that is a separate sweep.
- Findings M1 and M2 were confirmed by throwaway integration tests that
were removed after reproduction; the regression tests for the fixes
should be added to the permanent suite.
+1761
View File
File diff suppressed because it is too large. Load diff
+521
View File
@@ -0,0 +1,521 @@
//! BAST meta-schema — the standard JSON Schema (Draft 2020-12) that
//! validates the *structure* of BAST documents (is a document
//! well-formed?).
//!
//! This is distinct from the BAST-native validator (step 5), which
//! validates *binary data* against a BAST document (are the bytes a
//! valid instance?). See
//! [`docs/architecture/bast-format.md`](../docs/architecture/bast-format.md)
//! for the normative spec.
//!
//! The meta-schema is embedded at compile time via the `serde_json::json!`
//! macro — no I/O, no feature flags, wasm-clean. It is also published at
//! `https://alk.dev/bast/v1/schema` (the `$id`). Consumers and editors can
//! validate BAST documents against it with any standard JSON Schema
//! validator:
//!
//! ```ignore
//! jsonschema::options()
//! .build(&alktype::BAST_META_SCHEMA)?
//! .validate(&bast_doc)?;
//! ```
use serde_json::{json, Value};
use std::sync::LazyLock;
/// The BAST v1 meta-schema, as a `serde_json::Value`.
///
/// A BAST document is valid against this schema iff it is well-formed
/// (correct `$defs` shape, known `kind` strings, required properties
/// present, no additional properties). Value-domain constraints
/// (`maxLength`, enum index bounds, etc.) are enforced by the
/// BAST-native validator, not this meta-schema.
///
/// Built lazily on first access via `serde_json::json!` (the `json!`
/// macro allocates, so it can't be a `const`). The parsed `Value` is
/// then reused for every subsequent call — `&'static` via `LazyLock`.
pub static BAST_META_SCHEMA: LazyLock<Value> = LazyLock::new(|| {
json!({
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://alk.dev/bast/v1/schema",
"title": "Binary Abstract Syntax Tree (BAST) v1",
"description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
"type": "object",
"properties": {
"$defs": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeDef" }
}
},
"required": ["$defs"],
"$defs": {
"TypeDef": {
"oneOf": [
{ "$ref": "#/$defs/StructDef" },
{ "$ref": "#/$defs/UnionDef" },
{ "$ref": "#/$defs/EnumDef" }
]
},
"StructDef": {
"type": "object",
"properties": {
"kind": { "const": "struct" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
}
},
"required": ["kind", "fields"],
"additionalProperties": false
},
"FieldDef": {
"type": "object",
"properties": {
"name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
"kind": { "$ref": "#/$defs/TypeRef" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"encoding": { "enum": ["length-prefixed", "offset-indirect"] },
"maxLength": { "type": "integer", "minimum": 0 }
},
"required": ["name", "kind"],
"additionalProperties": false
},
"TypeRef": {
"oneOf": [
{
"description": "Primitive type",
"type": "string",
"enum": [
"int8", "int16", "int32", "int64",
"uint8", "uint16", "uint32", "uint64",
"float32", "float64",
"bool", "string", "bytes"
]
},
{
"description": "Reference to a named $defs entry",
"type": "object",
"properties": {
"$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
},
"required": ["$ref"],
"additionalProperties": false
},
{
"description": "Array type (fixed-size only in v1 — count is required)",
"type": "object",
"properties": {
"kind": { "const": "array" },
"element": { "$ref": "#/$defs/TypeRef" },
"count": { "type": "integer", "minimum": 0 }
},
"required": ["kind", "element", "count"],
"additionalProperties": false
},
{
"description": "Record (string-keyed map) type",
"type": "object",
"properties": {
"kind": { "const": "record" },
"values": { "$ref": "#/$defs/TypeRef" }
},
"required": ["kind", "values"],
"additionalProperties": false
},
{
"description": "Inline struct type",
"$ref": "#/$defs/StructDef"
},
{
"description": "Inline union type",
"$ref": "#/$defs/UnionDef"
},
{
"description": "Inline enum type",
"$ref": "#/$defs/EnumDef"
}
]
},
"UnionDef": {
"type": "object",
"properties": {
"kind": { "const": "union" },
"endian": { "enum": ["little", "big"] },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
},
"discriminator": {
"oneOf": [
{
"type": "object",
"properties": {
"kind": { "const": "byte" },
"offset": { "type": "integer", "minimum": 0 },
"type": { "enum": ["uint8", "uint16", "uint32"] }
},
"required": ["kind", "offset", "type"],
"additionalProperties": false
},
{
"type": "object",
"properties": {
"kind": { "const": "field" },
"name": { "type": "string" }
},
"required": ["kind", "name"],
"additionalProperties": false
}
]
},
"mapping": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeRef" }
}
},
"required": ["kind", "discriminator", "mapping"],
"additionalProperties": false
},
"EnumDef": {
"type": "object",
"properties": {
"kind": { "const": "enum" },
"values": {
"type": "array",
"items": { "type": "string" },
"minItems": 1
}
},
"required": ["kind", "values"],
"additionalProperties": false
}
}
})
});
/// Validate a BAST document against the meta-schema.
///
/// Returns `Ok(())` if the document is well-formed, or
/// [`crate::error::AlkTypeError::Schema`] with the first validation error
/// otherwise.
/// This is the structural gate [`crate::AlkTypeEngine::compile`] applies
/// before parsing — it catches malformed annotations (unknown `endian`/
/// `encoding` strings, non-integer `align`/`maxLength`, missing required
/// properties) that the parser would otherwise silently tolerate.
pub fn validate_bast_doc(doc: &Value) -> Result<(), crate::error::AlkTypeError> {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.map_err(|e| {
crate::error::AlkTypeError::Schema(format!("BAST meta-schema failed to build: {e}"))
})?;
validator.validate(doc).map_err(|e| {
crate::error::AlkTypeError::Schema(format!("BAST document is not well-formed: {e}"))
})
}
#[cfg(test)]
mod tests {
use super::BAST_META_SCHEMA;
#[test]
fn meta_schema_has_correct_id_and_draft() {
assert_eq!(
BAST_META_SCHEMA["$schema"],
"https://json-schema.org/draft/2020-12/schema"
);
assert_eq!(
BAST_META_SCHEMA["$id"],
"https://alk.dev/bast/v1/schema"
);
}
#[test]
fn meta_schema_requires_defs() {
let required = BAST_META_SCHEMA["required"].as_array().expect("required is array");
assert!(required.iter().any(|v| v == "$defs"));
}
#[test]
fn meta_schema_covers_struct_union_enum_type_defs() {
let type_def_oneof = BAST_META_SCHEMA["$defs"]["TypeDef"]["oneOf"]
.as_array()
.expect("TypeDef.oneOf is array");
let refs: Vec<&str> = type_def_oneof
.iter()
.map(|v| v["$ref"].as_str().expect("ref string"))
.collect();
assert!(refs.contains(&"#/$defs/StructDef"));
assert!(refs.contains(&"#/$defs/UnionDef"));
assert!(refs.contains(&"#/$defs/EnumDef"));
}
#[test]
fn meta_schema_typeref_enumerates_all_13_primitives() {
let typeref_oneof = BAST_META_SCHEMA["$defs"]["TypeRef"]["oneOf"]
.as_array()
.expect("TypeRef.oneOf is array");
let primitive_arm = typeref_oneof
.iter()
.find(|v| v["description"] == "Primitive type")
.expect("primitive arm present");
let primitives = primitive_arm["enum"]
.as_array()
.expect("primitive enum is array");
let strs: Vec<&str> = primitives.iter().map(|v| v.as_str().unwrap()).collect();
assert_eq!(strs.len(), 13);
for expected in [
"int8", "int16", "int32", "int64",
"uint8", "uint16", "uint32", "uint64",
"float32", "float64",
"bool", "string", "bytes",
] {
assert!(strs.contains(&expected), "missing primitive {expected}");
}
}
#[test]
fn meta_schema_validates_minimal_bast_document() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Point": {
"kind": "struct",
"fields": [
{ "name": "x", "kind": "uint32" },
{ "name": "y", "kind": "uint32" }
]
}
}
});
assert!(validator.validate(&doc).is_ok(), "valid doc rejected");
}
#[test]
fn meta_schema_rejects_missing_defs() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({ "type": "object" });
assert!(validator.validate(&doc).is_err(), "missing $defs accepted");
}
#[test]
fn meta_schema_rejects_unknown_kind() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Bad": {
"kind": "mystery",
"fields": []
}
}
});
assert!(validator.validate(&doc).is_err(), "unknown kind accepted");
}
#[test]
fn meta_schema_rejects_additional_properties_on_struct() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Bad": {
"kind": "struct",
"fields": [],
"bogus": true
}
}
});
assert!(validator.validate(&doc).is_err(), "additional property accepted");
}
#[test]
fn meta_schema_accepts_union_with_byte_discriminator() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Msg": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": {
"1": { "$ref": "#/$defs/A" },
"2": { "$ref": "#/$defs/B" }
}
},
"A": { "kind": "struct", "fields": [] },
"B": { "kind": "struct", "fields": [] }
}
});
assert!(validator.validate(&doc).is_ok(), "byte-discriminator union rejected");
}
#[test]
fn meta_schema_accepts_enum() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Status": {
"kind": "enum",
"values": ["Ok", "Err"]
}
}
});
assert!(validator.validate(&doc).is_ok(), "enum rejected");
}
#[test]
fn meta_schema_rejects_empty_enum_values() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Bad": { "kind": "enum", "values": [] }
}
});
assert!(validator.validate(&doc).is_err(), "empty enum accepted");
}
#[test]
fn meta_schema_accepts_array_with_count() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Vec3": {
"kind": "struct",
"fields": [
{ "name": "data", "kind": { "kind": "array", "element": "float32", "count": 3 } }
]
}
}
});
assert!(validator.validate(&doc).is_ok(), "array with count rejected");
}
#[test]
fn meta_schema_rejects_array_without_count() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Bad": {
"kind": "struct",
"fields": [
{ "name": "data", "kind": { "kind": "array", "element": "float32" } }
]
}
}
});
assert!(validator.validate(&doc).is_err(), "array without count accepted");
}
#[test]
fn meta_schema_accepts_record() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Map": {
"kind": "struct",
"fields": [
{ "name": "entries", "kind": { "kind": "record", "values": "string" } }
]
}
}
});
assert!(validator.validate(&doc).is_ok(), "record rejected");
}
#[test]
fn meta_schema_accepts_ref_object() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Outer": {
"kind": "struct",
"fields": [
{ "name": "inner", "kind": { "$ref": "#/$defs/Inner" } }
]
},
"Inner": { "kind": "struct", "fields": [] }
}
});
assert!(validator.validate(&doc).is_ok(), "ref object rejected");
}
#[test]
fn meta_schema_rejects_malformed_ref() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"Bad": {
"kind": "struct",
"fields": [
{ "name": "inner", "kind": { "$ref": "Inner" } }
]
}
}
});
assert!(validator.validate(&doc).is_err(), "bare-name ref accepted");
}
#[test]
fn meta_schema_accepts_inline_struct_field() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{ "name": "inner", "kind": {
"kind": "struct",
"fields": [ { "name": "x", "kind": "uint16" } ]
} }
]
}
}
});
assert!(validator.validate(&doc).is_ok(), "inline struct field rejected");
}
#[test]
fn meta_schema_accepts_inline_struct_in_union_mapping() {
let validator = jsonschema::options()
.build(&BAST_META_SCHEMA)
.expect("meta-schema compiles");
let doc = serde_json::json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": {
"5": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
}
}
}
});
assert!(validator.validate(&doc).is_ok(), "inline struct in union mapping rejected");
}
}
+747
View File
@@ -0,0 +1,747 @@
//! BAST-native validator — the `validate_bytes` validation step.
//!
//! A recursive walker over the BAST typed tree ([`crate::bast::BastDoc`]/
//! [`crate::bast::BastType`]) that checks the value-domain constraints the
//! materializer ([`crate::materialize`]) does NOT check. This replaces the
//! v0.1.0 `jsonschema` custom-keyword validators on the bytes path
//! (D-BAST-006) with a flat match — no factories, no trait objects, no
//! sub-validator pre-computation.
//!
//! ## What the materializer already guarantees
//!
//! By construction, a materialized [`serde_json::Value`] is structurally
//! correct: all declared fields are present, types are correct, bounds
//! are checked (via [`crate::data_access`]), UTF-8 is valid, the boolean
//! byte is 0 or 1, and the union discriminator is in the mapping. The
//! validator only needs to enforce the **value-domain constraints
//! expressed in the BAST document** — the ones the materializer can't
//! see from the bytes alone.
//!
//! ## Constraint table
//!
//! See [bast-format.md §Validation Model](../../docs/architecture/bast-format.md#validation-model).
//! The arms of the `validate_typeref` walker implement the table:
//!
//! | Constraint | Validator arm |
//! |---|---|
//! | Integer range (Int8..Uint64) | `validate_int`/`validate_uint` |
//! | Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` |
//! | Float finiteness (Float32/64) | `validate_float` |
//! | String `maxLength` (byte length) | `check_string` reads [`BastField::max_length`](crate::bast::BastField::max_length) |
//! | Bytes `maxLength` (array length) | `check_bytes` |
//! | Enum index bounds | `validate_enum` — **fixes the v0.1.0 dead constraint** |
//! | Union variant dispatch | `validate_union` reads `__discriminator`, resolves the variant, recurses |
//! | Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
//! | Array count | `validate_array` checks `arr.len() == count` and recurses per element |
//! | Record values | `validate_record` recurses into each value's `values` type |
//! | Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
//!
//! ## Error payload (D-BAST-009)
//!
//! The bytes path no longer touches `jsonschema` for validation, but the
//! error variant retains the `jsonschema::ValidationError<'static>` type
//! for uniformity with the `validate_json` path. Errors are constructed
//! via [`jsonschema::ValidationError::custom`] so consumers handle one
//! `AlkTypeError::Validation` match arm for both paths.
//!
//! ## Untrusted input
//!
//! Every walk path returns [`AlkTypeError::Schema`] on a malformed BAST
//! document, never `panic!`/`unreachable!`/`unwrap` (AGENTS.md §3). The
//! typed tree is parsed once by [`BastDoc::new`](crate::bast::BastDoc::new);
//! lazy `$ref` resolution via [`BastDoc::resolve_typeref`] surfaces a
//! `Schema` error for dangling refs.
use crate::bast::{
BastArray, BastDefKind, BastDoc, BastEnum, BastField, BastRecord, BastStruct, BastType,
BastUnion,
};
use crate::error::AlkTypeError;
use crate::schema::AlkTypeKind;
use serde_json::Value;
const DISCRIMINATOR_KEY: &str = "__discriminator";
/// Validate a materialized `value` against the BAST root type.
///
/// This is the `validate_bytes` validation step: after
/// [`crate::materialize::materialize_packed`]/
/// [`crate::materialize::materialize_aligned`] produces a `Value` tree
/// from the bytes, this walker enforces the value-domain constraints
/// expressed in the BAST document. The materializer already guarantees
/// structural correctness; the validator only checks what the bytes
/// alone can't tell you (integer ranges, `maxLength`,
/// enum index bounds, union variant constraints).
///
/// Returns `Err(AlkTypeError::Validation(...))` on the first violated
/// constraint, or `Err(AlkTypeError::Schema(...))` if the BAST document
/// is malformed (a dangling `$ref`, a missing `values` array, etc.).
pub fn validate_value(doc: &BastDoc<'_>, value: &Value) -> Result<(), AlkTypeError> {
let root_def = doc.root_def();
match root_def.kind() {
BastDefKind::Struct(s) => validate_struct(doc, s, value, ""),
BastDefKind::Union(u) => validate_union(doc, u, value, ""),
BastDefKind::Enum(e) => validate_enum(value, "", e),
}
}
/// Recursively validate `value` against `ty`, resolving `$ref`s lazily.
///
/// `field` is the owning field (for `maxLength`/annotations) — `None` for
/// synthetic contexts (array elements, record values, the root). The
/// validator reads [`BastField::max_length`] only for `String`/`Bytes`
/// arms, so passing a synthetic field (which has `max_length == None`) is
/// correct for those contexts.
fn validate_typeref(
doc: &BastDoc<'_>,
ty: &BastType<'_>,
field: Option<&BastField<'_>>,
value: &Value,
path: &str,
) -> Result<(), AlkTypeError> {
let resolved = doc.resolve_typeref(ty)?;
match &resolved {
BastType::Primitive(AlkTypeKind::Int8) => validate_int(value, path, -128, 127),
BastType::Primitive(AlkTypeKind::Int16) => validate_int(value, path, -32768, 32767),
BastType::Primitive(AlkTypeKind::Int32) => {
validate_int(value, path, -2147483648, 2147483647)
}
BastType::Primitive(AlkTypeKind::Int64) => validate_int64(value, path),
BastType::Primitive(AlkTypeKind::Uint8) => validate_uint(value, path, 255),
BastType::Primitive(AlkTypeKind::Uint16) => validate_uint(value, path, 65535),
BastType::Primitive(AlkTypeKind::Uint32) => validate_uint(value, path, 4294967295),
BastType::Primitive(AlkTypeKind::Uint64) => validate_uint64(value, path),
BastType::Primitive(AlkTypeKind::Float32) => validate_float(value, path),
BastType::Primitive(AlkTypeKind::Float64) => validate_float(value, path),
BastType::Primitive(AlkTypeKind::Boolean) => validate_bool(value, path),
BastType::Primitive(AlkTypeKind::String) => {
check_string(value, path, field.and_then(|f| f.max_length()))
}
BastType::Primitive(AlkTypeKind::Bytes) => {
check_bytes(value, path, field.and_then(|f| f.max_length()))
}
BastType::Enum(e) => validate_enum(value, path, e),
BastType::Struct(s) => validate_struct(doc, s, value, path),
BastType::Union(u) => validate_union(doc, u, value, path),
BastType::Array(a) => validate_array(doc, a, value, path),
BastType::Record(r) => validate_record(doc, r, value, path),
BastType::Ref(_) => Err(AlkTypeError::Schema(format!(
"bast_validation: unresolved $ref at {path} (resolve_typeref should have deref'd it)"
))),
BastType::Primitive(k) => Err(AlkTypeError::Schema(format!(
"bast_validation: unsupported primitive kind {k} at {path}"
))),
}
}
fn validate_int(value: &Value, path: &str, min: i64, max: i64) -> Result<(), AlkTypeError> {
let n = value
.as_i64()
.ok_or_else(|| validation_err(path, "expected an integer"))?;
if n < min || n > max {
return Err(validation_err(
path,
format!("integer {n} out of range [{min}, {max}]"),
));
}
Ok(())
}
fn validate_int64(value: &Value, path: &str) -> Result<(), AlkTypeError> {
if value.as_i64().is_none() {
return Err(validation_err(path, "expected an i64 integer"));
}
Ok(())
}
fn validate_uint(value: &Value, path: &str, max: u64) -> Result<(), AlkTypeError> {
let n = value
.as_u64()
.ok_or_else(|| validation_err(path, "expected a non-negative integer"))?;
if n > max {
return Err(validation_err(path, format!("integer {n} exceeds {max}")));
}
Ok(())
}
fn validate_uint64(value: &Value, path: &str) -> Result<(), AlkTypeError> {
if value.as_u64().is_none() {
return Err(validation_err(path, "expected a u64 integer"));
}
Ok(())
}
fn validate_float(value: &Value, path: &str) -> Result<(), AlkTypeError> {
match value {
Value::Number(n) => {
let f = n
.as_f64()
.ok_or_else(|| validation_err(path, "expected a number"))?;
if !f.is_finite() {
return Err(validation_err(path, "expected a finite number"));
}
Ok(())
}
_ => Err(validation_err(path, "expected a number")),
}
}
fn validate_bool(value: &Value, path: &str) -> Result<(), AlkTypeError> {
match value {
Value::Bool(_) => Ok(()),
_ => Err(validation_err(path, "expected a boolean")),
}
}
fn check_string(value: &Value, path: &str, max_length: Option<usize>) -> Result<(), AlkTypeError> {
let s = value
.as_str()
.ok_or_else(|| validation_err(path, "expected a string"))?;
if let Some(max) = max_length {
if s.len() > max {
return Err(validation_err(
path,
format!("string byte length {} exceeds maxLength {max}", s.len()),
));
}
}
Ok(())
}
fn check_bytes(value: &Value, path: &str, max_length: Option<usize>) -> Result<(), AlkTypeError> {
let arr = value
.as_array()
.ok_or_else(|| validation_err(path, "expected an array of u8 for bytes"))?;
if let Some(max) = max_length {
if arr.len() > max {
return Err(validation_err(
path,
format!("bytes array length {} exceeds maxLength {max}", arr.len()),
));
}
}
for (i, entry) in arr.iter().enumerate() {
let n = entry.as_u64().ok_or_else(|| {
validation_err(
path,
format!("bytes array entry {i} is not a non-negative integer"),
)
})?;
if n > 255 {
return Err(validation_err(
path,
format!("bytes array entry {i} = {n} is not a u8 (0..=255)"),
));
}
}
Ok(())
}
/// Enum validation on the bytes path: the materializer emits a numeric
/// index (`Value::Number`), and the constraint is that the index is
/// within the `values` array bounds (`0..len-1`).
///
/// This is the **fix for the v0.1.0 dead constraint**: the built-in
/// `enum` keyword checked string membership, but the materializer emitted
/// a numeric index that never matched — so out-of-bounds enum indices
/// silently passed. The BAST-native validator checks the index bounds
/// directly.
fn validate_enum(value: &Value, path: &str, enum_def: &BastEnum<'_>) -> Result<(), AlkTypeError> {
let idx = value
.as_u64()
.ok_or_else(|| validation_err(path, "expected a non-negative integer enum index"))?;
let len = enum_def.values().len() as u64;
if idx >= len {
return Err(validation_err(
path,
format!("enum index {idx} out of bounds (values has {len} entries)"),
));
}
Ok(())
}
fn validate_struct(
doc: &BastDoc<'_>,
struct_def: &BastStruct<'_>,
value: &Value,
path: &str,
) -> Result<(), AlkTypeError> {
let obj = value
.as_object()
.ok_or_else(|| validation_err(path, "expected an object"))?;
for field in struct_def.fields() {
let name = field.name();
let field_path = if path.is_empty() {
name.to_string()
} else {
format!("{path}.{name}")
};
let field_value = obj.get(name).ok_or_else(|| {
validation_err(&field_path, format!("missing field {name:?}"))
})?;
validate_typeref(doc, field.ty(), Some(field), field_value, &field_path)?;
}
Ok(())
}
/// Union validation: read `__discriminator`, look up the variant
/// [`BastType`] in the union's `mapping`, and recurse into the variant.
///
/// This recovers OQ-008 per-variant constraint enforcement (e.g.
/// `maxLength` on a `bytes` field inside a variant struct) without
/// custom keywords — the recursion walks the variant's BAST definition
/// and enforces every field constraint it declares.
fn validate_union(
doc: &BastDoc<'_>,
union_def: &BastUnion<'_>,
value: &Value,
path: &str,
) -> Result<(), AlkTypeError> {
let obj = value
.as_object()
.ok_or_else(|| validation_err(path, "expected an object for union"))?;
let disc = obj.get(DISCRIMINATOR_KEY).ok_or_else(|| {
validation_err(path, "union instance is missing the '__discriminator' field")
})?;
let key = match disc {
Value::String(s) => s.clone(),
Value::Number(n) => n.to_string(),
_ => {
return Err(validation_err(
path,
"union '__discriminator' must be a string or number",
));
}
};
let variant_ty = union_def.variant_for(&key).ok_or_else(|| {
validation_err(path, format!("union discriminator value '{key}' not in mapping"))
})?;
validate_typeref(doc, variant_ty, None, value, path)
}
fn validate_array(
doc: &BastDoc<'_>,
array_def: &BastArray<'_>,
value: &Value,
path: &str,
) -> Result<(), AlkTypeError> {
let arr = value
.as_array()
.ok_or_else(|| validation_err(path, "expected an array"))?;
let count = array_def.count();
if arr.len() != count {
return Err(validation_err(
path,
format!("array length {} does not match declared count {count}", arr.len()),
));
}
let element_ty = array_def.element();
for (i, item) in arr.iter().enumerate() {
let item_path = format!("{path}[{i}]");
validate_typeref(doc, element_ty, None, item, &item_path)?;
}
Ok(())
}
fn validate_record(
doc: &BastDoc<'_>,
record_def: &BastRecord<'_>,
value: &Value,
path: &str,
) -> Result<(), AlkTypeError> {
let obj = value
.as_object()
.ok_or_else(|| validation_err(path, "expected an object for record"))?;
let values_ty = record_def.values();
for (k, v) in obj.iter() {
let entry_path = format!("{path}[{k}]");
validate_typeref(doc, values_ty, None, v, &entry_path)?;
}
Ok(())
}
/// Construct a `Validation` error from a path + reason string. The
/// payload is a `jsonschema::ValidationError::custom` so the variant
/// type stays uniform with the `validate_json` path (D-BAST-009).
fn validation_err(path: &str, reason: impl Into<String>) -> AlkTypeError {
let msg = if path.is_empty() {
reason.into()
} else {
format!("{path}: {}", reason.into())
};
AlkTypeError::Validation(jsonschema::ValidationError::custom(msg))
}
#[cfg(test)]
mod tests {
use super::*;
use serde_json::json;
fn doc_from<'a>(root: &'a Value, name: &'a str) -> BastDoc<'a> {
BastDoc::new(root, name).expect("bast doc")
}
fn u32_le(n: u32) -> Vec<u8> {
n.to_le_bytes().to_vec()
}
fn prefixed_str_le(s: &str) -> Vec<u8> {
let mut buf = u32_le(s.len() as u32);
buf.extend_from_slice(s.as_bytes());
buf
}
fn materialize_and_validate(
doc: &BastDoc<'_>,
buffer: &[u8],
) -> Result<(), AlkTypeError> {
let value = crate::materialize::materialize_packed(doc, buffer)?;
validate_value(doc, &value)
}
// ----- Integer ranges --------------------------------------------------
#[test]
fn uint32_valid_boundary_passes() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "id", "kind": "uint32" }
] } }
});
let d = doc_from(&root, "S");
assert!(materialize_and_validate(&d, &u32_le(0)).is_ok());
assert!(materialize_and_validate(&d, &u32_le(0xFFFF_FFFF)).is_ok());
}
#[test]
fn int8_boundary_passes() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "v", "kind": "int8" }
] } }
});
let d = doc_from(&root, "S");
assert!(materialize_and_validate(&d, &[127u8]).is_ok());
assert!(materialize_and_validate(&d, &[128u8]).is_ok());
}
#[test]
fn validate_value_rejects_uint32_out_of_range() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "id", "kind": "uint32" }
] } }
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"id": -1})).is_err());
assert!(validate_value(&d, &json!({"id": 4294967296u64})).is_err());
assert!(validate_value(&d, &json!({"id": "x"})).is_err());
}
#[test]
fn validate_value_rejects_int8_out_of_range() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "v", "kind": "int8" }
] } }
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"v": 0})).is_ok());
assert!(validate_value(&d, &json!({"v": 127})).is_ok());
assert!(validate_value(&d, &json!({"v": -128})).is_ok());
assert!(validate_value(&d, &json!({"v": 128})).is_err());
assert!(validate_value(&d, &json!({"v": -129})).is_err());
}
#[test]
fn validate_value_int64_and_uint64() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "a", "kind": "int64" },
{ "name": "b", "kind": "uint64" }
] } }
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"a": 0, "b": 0})).is_ok());
assert!(validate_value(&d, &json!({"a": 9223372036854775807i64, "b": 18446744073709551615u64})).is_ok());
assert!(validate_value(&d, &json!({"a": "x", "b": 0})).is_err());
assert!(validate_value(&d, &json!({"a": 0, "b": -1})).is_err());
}
// ----- maxLength -------------------------------------------------------
#[test]
fn string_max_length_enforced() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "name", "kind": "string", "maxLength": 3 }
] } }
});
let d = doc_from(&root, "S");
assert!(materialize_and_validate(&d, &prefixed_str_le("hi")).is_ok());
let err = materialize_and_validate(&d, &prefixed_str_le("hello")).unwrap_err();
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
}
#[test]
fn bytes_max_length_enforced_array_form() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "blob", "kind": "bytes", "maxLength": 2 }
] } }
});
let d = doc_from(&root, "S");
let mut buf = u32_le(2);
buf.extend_from_slice(&[0xAA, 0xBB]);
assert!(materialize_and_validate(&d, &buf).is_ok());
let mut buf = u32_le(3);
buf.extend_from_slice(&[0xAA, 0xBB, 0xCC]);
let err = materialize_and_validate(&d, &buf).unwrap_err();
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
}
#[test]
fn bytes_array_entry_out_of_range_rejected() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "blob", "kind": "bytes" }
] } }
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"blob": [65, 256]})).is_err());
assert!(validate_value(&d, &json!({"blob": [65, "x"]})).is_err());
assert!(validate_value(&d, &json!({"blob": [65, 255]})).is_ok());
}
// ----- Enum index bounds (the v0.1.0 dead-constraint fix) -------------
#[test]
fn enum_index_in_bounds_passes() {
let root = json!({
"$defs": {
"S": { "kind": "struct", "fields": [
{ "name": "status", "kind": { "$ref": "#/$defs/StatusCode" } }
] },
"StatusCode": { "kind": "enum", "values": ["Ok", "Error", "Pending"] }
}
});
let d = doc_from(&root, "S");
assert!(materialize_and_validate(&d, &u32_le(0)).is_ok());
assert!(materialize_and_validate(&d, &u32_le(2)).is_ok());
}
#[test]
fn enum_index_out_of_bounds_rejected() {
let root = json!({
"$defs": {
"S": { "kind": "struct", "fields": [
{ "name": "status", "kind": { "$ref": "#/$defs/StatusCode" } }
] },
"StatusCode": { "kind": "enum", "values": ["Ok", "Error"] }
}
});
let d = doc_from(&root, "S");
let err = materialize_and_validate(&d, &u32_le(5)).unwrap_err();
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
}
// ----- Union variant dispatch (OQ-008) --------------------------------
#[test]
fn union_byte_disc_max_length_inside_variant_enforced() {
let root = json!({
"$defs": {
"Wrapper": { "kind": "struct", "fields": [
{ "name": "payload", "kind": { "$ref": "#/$defs/Packet" } }
] },
"Packet": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": {
"1": { "$ref": "#/$defs/Ack" },
"2": { "$ref": "#/$defs/Data" }
}
},
"Ack": { "kind": "struct", "fields": [ { "name": "code", "kind": "uint8" } ] },
"Data": { "kind": "struct", "fields": [
{ "name": "blob", "kind": "bytes", "maxLength": 2 }
] }
}
});
let d = doc_from(&root, "Wrapper");
let mut buf = vec![2u8];
buf.extend_from_slice(&u32_le(2));
buf.extend_from_slice(&[0xAA, 0xBB]);
assert!(materialize_and_validate(&d, &buf).is_ok());
let mut buf = vec![2u8];
buf.extend_from_slice(&u32_le(3));
buf.extend_from_slice(&[0xAA, 0xBB, 0xCC]);
let err = materialize_and_validate(&d, &buf).unwrap_err();
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
let buf = vec![1u8, 7u8];
assert!(materialize_and_validate(&d, &buf).is_ok());
}
#[test]
fn union_field_disc_max_length_inside_variant_enforced() {
let root = json!({
"$defs": {
"Wrapper": { "kind": "struct", "fields": [
{ "name": "payload", "kind": { "$ref": "#/$defs/Event" } }
] },
"Event": {
"kind": "union",
"discriminator": { "kind": "field", "name": "type" },
"fields": [ { "name": "type", "kind": "string" } ],
"mapping": { "data": { "$ref": "#/$defs/Data" } }
},
"Data": { "kind": "struct", "fields": [
{ "name": "blob", "kind": "bytes", "maxLength": 1 }
] }
}
});
let d = doc_from(&root, "Wrapper");
let mut buf = prefixed_str_le("data");
buf.extend_from_slice(&u32_le(2));
buf.extend_from_slice(&[0xFF, 0xFE]);
let err = materialize_and_validate(&d, &buf).unwrap_err();
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
}
#[test]
fn validate_value_union_dispatch_on_value() {
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": {
"5": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] }
}
}
}
});
let d = doc_from(&root, "U");
assert!(validate_value(&d, &json!({"__discriminator": 5, "id": 42})).is_ok());
assert!(validate_value(&d, &json!({"__discriminator": 5, "id": -1})).is_err());
assert!(validate_value(&d, &json!({"__discriminator": 99, "id": 42})).is_err());
assert!(validate_value(&d, &json!({"id": 42})).is_err());
assert!(validate_value(&d, &json!("not-object")).is_err());
}
// ----- Array / Record --------------------------------------------------
#[test]
fn array_count_mismatch_rejected() {
let root = json!({
"$defs": {
"S": { "kind": "struct", "fields": [
{ "name": "pts", "kind": { "kind": "array", "element": "uint16", "count": 2 } }
] }
}
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"pts": [1, 2]})).is_ok());
assert!(validate_value(&d, &json!({"pts": [1, 2, 3]})).is_err());
assert!(validate_value(&d, &json!({"pts": [1]})).is_err());
}
#[test]
fn array_of_structs_validates() {
let root = json!({
"$defs": {
"S": { "kind": "struct", "fields": [
{ "name": "points", "kind": {
"kind": "array",
"element": { "$ref": "#/$defs/Point" },
"count": 2
} }
] },
"Point": { "kind": "struct", "fields": [
{ "name": "x", "kind": "uint16" },
{ "name": "y", "kind": "uint16" }
] }
}
});
let d = doc_from(&root, "S");
let mut buf = Vec::new();
buf.extend_from_slice(&1u16.to_le_bytes());
buf.extend_from_slice(&2u16.to_le_bytes());
buf.extend_from_slice(&3u16.to_le_bytes());
buf.extend_from_slice(&4u16.to_le_bytes());
assert!(materialize_and_validate(&d, &buf).is_ok());
}
#[test]
fn record_values_validated() {
let root = json!({
"$defs": {
"S": { "kind": "struct", "fields": [
{ "name": "counts", "kind": { "kind": "record", "values": "uint32" } }
] }
}
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"counts": {"a": 1, "b": 2}})).is_ok());
assert!(validate_value(&d, &json!({"counts": {"a": -1}})).is_err());
assert!(validate_value(&d, &json!({"counts": "not-object"})).is_err());
}
// ----- Struct missing field -------------------------------------------
#[test]
fn validate_value_struct_missing_field_rejected() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "id", "kind": "uint32" },
{ "name": "flag", "kind": "uint8" }
] } }
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"id": 1, "flag": 0})).is_ok());
let err = validate_value(&d, &json!({"id": 1})).unwrap_err();
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
}
// ----- Untrusted schema: error, not panic -----------------------------
#[test]
fn short_buffer_rejected_with_access_error() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "id", "kind": "uint32" }
] } }
});
let d = doc_from(&root, "S");
let err = materialize_and_validate(&d, &[0u8; 2]).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
}
#[test]
fn float_finiteness_checked() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "f", "kind": "float32" }
] } }
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"f": 3.5})).is_ok());
assert!(validate_value(&d, &json!({"f": 0})).is_ok());
assert!(validate_value(&d, &json!({"f": "x"})).is_err());
}
#[test]
fn bool_validated() {
let root = json!({
"$defs": { "S": { "kind": "struct", "fields": [
{ "name": "flag", "kind": "bool" }
] } }
});
let d = doc_from(&root, "S");
assert!(validate_value(&d, &json!({"flag": true})).is_ok());
assert!(validate_value(&d, &json!({"flag": false})).is_ok());
assert!(validate_value(&d, &json!({"flag": "yes"})).is_err());
}
}
+951 -364
View File
File diff suppressed because it is too large. Load diff
+131 -38
View File
@@ -1,4 +1,4 @@
//! Data access layer: primitive read/write functions for all 19 AlkType
//! Data access layer: primitive read/write functions for all 18 AlkType
//! kinds with endianness support, bounds checking, and zero-copy access.
//!
//! These are the building blocks used by the layout types ([`crate::offset_map`],
@@ -294,24 +294,23 @@ pub fn write_bytes(
}
// ---------------------------------------------------------------------------
// Variable-length read (offset indirection)
// Variable-length read/write (offset indirection)
// ---------------------------------------------------------------------------
/// Read an offset-indirect string.
///
/// The 8-byte struct at `buffer[offset..offset+8]` is
/// `{ data_offset: u32, data_length: u32 }` (endian-aware). The actual UTF-8
/// bytes live in `data_region[data_offset..data_offset+data_length]`. Returns
/// a `&'a str` borrowing from `data_region`. Invalid UTF-8 produces
/// [`AlkTypeError::Access`].
/// bytes live at `buffer[data_offset..data_offset+data_length]` — offsets are
/// absolute into the same buffer (safetensors-style). Returns a `&'a str`
/// borrowing from `buffer`. Invalid UTF-8 produces [`AlkTypeError::Access`].
pub fn read_string_indirect<'a>(
buffer: &'a [u8],
offset: usize,
data_region: &'a [u8],
field_path: &str,
endian: Endian,
) -> Result<&'a str, AlkTypeError> {
let bytes = read_bytes_indirect(buffer, offset, data_region, field_path, endian)?;
let bytes = read_bytes_indirect(buffer, offset, field_path, endian)?;
std::str::from_utf8(bytes).map_err(|e| {
access_err(
field_path,
@@ -324,34 +323,97 @@ pub fn read_string_indirect<'a>(
///
/// The 8-byte struct at `buffer[offset..offset+8]` is
/// `{ data_offset: u32, data_length: u32 }` (endian-aware). Returns a
/// `&'a [u8]` slice of `data_region[data_offset..data_offset+data_length]`.
/// `&'a [u8]` slice of `buffer[data_offset..data_offset+data_length]` —
/// offsets are absolute into the same buffer.
pub fn read_bytes_indirect<'a>(
buffer: &'a [u8],
offset: usize,
data_region: &'a [u8],
field_path: &str,
endian: Endian,
) -> Result<&'a [u8], AlkTypeError> {
let struct_end = offset
.checked_add(8)
.ok_or_else(|| access_err(field_path, format!("offset {offset} + 8 overflows usize")))?;
check_bounds(buffer.len(), offset, struct_end, field_path)?;
let off_bytes: [u8; U32_SIZE] = buffer[offset..offset + U32_SIZE]
.try_into()
.map_err(|_| access_err(field_path, "internal: try_into failed for data_offset"))?;
let len_bytes: [u8; U32_SIZE] = buffer[offset + U32_SIZE..offset + 8]
.try_into()
.map_err(|_| access_err(field_path, "internal: try_into failed for data_length"))?;
let data_offset = u32_from(off_bytes, endian) as usize;
let data_length = u32_from(len_bytes, endian) as usize;
let data_offset = read_u32(buffer, offset, field_path, endian)? as usize;
let len_offset = offset.checked_add(U32_SIZE).ok_or_else(|| {
access_err(
field_path,
format!("offset {offset} + {U32_SIZE} overflows usize"),
)
})?;
let data_length = read_u32(buffer, len_offset, field_path, endian)? as usize;
let data_end = data_offset.checked_add(data_length).ok_or_else(|| {
access_err(
field_path,
format!("data_offset {data_offset} + data_length {data_length} overflows usize"),
)
})?;
check_bounds(data_region.len(), data_offset, data_end, field_path)?;
Ok(&data_region[data_offset..data_end])
check_bounds(buffer.len(), data_offset, data_end, field_path)?;
Ok(&buffer[data_offset..data_end])
}
/// Write an offset-indirect string.
///
/// Writes the `{ data_offset: u32, data_length: u32 }` pair at `offset` and
/// the UTF-8 bytes at `data_offset` (absolute into the same buffer). Returns
/// the number of bytes written for the pair (`8`).
pub fn write_string_indirect(
buffer: &mut [u8],
offset: usize,
data_offset: usize,
value: &str,
field_path: &str,
endian: Endian,
) -> Result<usize, AlkTypeError> {
write_bytes_indirect(buffer, offset, data_offset, value.as_bytes(), field_path, endian)
}
/// Write offset-indirect raw bytes.
///
/// Writes the `{ data_offset: u32, data_length: u32 }` pair at `offset` and
/// the raw bytes at `data_offset` (absolute into the same buffer). Returns
/// the number of bytes written for the pair (`8`).
pub fn write_bytes_indirect(
buffer: &mut [u8],
offset: usize,
data_offset: usize,
value: &[u8],
field_path: &str,
endian: Endian,
) -> Result<usize, AlkTypeError> {
let data_len = value.len();
let data_len_u32 = u32::try_from(data_len).map_err(|_| {
access_err(
field_path,
format!("data length {data_len} exceeds u32::MAX (length field width)"),
)
})?;
let data_offset_u32 = u32::try_from(data_offset).map_err(|_| {
access_err(
field_path,
format!("data offset {data_offset} exceeds u32::MAX (offset field width)"),
)
})?;
write_u32(buffer, offset, data_offset_u32, field_path, endian)?;
let len_offset = offset.checked_add(U32_SIZE).ok_or_else(|| {
access_err(
field_path,
format!("offset {offset} + {U32_SIZE} overflows usize"),
)
})?;
write_u32(buffer, len_offset, data_len_u32, field_path, endian)?;
let data_end = data_offset.checked_add(data_len).ok_or_else(|| {
access_err(
field_path,
format!("data offset {data_offset} + length {data_len} overflows usize"),
)
})?;
check_bounds(buffer.len(), data_offset, data_end, field_path)?;
let dest = buffer.get_mut(data_offset..data_end).ok_or_else(|| {
access_err(
field_path,
format!("mutable data slice [{data_offset}..{data_end}) unavailable"),
)
})?;
dest.copy_from_slice(value);
Ok(8)
}
#[cfg(test)]
@@ -583,39 +645,70 @@ mod tests {
#[test]
fn read_string_indirect_round_trip() {
let data_region = b"the quick brown fox";
let mut index = [0u8; 8];
write_u32(&mut index, 0, 4, "idx.off", LE).unwrap();
write_u32(&mut index, 4, 11, "idx.len", LE).unwrap();
let s = read_string_indirect(&index, 0, data_region, "msg", LE).unwrap();
let data = b"the quick brown fox";
let mut full = vec![0u8; 8 + data.len()];
write_u32(&mut full, 0, 12, "idx.off", LE).unwrap();
write_u32(&mut full, 4, 11, "idx.len", LE).unwrap();
full[8..].copy_from_slice(data);
let s = read_string_indirect(&full, 0, "msg", LE).unwrap();
assert_eq!(s, "quick brown");
}
#[test]
fn read_bytes_indirect_round_trip() {
let data_region: &[u8] = b"HEADERbody-payloadTAIL";
let mut index = [0u8; 8];
write_u32(&mut index, 0, 6, "idx.off", BE).unwrap();
write_u32(&mut index, 4, 12, "idx.len", BE).unwrap();
let bytes = read_bytes_indirect(&index, 0, data_region, "blob", BE).unwrap();
let data: &[u8] = b"HEADERbody-payloadTAIL";
let mut full = vec![0u8; 8 + data.len()];
write_u32(&mut full, 0, 14, "idx.off", BE).unwrap();
write_u32(&mut full, 4, 12, "idx.len", BE).unwrap();
full[8..].copy_from_slice(data);
let bytes = read_bytes_indirect(&full, 0, "blob", BE).unwrap();
assert_eq!(bytes, b"body-payload");
}
#[test]
fn write_bytes_indirect_round_trip() {
let mut buf = vec![0u8; 8 + 11];
let written = write_bytes_indirect(&mut buf, 0, 8, b"quick brown", "blob", LE).unwrap();
assert_eq!(written, 8);
assert_eq!(&buf[0..4], &8u32.to_le_bytes());
assert_eq!(&buf[4..8], &11u32.to_le_bytes());
assert_eq!(&buf[8..19], b"quick brown");
let bytes = read_bytes_indirect(&buf, 0, "blob", LE).unwrap();
assert_eq!(bytes, b"quick brown");
}
#[test]
fn write_string_indirect_round_trip() {
let mut buf = vec![0u8; 8 + 5];
let written = write_string_indirect(&mut buf, 0, 8, "hello", "msg", BE).unwrap();
assert_eq!(written, 8);
assert_eq!(&buf[0..4], &8u32.to_be_bytes());
assert_eq!(&buf[4..8], &5u32.to_be_bytes());
assert_eq!(&buf[8..13], b"hello");
let s = read_string_indirect(&buf, 0, "msg", BE).unwrap();
assert_eq!(s, "hello");
}
#[test]
fn read_bytes_indirect_bounds_failure_on_index() {
let buf = [0u8; 4];
let data_region = b"anything";
let err = read_bytes_indirect(&buf, 0, data_region, "blob", LE).unwrap_err();
let err = read_bytes_indirect(&buf, 0, "blob", LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }));
}
#[test]
fn read_bytes_indirect_bounds_failure_on_data_region() {
fn read_bytes_indirect_bounds_failure_on_data() {
let mut buf = [0u8; 8];
write_u32(&mut buf, 0, 100, "idx.off", LE).unwrap();
write_u32(&mut buf, 4, 10, "idx.len", LE).unwrap();
let data_region = b"too short";
let err = read_bytes_indirect(&buf, 0, data_region, "blob", LE).unwrap_err();
let err = read_bytes_indirect(&buf, 0, "blob", LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }));
}
#[test]
fn write_bytes_indirect_bounds_failure_on_data() {
let mut buf = vec![0u8; 8];
let err = write_bytes_indirect(&mut buf, 0, 100, b"hello", "blob", LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }));
}
+764 -337
View File
File diff suppressed because it is too large. Load diff
+3 -3
View File
@@ -1,6 +1,6 @@
//! Error types for the alktype engine.
//!
//! Decided in ADR-098: a single `AlkTypeError` enum covers all error
//! Decided in ADR-004: a single `AlkTypeError` enum covers all error
//! conditions across the engine's three phases (schema parsing, offset
//! computation, read/write) plus validation.
@@ -9,8 +9,8 @@ use std::fmt;
/// Errors produced by the alktype engine across all phases.
#[derive(Debug)]
pub enum AlkTypeError {
/// Schema parsing errors — invalid JSON, missing required keywords,
/// unknown `AlkType:*` kinds, malformed annotations.
/// Schema parsing errors — invalid JSON, missing required properties,
/// unknown BAST kinds, malformed annotations.
Schema(String),
/// Offset computation errors — field not found, type not supported
+609 -860
View File
File diff suppressed because it is too large. Load diff
+29 -15
View File
@@ -1,33 +1,46 @@
//! alktype: The binary struct engine.
//!
//! Takes a JSON Schema with `AlkType:*` custom keywords and produces
//! Takes a BAST (Binary Abstract Syntax Tree) document and produces
//! an offset map, read/write functions, and validation — all driven
//! by the schema. The schema is the format definition; the engine is
//! generic.
//!
//! ## Architecture
//!
//! - **Schema layer** ([`schema`]): AlkType kind detection, annotation
//! parsing, `$ref` normalization, endianness.
//! - **BAST parser** ([`bast`]): Typed tree over a BAST document —
//! `BastDoc`/`BastDef`/`BastStruct`/`BastField`/`BastType`/etc.
//! Borrows from the source `serde_json::Value` without cloning field
//! data.
//! - **Layout engine** ([`offset_map`], [`layout_builder`],
//! [`sequential_reader`]): Two layout modes — aligned static for
//! mmap-friendly formats, packed sequential for protocol wire formats.
//! All three consume the BAST typed tree.
//! - **Data access** ([`data_access`]): Typed read/write at computed
//! offsets, zero-copy for fixed-size types.
//! - **TUnion dispatch** ([`tunion`]): Byte-offset and field-name
//! discriminator dispatch.
//! - **Validation** ([`validation`]): Custom keyword validators for all
//! 19 `AlkType:*` kinds, delegated to the `jsonschema` crate.
//! - **Builder** ([`builder`]): Fluent Rust API for constructing alktype
//! JSON Schemas at runtime, producing `serde_json::Value` (ADR-009).
//! discriminator dispatch over `BastUnion`.
//! - **Validation** ([`validation`], [`bast_validation`]): two
//! validators for two paths. `bast_validation` is the BAST-native
//! validator for `validate_bytes` — a recursive walker over the BAST
//! type tree (D-BAST-006). `validation` builds a standard
//! `jsonschema::Validator` from a consumer-provided JSON Schema for
//! `validate_json` / `is_valid_json` (D-BAST-007) — no custom keywords,
//! no BAST involvement.
//! - **Builder** ([`builder`]): Fluent Rust API for constructing BAST
//! documents (binary layout) and standard JSON Schemas (JSON
//! validation) at runtime, producing `serde_json::Value` (ADR-009,
//! D-BAST-008). `struct_()` → BAST, `object()` → standard JSON Schema.
//! - **Materialize** ([`materialize`]): Materialize a `serde_json::Value`
//! tree from a binary buffer by walking the schema. Used by
//! tree from a binary buffer by walking the BAST typed tree. Used by
//! `AlkTypeEngine::validate_bytes` (ADR-010).
//! - **Engine** ([`engine`]): `AlkTypeEngine` — the compiled form of a
//! schema, combining layout and validation.
//! BAST document, combining layout and validation.
#[macro_use]
mod macros;
pub mod bast;
pub mod bast_meta;
pub mod bast_validation;
pub mod builder;
pub mod data_access;
pub mod engine;
@@ -40,16 +53,17 @@ pub mod sequential_reader;
pub mod tunion;
pub mod validation;
pub use bast_meta::BAST_META_SCHEMA;
pub use bast::{
BastArray, BastDef, BastDefKind, BastDiscriminator, BastDoc, BastEnum, BastField, BastRef,
BastRecord, BastStruct, BastType, BastUnion,
};
pub use builder::{Definitions, Discriminator, Schema};
pub use engine::{LayoutMode, AlkTypeEngine};
pub use error::AlkTypeError;
pub use layout_builder::{FieldPosition, LayoutBuilder, PackedLayout};
pub use offset_map::{ByteRange, OffsetMap};
pub use schema::{
get_alktype_kind_loose, get_alktype_kind_loose_enum, inline_union_variant_refs, normalize_refs,
parse_align, parse_discriminator, parse_encoding, parse_endian, parse_max_length, resolve_ref,
resolve_ref_or_inline, DiscriminatorKind, Endian, AlkTypeKind, VariableEncoding,
};
pub use schema::{Endian, AlkTypeKind, VariableEncoding};
pub use sequential_reader::{FieldValue, SequentialReader};
pub use tunion::UnionDispatch;
pub use validation::build_validator;
+5 -169
View File
@@ -1,173 +1,9 @@
//! Macros for generating repetitive code across the 19 AlkType kinds.
//! Macros for generating repetitive read/write code across the
//! endian-sensitive fixed-size kinds.
//!
//! These macros eliminate boilerplate in validation, data access, and
//! dispatch. Each macro takes a compact specification and generates the
//! full implementation, ensuring consistency across all types.
// ---------------------------------------------------------------------------
// Validation macros
// ---------------------------------------------------------------------------
/// Generate a signed integer validator struct and its factory closure.
#[macro_export]
macro_rules! define_int_validator {
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $min:literal, $max:literal) => {
struct $validator_struct;
impl jsonschema::Keyword for $validator_struct {
fn validate<'i>(
&self,
instance: &'i serde_json::Value,
) -> Result<(), jsonschema::ValidationError<'i>> {
match instance.as_i64() {
Some(n) if ($min..=$max).contains(&n) => Ok(()),
_ => Err(jsonschema::ValidationError::custom(concat!(
"expected an integer in range [",
stringify!($min),
", ",
stringify!($max),
"]"
))),
}
}
fn is_valid(&self, instance: &serde_json::Value) -> bool {
instance
.as_i64()
.is_some_and(|n| ($min..=$max).contains(&n))
}
}
fn $factory_fn<'a>(
_parent: &'a serde_json::Map<String, serde_json::Value>,
value: &'a serde_json::Value,
_path: jsonschema::paths::Location,
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
if value.as_bool() == Some(true) {
Ok(Box::new($validator_struct))
} else {
Err(jsonschema::ValidationError::schema(concat!(
$keyword,
" must be set to true"
)))
}
}
};
}
/// Generate an unsigned integer validator struct and its factory closure.
#[macro_export]
macro_rules! define_uint_validator {
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $max:literal) => {
struct $validator_struct;
impl jsonschema::Keyword for $validator_struct {
fn validate<'i>(
&self,
instance: &'i serde_json::Value,
) -> Result<(), jsonschema::ValidationError<'i>> {
match instance.as_u64() {
Some(n) if n <= $max => Ok(()),
_ => Err(jsonschema::ValidationError::custom(concat!(
"expected an unsigned integer in range [0, ",
stringify!($max),
"]"
))),
}
}
fn is_valid(&self, instance: &serde_json::Value) -> bool {
instance.as_u64().is_some_and(|n| n <= $max)
}
}
fn $factory_fn<'a>(
_parent: &'a serde_json::Map<String, serde_json::Value>,
value: &'a serde_json::Value,
_path: jsonschema::paths::Location,
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
if value.as_bool() == Some(true) {
Ok(Box::new($validator_struct))
} else {
Err(jsonschema::ValidationError::schema(concat!(
$keyword,
" must be set to true"
)))
}
}
};
}
/// Generate a float validator struct and its factory closure.
#[macro_export]
macro_rules! define_float_validator {
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $error_msg:literal) => {
struct $validator_struct;
impl jsonschema::Keyword for $validator_struct {
fn validate<'i>(
&self,
instance: &'i serde_json::Value,
) -> Result<(), jsonschema::ValidationError<'i>> {
match instance.as_f64() {
Some(f) if f.is_finite() => Ok(()),
_ => Err(jsonschema::ValidationError::custom($error_msg)),
}
}
fn is_valid(&self, instance: &serde_json::Value) -> bool {
instance.as_f64().is_some_and(|f| f.is_finite())
}
}
fn $factory_fn<'a>(
_parent: &'a serde_json::Map<String, serde_json::Value>,
value: &'a serde_json::Value,
_path: jsonschema::paths::Location,
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
if value.as_bool() == Some(true) {
Ok(Box::new($validator_struct))
} else {
Err(jsonschema::ValidationError::schema(concat!(
$keyword,
" must be set to true"
)))
}
}
};
}
/// Generate a simple type-check validator (object/array/boolean) and its factory.
#[macro_export]
macro_rules! define_type_validator {
($validator_struct:ident, $factory_fn:ident, $keyword:literal, $check_method:ident, $error_msg:literal) => {
struct $validator_struct;
impl jsonschema::Keyword for $validator_struct {
fn validate<'i>(
&self,
instance: &'i serde_json::Value,
) -> Result<(), jsonschema::ValidationError<'i>> {
if instance.$check_method() {
Ok(())
} else {
Err(jsonschema::ValidationError::custom($error_msg))
}
}
fn is_valid(&self, instance: &serde_json::Value) -> bool {
instance.$check_method()
}
}
fn $factory_fn<'a>(
_parent: &'a serde_json::Map<String, serde_json::Value>,
value: &'a serde_json::Value,
_path: jsonschema::paths::Location,
) -> Result<Box<dyn jsonschema::Keyword>, jsonschema::ValidationError<'a>> {
if value.as_bool() == Some(true) {
Ok(Box::new($validator_struct))
} else {
Err(jsonschema::ValidationError::schema(concat!(
$keyword,
" must be set to true"
)))
}
}
};
}
//! These macros eliminate boilerplate in data access. Each macro takes a
//! compact specification and generates the full implementation, ensuring
//! consistency across all types.
// ---------------------------------------------------------------------------
// Data access macros
+859 -644
View File
File diff suppressed because it is too large. Load diff
+407 -419
View File
File diff suppressed because it is too large. Load diff
+152 -791
View File
File diff suppressed because it is too large. Load diff
+757 -852
View File
File diff suppressed because it is too large. Load diff
+173 -275
View File
@@ -1,18 +1,18 @@
//! TUnion discriminator dispatch (ADR-097 §4).
//! TUnion discriminator dispatch (ADR-003 §4).
//!
//! TUnion supports two discriminator kinds: byte-offset (protocol
//! dispatch, e.g., SFTP type bytes) and field-name (typedef.ts string
//! pattern). This module reads the discriminator value from a byte
//! buffer, looks up the variant schema in the union's `mapping`, and
//! buffer, looks up the variant type in the union's `mapping`, and
//! reports the offset where the variant struct begins.
//!
//! All reads go through [`crate::data_access`] so bounds checks and
//! endianness handling are uniform with the rest of the engine.
use crate::bast::{BastDiscriminator, BastType, BastUnion};
use crate::data_access::{read_enum, read_string, read_u16, read_u32, read_u8};
use crate::error::AlkTypeError;
use crate::schema::{get_alktype_kind, parse_discriminator, DiscriminatorKind, Endian, AlkTypeKind, DISCRIMINATOR_PATH, U32_SIZE};
use serde_json::Value;
use crate::schema::{AlkTypeKind, Endian, DISCRIMINATOR_PATH, U32_SIZE};
const STRING_PREFIX_SIZE: usize = 4;
@@ -39,21 +39,19 @@ pub struct UnionDispatch {
///
/// # Errors
///
/// - [`AlkTypeError::Schema`] if the discriminator annotation is missing
/// or malformed, or if the discriminator `type` is not one of
/// `AlkType:Uint8` / `AlkType:Uint16` / `AlkType:Uint32`.
/// - [`AlkTypeError::Schema`] if the union does not have a byte-offset
/// discriminator.
/// - [`AlkTypeError::Access`] if the buffer is too short to contain the
/// discriminator, or if the read value is not present in the union's
/// `mapping`.
pub fn read_byte_discriminator(
buffer: &[u8],
union_schema: &Value,
union_node: &BastUnion<'_>,
endian: Endian,
) -> Result<UnionDispatch, AlkTypeError> {
let disc = parse_discriminator(union_schema)?;
let (offset, disc_type) = match disc {
DiscriminatorKind::Byte { offset, disc_type } => (offset, disc_type),
DiscriminatorKind::Field { .. } => {
let (offset, disc_type) = match union_node.discriminator() {
BastDiscriminator::Byte { offset, disc_type } => (*offset, *disc_type),
BastDiscriminator::Field { .. } => {
return Err(AlkTypeError::Schema(
"read_byte_discriminator requires a byte-offset discriminator".to_string(),
));
@@ -75,7 +73,7 @@ pub fn read_byte_discriminator(
};
let key = disc_value.to_string();
verify_mapping_key(union_schema, &key, DISCRIMINATOR_PATH, &key)?;
verify_mapping_key(union_node, &key, DISCRIMINATOR_PATH, &key)?;
let variant_offset =
offset
@@ -105,56 +103,44 @@ pub fn read_byte_discriminator(
///
/// # Errors
///
/// - [`AlkTypeError::Schema`] if the discriminator annotation is missing
/// or malformed, the discriminator field is not declared in
/// `properties`, the field has no `AlkType:*` kind, or the field's
/// kind is not one of `AlkType:String` / `AlkType:Uint8` /
/// `AlkType:Enum`.
/// - [`AlkTypeError::Schema`] if the union does not have a field-name
/// discriminator, the discriminator field is not declared in
/// `fields`, or the field's kind is not one of `string` / `uint8` /
/// `enum`.
/// - [`AlkTypeError::Access`] if the buffer is too short to contain the
/// discriminator field, or if the read value is not present in the
/// union's `mapping`.
pub fn read_field_discriminator(
buffer: &[u8],
union_schema: &Value,
union_node: &BastUnion<'_>,
disc_field_offset: usize,
endian: Endian,
) -> Result<UnionDispatch, AlkTypeError> {
let disc = parse_discriminator(union_schema)?;
let name = match disc {
DiscriminatorKind::Field { name } => name,
DiscriminatorKind::Byte { .. } => {
let name = match union_node.discriminator() {
BastDiscriminator::Field { name } => *name,
BastDiscriminator::Byte { .. } => {
return Err(AlkTypeError::Schema(
"read_field_discriminator requires a field-name discriminator".to_string(),
));
}
};
let field_schema = union_schema
.get("properties")
.and_then(Value::as_object)
.and_then(|props| props.get(&name))
.ok_or_else(|| {
AlkTypeError::Schema(format!(
"discriminator field '{name}' not found in union properties"
))
})?;
let disc_field = union_node.fields().iter().find(|f| f.name() == name).ok_or_else(|| {
AlkTypeError::Schema(format!(
"discriminator field '{name}' not found in union fields"
))
})?;
let kind = get_alktype_kind(field_schema)
.and_then(|s| s.parse::<AlkTypeKind>().ok())
.ok_or_else(|| {
AlkTypeError::Schema(format!(
"discriminator field '{name}' has no AlkType:* kind"
))
})?;
let kind = disc_field.ty().alk_kind();
let (key, discriminator_field_size) = match kind {
AlkTypeKind::String => {
let s = read_string(buffer, disc_field_offset, &name, endian)?;
let s = read_string(buffer, disc_field_offset, name, endian)?;
let size =
STRING_PREFIX_SIZE
.checked_add(s.len())
.ok_or_else(|| AlkTypeError::Access {
field_path: name.clone(),
field_path: name.to_string(),
reason: format!(
"string prefix {STRING_PREFIX_SIZE} + data length {} overflows usize",
s.len()
@@ -163,11 +149,11 @@ pub fn read_field_discriminator(
(s.to_string(), size)
}
AlkTypeKind::Uint8 => {
let v = read_u8(buffer, disc_field_offset, &name)?;
let v = read_u8(buffer, disc_field_offset, name)?;
(v.to_string(), 1)
}
AlkTypeKind::Enum => {
let v = read_enum(buffer, disc_field_offset, &name, endian)?;
let v = read_enum(buffer, disc_field_offset, name, endian)?;
(v.to_string(), U32_SIZE)
}
other => {
@@ -177,12 +163,12 @@ pub fn read_field_discriminator(
}
};
verify_mapping_key(union_schema, &key, &name, &key)?;
verify_mapping_key(union_node, &key, name, &key)?;
let variant_offset = disc_field_offset
.checked_add(discriminator_field_size)
.ok_or_else(|| AlkTypeError::Access {
field_path: name.clone(),
field_path: name.to_string(),
reason: format!(
"disc_field_offset {disc_field_offset} + discriminator_field_size {discriminator_field_size} overflows usize"
),
@@ -195,63 +181,39 @@ pub fn read_field_discriminator(
})
}
/// Look up a variant schema from the union's mapping.
/// Look up a variant type from the union's mapping.
///
/// Returns the variant schema. Inline schemas are returned directly.
/// `$ref` pointers of the form `"#/$defs/<name>"` are resolved against
/// the `union_schema`'s own `$defs` block (when the union schema is the
/// schema root). For nested unions whose `$defs` live on an ancestor,
/// the caller (typically `AlkTypeEngine::compile`) is expected to
/// resolve refs before reaching this function, or to inline the
/// variant schemas into the mapping at load time.
/// Returns the variant [`BastType`]. Inline struct/union/enum types are
/// returned directly; `$ref` pointers are returned as
/// [`BastType::Ref`] — the caller resolves them via
/// [`crate::bast::BastDoc::resolve_typeref`] when a concrete definition is needed.
///
/// # Errors
///
/// - [`AlkTypeError::Schema`] if the union has no `mapping` object, the
/// `key` is not present, a `$ref` is malformed, or a `$ref` cannot be
/// resolved against the union schema's own `$defs`.
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str) -> Result<&'a Value, AlkTypeError> {
let mapping = union_schema
.get("mapping")
.and_then(Value::as_object)
.ok_or_else(|| AlkTypeError::Schema("union is missing 'mapping' object".to_string()))?;
let variant = mapping
.get(key)
.ok_or_else(|| AlkTypeError::Schema(format!("unknown mapping key: {key}")))?;
let ref_str = match variant.get("$ref").and_then(Value::as_str) {
Some(r) => r,
None => return Ok(variant),
};
let pointer = ref_str
.strip_prefix('#')
.ok_or_else(|| AlkTypeError::Schema(format!("unsupported $ref form: {ref_str}")))?;
let resolved = resolve_json_pointer(union_schema, pointer).ok_or_else(|| {
AlkTypeError::Schema(format!(
"cannot resolve $ref {ref_str} against union schema; ensure refs are inlined or the union schema contains $defs"
))
})?;
Ok(resolved)
/// - [`AlkTypeError::Schema`] if the union has no `mapping` entries or
/// the `key` is not present.
pub fn resolve_variant<'a>(
union_node: &'a BastUnion<'a>,
key: &str,
) -> Result<&'a BastType<'a>, AlkTypeError> {
union_node
.variant_for(key)
.ok_or_else(|| AlkTypeError::Schema(format!("unknown mapping key: {key}")))
}
/// Get the discriminator size in bytes for a byte-offset discriminator.
///
/// Returns 1 for `AlkType:Uint8`, 2 for `AlkType:Uint16`, and 4 for
/// `AlkType:Uint32`. Field-name discriminators have no fixed size and
/// produce a [`AlkTypeError::Schema`].
/// Returns 1 for `uint8`, 2 for `uint16`, and 4 for `uint32`.
/// Field-name discriminators have no fixed size and produce a
/// [`AlkTypeError::Schema`].
///
/// # Errors
///
/// - [`AlkTypeError::Schema`] if the discriminator annotation is
/// missing/malformed, the discriminator `type` is unsupported, or the
/// discriminator is a field-name discriminator.
pub fn discriminator_size(union_schema: &Value) -> Result<usize, AlkTypeError> {
let disc = parse_discriminator(union_schema)?;
match disc {
DiscriminatorKind::Byte { disc_type, .. } => match disc_type {
/// - [`AlkTypeError::Schema`] if the discriminator is a field-name
/// discriminator.
pub fn discriminator_size(union_node: &BastUnion<'_>) -> Result<usize, AlkTypeError> {
match union_node.discriminator() {
BastDiscriminator::Byte { disc_type, .. } => match disc_type {
AlkTypeKind::Uint8 => Ok(1),
AlkTypeKind::Uint16 => Ok(2),
AlkTypeKind::Uint32 => Ok(4),
@@ -259,24 +221,19 @@ pub fn discriminator_size(union_schema: &Value) -> Result<usize, AlkTypeError> {
"unsupported byte discriminator type: {other}"
))),
},
DiscriminatorKind::Field { .. } => Err(AlkTypeError::Schema(
BastDiscriminator::Field { .. } => Err(AlkTypeError::Schema(
"field-name discriminator has no fixed size".to_string(),
)),
}
}
fn verify_mapping_key(
union_schema: &Value,
union_node: &BastUnion<'_>,
key: &str,
field_path: &str,
raw_value: &str,
) -> Result<(), AlkTypeError> {
let in_mapping = union_schema
.get("mapping")
.and_then(Value::as_object)
.map(|m| m.contains_key(key))
.unwrap_or(false);
if in_mapping {
if union_node.variant_for(key).is_some() {
Ok(())
} else {
Err(AlkTypeError::Access {
@@ -286,84 +243,64 @@ fn verify_mapping_key(
}
}
fn resolve_json_pointer<'a>(root: &'a Value, pointer: &str) -> Option<&'a Value> {
if pointer.is_empty() {
return Some(root);
}
let trimmed = pointer.strip_prefix('/')?;
let mut current = root;
for unescaped in trimmed.split('/') {
let segment = unescape_json_pointer_token(unescaped)?;
current = current.get(&segment)?;
}
Some(current)
}
fn unescape_json_pointer_token(token: &str) -> Option<String> {
let mut out = String::with_capacity(token.len());
let mut chars = token.chars();
while let Some(c) = chars.next() {
match c {
'~' => match chars.next() {
Some('0') => out.push('~'),
Some('1') => out.push('/'),
_ => return None,
},
other => out.push(other),
}
}
Some(out)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::bast::BastDoc;
use serde_json::json;
const LE: Endian = Endian::Little;
const BE: Endian = Endian::Big;
fn byte_union_schema(offset: usize, disc_type: &str) -> Value {
fn doc_union<'a>(root: &'a serde_json::Value, name: &'a str) -> BastUnion<'a> {
let doc = BastDoc::new(root, name).expect("bast doc");
match doc.root_def().kind() {
crate::bast::BastDefKind::Union(u) => u.clone(),
_ => panic!("root must be a union"),
}
}
fn byte_union_root(offset: usize, disc_type: &str) -> serde_json::Value {
json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "offset": offset, "type": disc_type},
"mapping": {
"5": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}},
"6": {"AlkType:Struct": true, "properties": {"len": {"AlkType:Uint16": true}}}
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": offset, "type": disc_type },
"mapping": {
"5": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] },
"6": { "kind": "struct", "fields": [ { "name": "len", "kind": "uint16" } ] }
}
}
}
})
}
fn field_union_schema(field_name: &str, field_kind: &str) -> Value {
let field_schema = match field_kind {
"AlkType:Enum" => json!({
"AlkType:Enum": true,
"enum": ["read", "write"]
}),
_ => json!({field_kind: true}),
};
fn field_union_root(field_name: &str, field_kind: &str) -> serde_json::Value {
let (key_a, key_b) = match field_kind {
"AlkType:String" => ("read", "write"),
"string" => ("read", "write"),
_ => ("0", "1"),
};
json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": field_name},
"properties": {
field_name: field_schema
},
"mapping": {
key_a: {"AlkType:Struct": true, "properties": {"n": {"AlkType:Uint32": true}}},
key_b: {"AlkType:Struct": true, "properties": {"m": {"AlkType:Uint16": true}}}
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "field", "name": field_name },
"fields": [ { "name": field_name, "kind": field_kind } ],
"mapping": {
key_a: { "kind": "struct", "fields": [ { "name": "n", "kind": "uint32" } ] },
key_b: { "kind": "struct", "fields": [ { "name": "m", "kind": "uint16" } ] }
}
}
}
})
}
#[test]
fn read_byte_discriminator_uint8_default_offset() {
let schema = byte_union_schema(0, "AlkType:Uint8");
let root = byte_union_root(0, "uint8");
let u = doc_union(&root, "U");
let buf = [5u8, 0xAA, 0xBB, 0xCC];
let d = read_byte_discriminator(&buf, &schema, LE).expect("read");
let d = read_byte_discriminator(&buf, &u, LE).expect("read");
assert_eq!(d.key, "5");
assert_eq!(d.variant_offset, 1);
assert_eq!(d.discriminator_size, 1);
@@ -371,19 +308,21 @@ mod tests {
#[test]
fn read_byte_discriminator_uint8_big_endian() {
let schema = byte_union_schema(0, "AlkType:Uint8");
let root = byte_union_root(0, "uint8");
let u = doc_union(&root, "U");
let buf = [6u8];
let d = read_byte_discriminator(&buf, &schema, BE).expect("read");
let d = read_byte_discriminator(&buf, &u, BE).expect("read");
assert_eq!(d.key, "6");
assert_eq!(d.variant_offset, 1);
}
#[test]
fn read_byte_discriminator_uint16_little_endian() {
let schema = byte_union_schema(2, "AlkType:Uint16");
let root = byte_union_root(2, "uint16");
let u = doc_union(&root, "U");
let mut buf = vec![0u8; 4];
buf[2..4].copy_from_slice(&5u16.to_le_bytes());
let d = read_byte_discriminator(&buf, &schema, LE).expect("read");
let d = read_byte_discriminator(&buf, &u, LE).expect("read");
assert_eq!(d.key, "5");
assert_eq!(d.variant_offset, 4);
assert_eq!(d.discriminator_size, 2);
@@ -391,19 +330,21 @@ mod tests {
#[test]
fn read_byte_discriminator_uint16_big_endian() {
let schema = byte_union_schema(0, "AlkType:Uint16");
let root = byte_union_root(0, "uint16");
let u = doc_union(&root, "U");
let buf = [0x00, 0x06, 0xAA, 0xBB];
let d = read_byte_discriminator(&buf, &schema, BE).expect("read");
let d = read_byte_discriminator(&buf, &u, BE).expect("read");
assert_eq!(d.key, "6");
assert_eq!(d.variant_offset, 2);
}
#[test]
fn read_byte_discriminator_uint32_little_endian() {
let schema = byte_union_schema(0, "AlkType:Uint32");
let root = byte_union_root(0, "uint32");
let u = doc_union(&root, "U");
let mut buf = vec![0u8; 8];
buf[0..4].copy_from_slice(&5u32.to_le_bytes());
let d = read_byte_discriminator(&buf, &schema, LE).expect("read");
let d = read_byte_discriminator(&buf, &u, LE).expect("read");
assert_eq!(d.key, "5");
assert_eq!(d.variant_offset, 4);
assert_eq!(d.discriminator_size, 4);
@@ -411,19 +352,21 @@ mod tests {
#[test]
fn read_byte_discriminator_uint32_big_endian() {
let schema = byte_union_schema(0, "AlkType:Uint32");
let root = byte_union_root(0, "uint32");
let u = doc_union(&root, "U");
let mut buf = vec![0u8; 8];
buf[0..4].copy_from_slice(&6u32.to_be_bytes());
let d = read_byte_discriminator(&buf, &schema, BE).expect("read");
let d = read_byte_discriminator(&buf, &u, BE).expect("read");
assert_eq!(d.key, "6");
assert_eq!(d.variant_offset, 4);
}
#[test]
fn read_byte_discriminator_unknown_value_is_access_error() {
let schema = byte_union_schema(0, "AlkType:Uint8");
let root = byte_union_root(0, "uint8");
let u = doc_union(&root, "U");
let buf = [99u8];
let err = read_byte_discriminator(&buf, &schema, LE).unwrap_err();
let err = read_byte_discriminator(&buf, &u, LE).unwrap_err();
match err {
AlkTypeError::Access { field_path, reason } => {
assert_eq!(field_path, DISCRIMINATOR_PATH);
@@ -435,29 +378,32 @@ mod tests {
#[test]
fn read_byte_discriminator_buffer_too_short_is_access_error() {
let schema = byte_union_schema(4, "AlkType:Uint32");
let root = byte_union_root(4, "uint32");
let u = doc_union(&root, "U");
let buf = [0u8; 2];
let err = read_byte_discriminator(&buf, &schema, LE).unwrap_err();
let err = read_byte_discriminator(&buf, &u, LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }));
}
#[test]
fn read_byte_discriminator_field_kind_is_schema_error() {
let schema = field_union_schema("type", "AlkType:String");
let root = field_union_root("type", "string");
let u = doc_union(&root, "U");
let buf = [0u8; 16];
let err = read_byte_discriminator(&buf, &schema, LE).unwrap_err();
let err = read_byte_discriminator(&buf, &u, LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn read_field_discriminator_string() {
let schema = field_union_schema("type", "AlkType:String");
let root = field_union_root("type", "string");
let u = doc_union(&root, "U");
let mut buf = vec![0u8; 32];
let value = "read";
let len_bytes = (value.len() as u32).to_le_bytes();
buf[0..4].copy_from_slice(&len_bytes);
buf[4..4 + value.len()].copy_from_slice(value.as_bytes());
let d = read_field_discriminator(&buf, &schema, 0, LE).expect("read");
let d = read_field_discriminator(&buf, &u, 0, LE).expect("read");
assert_eq!(d.key, "read");
assert_eq!(d.variant_offset, 4 + value.len());
assert_eq!(d.discriminator_size, 4 + value.len());
@@ -465,45 +411,37 @@ mod tests {
#[test]
fn read_field_discriminator_uint8() {
let schema = field_union_schema("type", "AlkType:Uint8");
let root = field_union_root("type", "uint8");
let u = doc_union(&root, "U");
let mut buf = vec![0u8; 8];
buf[0] = 0;
let d = read_field_discriminator(&buf, &schema, 0, LE).expect("read");
let d = read_field_discriminator(&buf, &u, 0, LE).expect("read");
assert_eq!(d.key, "0");
assert_eq!(d.variant_offset, 1);
assert_eq!(d.discriminator_size, 1);
}
#[test]
fn read_field_discriminator_enum() {
let schema = field_union_schema("type", "AlkType:Enum");
let mut buf = vec![0u8; 8];
buf[0..4].copy_from_slice(&0u32.to_le_bytes());
let d = read_field_discriminator(&buf, &schema, 0, LE).expect("read");
assert_eq!(d.key, "0");
assert_eq!(d.variant_offset, 4);
assert_eq!(d.discriminator_size, 4);
}
#[test]
fn read_field_discriminator_string_big_endian() {
let schema = field_union_schema("type", "AlkType:String");
let root = field_union_root("type", "string");
let u = doc_union(&root, "U");
let mut buf = vec![0u8; 32];
let value = "write";
let len_bytes = (value.len() as u32).to_be_bytes();
buf[0..4].copy_from_slice(&len_bytes);
buf[4..4 + value.len()].copy_from_slice(value.as_bytes());
let d = read_field_discriminator(&buf, &schema, 0, BE).expect("read");
let d = read_field_discriminator(&buf, &u, 0, BE).expect("read");
assert_eq!(d.key, "write");
assert_eq!(d.variant_offset, 4 + value.len());
}
#[test]
fn read_field_discriminator_unknown_value_is_access_error() {
let schema = field_union_schema("type", "AlkType:Uint8");
let root = field_union_root("type", "uint8");
let u = doc_union(&root, "U");
let mut buf = vec![0u8; 8];
buf[0] = 99;
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
let err = read_field_discriminator(&buf, &u, 0, LE).unwrap_err();
match err {
AlkTypeError::Access { field_path, reason } => {
assert_eq!(field_path, "type");
@@ -515,131 +453,91 @@ mod tests {
#[test]
fn read_field_discriminator_field_not_found_is_schema_error() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "missing"},
"properties": {"other": {"AlkType:Uint8": true}},
"mapping": {"5": {"AlkType:Struct": true}}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "field", "name": "missing" },
"fields": [ { "name": "other", "kind": "uint8" } ],
"mapping": { "5": { "kind": "struct", "fields": [] } }
}
}
});
let u = doc_union(&root, "U");
let buf = [0u8; 4];
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn read_field_discriminator_no_alktype_kind_is_schema_error() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "type"},
"properties": {"type": {"type": "string"}},
"mapping": {"read": {"AlkType:Struct": true}}
});
let buf = [0u8; 4];
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
let err = read_field_discriminator(&buf, &u, 0, LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn read_field_discriminator_unsupported_kind_is_schema_error() {
let schema = field_union_schema("type", "AlkType:Float32");
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "field", "name": "type" },
"fields": [ { "name": "type", "kind": "float32" } ],
"mapping": { "read": { "kind": "struct", "fields": [] } }
}
}
});
let u = doc_union(&root, "U");
let buf = [0u8; 8];
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
let err = read_field_discriminator(&buf, &u, 0, LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn read_field_discriminator_byte_kind_is_schema_error() {
let schema = byte_union_schema(0, "AlkType:Uint8");
let root = byte_union_root(0, "uint8");
let u = doc_union(&root, "U");
let buf = [5u8];
let err = read_field_discriminator(&buf, &schema, 0, LE).unwrap_err();
let err = read_field_discriminator(&buf, &u, 0, LE).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn resolve_variant_inline_schema() {
let schema = byte_union_schema(0, "AlkType:Uint8");
let variant = resolve_variant(&schema, "5").expect("resolve");
assert_eq!(
variant.get("AlkType:Struct").and_then(Value::as_bool),
Some(true)
);
}
#[test]
fn resolve_variant_ref_against_own_defs() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte"},
"mapping": {
"5": {"$ref": "#/$defs/Read"}
},
"$defs": {
"Read": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}}
}
});
let variant = resolve_variant(&schema, "5").expect("resolve");
assert_eq!(
variant.get("AlkType:Struct").and_then(Value::as_bool),
Some(true)
);
fn resolve_variant_returns_variant_type() {
let root = byte_union_root(0, "uint8");
let u = doc_union(&root, "U");
let variant = resolve_variant(&u, "5").expect("resolve");
assert!(matches!(variant, BastType::Struct(_)));
}
#[test]
fn resolve_variant_unknown_key_is_schema_error() {
let schema = byte_union_schema(0, "AlkType:Uint8");
let err = resolve_variant(&schema, "999").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn resolve_variant_missing_mapping_is_schema_error() {
let schema = json!({"AlkType:Union": true, "discriminator": {"kind": "byte"}});
let err = resolve_variant(&schema, "5").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn resolve_variant_unresolvable_ref_is_schema_error() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte"},
"mapping": {
"5": {"$ref": "#/$defs/Read"}
}
});
let err = resolve_variant(&schema, "5").unwrap_err();
let root = byte_union_root(0, "uint8");
let u = doc_union(&root, "U");
let err = resolve_variant(&u, "999").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn discriminator_size_uint8() {
let schema = byte_union_schema(0, "AlkType:Uint8");
assert_eq!(discriminator_size(&schema).unwrap(), 1);
let root = byte_union_root(0, "uint8");
let u = doc_union(&root, "U");
assert_eq!(discriminator_size(&u).unwrap(), 1);
}
#[test]
fn discriminator_size_uint16() {
let schema = byte_union_schema(0, "AlkType:Uint16");
assert_eq!(discriminator_size(&schema).unwrap(), 2);
let root = byte_union_root(0, "uint16");
let u = doc_union(&root, "U");
assert_eq!(discriminator_size(&u).unwrap(), 2);
}
#[test]
fn discriminator_size_uint32() {
let schema = byte_union_schema(0, "AlkType:Uint32");
assert_eq!(discriminator_size(&schema).unwrap(), 4);
let root = byte_union_root(0, "uint32");
let u = doc_union(&root, "U");
assert_eq!(discriminator_size(&u).unwrap(), 4);
}
#[test]
fn discriminator_size_field_kind_is_schema_error() {
let schema = field_union_schema("type", "AlkType:String");
let err = discriminator_size(&schema).unwrap_err();
let root = field_union_root("type", "string");
let u = doc_union(&root, "U");
let err = discriminator_size(&u).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn discriminator_size_missing_discriminator_is_schema_error() {
let schema = json!({"AlkType:Union": true});
let err = discriminator_size(&schema).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
}
}
+101 -1106
View File
File diff suppressed because it is too large. Load diff
+161 -170
View File
@@ -4,27 +4,34 @@
//! accessors, validation convenience methods, and the aligned-mode
//! `read_field` / `write_field` round-trip for the fixed-size primitive
//! kinds and length-prefixed `String` / `Bytes`.
//!
//! All schemas are BAST documents (`{ "$defs": { ... } }` with `kind`-
//! based vocabulary).
use alktype::*;
use serde_json::json;
fn mixed_fixed_struct_schema() -> serde_json::Value {
fn mixed_fixed_struct_doc() -> serde_json::Value {
json!({
"AlkType:Struct": true,
"endian": "little",
"properties": {
"flag": { "AlkType:Uint8": true },
"id": { "AlkType:Uint32": true },
"score": { "AlkType:Float32": true },
"tag": { "AlkType:String": true }
"$defs": {
"S": {
"kind": "struct",
"endian": "little",
"fields": [
{ "name": "flag", "kind": "uint8" },
{ "name": "id", "kind": "uint32" },
{ "name": "score", "kind": "float32" },
{ "name": "tag", "kind": "string" }
]
}
}
})
}
#[test]
fn compile_aligned_builds_engine_with_offset_map() -> Result<(), AlkTypeError> {
let mut schema = mixed_fixed_struct_schema();
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let doc = mixed_fixed_struct_doc();
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
assert_eq!(engine.mode(), LayoutMode::Aligned);
assert!(engine.offset_map().is_some());
assert!(engine.layout_builder().is_none());
@@ -34,8 +41,8 @@ fn compile_aligned_builds_engine_with_offset_map() -> Result<(), AlkTypeError> {
#[test]
fn compile_packed_builds_engine_with_builder_and_reader() -> Result<(), AlkTypeError> {
let mut schema = mixed_fixed_struct_schema();
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
let doc = mixed_fixed_struct_doc();
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
assert_eq!(engine.mode(), LayoutMode::Packed);
assert!(engine.offset_map().is_none());
assert!(engine.layout_builder().is_some());
@@ -44,147 +51,96 @@ fn compile_packed_builds_engine_with_builder_and_reader() -> Result<(), AlkTypeE
}
#[test]
fn compile_normalizes_bare_name_refs() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"child": { "$ref": "Child" }
},
fn compile_resolves_ref_fields() -> Result<(), AlkTypeError> {
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{ "name": "child", "kind": { "$ref": "#/$defs/Child" } }
]
},
"Child": {
"AlkType:Struct": true,
"properties": { "x": { "AlkType:Uint8": true } }
"kind": "struct",
"fields": [ { "name": "x", "kind": "uint8" } ]
}
}
});
let _engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
assert_eq!(
schema["properties"]["child"]["$ref"],
json!("#/$defs/Child")
);
let _engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
Ok(())
}
#[test]
fn compile_leaves_full_pointer_refs_unchanged() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"child": { "$ref": "#/$defs/Child" }
},
"$defs": {
"Child": {
"AlkType:Struct": true,
"properties": { "x": { "AlkType:Uint8": true } }
}
}
});
let _engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
assert_eq!(
schema["properties"]["child"]["$ref"],
json!("#/$defs/Child")
);
Ok(())
fn compile_returns_schema_error_when_no_defs() {
let doc = json!({ "type": "object", "properties": {} });
let err = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn compile_returns_schema_error_when_no_alktype_kind() {
let mut schema = json!({ "type": "object", "properties": {} });
let err = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned).unwrap_err();
fn compile_returns_schema_error_for_missing_root() {
let doc = json!({ "$defs": { "Other": { "kind": "struct", "fields": [] } } });
let err = AlkTypeEngine::compile(&doc, "Missing", LayoutMode::Aligned, None).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn endian_parsed_from_schema_big() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"endian": "big",
"properties": { "id": { "AlkType:Uint32": true } }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"endian": "big",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
assert_eq!(engine.endian(), Endian::Big);
Ok(())
}
#[test]
fn endian_defaults_to_little() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": { "id": { "AlkType:Uint32": true } }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
assert_eq!(engine.endian(), Endian::Little);
Ok(())
}
#[test]
fn validate_json_accepts_valid_instance() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"type": "object",
"properties": {
"id": { "AlkType:Uint32": true, "type": "integer" }
},
"required": ["id"]
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
assert!(engine.validate_json(&json!({"id": 42})).is_ok());
Ok(())
}
#[test]
fn validate_json_rejects_invalid_instance() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"type": "object",
"properties": {
"id": { "AlkType:Uint32": true, "type": "integer" }
},
"required": ["id"]
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let err = engine.validate_json(&json!({"id": -1})).unwrap_err();
assert!(matches!(err, AlkTypeError::Validation(_)), "got {err:?}");
Ok(())
}
#[test]
fn is_valid_json_returns_bool() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"type": "object",
"properties": {
"id": { "AlkType:Uint32": true, "type": "integer" }
},
"required": ["id"]
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
assert!(engine.is_valid_json(&json!({"id": 42})));
assert!(!engine.is_valid_json(&json!({"id": -1})));
Ok(())
}
#[test]
fn read_write_aligned_round_trips_all_fixed_size_kinds() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"endian": "little",
"properties": {
"i8": { "AlkType:Int8": true },
"u8": { "AlkType:Uint8": true },
"i16": { "AlkType:Int16": true },
"u16": { "AlkType:Uint16": true },
"i32": { "AlkType:Int32": true },
"u32": { "AlkType:Uint32": true },
"i64": { "AlkType:Int64": true },
"u64": { "AlkType:Uint64": true },
"f32": { "AlkType:Float32": true },
"f64": { "AlkType:Float64": true },
"b": { "AlkType:Boolean": true },
"e": { "AlkType:Enum": true }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"endian": "little",
"fields": [
{ "name": "i8", "kind": "int8" },
{ "name": "u8", "kind": "uint8" },
{ "name": "i16", "kind": "int16" },
{ "name": "u16", "kind": "uint16" },
{ "name": "i32", "kind": "int32" },
{ "name": "u32", "kind": "uint32" },
{ "name": "i64", "kind": "int64" },
{ "name": "u64", "kind": "uint64" },
{ "name": "f32", "kind": "float32" },
{ "name": "f64", "kind": "float64" },
{ "name": "b", "kind": "bool" },
{ "name": "e", "kind": { "$ref": "#/$defs/E" } }
]
},
"E": { "kind": "enum", "values": ["A", "B", "C"] }
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
let mut buffer = vec![0u8; offset_map.total_size()];
@@ -230,13 +186,15 @@ fn read_write_aligned_round_trips_all_fixed_size_kinds() -> Result<(), AlkTypeEr
#[test]
fn read_write_aligned_round_trips_string() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"name": { "AlkType:String": true }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "name", "kind": "string" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
let mut buffer = vec![0u8; offset_map.total_size() + 64];
engine.write_field(&mut buffer, "name", &FieldValue::String("hello world"))?;
@@ -249,13 +207,15 @@ fn read_write_aligned_round_trips_string() -> Result<(), AlkTypeError> {
#[test]
fn read_write_aligned_round_trips_bytes() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"blob": { "AlkType:Bytes": true }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "blob", "kind": "bytes" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
let payload = b"the quick brown fox".to_vec();
let mut buffer = vec![0u8; offset_map.total_size() + payload.len()];
@@ -269,11 +229,15 @@ fn read_write_aligned_round_trips_bytes() -> Result<(), AlkTypeError> {
#[test]
fn read_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": { "id": { "AlkType:Uint32": true } }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
let buffer = [0u8; 4];
let err = engine.read_field(&buffer, "id").unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
@@ -282,11 +246,15 @@ fn read_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError>
#[test]
fn write_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": { "id": { "AlkType:Uint32": true } }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Packed, None)?;
let mut buffer = [0u8; 4];
let err = engine
.write_field(&mut buffer, "id", &FieldValue::U32(1))
@@ -297,11 +265,15 @@ fn write_field_returns_access_error_in_packed_mode() -> Result<(), AlkTypeError>
#[test]
fn read_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": { "id": { "AlkType:Uint32": true } }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let buffer = [0u8; 8];
let err = engine.read_field(&buffer, "missing").unwrap_err();
assert!(matches!(err, AlkTypeError::Offset { .. }), "got {err:?}");
@@ -310,11 +282,15 @@ fn read_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeError
#[test]
fn write_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": { "id": { "AlkType:Uint32": true } }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let mut buffer = [0u8; 8];
let err = engine
.write_field(&mut buffer, "missing", &FieldValue::U32(1))
@@ -324,30 +300,38 @@ fn write_field_returns_offset_error_for_missing_path() -> Result<(), AlkTypeErro
}
#[test]
fn read_field_returns_access_error_for_composite_types() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"vals": {
"AlkType:Array": true,
"items": { "AlkType:Uint32": true }
fn read_field_returns_error_for_composite_types() -> Result<(), AlkTypeError> {
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{ "name": "vals", "kind": { "kind": "array", "element": "uint32", "count": 2 } }
]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let buffer = [0u8; 8];
let err = engine.read_field(&buffer, "vals").unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
assert!(
matches!(err, AlkTypeError::Access { .. } | AlkTypeError::Offset { .. }),
"got {err:?}"
);
Ok(())
}
#[test]
fn write_field_returns_access_error_for_composite_value() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": { "id": { "AlkType:Uint32": true } }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let mut buffer = [0u8; 8];
let err = engine
.write_field(&mut buffer, "id", &FieldValue::Struct { start: 0, end: 4 })
@@ -358,19 +342,26 @@ fn write_field_returns_access_error_for_composite_value() -> Result<(), AlkTypeE
#[test]
fn read_field_aligned_reads_nested_struct_byte_range() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"header": {
"AlkType:Struct": true,
"properties": {
"version": { "AlkType:Uint8": true },
"magic": { "AlkType:Uint32": true }
}
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{
"name": "header",
"kind": {
"kind": "struct",
"fields": [
{ "name": "version", "kind": "uint8" },
{ "name": "magic", "kind": "uint32" }
]
}
}
]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let offset_map = engine.offset_map().expect("aligned mode");
let mut buffer = vec![0u8; offset_map.total_size()];
@@ -386,4 +377,4 @@ fn read_field_aligned_reads_nested_struct_byte_range() -> Result<(), AlkTypeErro
FieldValue::U32(0xCAFEBABE)
);
Ok(())
}
}
+163 -138
View File
@@ -2,10 +2,13 @@
//!
//! Exercises the `AlkTypeError` variants across the crate:
//! `Access` (buffer too short, invalid UTF-8, invalid boolean byte,
//! unknown discriminator value), `Schema` (missing AlkType kind,
//! malformed discriminator annotation), and `Offset` (missing
//! variable-length field size in `LayoutBuilder::build`).
//! unknown discriminator value), `Schema` (missing root, malformed
//! discriminator annotation), and `Offset` (missing variable-length
//! field size in `LayoutBuilder::build`).
//!
//! All schemas are BAST documents.
use alktype::bast::BastDoc;
use alktype::data_access;
use alktype::tunion;
use alktype::*;
@@ -149,128 +152,145 @@ fn write_string_buffer_too_short_returns_access_error() {
}
#[test]
fn compile_missing_alktype_kind_returns_schema_error() {
let mut schema = json!({ "type": "object", "properties": {} });
let err = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned).unwrap_err();
fn compile_missing_defs_returns_schema_error() {
let doc = json!({ "type": "object", "properties": {} });
let err = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn offset_map_compute_missing_alktype_kind_returns_schema_error() {
let schema = json!({ "type": "object", "properties": {} });
let err = OffsetMap::compute(&schema).unwrap_err();
fn compile_missing_root_returns_schema_error() {
let doc = json!({ "$defs": { "Other": { "kind": "struct", "fields": [] } } });
let err = AlkTypeEngine::compile(&doc, "Missing", LayoutMode::Aligned, None).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn offset_map_compute_non_struct_top_level_returns_schema_error() {
let schema = json!({ "AlkType:Uint32": true });
let err = OffsetMap::compute(&schema).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": { "1": { "$ref": "#/$defs/A" } }
},
"A": { "kind": "struct", "fields": [] }
}
});
let doc = BastDoc::new(&root, "U").expect("doc");
let err = OffsetMap::compute(&doc).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)));
}
#[test]
fn layout_builder_new_missing_alktype_kind_returns_schema_error() {
let schema = json!({ "type": "object", "properties": {} });
let err = LayoutBuilder::new(&schema).unwrap_err();
fn layout_builder_new_missing_root_returns_schema_error() {
let root = json!({ "$defs": { "Other": { "kind": "struct", "fields": [] } } });
let err = LayoutBuilder::new(&root, "Missing").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn layout_builder_new_non_struct_top_level_returns_schema_error() {
let schema = json!({ "AlkType:Uint32": true });
let err = LayoutBuilder::new(&schema).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn parse_discriminator_missing_returns_schema_error() {
let schema = json!({"AlkType:Union": true});
let err = parse_discriminator(&schema).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn parse_discriminator_field_missing_name_returns_schema_error() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field"}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": { "1": { "$ref": "#/$defs/A" } }
},
"A": { "kind": "struct", "fields": [] }
}
});
let err = parse_discriminator(&schema).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn parse_discriminator_unknown_kind_returns_schema_error() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "magic"}
});
let err = parse_discriminator(&schema).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn parse_discriminator_byte_invalid_type_returns_schema_error() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Float32"}
});
let err = parse_discriminator(&schema).unwrap_err();
let err = LayoutBuilder::new(&root, "U").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn read_byte_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
"mapping": {"5": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "type": "uint8" },
"mapping": {
"5": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
}
}
}
});
let doc = BastDoc::new(&root, "U")?;
let union_def = match doc.root_def().kind() {
alktype::bast::BastDefKind::Union(u) => u,
_ => unreachable!(),
};
let buffer = [99u8, 0x00, 0x00];
let err = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Little).unwrap_err();
let err = tunion::read_byte_discriminator(&buffer, union_def, Endian::Little).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
Ok(())
}
#[test]
fn read_byte_discriminator_buffer_too_short_returns_access_error() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "offset": 4, "type": "AlkType:Uint32"},
"mapping": {"5": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 4, "type": "uint32" },
"mapping": {
"5": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
}
}
}
});
let doc = BastDoc::new(&root, "U")?;
let union_def = match doc.root_def().kind() {
alktype::bast::BastDefKind::Union(u) => u,
_ => unreachable!(),
};
let buffer = [0u8; 2];
let err = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Little).unwrap_err();
let err = tunion::read_byte_discriminator(&buffer, union_def, Endian::Little).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
Ok(())
}
#[test]
fn read_field_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "type"},
"properties": {"type": {"AlkType:Uint8": true}},
"mapping": {"0": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "field", "name": "type" },
"fields": [ { "name": "type", "kind": "uint8" } ],
"mapping": {
"0": { "kind": "struct", "fields": [ { "name": "x", "kind": "uint8" } ] }
}
}
}
});
let doc = BastDoc::new(&root, "U")?;
let union_def = match doc.root_def().kind() {
alktype::bast::BastDefKind::Union(u) => u,
_ => unreachable!(),
};
let mut buffer = vec![0u8; 8];
buffer[0] = 99;
let err =
tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little).unwrap_err();
tunion::read_field_discriminator(&buffer, union_def, 0, Endian::Little).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
Ok(())
}
#[test]
fn layout_builder_missing_var_size_returns_offset_error() {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"name": { "AlkType:String": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "name", "kind": "string" } ]
}
}
});
let builder = LayoutBuilder::new(&schema).expect("builder");
let builder = LayoutBuilder::new(&root, "S").expect("builder");
let empty: HashMap<String, usize> = HashMap::new();
let err = builder.build(&empty).unwrap_err();
match err {
@@ -285,39 +305,28 @@ fn layout_builder_missing_var_size_returns_offset_error() {
}
}
#[test]
fn layout_builder_missing_array_data_size_returns_offset_error() {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"vals": {
"AlkType:Array": true,
"items": { "AlkType:Uint32": true }
}
}
});
let builder = LayoutBuilder::new(&schema).expect("builder");
let empty: HashMap<String, usize> = HashMap::new();
let err = builder.build(&empty).unwrap_err();
assert!(matches!(err, AlkTypeError::Offset { .. }), "got {err:?}");
}
#[test]
fn layout_builder_missing_discriminator_value_returns_offset_error() {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"payload": {
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
"mapping": {"5": {"$ref": "#/$defs/Read"}}
}
},
let root = json!({
"$defs": {
"Read": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}
"S": {
"kind": "struct",
"fields": [
{ "name": "payload", "kind": { "$ref": "#/$defs/Packet" } }
]
},
"Packet": {
"kind": "union",
"discriminator": { "kind": "byte", "type": "uint8" },
"mapping": { "5": { "$ref": "#/$defs/Read" } }
},
"Read": {
"kind": "struct",
"fields": [ { "name": "x", "kind": "uint8" } ]
}
}
});
let builder = LayoutBuilder::new(&schema).expect("builder");
let builder = LayoutBuilder::new(&root, "S").expect("builder");
let empty: HashMap<String, usize> = HashMap::new();
let err = builder.build(&empty).unwrap_err();
assert!(matches!(err, AlkTypeError::Offset { .. }), "got {err:?}");
@@ -325,20 +334,26 @@ fn layout_builder_missing_discriminator_value_returns_offset_error() {
#[test]
fn layout_builder_unknown_discriminator_value_returns_offset_error() {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"payload": {
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
"mapping": {"5": {"$ref": "#/$defs/Read"}}
}
},
let root = json!({
"$defs": {
"Read": {"AlkType:Struct": true, "properties": {"x": {"AlkType:Uint8": true}}}
"S": {
"kind": "struct",
"fields": [
{ "name": "payload", "kind": { "$ref": "#/$defs/Packet" } }
]
},
"Packet": {
"kind": "union",
"discriminator": { "kind": "byte", "type": "uint8" },
"mapping": { "5": { "$ref": "#/$defs/Read" } }
},
"Read": {
"kind": "struct",
"fields": [ { "name": "x", "kind": "uint8" } ]
}
}
});
let builder = LayoutBuilder::new(&schema).expect("builder");
let builder = LayoutBuilder::new(&root, "S").expect("builder");
let mut vs = HashMap::new();
vs.insert("payload.__discriminator".to_string(), 99);
let err = builder.build(&vs).unwrap_err();
@@ -352,64 +367,74 @@ fn layout_builder_unknown_discriminator_value_returns_offset_error() {
#[test]
fn sequential_reader_buffer_too_short_returns_access_error() {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"id": { "AlkType:Uint32": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "id", "kind": "uint32" } ]
}
}
});
let buffer = [0u8; 2];
let mut reader = SequentialReader::new(&schema).unwrap();
let mut reader = SequentialReader::new(&root, "S").unwrap();
let err = reader.read_next(&buffer).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
}
#[test]
fn sequential_reader_unknown_field_returns_schema_error() {
let schema = json!({
"AlkType:Struct": true,
"properties": { "a": { "AlkType:Uint8": true } }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "a", "kind": "uint8" } ]
}
}
});
let buffer = [0u8; 4];
let mut reader = SequentialReader::new(&schema).unwrap();
let mut reader = SequentialReader::new(&root, "S").unwrap();
let err = reader.read_field(&buffer, "missing").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn sequential_reader_new_non_struct_returns_schema_error() {
let schema = json!({ "AlkType:Uint32": true });
let err = SequentialReader::new(&schema).unwrap_err();
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": { "1": { "$ref": "#/$defs/A" } }
},
"A": { "kind": "struct", "fields": [] }
}
});
let err = SequentialReader::new(&root, "U").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn read_string_indirect_data_region_too_short_returns_access_error() {
fn read_string_indirect_data_too_short_returns_access_error() {
let mut index = [0u8; 8];
let _ = data_access::write_u32(&mut index, 0, 100, "idx.off", Endian::Little);
let _ = data_access::write_u32(&mut index, 4, 10, "idx.len", Endian::Little);
let data_region = b"too short";
let err = data_access::read_bytes_indirect(&index, 0, data_region, "blob", Endian::Little)
.unwrap_err();
let err = data_access::read_bytes_indirect(&index, 0, "blob", Endian::Little).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
}
#[test]
fn read_bytes_indirect_index_too_short_returns_access_error() {
let buffer = [0u8; 4];
let data_region = b"anything";
let err = data_access::read_bytes_indirect(&buffer, 0, data_region, "blob", Endian::Little)
.unwrap_err();
let err = data_access::read_bytes_indirect(&buffer, 0, "blob", Endian::Little).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
}
#[test]
fn read_string_indirect_invalid_utf8_returns_access_error() {
let data_region: &[u8] = &[0xFF, 0xFE, 0xFD];
let mut index = [0u8; 8];
let _ = data_access::write_u32(&mut index, 0, 0, "idx.off", Endian::Little);
let _ = data_access::write_u32(&mut index, 4, 3, "idx.len", Endian::Little);
let err = data_access::read_string_indirect(&index, 0, data_region, "name", Endian::Little)
.unwrap_err();
let mut buf = vec![0u8; 8 + 3];
let _ = data_access::write_u32(&mut buf, 0, 8, "idx.off", Endian::Little);
let _ = data_access::write_u32(&mut buf, 4, 3, "idx.len", Endian::Little);
buf[8..11].copy_from_slice(&[0xFF, 0xFE, 0xFD]);
let err = data_access::read_string_indirect(&buf, 0, "name", Endian::Little).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
}
}
+202 -134
View File
@@ -8,7 +8,12 @@
//! walks. Each test writes values to a buffer at computed offsets and
//! reads them back, asserting both the values and (where applicable)
//! the byte positions.
//!
//! All schemas are BAST documents (`{ "$defs": { ... } }` with `kind`-
//! based vocabulary). The root type name is passed to `OffsetMap::compute`
//! / `LayoutBuilder::new` / `SequentialReader::new` / `AlkTypeEngine::compile`.
use alktype::bast::BastDoc;
use alktype::data_access;
use alktype::tunion;
use alktype::*;
@@ -21,16 +26,21 @@ fn var_sizes(pairs: &[(&str, usize)]) -> HashMap<String, usize> {
#[test]
fn fixed_size_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"id": { "AlkType:Uint32": true },
"score": { "AlkType:Float32": true },
"flag": { "AlkType:Uint8": true },
"count": { "AlkType:Uint16": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{ "name": "id", "kind": "uint32" },
{ "name": "score", "kind": "float32" },
{ "name": "flag", "kind": "uint8" },
{ "name": "count", "kind": "uint16" }
]
}
}
});
let offset_map = OffsetMap::compute(&schema)?;
let doc = BastDoc::new(&root, "S")?;
let offset_map = OffsetMap::compute(&doc)?;
let mut buffer = vec![0u8; offset_map.total_size()];
let id_range = offset_map.get("id").expect("id range");
@@ -64,16 +74,20 @@ fn fixed_size_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
#[test]
fn fixed_size_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"endian": "little",
"properties": {
"id": { "AlkType:Uint32": true },
"score": { "AlkType:Float32": true },
"flag": { "AlkType:Uint8": true }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"endian": "little",
"fields": [
{ "name": "id", "kind": "uint32" },
{ "name": "score", "kind": "float32" },
{ "name": "flag", "kind": "uint8" }
]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
let mut buffer = vec![0u8; offset_map.total_size()];
@@ -107,13 +121,15 @@ fn string_round_trip_via_data_access() -> Result<(), AlkTypeError> {
#[test]
fn string_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"name": { "AlkType:String": true }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [ { "name": "name", "kind": "string" } ]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let offset_map = engine.offset_map().expect("aligned mode has offset_map");
let mut buffer = vec![0u8; offset_map.total_size() + 64];
engine.write_field(&mut buffer, "name", &FieldValue::String("hello"))?;
@@ -141,20 +157,28 @@ fn bytes_round_trip_via_data_access() -> Result<(), AlkTypeError> {
#[test]
fn nested_struct_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"header": {
"AlkType:Struct": true,
"properties": {
"version": { "AlkType:Uint32": true },
"magic": { "AlkType:Uint32": true }
}
},
"payload": { "AlkType:Bytes": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{
"name": "header",
"kind": {
"kind": "struct",
"fields": [
{ "name": "version", "kind": "uint32" },
{ "name": "magic", "kind": "uint32" }
]
}
},
{ "name": "payload", "kind": "bytes" }
]
}
}
});
let offset_map = OffsetMap::compute(&schema)?;
let doc = BastDoc::new(&root, "S")?;
let offset_map = OffsetMap::compute(&doc)?;
let header_version = offset_map.get("header.version").expect("header.version");
let header_magic = offset_map.get("header.magic").expect("header.magic");
@@ -210,20 +234,27 @@ fn nested_struct_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
#[test]
fn nested_struct_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
let mut schema = json!({
"AlkType:Struct": true,
"properties": {
"header": {
"AlkType:Struct": true,
"properties": {
"version": { "AlkType:Uint8": true },
"flags": { "AlkType:Uint8": true }
}
},
"payload_len": { "AlkType:Uint32": true }
let doc = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{
"name": "header",
"kind": {
"kind": "struct",
"fields": [
{ "name": "version", "kind": "uint8" },
{ "name": "flags", "kind": "uint8" }
]
}
},
{ "name": "payload_len", "kind": "uint32" }
]
}
}
});
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Aligned)?;
let engine = AlkTypeEngine::compile(&doc, "S", LayoutMode::Aligned, None)?;
let offset_map = engine.offset_map().expect("aligned mode");
assert_eq!(offset_map.get("header.version").unwrap().start, 0);
@@ -252,17 +283,21 @@ fn nested_struct_round_trip_via_engine_aligned() -> Result<(), AlkTypeError> {
#[test]
fn big_endian_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Struct": true,
"endian": "big",
"properties": {
"id": { "AlkType:Uint32": true },
"offset": { "AlkType:Float64": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "id", "kind": "uint32" },
{ "name": "offset", "kind": "float64" }
]
}
}
});
let offset_map = OffsetMap::compute(&schema)?;
let endian = Endian::from_schema(&schema);
assert_eq!(endian, Endian::Big);
let doc = BastDoc::new(&root, "S")?;
let offset_map = OffsetMap::compute(&doc)?;
let endian = Endian::Big;
let id_range = offset_map.get("id").expect("id");
let offset_range = offset_map.get("offset").expect("offset");
@@ -290,14 +325,19 @@ fn big_endian_round_trip_via_offset_map() -> Result<(), AlkTypeError> {
#[test]
fn alignment_padding_round_trip_u8_then_u32() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Struct": true,
"properties": {
"flag": { "AlkType:Uint8": true },
"id": { "AlkType:Uint32": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"fields": [
{ "name": "flag", "kind": "uint8" },
{ "name": "id", "kind": "uint32" }
]
}
}
});
let offset_map = OffsetMap::compute(&schema)?;
let doc = BastDoc::new(&root, "S")?;
let offset_map = OffsetMap::compute(&doc)?;
let flag_range = offset_map.get("flag").expect("flag");
let id_range = offset_map.get("id").expect("id");
@@ -335,16 +375,20 @@ fn alignment_padding_round_trip_u8_then_u32() -> Result<(), AlkTypeError> {
#[test]
fn packed_layout_round_trip_via_layout_builder() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Struct": true,
"endian": "little",
"properties": {
"flag": { "AlkType:Uint8": true },
"id": { "AlkType:Uint32": true },
"payload": { "AlkType:String": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"endian": "little",
"fields": [
{ "name": "flag", "kind": "uint8" },
{ "name": "id", "kind": "uint32" },
{ "name": "payload", "kind": "string" }
]
}
}
});
let builder = LayoutBuilder::new(&schema)?;
let builder = LayoutBuilder::new(&root, "S")?;
let layout = builder.build(&var_sizes(&[("payload", 10)]))?;
let flag_pos = layout.get("flag").expect("flag");
@@ -392,16 +436,20 @@ fn packed_layout_round_trip_via_layout_builder() -> Result<(), AlkTypeError> {
#[test]
fn sequential_reader_round_trip_packed_buffer() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Struct": true,
"endian": "little",
"properties": {
"id": { "AlkType:Uint8": true },
"name": { "AlkType:String": true },
"tail": { "AlkType:Uint8": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"endian": "little",
"fields": [
{ "name": "id", "kind": "uint8" },
{ "name": "name", "kind": "string" },
{ "name": "tail", "kind": "uint8" }
]
}
}
});
let builder = LayoutBuilder::new(&schema)?;
let builder = LayoutBuilder::new(&root, "S")?;
let payload = "hello";
let layout = builder.build(&var_sizes(&[("name", payload.len())]))?;
@@ -411,7 +459,7 @@ fn sequential_reader_round_trip_packed_buffer() -> Result<(), AlkTypeError> {
let after = 1 + 4 + payload.len();
data_access::write_u8(&mut buffer, after, 99, "tail")?;
let mut reader = SequentialReader::new(&schema)?;
let mut reader = SequentialReader::new(&root, "S")?;
assert_eq!(reader.endian(), Endian::Little);
assert_eq!(reader.position(), 0);
@@ -436,13 +484,17 @@ fn sequential_reader_round_trip_packed_buffer() -> Result<(), AlkTypeError> {
#[test]
fn sequential_reader_read_field_walks_preceding_fields() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Struct": true,
"endian": "little",
"properties": {
"a": { "AlkType:Uint8": true },
"b": { "AlkType:Uint32": true },
"c": { "AlkType:Uint8": true }
let root = json!({
"$defs": {
"S": {
"kind": "struct",
"endian": "little",
"fields": [
{ "name": "a", "kind": "uint8" },
{ "name": "b", "kind": "uint32" },
{ "name": "c", "kind": "uint8" }
]
}
}
});
let mut buffer = vec![0u8; 16];
@@ -450,7 +502,7 @@ fn sequential_reader_read_field_walks_preceding_fields() -> Result<(), AlkTypeEr
data_access::write_u32(&mut buffer, 1, 0xDEADBEEF, "b", Endian::Little)?;
data_access::write_u8(&mut buffer, 5, 9, "c")?;
let mut reader = SequentialReader::new(&schema)?;
let mut reader = SequentialReader::new(&root, "S")?;
let value = reader.read_field(&buffer, "c")?;
assert_eq!(value, FieldValue::U8(9));
assert_eq!(reader.position(), 6);
@@ -463,74 +515,90 @@ fn sequential_reader_read_field_walks_preceding_fields() -> Result<(), AlkTypeEr
#[test]
fn tunion_byte_offset_discriminator_dispatch() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "AlkType:Uint8"
},
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" }
},
let root = json!({
"$defs": {
"Read": {
"AlkType:Struct": true,
"properties": {
"handle": { "AlkType:Uint32": true },
"length": { "AlkType:Uint32": true }
"Packet": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" }
}
},
"Read": {
"kind": "struct",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
},
"Write": {
"AlkType:Struct": true,
"properties": {
"handle": { "AlkType:Uint32": true },
"length": { "AlkType:Uint32": true },
"data": { "AlkType:Uint32": true }
}
"kind": "struct",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" },
{ "name": "data", "kind": "uint32" }
]
}
}
});
let doc = BastDoc::new(&root, "Packet")?;
let union_def = match doc.root_def().kind() {
alktype::bast::BastDefKind::Union(u) => u,
_ => unreachable!(),
};
let mut buffer = vec![0u8; 32];
buffer[0] = 5;
data_access::write_u32(&mut buffer, 1, 0x01020304, "Read.handle", Endian::Big)?;
data_access::write_u32(&mut buffer, 5, 4096, "Read.length", Endian::Big)?;
let dispatch = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big)?;
let dispatch = tunion::read_byte_discriminator(&buffer, union_def, Endian::Big)?;
assert_eq!(dispatch.key, "5");
assert_eq!(dispatch.variant_offset, 1);
assert_eq!(dispatch.discriminator_size, 1);
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
assert_eq!(
variant
.get("AlkType:Struct")
.and_then(serde_json::Value::as_bool),
Some(true)
);
let variant = tunion::resolve_variant(union_def, &dispatch.key)?;
match variant {
alktype::bast::BastType::Ref(r) => assert_eq!(r.name(), "Read"),
other => panic!("expected Ref to Read, got {other:?}"),
}
Ok(())
}
#[test]
fn tunion_byte_offset_discriminator_size_lookup() -> Result<(), AlkTypeError> {
let u8_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
"mapping": {}
});
let u16_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint16"},
"mapping": {}
});
let u32_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint32"},
"mapping": {}
});
assert_eq!(tunion::discriminator_size(&u8_schema)?, 1);
assert_eq!(tunion::discriminator_size(&u16_schema)?, 2);
assert_eq!(tunion::discriminator_size(&u32_schema)?, 4);
fn union_with(disc_type: &str) -> serde_json::Value {
json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "type": disc_type },
"mapping": { "1": { "$ref": "#/$defs/A" } }
},
"A": { "kind": "struct", "fields": [] }
}
})
}
let u8_root = union_with("uint8");
let u16_root = union_with("uint16");
let u32_root = union_with("uint32");
let u8_doc = BastDoc::new(&u8_root, "U")?;
let u16_doc = BastDoc::new(&u16_root, "U")?;
let u32_doc = BastDoc::new(&u32_root, "U")?;
let u8_union = match u8_doc.root_def().kind() {
alktype::bast::BastDefKind::Union(u) => u,
_ => unreachable!(),
};
let u16_union = match u16_doc.root_def().kind() {
alktype::bast::BastDefKind::Union(u) => u,
_ => unreachable!(),
};
let u32_union = match u32_doc.root_def().kind() {
alktype::bast::BastDefKind::Union(u) => u,
_ => unreachable!(),
};
assert_eq!(tunion::discriminator_size(u8_union)?, 1);
assert_eq!(tunion::discriminator_size(u16_union)?, 2);
assert_eq!(tunion::discriminator_size(u32_union)?, 4);
Ok(())
}
}
+186 -165
View File
@@ -6,103 +6,116 @@
//! correct mapping key and variant offset, that `resolve_variant`
//! follows `$ref` pointers, and that `discriminator_size` reports the
//! right fixed sizes.
//!
//! All schemas are BAST documents.
use alktype::bast::{BastDefKind, BastDoc, BastType, BastUnion};
use alktype::data_access;
use alktype::tunion;
use alktype::{Endian, AlkTypeError};
use serde_json::json;
fn sftp_like_byte_union() -> serde_json::Value {
fn sftp_like_byte_union_doc() -> serde_json::Value {
json!({
"AlkType:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "AlkType:Uint8"
},
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" }
},
"$defs": {
"Read": {
"AlkType:Struct": true,
"properties": {
"handle": { "AlkType:Uint32": true },
"length": { "AlkType:Uint32": true }
"Packet": {
"kind": "union",
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "uint8"
},
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" }
}
},
"Read": {
"kind": "struct",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
},
"Write": {
"AlkType:Struct": true,
"properties": {
"handle": { "AlkType:Uint32": true },
"length": { "AlkType:Uint32": true },
"data": { "AlkType:Uint32": true }
}
"kind": "struct",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" },
{ "name": "data", "kind": "uint32" }
]
}
}
})
}
fn union_of<'a>(root: &'a serde_json::Value, name: &'a str) -> BastUnion<'a> {
let doc = BastDoc::new(root, name).expect("bast doc");
match doc.root_def().kind() {
BastDefKind::Union(u) => u.clone(),
_ => panic!("root must be a union"),
}
}
#[test]
fn read_byte_discriminator_uint8_dispatches_to_read() -> Result<(), AlkTypeError> {
let union_schema = sftp_like_byte_union();
let root = sftp_like_byte_union_doc();
let union_def = union_of(&root, "Packet");
let mut buffer = vec![0u8; 16];
buffer[0] = 5;
data_access::write_u32(&mut buffer, 1, 0x01020304, "Read.handle", Endian::Big)?;
let dispatch = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big)?;
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big)?;
assert_eq!(dispatch.key, "5");
assert_eq!(dispatch.variant_offset, 1);
assert_eq!(dispatch.discriminator_size, 1);
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
assert_eq!(
variant
.get("AlkType:Struct")
.and_then(serde_json::Value::as_bool),
Some(true)
);
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
match variant {
BastType::Ref(r) => assert_eq!(r.name(), "Read"),
other => panic!("expected Ref to Read, got {other:?}"),
}
Ok(())
}
#[test]
fn read_byte_discriminator_uint8_dispatches_to_write() -> Result<(), AlkTypeError> {
let union_schema = sftp_like_byte_union();
let root = sftp_like_byte_union_doc();
let union_def = union_of(&root, "Packet");
let mut buffer = vec![0u8; 16];
buffer[0] = 6;
data_access::write_u32(&mut buffer, 1, 0xDEADBEEF, "Write.handle", Endian::Big)?;
let dispatch = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big)?;
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big)?;
assert_eq!(dispatch.key, "6");
assert_eq!(dispatch.variant_offset, 1);
assert_eq!(dispatch.discriminator_size, 1);
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
let props = variant
.get("properties")
.and_then(serde_json::Value::as_object)
.expect("variant has properties");
assert!(props.contains_key("data"));
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
match variant {
BastType::Ref(r) => assert_eq!(r.name(), "Write"),
other => panic!("expected Ref to Write, got {other:?}"),
}
Ok(())
}
#[test]
fn read_byte_discriminator_uint16_little_endian() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Union": true,
"discriminator": {
"kind": "byte",
"offset": 2,
"type": "AlkType:Uint16"
},
"mapping": {
"5": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 2, "type": "uint16" },
"mapping": {
"5": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] }
}
}
}
});
let union_def = union_of(&root, "U");
let mut buffer = vec![0u8; 16];
buffer[2..4].copy_from_slice(&5u16.to_le_bytes());
let dispatch = tunion::read_byte_discriminator(&buffer, &schema, Endian::Little)?;
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Little)?;
assert_eq!(dispatch.key, "5");
assert_eq!(dispatch.variant_offset, 4);
assert_eq!(dispatch.discriminator_size, 2);
@@ -111,20 +124,21 @@ fn read_byte_discriminator_uint16_little_endian() -> Result<(), AlkTypeError> {
#[test]
fn read_byte_discriminator_uint32_big_endian() -> Result<(), AlkTypeError> {
let schema = json!({
"AlkType:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "AlkType:Uint32"
},
"mapping": {
"101": {"AlkType:Struct": true, "properties": {"id": {"AlkType:Uint32": true}}}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint32" },
"mapping": {
"101": { "kind": "struct", "fields": [ { "name": "id", "kind": "uint32" } ] }
}
}
}
});
let union_def = union_of(&root, "U");
let mut buffer = vec![0u8; 16];
buffer[0..4].copy_from_slice(&101u32.to_be_bytes());
let dispatch = tunion::read_byte_discriminator(&buffer, &schema, Endian::Big)?;
let dispatch = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big)?;
assert_eq!(dispatch.key, "101");
assert_eq!(dispatch.variant_offset, 4);
assert_eq!(dispatch.discriminator_size, 4);
@@ -133,191 +147,198 @@ fn read_byte_discriminator_uint32_big_endian() -> Result<(), AlkTypeError> {
#[test]
fn read_byte_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
let union_schema = sftp_like_byte_union();
let root = sftp_like_byte_union_doc();
let union_def = union_of(&root, "Packet");
let buffer = [99u8, 0x00, 0x00, 0x00];
let err = tunion::read_byte_discriminator(&buffer, &union_schema, Endian::Big).unwrap_err();
let err = tunion::read_byte_discriminator(&buffer, &union_def, Endian::Big).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
Ok(())
}
#[test]
fn read_field_discriminator_string_dispatches_to_read() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "type"},
"properties": {
"type": { "AlkType:String": true }
},
"mapping": {
"read": {"$ref": "#/$defs/Read"},
"write": {"$ref": "#/$defs/Write"}
},
let root = json!({
"$defs": {
"Read": {
"AlkType:Struct": true,
"properties": {
"handle": { "AlkType:Uint32": true },
"length": { "AlkType:Uint32": true }
"Event": {
"kind": "union",
"discriminator": { "kind": "field", "name": "type" },
"fields": [ { "name": "type", "kind": "string" } ],
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
},
"Read": {
"kind": "struct",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
},
"Write": {
"AlkType:Struct": true,
"properties": {
"handle": { "AlkType:Uint32": true },
"data": { "AlkType:Bytes": true }
}
"kind": "struct",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "data", "kind": "bytes" }
]
}
}
});
let union_def = union_of(&root, "Event");
let value = "read";
let mut buffer = vec![0u8; 32];
data_access::write_string(&mut buffer, 0, value, "type", Endian::Little)?;
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
assert_eq!(dispatch.key, "read");
assert_eq!(dispatch.variant_offset, 4 + value.len());
assert_eq!(dispatch.discriminator_size, 4 + value.len());
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
assert_eq!(
variant
.get("AlkType:Struct")
.and_then(serde_json::Value::as_bool),
Some(true)
);
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
match variant {
BastType::Ref(r) => assert_eq!(r.name(), "Read"),
other => panic!("expected Ref to Read, got {other:?}"),
}
Ok(())
}
#[test]
fn read_field_discriminator_string_dispatches_to_write() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "type"},
"properties": {
"type": { "AlkType:String": true }
},
"mapping": {
"read": {"$ref": "#/$defs/Read"},
"write": {"$ref": "#/$defs/Write"}
},
let root = json!({
"$defs": {
"Event": {
"kind": "union",
"discriminator": { "kind": "field", "name": "type" },
"fields": [ { "name": "type", "kind": "string" } ],
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
},
"Read": {
"AlkType:Struct": true,
"properties": {"x": {"AlkType:Uint8": true}}
"kind": "struct",
"fields": [ { "name": "x", "kind": "uint8" } ]
},
"Write": {
"AlkType:Struct": true,
"properties": {"y": {"AlkType:Uint16": true}}
"kind": "struct",
"fields": [ { "name": "y", "kind": "uint16" } ]
}
}
});
let union_def = union_of(&root, "Event");
let value = "write";
let mut buffer = vec![0u8; 32];
data_access::write_string(&mut buffer, 0, value, "type", Endian::Little)?;
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
assert_eq!(dispatch.key, "write");
assert_eq!(dispatch.variant_offset, 4 + value.len());
let variant = tunion::resolve_variant(&union_schema, &dispatch.key)?;
let props = variant
.get("properties")
.and_then(serde_json::Value::as_object)
.expect("variant has properties");
assert!(props.contains_key("y"));
assert!(!props.contains_key("x"));
let variant = tunion::resolve_variant(&union_def, &dispatch.key)?;
match variant {
BastType::Ref(r) => assert_eq!(r.name(), "Write"),
other => panic!("expected Ref to Write, got {other:?}"),
}
Ok(())
}
#[test]
fn read_field_discriminator_uint8_field() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "tag"},
"properties": {
"tag": { "AlkType:Uint8": true }
},
"mapping": {
"0": {"AlkType:Struct": true, "properties": {"a": {"AlkType:Uint32": true}}},
"1": {"AlkType:Struct": true, "properties": {"b": {"AlkType:Uint16": true}}}
let root = json!({
"$defs": {
"Event": {
"kind": "union",
"discriminator": { "kind": "field", "name": "tag" },
"fields": [ { "name": "tag", "kind": "uint8" } ],
"mapping": {
"0": { "kind": "struct", "fields": [ { "name": "a", "kind": "uint32" } ] },
"1": { "kind": "struct", "fields": [ { "name": "b", "kind": "uint16" } ] }
}
}
}
});
let union_def = union_of(&root, "Event");
let mut buffer = vec![0u8; 8];
buffer[0] = 0;
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
assert_eq!(dispatch.key, "0");
assert_eq!(dispatch.variant_offset, 1);
assert_eq!(dispatch.discriminator_size, 1);
buffer[0] = 1;
let dispatch = tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little)?;
let dispatch = tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little)?;
assert_eq!(dispatch.key, "1");
Ok(())
}
#[test]
fn read_field_discriminator_unknown_value_returns_access_error() -> Result<(), AlkTypeError> {
let union_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "tag"},
"properties": {
"tag": { "AlkType:Uint8": true }
},
"mapping": {
"0": {"AlkType:Struct": true, "properties": {"a": {"AlkType:Uint32": true}}}
let root = json!({
"$defs": {
"Event": {
"kind": "union",
"discriminator": { "kind": "field", "name": "tag" },
"fields": [ { "name": "tag", "kind": "uint8" } ],
"mapping": {
"0": { "kind": "struct", "fields": [ { "name": "a", "kind": "uint32" } ] }
}
}
}
});
let union_def = union_of(&root, "Event");
let mut buffer = vec![0u8; 8];
buffer[0] = 99;
let err =
tunion::read_field_discriminator(&buffer, &union_schema, 0, Endian::Little).unwrap_err();
tunion::read_field_discriminator(&buffer, &union_def, 0, Endian::Little).unwrap_err();
assert!(matches!(err, AlkTypeError::Access { .. }), "got {err:?}");
Ok(())
}
#[test]
fn discriminator_size_returns_correct_values() -> Result<(), AlkTypeError> {
let u8_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint8"},
"mapping": {}
});
let u16_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint16"},
"mapping": {}
});
let u32_schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "byte", "type": "AlkType:Uint32"},
"mapping": {}
});
assert_eq!(tunion::discriminator_size(&u8_schema)?, 1);
assert_eq!(tunion::discriminator_size(&u16_schema)?, 2);
assert_eq!(tunion::discriminator_size(&u32_schema)?, 4);
fn union_with(disc_type: &str) -> serde_json::Value {
json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "byte", "type": disc_type },
"mapping": { "1": { "$ref": "#/$defs/A" } }
},
"A": { "kind": "struct", "fields": [] }
}
})
}
let u8_root = union_with("uint8");
let u16_root = union_with("uint16");
let u32_root = union_with("uint32");
let u8_union = union_of(&u8_root, "U");
let u16_union = union_of(&u16_root, "U");
let u32_union = union_of(&u32_root, "U");
assert_eq!(tunion::discriminator_size(&u8_union)?, 1);
assert_eq!(tunion::discriminator_size(&u16_union)?, 2);
assert_eq!(tunion::discriminator_size(&u32_union)?, 4);
Ok(())
}
#[test]
fn discriminator_size_field_kind_returns_schema_error() {
let schema = json!({
"AlkType:Union": true,
"discriminator": {"kind": "field", "name": "type"},
"properties": {"type": {"AlkType:Uint8": true}},
"mapping": {}
let root = json!({
"$defs": {
"U": {
"kind": "union",
"discriminator": { "kind": "field", "name": "type" },
"fields": [ { "name": "type", "kind": "uint8" } ],
"mapping": { "0": { "kind": "struct", "fields": [] } }
}
}
});
let err = tunion::discriminator_size(&schema).unwrap_err();
let union_def = union_of(&root, "U");
let err = tunion::discriminator_size(&union_def).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn resolve_variant_returns_schema_error_for_unknown_key() {
let union_schema = sftp_like_byte_union();
let err = tunion::resolve_variant(&union_schema, "999").unwrap_err();
let root = sftp_like_byte_union_doc();
let union_def = union_of(&root, "Packet");
let err = tunion::resolve_variant(&union_def, "999").unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
#[test]
fn parse_discriminator_missing_returns_schema_error() {
let schema = json!({"AlkType:Union": true});
let err = alktype::parse_discriminator(&schema).unwrap_err();
assert!(matches!(err, AlkTypeError::Schema(_)), "got {err:?}");
}
}