From 62270b03cab701ff6f4fd2bfafe0f9f6defa02a0 Mon Sep 17 00:00:00 2001 From: "glm-5.2" Date: Sat, 15 Aug 2026 14:03:21 +0000 Subject: [PATCH] Sync architecture docs and ADRs to BAST pivot (steps 9-10) Step 9 (convert tests to BAST format) was a no-op: steps 4-8 converted the tests as they went. The only remaining reference in src/tests was the intentional rejection test at src/schema.rs:462 (asserting the old keyword form is rejected). Full suite passes: 389 tests (312 lib + 77 integration). Step 10 (sync architecture docs and ADRs): Descriptive docs rewritten/updated for BAST: - schema-layer.md: rewritten for the BAST parser (BastDoc/BastDef/ BastType typed tree, AlkTypeKind enum with to_bast_str/from_bast_str, what was removed). Points at bast-format.md for the normative format. - validation.md: rewritten for the two-validator model (bast_validation for validate_bytes, standard jsonschema for validate_json). Documents the repurposed build_validator, the AlkTypeError::Validation uniform payload (D-BAST-009), and what is removed. - builder.md: updated all output examples to BAST JSON (struct_() -> { kind: struct, fields: [...] }; object() -> standard JSON Schema). Documents build_doc, count(), and the field-name union fields requirement (D-BAST-005). - overview.md: updated for BAST (what/why, schema-is-the-format table, dependencies, architecture pointers, design decisions table). - README.md (architecture index): updated document table, ADR table (new ADR-BAST + ADR-VAL-SPLIT, superseded ADR-001), OQ table (OQ-007/OQ-008 resolutions updated for BAST-native validator), and key design principles (#1, #2, #7, #10 reworded for BAST). - data-access.md: updated tunion function signatures to BastUnion and the variant resolution to return BastType (resolve_typeref for refs). - layout-engine.md: updated construct signatures (LayoutBuilder::new(bast_doc, root_name), OffsetMap::compute(&doc), SequentialReader::new(bast_doc, root_name)), the recursive-walk description (BAST typed tree), and composite-kind headings (TStruct/TUnion/TArray -> struct/union/array). Added D-BAST-004 note on array count requirement. New ADRs: - ADR-BAST (bast-bast-format.md): the BAST format, meta-schema, //kind vocabulary, design principles, what is removed, the enum index bounds bug fix. Supersedes ADR-001's format-specific content; records D-BAST-001..009. - ADR-VAL-SPLIT (val-split-two-validator-model.md): the two-validator model (BAST-native for validate_bytes, standard jsonschema for validate_json), the repurposed build_validator, the uniform AlkTypeError::Validation payload. Refines ADR-004's validation strategy and ADR-010's validation step; records D-BAST-006/007/009. Amended ADRs (supersession/amendment notes added; original decision text preserved as historical record): - ADR-001: format-specific content superseded by ADR-BAST; purpose/scope and schema-is-the-format principle retained. - ADR-002: unchanged under the pivot; one-line note that the input format changed but the modes didn't. - ADR-003: annotation semantics retained; annotation location moved to BAST type-level properties (amended by ADR-BAST). - ADR-004: AlkTypeError enum retained (D-BAST-009); validation strategy section refined by ADR-VAL-SPLIT. - ADR-009: builder API surface retained; build() output format amended to BAST / standard JSON Schema by ADR-BAST (D-BAST-008). - ADR-010: validate_bytes two-step concept retained; validation step amended to the BAST-native validator by ADR-VAL-SPLIT. Other: - Cargo.toml description: JSON Schema with AlkType:* custom keywords -> BAST document. - bast-pivot.md research record: status draft -> implemented, with a pointer to the ADRs that superseded its decisions. - bast-implementation.md plan: status draft -> complete, with a note that step 9 was a no-op and step 10 is this commit. - open-questions.md: OQ-006/OQ-007/OQ-008 resolutions updated for the BAST-native validator. - questions/008-unionvalidator-variant-dispatch.md: added a post-BAST-pivot note pointing to the current bast_validation implementation; v0.1.0 resolution text preserved as historical record. Verification: - cargo test --release: 389 pass (312 lib + 77 integration) - cargo clippy --all-targets -- -D warnings: clean - cargo doc --no-deps: clean - cross-reference check: every relative link in the new/updated docs resolves (verified by script). --- Cargo.toml | 2 +- docs/architecture/README.md | 128 ++-- docs/architecture/builder.md | 408 ++++++----- docs/architecture/data-access.md | 41 +- ...alktype-purpose-scope-jsonschema-engine.md | 14 +- .../002-two-layout-modes-packed-vs-aligned.md | 9 +- .../decisions/003-schema-annotations.md | 13 +- .../004-error-handling-validation-strategy.md | 16 +- .../architecture/decisions/009-builder-api.md | 14 +- ...0-generalized-validation-validate-bytes.md | 18 +- .../decisions/bast-bast-format.md | 316 +++++++++ .../val-split-two-validator-model.md | 229 +++++++ docs/architecture/layout-engine.md | 37 +- docs/architecture/open-questions.md | 34 +- docs/architecture/overview.md | 191 +++--- .../008-unionvalidator-variant-dispatch.md | 14 + docs/architecture/schema-layer.md | 633 ++++++------------ docs/architecture/validation.md | 558 ++++++++------- docs/plans/bast-implementation.md | 19 +- docs/research/bast-pivot.md | 23 +- 20 files changed, 1642 insertions(+), 1075 deletions(-) create mode 100644 docs/architecture/decisions/bast-bast-format.md create mode 100644 docs/architecture/decisions/val-split-two-validator-model.md diff --git a/Cargo.toml b/Cargo.toml index 7ace494..dfdfb8f 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -4,7 +4,7 @@ version = "0.1.0" edition = "2021" rust-version = "1.85" license = "MIT OR Apache-2.0" -description = "Binary struct engine: takes a JSON Schema with AlkType:* custom keywords and produces an offset map, read/write functions, and validation" +description = "Binary struct engine: takes a BAST (Binary Abstract Syntax Tree) document and produces an offset map, read/write functions, and validation" repository = "https://git.alk.dev/alkdev/alktype" readme = "README.md" keywords = ["binary", "jsonschema", "wire-format", "serialization", "layout"] diff --git a/docs/architecture/README.md b/docs/architecture/README.md index ce4609f..8c082de 100644 --- a/docs/architecture/README.md +++ b/docs/architecture/README.md @@ -1,12 +1,12 @@ --- -status: draft -last_updated: 2026-08-11 +status: accepted +last_updated: 2026-08-15 --- # alktype -The binary struct engine: a small Rust crate that takes a JSON Schema -with `AlkType:*` custom keywords and produces an offset map, read/write +The binary struct engine: a small Rust crate that takes a BAST (Binary +Abstract Syntax Tree) document and produces an offset map, read/write functions, and validation — all driven by the schema. The schema is the format definition; the engine is generic. @@ -14,68 +14,73 @@ format definition; the engine is generic. | Document | Status | Description | |----------|--------|-------------| -| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries | -| [schema-layer.md](schema-layer.md) | draft | The 19 `AlkType:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations. *(Current v0.1.0 format; will be rewritten for BAST — see [bast-format.md](bast-format.md).)* | -| [bast-format.md](bast-format.md) | draft | **Target schema format.** BAST (Binary Abstract Syntax Tree): meta-schema, TypeRef, examples, the two-validator model. Supersedes the format-spec content of schema-layer.md when the [BAST pivot](../plans/bast-implementation.md) lands. | +| [overview.md](overview.md) | accepted | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries | +| [`bast-format.md`](bast-format.md) | accepted | **Normative BAST format specification.** Meta-schema, TypeRef, TypeDef shapes (Struct/Union/Enum/FieldDef), examples, validation model. The format the engine consumes. | +| [schema-layer.md](schema-layer.md) | accepted | The BAST parser (`src/bast.rs`) — the typed tree (`BastDoc`/`BastDef`/`BastType`/…) every engine module walks, the 19 BAST kinds, the `AlkTypeKind` enum, and the foundational annotation types. | | [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling | | [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading | -| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine`; `validate_bytes` for binary buffers (ADR-010). *(Current v0.1.0 validation; will be rewritten for the validator split — see [bast-format.md §Validation Model](bast-format.md#validation-model).)* | -| [builder.md](builder.md) | draft | Fluent Rust API for constructing alktype JSON Schemas at runtime, producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema (ADR-009) | +| [validation.md](validation.md) | accepted | The two-validator model (BAST-native for `validate_bytes`, standard `jsonschema` for `validate_json`), `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine` as the compiled form of a BAST document (ADR-010, ADR-VAL-SPLIT). | +| [builder.md](builder.md) | accepted | Fluent Rust API for constructing BAST documents (`struct_()`) and standard JSON Schemas (`object()`) at runtime, producing `serde_json::Value` (ADR-009, D-BAST-008). | ### In-progress work | Document | Status | Description | |----------|--------|-------------| -| [BAST pivot — research record](../research/bast-pivot.md) | draft | Motivation, POC scope and result, decisions D-BAST-001..009, risks for the BAST format pivot | -| [BAST pivot — implementation plan](../plans/bast-implementation.md) | draft | Ordered implementation steps, the public-API semver contract, and the ADR-sync checklist for the BAST pivot | +| [BAST pivot — research record](../research/bast-pivot.md) | accepted | Motivation, POC scope and result, decisions D-BAST-001..009, risks for the BAST format pivot. Implemented in steps 1–10. | +| [BAST pivot — implementation plan](../plans/bast-implementation.md) | accepted | Ordered implementation steps, the public-API semver contract, and the ADR-sync checklist for the BAST pivot. Steps 1–10 complete. | ## Applicable ADRs | ADR | Title | Relevance | |-----|-------|-----------| -| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries | -| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` | -| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations | -| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors | +| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries. *Format-specific content superseded by ADR-BAST; purpose/scope retained.* | +| [BAST](decisions/bast-bast-format.md) | BAST (Binary Abstract Syntax Tree) as the Schema Format | The BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary. Supersedes ADR-001's format-specific content; records D-BAST-001..009. | +| [VAL-SPLIT](decisions/val-split-two-validator-model.md) | Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON | `validate_bytes` uses the BAST-native validator; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema. Records D-BAST-006/007/009. | +| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` (format-agnostic — input format changed, modes didn't) | +| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Annotation *semantics* (carry forward unchanged); annotation *location* moved to BAST type-level properties under the pivot | +| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum (shape unchanged, D-BAST-009); load-time build, access-time check; field-path-carrying errors. *Validation-strategy section refined by ADR-VAL-SPLIT.* | | [005](decisions/005-int64-uint64-first-class-kinds.md) | Int64/Uint64 as First-Class Kinds | 64-bit integers (SFTP offsets, metatensor data_offsets); JSON precision caveat | | [006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode | Prevents silent data corruption (inline variable data clobbering subsequent fields) | | [007](decisions/007-packed-mode-read-factory.md) | Packed-Mode Read API — Engine as SequentialReader Factory | `engine.sequential_reader()` returns an owned reader, not a reference | | [008](decisions/008-reject-tunion-in-aligned-mode.md) | Reject TUnion in Aligned Mode for v1 | Unions are the protocol pattern; aligned-mode union semantics were broken | -| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 | -| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate; two methods on one struct, not a trait | +| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003. *Output format amended to BAST / standard JSON Schema by ADR-BAST.* | +| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate. *Validation step amended to the BAST-native validator by ADR-VAL-SPLIT.* | ## Relevant Open Questions | OQ | Title | Status | Relevance | |----|-------|--------|-----------| -| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it | +| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it. BAST arrays require `count` in v1 (D-BAST-004), aligning with this deferral. | | OQ-002 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case | -| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Resolved in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) | -| OQ-004 | `Discriminator::Field` name — `&str` or `String` | open | Builder API ownership question; resolve before the SFTP Packet POC's field-name discriminator path | -| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | open | Blocks the SFTP Packet `validate_bytes` POC (next round) | -| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | open | Documentation fix in builder.md; the engine requires `AlkType:Struct` at the top level | -| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | open | Blocks the SFTP use case for `validate_bytes` (binary `handle`/`data` fields) | +| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Shipped in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) | +| OQ-004 | `Discriminator::Field` name — `&str` or `String` | resolved | `String`, for ownership simplicity | +| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | resolved | Both kinds return `{ "__discriminator": , ...variant-fields }` | +| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | resolved | [builder.md](builder.md) Example 3 wraps the union in a `Schema::struct_().field("payload", ...)` | +| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | resolved | Array of u8: materializer produces `Value::Array` of `Value::Number`; BAST-native validator accepts both `Value::String` and `Value::Array` | +| OQ-008 | `UnionValidator` variant dispatch | resolved | BAST-native validator recurses into the selected variant's BAST definition on `__discriminator` lookup — no custom keywords, no `inline_union_variant_refs` | ## Key Design Principles -1. **The schema is the format.** A JSON Schema with `AlkType:*` custom - keywords is both the validation spec and the layout spec. No separate - format definition, no separate parser, no separate validator. One - schema, three uses: validate, compute offsets, access data. See - [overview.md](overview.md) and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md). +1. **The schema is the format.** A BAST document is both the layout + spec and the validation spec for bytes. No separate format + definition, no separate parser, no separate validator. One schema, + three uses: validate, compute offsets, access data. See + [overview.md](overview.md), [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md), + and [ADR-BAST](decisions/bast-bast-format.md). -2. **jsonschema is the validation engine, not a custom engine.** The - `jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with - custom keyword support. The novel code is the offset computation, not - the validation. This eliminates ~14,000 lines of hand-rolled schema - engines (typebox-rs, the @alkdev/alktype prototype). See [schema-layer.md](schema-layer.md) - and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md). +2. **BAST is a JSON Schema dialect, not a custom format.** A BAST + document is valid JSON conforming to the BAST meta-schema (a + standard Draft 2020-12 JSON Schema). Any JSON Schema validator can + check whether a BAST document is well-formed; editors with JSON + Schema support provide autocomplete for free. See + [`bast-format.md`](bast-format.md) and + [ADR-BAST](decisions/bast-bast-format.md). 3. **Two layout modes for two use cases.** Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP, channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats - (metatensor). The consumer selects the mode; the schema is the same. - See [layout-engine.md](layout-engine.md) and + (metatensor). The consumer selects the mode; the BAST document is the + same. See [layout-engine.md](layout-engine.md) and [ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md). 4. **Variable-length types default to inline length-prefixing.** @@ -92,38 +97,41 @@ format definition; the engine is generic. [ADR-003](decisions/003-schema-annotations.md). 6. **Endianness is per-schema, default little-endian.** The engine reads - the `"endian"` annotation and byte-swaps accordingly. SFTP consumers - specify `"endian": "big"`. See [layout-engine.md](layout-engine.md) - and [ADR-003](decisions/003-schema-annotations.md). + the struct-level `"endian"` annotation and byte-swaps accordingly. + SFTP consumers specify `"endian": "big"`. See + [layout-engine.md](layout-engine.md) and + [ADR-003](decisions/003-schema-annotations.md). -7. **Validation is opt-in, built once at load time.** The jsonschema - validator is compiled once at schema load time. Access-time validation - is a fast `is_valid()` check. High-throughput paths can skip - validation; security-sensitive paths can validate every frame. See +7. **Two validators for two input types.** `validate_bytes(&[u8])` uses + the BAST-native validator (a recursive walker over the BAST type + tree — no `jsonschema` involvement). `validate_json(&Value)` uses a + standard `jsonschema::Validator` from a consumer-provided JSON Schema + (BAST is not involved — BAST describes bytes, not JSON shape). One + `AlkTypeError::Validation` variant covers both (D-BAST-009). See [validation.md](validation.md) and - [ADR-004](decisions/004-error-handling-validation-strategy.md). + [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md). 8. **Not a serialization framework.** The alktype engine is not a general-purpose serde replacement. It operates on raw byte buffers at - computed offsets — no intermediate `Value` tree, no reflection, no - dynamic dispatch per field. For JSON data, use serde. For binary data - with a known schema, use alktype. See [overview.md](overview.md) and + computed offsets — no reflection, no dynamic dispatch per field. For + JSON data, use serde. For binary data with a known BAST document, use + alktype. See [overview.md](overview.md) and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md). -9. **Schemas can be built at runtime from Rust (v0.1.0).** A fluent - builder API produces `serde_json::Value` for both AlkType-kind - schemas and standard JSON Schema, covering alkcall's two roles +9. **Schemas can be built at runtime from Rust.** A fluent builder API + produces `serde_json::Value` for both BAST documents (`struct_()`) and + standard JSON Schemas (`object()`), covering alkcall's two roles (binary layout + JSON payloads) from one module. The builder is - additive — consumers with static schemas continue to load JSON. - See [builder.md](builder.md) and [ADR-009](decisions/009-builder-api.md). + additive — consumers with static BAST documents continue to load + JSON. See [builder.md](builder.md) and + [ADR-009](decisions/009-builder-api.md). -10. **Two validation entry points, one engine (v0.1.0).** - `validate_json(&Value)` for already-parsed JSON (call's payloads); - `validate_bytes(&[u8])` for binary buffers (channels' chunk header). - Same underlying `jsonschema` validator; the bytes path materializes - a `Value` tree via the layout engine, then validates. See - [validation.md](validation.md) and - [ADR-010](decisions/010-generalized-validation-validate-bytes.md). +10. **Two validation entry points, one engine.** `validate_json(&Value)` + for already-parsed JSON (call's payloads); `validate_bytes(&[u8])` + for binary buffers (channels' chunk header). Different validators, + one `AlkTypeError::Validation` variant. See [validation.md](validation.md), + [ADR-010](decisions/010-generalized-validation-validate-bytes.md), + and [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md). ## References @@ -143,4 +151,4 @@ format definition; the engine is generic. > **Note**: The research findings, POC code, and prior-attempt paths above > refer to the parent `@alkdev/alknet` workspace where this crate originated. > They are preserved here as historical context for the architectural -> decisions; the artifacts themselves are not part of this standalone repo. +> decisions; the artifacts themselves are not part of this standalone repo. \ No newline at end of file diff --git a/docs/architecture/builder.md b/docs/architecture/builder.md index 1294909..49504d7 100644 --- a/docs/architecture/builder.md +++ b/docs/architecture/builder.md @@ -1,29 +1,41 @@ --- -status: draft -last_updated: 2026-08-11 +status: accepted +last_updated: 2026-08-15 --- # alktype — Builder API -The builder layer: a fluent Rust API for constructing alktype JSON -Schemas (both `AlkType:*`-bearing binary-layout schemas and plain -JSON-Schema-only operation payload schemas) at runtime, producing -`serde_json::Value`. Decided in [ADR-009](decisions/009-builder-api.md); -resolves [OQ-003](questions/003-builder-api-for-schema-construction.md). +The builder layer: a fluent Rust API for constructing BAST documents +(binary-layout schemas) and standard JSON Schemas (JSON-validation +schemas) at runtime, producing `serde_json::Value`. Decided in +[ADR-009](decisions/009-builder-api.md); resolves +[OQ-003](questions/003-builder-api-for-schema-construction.md). The +two-output-format split is D-BAST-008, recorded in +[ADR-BAST](decisions/bast-bast-format.md). ## What The `builder` module provides a single `Schema` builder type and a `Definitions` helper for named `$defs`. The builder's `.build()` method -returns a `serde_json::Value` — the same form alktype already consumes -via `AlkTypeEngine::compile` (for `AlkType:*` schemas) and the same form -`OperationSpec.input_schema` / `output_schema` / `error_schemas` hold -(for plain JSON Schema, no `AlkType:*` kinds). +returns a `serde_json::Value` — one of two forms depending on the +constructor used (D-BAST-008): + +- **BAST JSON** (binary layout) — `Schema::struct_().field(...).build()` + produces a BAST TypeDef (`{ "kind": "struct", "fields": [...] }`). + Primitive constructors produce bare TypeRef strings (`"uint32"`). + Feed to [`AlkTypeEngine::compile`](validation.md) (the binary-layout + path) → `validate_bytes`. +- **Standard JSON Schema** (JSON validation) — `Schema::object().field(...)` + produces `{ "type": "object", "properties": {...}, "required": [...] }`. + No BAST `kind`, no custom keywords — a plain JSON Schema. Feed to a + standard `jsonschema::Validator` (or `AlkTypeEngine::compile` with a + JSON Schema for the `validate_json` path, D-BAST-007). The builder covers: -- All 19 `AlkType:*` kinds (binary-layout schemas) — see - [schema-layer.md](schema-layer.md) for the kinds. +- All 19 BAST kinds (binary-layout schemas) — see + [schema-layer.md](schema-layer.md) for the kinds and + [`bast-format.md`](bast-format.md) for the format. - All standard JSON Schema keywords needed for operation payload schemas: `type`, `properties`, `required`, `items`, `enum`, `format`, `additionalProperties`, `minimum`, `maximum`, `minItems`, @@ -39,16 +51,20 @@ alktype's first consumer and needs to build schemas at runtime from Rust code, for two roles: 1. **Binary layout schemas** (channels' 8-byte chunk header, future - binary call frames) — `AlkType:*` schemas, fed to - `AlkTypeEngine::compile` (packed mode, big-endian). + binary call frames) — BAST documents, fed to + `AlkTypeEngine::compile` (packed mode, big-endian) → + `validate_bytes`. 2. **JSON payload schemas** (call's `OperationSpec.input_schema` / - `output_schema` / `error_schemas`) — plain JSON Schema, no - `AlkType:*` kinds, validated via the standard `jsonschema` validator. + `output_schema` / `error_schemas`) — plain JSON Schema, no BAST + `kind`, validated via the standard `jsonschema` validator + (`AlkTypeEngine::compile` with a JSON Schema → `validate_json`). A single builder serving both roles means alkcall imports one module for schema construction. See [ADR-009](decisions/009-builder-api.md) for the decision rationale (why `Value` not a typed `Schema` enum, why -both AlkType and standard JSON Schema in one builder). +both BAST and standard JSON Schema in one builder) and +[ADR-BAST](decisions/bast-bast-format.md) for the two-output-format +decision (D-BAST-008). ## Architecture @@ -62,35 +78,42 @@ duplicating the JSON form that `AlkTypeEngine::compile`, ### Module placement `src/builder.rs`, re-exported from the crate root. The builder is a -peer of `schema.rs` (which parses schemas) and `engine.rs` (which +peer of `bast.rs` (which parses BAST documents) and `engine.rs` (which compiles them). The builder constructs; it does not parse or compile. ```rust // src/lib.rs (additions) pub mod builder; -pub use builder::{Schema, Definitions}; +pub use builder::{Schema, Definitions, Discriminator}; ``` -### Field order is load-bearing +### Field order is explicit -`serde_json` with `preserve_order` is already a dependency (ADR-001). -The builder's `Value` output uses `serde_json::Map` (which preserves -insertion order under `preserve_order`), so field declaration order in -the builder is the field order in the binary layout. This is critical -for packed mode (ADR-002) where field order determines offsets. +BAST struct fields are an ordered array (BAST design principle #4 — +see [`bast-format.md`](bast-format.md#design-principles)). The builder's +`struct_()`/`union_()` accumulates fields in call order and emits them +as the `fields` array on `.build()`. Field order in the builder is the +field order in the binary layout. This is critical for packed mode +(ADR-002) where field order determines offsets. + +(`serde_json`'s `preserve_order` feature remains a dependency, but +layout correctness no longer depends on it — the `fields` array makes +order explicit. `preserve_order` is still load-bearing for the +`mapping` object's iteration order and for `Definitions`' `$defs` +block, which the parser walks in document order.) ## Public API ### `Schema` builder -`Schema` is the single entry point. Constructors for each AlkType kind -and each standard JSON Schema type; setters for annotations and +`Schema` is the single entry point. Constructors for each BAST kind and +each standard JSON Schema type; setters for annotations and constraints; `.build()` produces `Value`. -#### AlkType kind constructors +#### BAST kind constructors One constructor per `AlkTypeKind` variant (see [schema-layer.md](schema-layer.md) -§"The 19 AlkType Kinds"): +§"The 19 BAST Kinds"): ```rust impl Schema { @@ -116,39 +139,41 @@ impl Schema { // Composite kinds pub fn struct_() -> Self; // fields added via .field() pub fn union_(disc: Discriminator) -> Self; // variants via .mapping() - pub fn array_of(element: Schema) -> Self; + pub fn array_of(element: Schema) -> Self; // .count() required for valid BAST (D-BAST-004) pub fn record_of(value: Schema) -> Self; } ``` -Each constructor sets the corresponding `"AlkType:": true` key. -For example, `Schema::uint32()` produces `{"AlkType:Uint32": true}`. +Primitive constructors produce the bare BAST TypeRef string on +`.build()`. For example, `Schema::uint32().build()` produces `"uint32"`. +Composite constructors produce the BAST object form. -**`enum_of`** sets both `"AlkType:Enum": true` and the standard -`"enum"` keyword with the provided values (declaration order is the -index order — see [schema-layer.md](schema-layer.md) §"TEnum binary -representation"): +**`enum_of`** produces a BAST enum TypeDef (`{ "kind": "enum", "values": +[...] }`); declaration order is the index order — see +[schema-layer.md](schema-layer.md) §"The 19 BAST Kinds"): ```rust -Schema::enum_of(&["read", "write", "execute"]) -// -> { "AlkType:Enum": true, "enum": ["read", "write", "execute"] } +Schema::enum_of(&["read", "write", "execute"]).build() +// -> { "kind": "enum", "values": ["read", "write", "execute"] } ``` **`array_of`** and **`record_of`** take the element/value schema as a -nested `Schema`: +nested `Schema`. `array_of` requires `.count(N)` for valid BAST +(D-BAST-004 — arrays of variable-length elements without a count are +deferred, aligning with OQ-001): ```rust -Schema::array_of(Schema::uint32()) -// -> { "AlkType:Array": true, "items": { "AlkType:Uint32": true } } +Schema::array_of(Schema::uint32()).count(3).build() +// -> { "kind": "array", "element": "uint32", "count": 3 } -Schema::record_of(Schema::float32()) -// -> { "AlkType:Record": true, "values": { "AlkType:Float32": true } } +Schema::record_of(Schema::float32()).build() +// -> { "kind": "record", "values": "float32" } ``` #### Standard JSON Schema type constructors -For plain JSON Schema (no `AlkType:*` kinds) — call's -`input_schema` / `output_schema` / `error_schemas`: +For plain JSON Schema (no BAST `kind`) — call's `input_schema` / +`output_schema` / `error_schemas`: ```rust impl Schema { @@ -163,22 +188,24 @@ impl Schema { } ``` -The `_` suffix disambiguates standard JSON Schema types from AlkType -kinds (`string` is the AlkType kind; `string_` is the standard JSON -Schema type — the AlkType kind constructor sets `"AlkType:String": -true`, the standard constructor sets `"type": "string"`). This is -deliberate: the two are distinct schema forms and the builder makes -the distinction visible at the call site. +The `_` suffix disambiguates standard JSON Schema types from BAST kinds +(`string` is the BAST primitive; `string_` is the standard JSON Schema +type — `string()` would produce `"string"` as a BAST TypeRef, +`string_()` produces `{ "type": "string" }` as a standard JSON Schema). +This is deliberate: the two are distinct schema forms and the builder +makes the distinction visible at the call site. #### Annotation setters -Annotation setters mirror ADR-003. Each setter is named after the -annotation it produces; calling the setter sets the corresponding JSON -key. Setters return `Self` for chaining. +Annotation setters mirror ADR-003 (semantics unchanged; location moved +to BAST type-level properties under the pivot — see +[ADR-BAST](decisions/bast-bast-format.md)). Each setter is named after +the annotation it produces; calling the setter sets the corresponding +JSON key. Setters return `Self` for chaining. ```rust impl Schema { - /// Schema-level endianness (ADR-003 §1). Default little. + /// Struct/union-level endianness (ADR-003 §1). Default little. pub fn endian(mut self, endian: Endian) -> Self; /// Struct or field alignment (ADR-003 §2). Struct-level sets the @@ -193,12 +220,19 @@ impl Schema { /// a variable-length type, reserves this many bytes (strategy 2). /// In packed mode, validation constraint only. pub fn max_length(mut self, max: usize) -> Self; + + /// Array count (D-BAST-004 — required for valid BAST arrays in v1). + pub fn count(mut self, count: usize) -> Self; } ``` `Endian` and `VariableEncoding` are re-exported from `schema.rs` (no -new types — the builder uses the existing enums). The setters produce -the exact JSON shapes from ADR-003: +new types — the builder uses the existing enums). When applied to a +struct, `endian`/`align` are struct-level; when the `Schema` is used as +a `.field()` argument, the builder extracts `endian`/`align`/`encoding`/ +`maxLength` and places them on the *field* object (BAST field-level +properties). The setters produce the exact BAST JSON shapes from +[`bast-format.md`](bast-format.md): ```rust Schema::struct_() @@ -207,31 +241,34 @@ Schema::struct_() .field("length", Schema::uint32()) .build() // -> { -// "AlkType:Struct": true, +// "kind": "struct", // "endian": "big", -// "properties": { -// "channel_id": { "AlkType:Uint32": true }, -// "length": { "AlkType:Uint32": true } -// } +// "fields": [ +// { "name": "channel_id", "kind": "uint32" }, +// { "name": "length", "kind": "uint32" } +// ] // } ``` #### Composite builders -`struct_()`, `union_()`, `array_of()`, `record_of()` are the -composite constructors. `struct_()` and `union_()` need additional -setters to populate their children: +`struct_()`, `union_()`, `array_of()`, `record_of()` are the composite +constructors. `struct_()` and `union_()` need additional setters to +populate their children: ```rust impl Schema { - /// Add a field to a struct (or object). Field order is load-bearing - /// for binary layouts (packed mode field order = byte order). + /// Add a field to a struct (or a field-name-discriminator union). + /// Field order is load-bearing for binary layouts (packed mode + /// field order = byte order — the `fields` array is ordered). /// The field's schema is built from the passed `Schema`. pub fn field(mut self, name: &str, field: Schema) -> Self; /// Mark fields as required (standard JSON Schema `required` keyword). - /// Can be called multiple times; required names accumulate. - /// Field names must have been added via `.field()`. + /// Only meaningful for `object()` (standard JSON Schema) — BAST + /// structs require all declared fields present (the validator + /// enforces this). Can be called multiple times; required names + /// accumulate. pub fn required(mut self, names: &[&str]) -> Self; /// Set the items schema for a standard `array` type. @@ -247,9 +284,11 @@ impl Schema { } ``` -**`field`** sets `properties[name] = field.build()`. Repeated calls -append. Field order in the built `Value` is the call order (because -`serde_json::Map` preserves insertion order under `preserve_order`). +**`field`** appends a `{ "name": ..., "kind": , ... }` +entry to the struct/union's `fields` array, extracting field-level +annotations (`endian`, `align`, `encoding`, `maxLength`) from the +passed `Schema`. Repeated calls append in order. Field order in the +built `Value` is the call order. **`required`** sets the standard JSON Schema `"required"` array. The builder does not check that the named fields exist (that's a @@ -260,11 +299,12 @@ accumulates names: ```rust Schema::object() - .field("path", Schema::string_()) + .field("path", Schema::string_().max_length(4096)) .field("offset", Schema::integer().minimum(0)) .field("length", Schema::integer().minimum(0)) .required(["path"]) .required(["offset", "length"]) + .build() // -> { // "type": "object", // "properties": { "path": {...}, "offset": {...}, "length": {...} }, @@ -280,35 +320,31 @@ For operation payload schemas (call's `input_schema` etc.): impl Schema { /// `minimum` (inclusive lower bound for numbers/integers). pub fn minimum(mut self, min: f64) -> Self; - /// `maximum` (inclusive upper bound for numbers/integers). pub fn maximum(mut self, max: f64) -> Self; - /// `minLength` (minimum string length). pub fn min_length(mut self, min: usize) -> Self; - /// `minItems` (minimum array length). pub fn min_items(mut self, min: usize) -> Self; - /// `maxItems` (maximum array length). pub fn max_items(mut self, max: usize) -> Self; - /// `format` (e.g. "date-time", "uri", "email"). pub fn format(mut self, fmt: &str) -> Self; - /// `title` (human-readable description). pub fn title(mut self, t: &str) -> Self; - /// `description` (human-readable description). pub fn description(mut self, d: &str) -> Self; } ``` These set the corresponding standard JSON Schema keywords. They apply -to both AlkType-kind schemas and standard JSON Schema type schemas -(e.g., `Schema::string().max_length(4096)` sets `maxLength`, which -serves as both a validation constraint and, in aligned mode, a -fixed-size reservation — ADR-003 §3). +to standard JSON Schema type schemas (e.g., +`Schema::string_().max_length(4096)` sets `maxLength`, which on the +`validate_json` path is a JSON-Schema validation constraint). On a BAST +schema, `max_length` also serves as the aligned-mode fixed-size +reservation (ADR-003 §3) and the packed-mode validation constraint +(enforced by the BAST-native validator — see +[validation.md](validation.md)). #### `.build()` @@ -343,21 +379,22 @@ to compose them. `from_value` wraps the `Value` so it can be passed to ### `Discriminator` for `union_()` `union_()` takes a `Discriminator` describing the union's dispatch -mechanism. This mirrors `schema.rs::DiscriminatorKind` but with a -builder-friendly shape (the kind enum is re-exported from `schema.rs`, -not duplicated): +mechanism. This mirrors `bast::BastDiscriminator` (the parser's typed +view) but with a builder-friendly shape: ```rust pub enum Discriminator { /// Byte-offset discriminator (ADR-003 §4 Kind A). - /// `offset` is the byte position; `disc_type` is the AlkType kind + /// `offset` is the byte position; `disc_type` is the BAST kind /// of the discriminator (Uint8/Uint16/Uint32). Byte { offset: usize, disc_type: AlkTypeKind, // restricted to Uint8/Uint16/Uint32 }, /// Field-name discriminator (ADR-003 §4 Kind B). - /// `name` is the field holding the discriminator value. + /// `name` is the field holding the discriminator value. The + /// discriminator field and any shared fields are declared via + /// `.field()` on the union builder. Field { name: String, }, @@ -376,9 +413,13 @@ let packet = Schema::union_(Discriminator::Byte { .mapping("101", Schema::ref_def("Status")) .build(); // -> { -// "AlkType:Union": true, -// "discriminator": { "kind": "byte", "offset": 0, "type": "AlkType:Uint8" }, -// "mapping": { "5": {"$ref":"#/$defs/Read"}, "6": {...}, "101": {...} } +// "kind": "union", +// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" }, +// "mapping": { +// "5": { "$ref": "#/$defs/Read" }, +// "6": { "$ref": "#/$defs/Write" }, +// "101": { "$ref": "#/$defs/Status" } +// } // } ``` @@ -386,22 +427,29 @@ let packet = Schema::union_(Discriminator::Byte { ```rust let event = Schema::union_(Discriminator::Field { name: "type" }) + .field("type", Schema::string()) .mapping("read", Schema::ref_def("Read")) .mapping("write", Schema::ref_def("Write")) .build(); // -> { -// "AlkType:Union": true, +// "kind": "union", // "discriminator": { "kind": "field", "name": "type" }, +// "fields": [ { "name": "type", "kind": "string" } ], // "mapping": { "read": {...}, "write": {...} } // } ``` +(Field-name-discriminator unions require a `fields` array declaring the +discriminator field — D-BAST-005. The builder emits `fields` only when +the discriminator is `Field` and at least one field was added.) + ### `Definitions` — named `$defs` for cross-reference `Definitions` is a helper for building named `$defs` that schemas can -`$ref` by name. This is the ergonomics win for alkcall's -`OperationSpec`, where input/output/error schemas reference shared -definitions (e.g., `FileNotFound`, `RateLimited`). +`$ref` by name, and for assembling a complete BAST document. This is +the ergonomics win for alkcall's `OperationSpec`, where +input/output/error schemas reference shared definitions (e.g., +`FileNotFound`, `RateLimited`). ```rust pub struct Definitions { /* ... */ } @@ -410,50 +458,67 @@ impl Definitions { pub fn new() -> Self; /// Define a named schema. Returns a `Schema` that produces - /// `{"$ref": "#/$defs/"}` — the JSON Pointer form that - /// `jsonschema` and `AlkTypeEngine::compile` expect (after - /// `normalize_refs`, which the engine runs at compile time). + /// `{"$ref": "#/$defs/"}` — the JSON Pointer form BAST + /// requires (no `normalize_refs` step; refs are always full + /// pointers). pub fn define(&mut self, name: &str, schema: Schema) -> Schema; /// Like `define`, but the schema is an existing `Value` (adopted /// via `Schema::from_value`). pub fn define_value(&mut self, name: &str, value: Value) -> Schema; - /// Produce the `{"$defs": { ... }}` object to merge into a - /// top-level schema. Call once at the end. + /// Produce the `{"$defs": { ... }}` object. pub fn build(self) -> Value; + + /// Build a complete BAST document with `root_name` as the root + /// type. The root schema is inserted into `$defs` alongside any + /// previously defined entries. The resulting `Value` is ready for + /// `AlkTypeEngine::compile(&doc, root_name, mode, ...)`. + pub fn build_doc(self, root_name: &str, root: Schema) -> Value; + + /// Merge the `$defs` into a top-level schema `Value`. If `top` + /// already has a `$defs` object, the definitions are merged into + /// it; otherwise a `$defs` key is inserted. For BAST documents, + /// prefer `build_doc` — it places the root type inside `$defs` + /// (where BAST requires it). + pub fn merge_into(self, top: &mut Value); } ``` -**Usage:** +**Usage (complete BAST document):** ```rust let mut defs = Definitions::new(); -let file_not_found = defs.define("FileNotFound", - Schema::object() - .field("path", Schema::string_()) - .field("errno", Schema::integer()) - .required(["path", "errno"]) -); +defs.define("Init", Schema::struct_().field("version", Schema::uint32())); +defs.define("Read", Schema::struct_() + .field("handle", Schema::bytes()) + .field("offset", Schema::uint64()) + .field("len", Schema::uint32())); -let rate_limited = defs.define("RateLimited", - Schema::object() - .field("retry_after_ms", Schema::integer().minimum(0)) - .required(["retry_after_ms"]) -); - -let read_file_error = Schema::object() - .field("code", Schema::string_()) - .field("details", Schema::any()) // one of the defined errors - .required(["code"]) - .build(); - -// Merge $defs into the top-level schema that references them -let mut top = Schema::object() - .field("error", read_file_error) - .build(); -top.as_object_mut().unwrap().insert("$defs".to_string(), defs.build()); +let doc = defs.build_doc("Packet", Schema::struct_() + .field("payload", Schema::union_(Discriminator::Byte { + offset: 0, + disc_type: AlkTypeKind::Uint8, + }) + .mapping("1", Schema::ref_def("Init")) + .mapping("5", Schema::ref_def("Read")))); +// -> { +// "$defs": { +// "Init": { "kind": "struct", "fields": [ { "name": "version", "kind": "uint32" } ] }, +// "Read": { "kind": "struct", "fields": [ ... ] }, +// "Packet": { "kind": "struct", "fields": [ +// { "name": "payload", "kind": { +// "kind": "union", +// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" }, +// "mapping": { "1": { "$ref": "#/$defs/Init" }, "5": { "$ref": "#/$defs/Read" } } +// } } +// ] } +// } +// } +// +// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None) +// then validate incoming frames via engine.validate_bytes(&frame). ``` `define` returns a `Schema` (the `$ref` to the definition), so it can @@ -483,31 +548,36 @@ For cases where the `Definitions::define` return value isn't handy ## Usage Examples -### Example 1: channels' 8-byte chunk header (binary layout) +### Example 1: channels' 8-byte chunk header (binary layout, BAST) ```rust -use alktype::{Schema, Endian}; +use alktype::{Schema, Endian, Definitions}; let chunk_header = Schema::struct_() .endian(Endian::Big) .field("channel_id", Schema::uint32()) - .field("length", Schema::uint32()) - .build(); + .field("length", Schema::uint32()); +// Build a complete BAST document (single-type — one $defs entry). +let doc = Definitions::new().build_doc("ChunkHeader", chunk_header); // -> { -// "AlkType:Struct": true, -// "endian": "big", -// "properties": { -// "channel_id": { "AlkType:Uint32": true }, -// "length": { "AlkType:Uint32": true } +// "$defs": { +// "ChunkHeader": { +// "kind": "struct", +// "endian": "big", +// "fields": [ +// { "name": "channel_id", "kind": "uint32" }, +// { "name": "length", "kind": "uint32" } +// ] +// } // } // } // -// Feed to AlkTypeEngine::compile(&mut chunk_header, LayoutMode::Packed) +// Feed to AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None) // then validate incoming frames via engine.validate_bytes(&frame). ``` -### Example 2: call's `OperationSpec` input schema (JSON payload) +### Example 2: call's `OperationSpec` input schema (JSON payload, standard JSON Schema) ```rust use alktype::Schema; @@ -530,18 +600,19 @@ let read_file_input = Schema::object() // } // // Stored in OperationSpec.input_schema; validated via the standard -// jsonschema validator (validate_json for parsed payloads, or via -// serde_json::from_slice then validate_json for wire frames). +// jsonschema validator (AlkTypeEngine::compile with Some(&read_file_input) +// for the validate_json path, or serde_json::from_slice then +// validate_json for wire frames). ``` ### Example 3: SFTP `Packet` union (binary layout, byte discriminator) The SFTP wire shape is `[type:u8][payload-struct]` — a struct with a -union payload field. The engine requires `AlkType:Struct` at the top -level (`OffsetMap::compute` / `SequentialReader::new` both enforce -this; a `Union` is a field type within a struct, not a top-level -schema). The builder constructs the union wrapped in a struct, and -`$defs` are merged into the top-level schema so `$ref`s resolve: +union payload field. The engine requires a struct at the root +(`OffsetMap::compute` / `SequentialReader::new` both enforce this; a +`Union` is a field type within a struct, not a top-level schema). The +builder constructs the union wrapped in a struct, and `$defs` are +placed inside the document via `build_doc` so `$ref`s resolve: ```rust use alktype::{Definitions, Discriminator, AlkTypeKind, Schema}; @@ -555,27 +626,21 @@ defs.define("Status", Schema::struct_().field("code", Schema::uint32()).field("m // A "Packet" is a struct with one field — the union. This mirrors // SFTP's wire shape: [type:u8][payload-struct]. -let mut packet = Schema::struct_() - .field( - "payload", - Schema::union_(Discriminator::Byte { - offset: 0, - disc_type: AlkTypeKind::Uint8, - }) - .mapping("1", Schema::ref_def("Init")) - .mapping("3", Schema::ref_def("Open")) - .mapping("5", Schema::ref_def("Read")) - .mapping("6", Schema::ref_def("Write")) - .mapping("101", Schema::ref_def("Status")), - ) - .build(); -// Merge $defs into the top-level schema so $refs resolve at compile time. -defs.merge_into(&mut packet); -// Feed to AlkTypeEngine::compile(&mut packet, LayoutMode::Packed) +let doc = defs.build_doc("Packet", Schema::struct_() + .field("payload", Schema::union_(Discriminator::Byte { + offset: 0, + disc_type: AlkTypeKind::Uint8, + }) + .mapping("1", Schema::ref_def("Init")) + .mapping("3", Schema::ref_def("Open")) + .mapping("5", Schema::ref_def("Read")) + .mapping("6", Schema::ref_def("Write")) + .mapping("101", Schema::ref_def("Status")))); +// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None) // then validate incoming frames via engine.validate_bytes(&frame). ``` -### Example 4: OperationSpec error schemas (named `$defs`) +### Example 4: OperationSpec error schemas (named `$defs`, standard JSON Schema) ```rust use alktype::{Definitions, Schema}; @@ -610,15 +675,19 @@ let op_errors = vec![ http_status: Some(429), }, ]; -// $defs is built once and stored alongside the OperationSpec +// `$defs` is built once and stored alongside the OperationSpec. +// (For the validate_json path, compile with Some(&defs.build()) as the +// json_schema argument — but typically OperationSpec schemas are +// validated directly via jsonschema, not via AlkTypeEngine.) ``` ## Design Decisions | Decision | ADR | Summary | |----------|-----|---------| -| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 | -| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation shapes the builder's setters produce | +| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 | +| BAST format + two output formats | [ADR-BAST](decisions/bast-bast-format.md) | `struct_()` → BAST, `object()` → standard JSON Schema (D-BAST-008) | +| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation semantics the builder's setters produce (location moved to BAST type-level properties) | | Load-time validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | The builder does not pre-validate; compile-time is the validation point | ## Open Questions @@ -634,12 +703,15 @@ and transitively on `Schema::union_`). See ## References - [ADR-009](decisions/009-builder-api.md) — the decision this spec implements +- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format and the + two-output-format decision (D-BAST-008) - [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) — scope boundaries this module extends; "schemas are JSON" principle - [ADR-003](decisions/003-schema-annotations.md) — the annotation - shapes the builder's setters produce -- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds the - builder's constructors produce + semantics the builder's setters produce +- [`bast-format.md`](bast-format.md) — the normative BAST format + specification (the output format for `struct_()`) +- [schema-layer.md](schema-layer.md) — the BAST kinds and parser - [validation.md](validation.md) — the validation layer that consumes builder output (via `AlkTypeEngine::compile`) - `@alkdev/alknet: docs/architecture/crates/call/operation-registry.md` diff --git a/docs/architecture/data-access.md b/docs/architecture/data-access.md index 3948431..dd9d4da 100644 --- a/docs/architecture/data-access.md +++ b/docs/architecture/data-access.md @@ -238,21 +238,23 @@ pub struct UnionDispatch { } ``` -After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)` -to get the variant schema, then reads the variant's fields at +After dispatch, the consumer calls `tunion::resolve_variant(union_node, &dispatch.key)` +to get the variant `BastType`, then reads the variant's fields at `dispatch.variant_offset` using the normal `data_access` functions (or a -fresh `SequentialReader` scoped to the variant). +fresh `SequentialReader` scoped to the variant). `$ref` variant types are +returned as `BastType::Ref`; the caller resolves them via +`BastDoc::resolve_typeref` when a concrete definition is needed. ### Byte-offset discriminator ```rust /// Read the discriminator value from a byte-offset TUnion. The discriminator -/// is a fixed-size integer (AlkType:Uint8/Uint16/Uint32) at a known byte -/// offset. Returns the mapping key (stringified integer) and the variant -/// struct offset. +/// is a fixed-size integer (uint8/uint16/uint32) at a known byte offset. +/// Returns the mapping key (stringified integer) and the variant struct +/// offset. pub fn read_byte_discriminator( buffer: &[u8], - union_schema: &Value, + union_node: &BastUnion<'_>, endian: Endian, ) -> Result; ``` @@ -268,11 +270,11 @@ starts at `offset + discriminator_size`. /// Read the discriminator value from a field-name TUnion. The /// discriminator is a named field within the struct — the consumer /// provides the field's computed offset (from the OffsetMap or -/// LayoutBuilder). Supports AlkType:String, Uint8, and Enum discriminator +/// LayoutBuilder). Supports string, uint8, and enum discriminator /// fields. pub fn read_field_discriminator( buffer: &[u8], - union_schema: &Value, + union_node: &BastUnion<'_>, disc_field_offset: usize, endian: Endian, ) -> Result; @@ -287,16 +289,19 @@ the variant's fields starting at the end of the discriminator field. ### Variant resolution ```rust -/// Look up a variant schema from the union's mapping. Inline schemas -/// are returned directly. $ref pointers of the form "#/$defs/" -/// are resolved against the union schema's own $defs block. -pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str) - -> Result<&'a Value, AlkTypeError>; +/// Look up a variant type from the union's mapping. Inline struct/ +/// union/enum types are returned directly; `$ref` pointers are +/// returned as `BastType::Ref` — the caller resolves them via +/// `BastDoc::resolve_typeref` when a concrete definition is needed. +pub fn resolve_variant<'a>( + union_node: &'a BastUnion<'a>, + key: &str, +) -> Result<&'a BastType<'a>, AlkTypeError>; -/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a +/// Get the discriminator's byte size (1/2/4 for uint8/16/32) for a /// byte-offset TUnion. Field-name discriminators have no fixed size -/// and produce a AlkTypeError::Schema. -pub fn discriminator_size(union_schema: &Value) -> Result; +/// and produce an `AlkTypeError::Schema`. +pub fn discriminator_size(union_node: &BastUnion<'_>) -> Result; ``` ### TUnion in the layout engines @@ -324,7 +329,7 @@ dispatch to the primitive `data_access` function for the field's kind. For aligned-mode access, `AlkTypeEngine::read_field(&buffer, "header.version")` returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds -the field's `AlkType:*` kind in the schema, and calls the matching +the field's `AlkTypeKind` in the BAST typed tree, and calls the matching `data_access::read_*` function. `write_field` is the mirror. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue` carrying a layout descriptor; the consumer recurses with a fresh reader diff --git a/docs/architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md b/docs/architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md index 466a1cc..dcb0457 100644 --- a/docs/architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md +++ b/docs/architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md @@ -1,7 +1,19 @@ # ADR-001: alktype — Purpose, Scope, and the jsonschema Engine ## Status -Accepted + +**Superseded (format-specific content) by +[ADR-BAST](bast-bast-format.md).** The crate's purpose, scope +boundaries, and the "schema is the format" principle are **retained and +strengthened** — BAST *is* the format. Only the *concrete format* +(custom-keyword JSON Schema → BAST) and the *validation strategy* +(single `jsonschema` custom-keyword validator → two-validator model) +are superseded: the format-specific content by ADR-BAST, the +validation-strategy content by +[ADR-VAL-SPLIT](val-split-two-validator-model.md). This ADR is kept as +the historical record of the v0.1.0 design and the purpose/scope +decision; read it alongside ADR-BAST and ADR-VAL-SPLIT for the current +state. ## Context diff --git a/docs/architecture/decisions/002-two-layout-modes-packed-vs-aligned.md b/docs/architecture/decisions/002-two-layout-modes-packed-vs-aligned.md index 175a2b9..5352f30 100644 --- a/docs/architecture/decisions/002-two-layout-modes-packed-vs-aligned.md +++ b/docs/architecture/decisions/002-two-layout-modes-packed-vs-aligned.md @@ -1,7 +1,14 @@ # ADR-002: Two Layout Modes — Packed Sequential vs Aligned Static ## Status -Accepted +Accepted — unchanged under the BAST pivot +([ADR-BAST](bast-bast-format.md)). Layout modes are format-agnostic: +the input format changed from custom-keyword JSON Schema to BAST, but +the two modes, their alignment/packing rules, and the +`LayoutBuilder`/`SequentialReader`/`OffsetMap` API did not. The layout +engines now walk the BAST typed tree ([`BastDoc`](../schema-layer.md)) +instead of raw JSON with `get_alktype_kind*`, but the offset +computation algorithm is identical. ## Context diff --git a/docs/architecture/decisions/003-schema-annotations.md b/docs/architecture/decisions/003-schema-annotations.md index c093275..9c2b84b 100644 --- a/docs/architecture/decisions/003-schema-annotations.md +++ b/docs/architecture/decisions/003-schema-annotations.md @@ -1,7 +1,18 @@ # ADR-003: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators ## Status -Accepted + +**Accepted (semantics); amended (location) by +[ADR-BAST](bast-bast-format.md).** The annotation *semantics* decided +here — endianness default, struct/field-level alignment, the three +variable-length encoding strategies, and the two TUnion discriminator +kinds — **carry forward unchanged** under the BAST pivot. Only the +annotation *location* moves: from v0.1.0's custom-keyword objects +(`{"AlkType:String": { "encoding": "..." }}`) to BAST type-level +properties (`{ "name": "handle", "kind": "string", "encoding": "..." }`). +The BAST shapes are normative in +[`bast-format.md`](../bast-format.md#variable-length-encoding); this +ADR is kept as the semantic reference. Read it alongside ADR-BAST. ## Context diff --git a/docs/architecture/decisions/004-error-handling-validation-strategy.md b/docs/architecture/decisions/004-error-handling-validation-strategy.md index 5ff1960..f84ef55 100644 --- a/docs/architecture/decisions/004-error-handling-validation-strategy.md +++ b/docs/architecture/decisions/004-error-handling-validation-strategy.md @@ -1,7 +1,21 @@ # ADR-004: Error Handling and Validation Strategy ## Status -Accepted + +**Accepted (error type); amended (validation strategy) by +[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The `AlkTypeError` +enum, its four variants, the load-time-build / access-time-check split, +and the field-path-carrying errors decided here are **retained +unchanged** under the BAST pivot (D-BAST-009 keeps +`Validation(jsonschema::ValidationError<'static>)`). The "validation +strategy" section — which described v0.1.0's single +`jsonschema`-custom-keyword validator for both paths — is **refined**: +the bytes path now uses the BAST-native validator +(`bast_validation`), the JSON path now uses a standard +`jsonschema::Validator` from a consumer-provided JSON Schema. See +[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the +two-validator model. This ADR is kept as the error-handling reference; +read it alongside ADR-VAL-SPLIT for the current validation strategy. ## Context diff --git a/docs/architecture/decisions/009-builder-api.md b/docs/architecture/decisions/009-builder-api.md index 65f65b9..faf2217 100644 --- a/docs/architecture/decisions/009-builder-api.md +++ b/docs/architecture/decisions/009-builder-api.md @@ -2,7 +2,19 @@ ## Status -Accepted +**Accepted (API surface); amended (output format) by +[ADR-BAST](bast-bast-format.md).** The fluent builder API, the +`Schema`/`Definitions`/`Discriminator` types, the constructor and +setter catalog, and the "produces `serde_json::Value`, not a typed +`Schema` enum" decision decided here are **retained unchanged** under +the BAST pivot. Only the `build()` *output format* changes: +`struct_()` now produces BAST JSON (`{ "kind": "struct", "fields": [...] }`) +instead of v0.1.0's custom-keyword JSON (`{ "AlkType:Struct": true, +"properties": {...} }`); `object()` continues to produce standard JSON +Schema. This is D-BAST-008, recorded in ADR-BAST. The builder examples +in [`builder.md`](../builder.md) reflect the current BAST output. This +ADR is kept as the API-surface decision; read it alongside ADR-BAST +for the output format. ## Context diff --git a/docs/architecture/decisions/010-generalized-validation-validate-bytes.md b/docs/architecture/decisions/010-generalized-validation-validate-bytes.md index 7819c57..5ed223d 100644 --- a/docs/architecture/decisions/010-generalized-validation-validate-bytes.md +++ b/docs/architecture/decisions/010-generalized-validation-validate-bytes.md @@ -2,7 +2,23 @@ ## Status -Accepted +**Accepted (two-step concept); amended (validation step) by +[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The +`validate_bytes(&[u8])` entry point, the "materialize `Value` from +bytes, then validate" two-step concept, the mode dispatch, the +field-path-carrying errors, and the "not a `Validator` trait / not +framing-aware / not a binary-payload validator for JSON-only schemas" +scope boundaries decided here are **retained unchanged** under the +BAST pivot. Only the validation *step's implementation* changes: the +materialized `Value` is validated by the **BAST-native validator** +(`bast_validation`) instead of v0.1.0's `jsonschema` custom-keyword +validator. The `jsonschema` crate is no longer touched on the bytes +path (it remains for the `validate_json` path and for BAST meta-schema +validation). The error payload type stays +`Validation(jsonschema::ValidationError<'static>)` (D-BAST-009). See +[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the +two-validator model. This ADR is kept as the `validate_bytes` decision; +read it alongside ADR-VAL-SPLIT for the current validation step. ## Context diff --git a/docs/architecture/decisions/bast-bast-format.md b/docs/architecture/decisions/bast-bast-format.md new file mode 100644 index 0000000..e922c0c --- /dev/null +++ b/docs/architecture/decisions/bast-bast-format.md @@ -0,0 +1,316 @@ +# ADR-BAST: BAST (Binary Abstract Syntax Tree) as the Schema Format + +## Status + +Accepted — supersedes the format-specific content of +[ADR-001](001-alktype-purpose-scope-jsonschema-engine.md). ADR-001's +purpose, scope, and "schema is the format" principle are retained and +strengthened; only the concrete format (custom-keyword JSON Schema → +BAST) is superseded by this ADR. + +## Context + +alktype v0.1.0 embedded binary layout information inside standard JSON +Schema documents via custom keywords: + +```json +{ + "AlkType:Struct": true, + "type": "object", + "properties": { + "channel_id": { "AlkType:Uint32": true, "type": "integer" }, + "length": { "AlkType:Uint32": true, "type": "integer" } + }, + "endian": "big" +} +``` + +This worked for the Rust engine — it walked the tree, detected +keywords, computed offsets. But it created friction for everything +outside Rust: + +1. **Cross-language consumption.** A Python, Go, or TypeScript consumer + that wanted to parse an alktype schema had to re-implement custom + keyword detection. The format was not self-describing — you needed + to know that `AlkType:Uint32` meant "4-byte little/big-endian + unsigned integer" out-of-band. +2. **No meta-schema.** The custom keywords were not part of any JSON + Schema dialect, so `jsonschema` itself could not validate an alktype + schema's *structure*. Editors had no autocomplete; a typo in + `AlkType:Uint32` (e.g., `AlkType:UINT32`) was a runtime engine + error, not a schema-validation error. +3. **Awkward composition.** The keyword-value shape (`true` vs an + annotation object) and the `normalize_refs` step needed to bridge + TypeBox's bare-name `$ref` output and `jsonschema`'s JSON Pointer + requirement were engine internals leaking into the format. +4. **Validator coupling.** The v0.1.0 bytes path validated by + registering 19 `jsonschema::Keyword` factories (~200 lines). The + custom keyword integration was the only way to enforce value-domain + constraints (integer ranges, `maxLength`, enum index bounds) on the + materialized `Value`. The built-in `enum` keyword checked string + membership, but the materializer emitted `Value::Number(index)` — + so out-of-bounds enum indices *silently passed* (a dead constraint). + +The engine's core logic (layout computation, data access, union +dispatch, two layout modes) was format-agnostic beneath the accessor +layer. A POC on branch `bast-validator-poc` (commit `f371fe4`, +`src/bast_poc.rs`) proved that a `kind`-based vocabulary with +`$defs`/`$ref`, a BAST-native validator, and lazy variant ref +resolution could replace the custom-keyword machinery end-to-end with +no loss of capability and a net reduction in code. The research record +is [`docs/research/bast-pivot.md`](../../research/bast-pivot.md); +decisions D-BAST-001 through D-BAST-009 are recorded there. + +## Decision + +**alktype's schema format is BAST (Binary Abstract Syntax Tree): a JSON +document that describes binary data layouts using a `kind`-based +vocabulary with `$defs`/`$ref` for composition.** BAST is itself a +valid JSON Schema instance (it has a meta-schema), making it +self-validating, editor-friendly, and trivially consumable from any +language with a JSON parser. + +The normative format specification is +[`docs/architecture/bast-format.md`](../bast-format.md) (meta-schema, +TypeRef, examples, validation model). This ADR records the decision and +its consequences; the spec records the shape. + +### Design principles + +1. **BAST is a JSON Schema instance.** A BAST document is valid JSON + that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON + Schema). Any JSON Schema validator can check whether a BAST document + is well-formed; editors with JSON Schema support provide autocomplete + and inline validation for free. +2. **`$defs`/`$ref` for composition.** Named type definitions live in a + top-level `$defs` block. `$ref` handles cross-references and union + variant references — the same pattern as JSON Schema's own `$defs` + and TypeBox's `Type.Module`. No custom reference resolution + mechanism. +3. **`kind`-based vocabulary.** Every type has a `kind` field whose + value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.). + This replaces the `AlkType:*` custom-keyword pattern with a flat, + easily-matched string. The 19 `AlkTypeKind` enum variants are + unchanged; `AlkTypeKind::from_bast_str`/`to_bast_str` map between + the enum and the lowercase BAST strings (D-BAST-002). +4. **Order is explicit.** Struct fields are an ordered array, not an + object with `properties`. Field order is unambiguous — no reliance + on `serde_json`'s `preserve_order` for correctness — and matches the + mental model of binary layouts. +5. **Annotations are type-level properties.** Endianness, alignment, + encoding, and discriminators are properties of the type definition + or field, not custom keywords on a separate schema object. Their + *semantics* carry forward unchanged from + [ADR-003](003-schema-annotations.md); only their *location* moves. + +### Document shape + +Every BAST document has the same top-level shape: + +```json +{ "$defs": { "": { ...TypeDef... }, ... } } +``` + +- The `$defs` block is **required** (D-BAST-003). Single-type documents + are a special case with one entry. +- The **root type name** is a required parameter to + `AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001). + Convention (first entry) is fragile and depends on JSON key order; an + explicit parameter is used instead. + +### TypeRef + +`TypeRef` is the central mechanism for referencing types. Four forms: +primitive string (`"uint32"`), `$ref` object +(`{ "$ref": "#/$defs/Read" }`), array object +(`{ "kind": "array", "element": "uint32", "count": 3 }`), and record +object (`{ "kind": "record", "values": "string" }`). + +The `$ref` form uses standard JSON Pointer syntax **restricted to +`#/$defs/`** — no external references, no fragment-only pointers, +no bare names. The restriction keeps resolution a single hash lookup +and eliminates the `normalize_refs` step the v0.1.0 engine needed for +TypeBox's bare-name refs. + +### Meta-schema + +The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that +validates the *structure* of BAST documents (is it well-formed?). It +lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded +in the crate as `BAST_META_SCHEMA` (re-exported from the crate root) for +offline use. A *different* validator — the BAST-native validator (see +[ADR-VAL-SPLIT](val-split-two-validator-model.md)) — validates *binary +data* against a BAST document (are the bytes a valid instance?). These +are different validators for different inputs. + +### The typed parser + +`src/bast.rs` parses a BAST document into a borrowed typed tree +(`BastDoc`/`BastDef`/`BastStruct`/`BastField`/`BastType`/`BastUnion`/ +`BastEnum`/`BastArray`/`BastRecord`/`BastRef`). Three consumers (layout +engines, materializer, BAST-native validator) walk the same tree, so a +typed view pays for itself. See +[`schema-layer.md`](../schema-layer.md) for the parser's surface and +[`bast-format.md`](../bast-format.md) for the format. + +Variant `$ref`s (union `mapping` entries) are resolved **lazily** by +the materializer and validator via `BastDoc::resolve_typeref` — no +compile-time inlining. + +### Untrusted input + +Every path that walks a BAST document returns +`Err(AlkTypeError::Schema)` on a malformed document, never +`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream +`alkcall` consumer accepts schemas from arbitrary internet peers in its +hub/spoke topology). + +### Bug fix: enum index bounds + +The v0.1.0 engine had a dead constraint on the bytes path — the +built-in `enum` keyword checked string membership, but the materializer +emitted `Value::Number(index)`, which never matched. The BAST-native +validator checks the materialized index against `values.len()` bounds, +fixing this. Net improvement, recorded as intended behavior in the +test suite. + +## What is removed + +Under the BAST pivot, the v0.1.0 custom-keyword machinery is removed: + +- All 19 `jsonschema::Keyword` implementations (~200 lines of validator + factories) — replaced by the BAST-native validator (~250 lines, a + flat match with no factories, no trait objects, no sub-validator + pre-computation). See [ADR-VAL-SPLIT](val-split-two-validator-model.md). +- `normalize_refs()` / `inline_union_variant_refs()` — BAST refs are + always `#/$defs/`; one hash lookup. Variant refs resolve lazily. +- `get_alktype_kind*` family — superseded by the parser's `kind`-string + dispatch. +- `parse_encoding`/`parse_align`/`parse_max_length`/`parse_endian`/ + `parse_discriminator` + `DiscriminatorKind` — replaced by the typed + `BastField`/`BastDiscriminator` views and the parser's internal + BAST-property-form copies. +- `resolve_ref`/`resolve_ref_or_inline` — replaced by + `BastDoc::lookup_def`/`resolve_typeref`. +- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`, + `BYTE_DISCRIMINATOR_TYPES` — replaced by `from_bast_str`/`to_bast_str` + and the parser's typed views. +- The custom-keyword `build_validator` path — `build_validator` is + repurposed to build a *standard* `jsonschema::Validator` from a + consumer-provided JSON Schema (no custom keywords). See + [ADR-VAL-SPLIT](val-split-two-validator-model.md). + +The `jsonschema` crate **remains a direct dependency** for +`validate_json` and for validating BAST documents against the BAST +meta-schema. The only thing removed is the custom keyword integration +path. + +## Consequences + +### Positive + +- **Self-describing, cross-language format.** A BAST document carries + its type vocabulary in a meta-schema'd JSON Schema instance. Any + language with a JSON parser and a JSON Schema validator can validate + BAST document structure without knowing alktype's Rust internals. + Editors with `$schema` support provide autocomplete and inline + validation for free. +- **Simpler `$ref` story.** One restricted form (`#/$defs/`), one + hash lookup, no normalization pass. TypeBox interop is a serialization + concern (TypeBox → BAST JSON), not an engine concern. +- **Explicit field order.** The `fields` array makes byte order + unambiguous — no reliance on `serde_json`'s `preserve_order` for + correctness (it remains a dependency for builder output and for + `mapping` iteration order, but layout correctness no longer depends + on it). +- **Architecture simplification.** The BAST-native validator is a flat + recursive match — no factories, no trait objects, no sub-validator + pre-computation, no `with_keyword` registration. ~200 lines of + custom-keyword validators become ~250 lines of straightforward + pattern matching. +- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed. +- **Wasm binary-size win.** The `validate_bytes` path no longer touches + `jsonschema` for validation (it still uses `jsonschema`'s + `ValidationError::custom` type for the error payload, per + D-BAST-009 — but no validator compilation, no keyword registration, + no sub-validators). + +### Negative + +- **Breaking change to the v0.1.0 public surface.** `compile`'s + signature changes (new `root_name` param, drops `&mut`, takes a BAST + document not a custom-keyword JSON Schema). `validate_json`/ + `is_valid_json` change contract (validate against a consumer-provided + JSON Schema, not the alktype schema). `Schema::build`/ + `Definitions::build` output format changes. The ~13 `schema::*` + helper re-exports are removed. `build_validator` is repurposed. The + crate is on crates.io at 0.1.0 with zero real consumers, so the bump + is free — but the contract is explicit (see the implementation plan's + Semver Contract table). +- **Two output formats from the builder.** `struct_()` → BAST, + `object()` → standard JSON Schema. The construction API is the same; + only the serialization differs. This is deliberate (D-BAST-008) but + is a thing consumers must learn. +- **`AlkTypeKind::Display` is backed by `to_bast_str`.** The v0.1.0 + `as_str`/`Display` rendered the `"AlkType:Uint8"` keyword; the new + `Display` renders the BAST canonical string (`"uint8"`). Error + messages across six modules surface the new name. This is the right + name to surface now, but it is a visible change in error output. + +## Scope Boundaries (What This Is Not) + +- **Not a replacement for JSON Schema for JSON validation.** A BAST + document cannot validate a JSON payload — it describes binary data + layouts and value-domain constraints for bytes. For JSON validation, + consumers use standard JSON Schema documents (which may be derived + from BAST via future codegen, or authored separately). See + [ADR-VAL-SPLIT](val-split-two-validator-model.md). +- **Not a code generator.** BAST is a data format, not a Rust source + generator. ADR-001's scope boundary stands. +- **Not a schema-evolution / Value system.** TypeBox's `Value.Diff`, + `Value.Migrate`, `Value.Convert` remain out of scope (ADR-001). +- **Not a framing format.** BAST describes one struct/union/enum + instance; it does not strip length prefixes or handle multi-frame + buffers. Framing stays in the consumer (ADR-010). + +## Decisions (D-BAST-001..009) + +The BAST format is grounded in decisions D-BAST-001 through D-BAST-009, +recorded in +[the pivot research record](../../research/bast-pivot.md#decisions). +Summary: + +| Decision | Summary | +|----------|---------| +| [D-BAST-001](../../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention | +| [D-BAST-002](../../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase | +| [D-BAST-003](../../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape | +| [D-BAST-004](../../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) | +| [D-BAST-005](../../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` | +| [D-BAST-006](../../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed | +| [D-BAST-007](../../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema | +| [D-BAST-008](../../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema | +| [D-BAST-009](../../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths | + +## References + +- [`bast-format.md`](../bast-format.md) — the normative BAST format + specification +- [`schema-layer.md`](../schema-layer.md) — the BAST parser + implementation +- [ADR-VAL-SPLIT](val-split-two-validator-model.md) — the two-validator + model (BAST-native for bytes, standard `jsonschema` for JSON) +- [ADR-001](001-alktype-purpose-scope-jsonschema-engine.md) — purpose, + scope, and the "schema is the format" principle (format-specific + content superseded by this ADR; purpose/scope retained) +- [ADR-003](003-schema-annotations.md) — annotation semantics (carry + forward unchanged; only location moves) +- [ADR-009](009-builder-api.md) — builder API (output format amended + to BAST / standard JSON Schema) +- [ADR-010](010-generalized-validation-validate-bytes.md) — + `validate_bytes` (validation step amended to the BAST-native + validator) +- [BAST pivot research record](../../research/bast-pivot.md) — + motivation, POC scope and result, decisions D-BAST-001..009, risks +- [BAST pivot implementation plan](../../plans/bast-implementation.md) + — ordered steps, semver contract, ADR-sync checklist \ No newline at end of file diff --git a/docs/architecture/decisions/val-split-two-validator-model.md b/docs/architecture/decisions/val-split-two-validator-model.md new file mode 100644 index 0000000..5716575 --- /dev/null +++ b/docs/architecture/decisions/val-split-two-validator-model.md @@ -0,0 +1,229 @@ +# ADR-VAL-SPLIT: Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON + +## Status + +Accepted — refines the "validation strategy" section of +[ADR-004](004-error-handling-validation-strategy.md) and the "validation +step" of [ADR-010](010-generalized-validation-validate-bytes.md) for +the BAST pivot. Records decisions D-BAST-006, D-BAST-007, and +D-BAST-009. + +## Context + +alktype v0.1.0 used a single validation mechanism — the `jsonschema` +crate with 19 custom keyword validators — for both the JSON path +(`validate_json(&Value)`) and the bytes path (`validate_bytes(&[u8])`). +The bytes path materialized a `serde_json::Value` tree from the buffer, +then ran the same `jsonschema::Validator` against it. + +Under the BAST pivot ([ADR-BAST](bast-bast-format.md)), the format +changed from custom-keyword JSON Schema to BAST, and the custom-keyword +integration was removed. This forced a re-evaluation of both validation +paths: + +1. **The bytes path.** BAST is the complete specification of the binary + format — it describes both the layout (how to read) and the + constraints (what values are valid). An external JSON Schema is not + needed for `validate_bytes`; the BAST document *is* the validation + spec for bytes. The natural validator is a recursive walker over the + BAST type tree that checks the value-domain constraints the + materializer does not (integer ranges, `maxLength`, timestamp shape, + enum index bounds, union variant constraints). The POC + (`bast-validator-poc` branch, `src/bast_poc.rs`) proved this out + end-to-end with 20 reference tests. + +2. **The JSON path.** BAST describes bytes, not JSON shape. A JSON + `Value` (e.g., an incoming JSON-RPC request) is the wrong input for + a BAST document; the right validator is a standard + `jsonschema::Validator` built from a standard JSON Schema document + the consumer provides. BAST is not involved on this path. This is + the path alkcall uses for its `OperationSpec` JSON validation. + +The two paths have different inputs (bytes vs JSON `Value`), different +schema sources (the BAST document vs a consumer-provided JSON Schema), +and different validators (a flat recursive match vs a compiled +`jsonschema::Validator`). But they share the same error variant — +`AlkTypeError::Validation` — so consumers handling both (alkcall uses +`validate_json` for channel 0 JSON-RPC and `validate_bytes` for binary +channels) match one arm. + +## Decision + +**alktype has two validators for two input types:** + +| Path | Input | Validator | Schema source | +|------|-------|-----------|---------------| +| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator (`bast_validation`) | The BAST document (binary layout + value constraints) | +| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema | + +### `validate_bytes` — BAST-native validator (D-BAST-006) + +`src/bast_validation.rs` is a recursive walker +(`validate_value(doc, &value)`) over the BAST typed tree +([`crate::bast::BastDoc`]/[`BastType`]). The materializer +(`src/materialize.rs`) produces a structurally-correct `Value` tree +from bytes (all declared fields present, types correct, bounds checked, +UTF-8 valid, discriminator in mapping, boolean byte 0 or 1). The +validator enforces only the **value-domain constraints expressed in the +BAST document** — the ones the materializer can't see from the bytes +alone: + +| Constraint | Validator arm | +|------------|---------------| +| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` | +| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` | +| Float finiteness (Float32/64) | `validate_float` | +| String `maxLength` (byte length) | `check_string` | +| Bytes `maxLength` (array length) | `check_bytes` (accepts `Value::String` and `Value::Array`) | +| RFC 3339 timestamp shape | `validate_timestamp` (non-strict, matching v0.1.0) | +| Enum index bounds | `validate_enum` — **fixes the v0.1.0 dead constraint** | +| Union variant dispatch | `validate_union` reads `__discriminator`, resolves the variant, recurses | +| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses | +| Array count | `validate_array` checks `arr.len() == count` and recurses per element | +| Record values | `validate_record` recurses into each value's `values` type | +| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) | + +The validator is a flat `match` — no factories, no trait objects, no +sub-validator pre-computation, no `jsonschema` involvement. ~250 lines +replace ~200 lines of v0.1.0 custom-keyword factories. + +No external JSON Schema is required. The BAST document is the complete +specification of the binary format. An optional external JSON Schema +can be layered on top for constraints BAST doesn't express (cross-field +consistency, regex patterns on string content) — additive, not +load-bearing. + +### `validate_json` — standard jsonschema (D-BAST-007) + +`validate_json(&Value)` / `is_valid_json(&Value)` validate a JSON +`Value` against a standard `jsonschema::Validator` compiled at +`AlkTypeEngine::compile` time from a consumer-provided JSON Schema +(`compile`'s `json_schema: Option<&Value>` parameter). No custom +keywords, no BAST involvement. The JSON Schema is independent of the +BAST document — BAST describes bytes, not JSON shape. It may be +authored separately or derived from BAST via future codegen. + +If no JSON Schema was supplied to `compile`, `validate_json` returns +`AlkTypeError::Schema` and `is_valid_json` returns `false`. + +`build_validator` (in `src/validation.rs`) is **repurposed**: it builds +a *standard* `jsonschema::Validator` from a plain JSON Schema (no +custom keywords). The engine calls it internally during `compile` when +`json_schema` is `Some`. Consumers that only need a one-off validator +may call `jsonschema::options().build(schema)` directly; `build_validator` +exists so the engine's error mapping (`jsonschema` build error → +`AlkTypeError::Schema`) is reused. The v0.1.0 custom-keyword +`build_validator` is removed. + +The `jsonschema` crate remains a direct dependency for this path and +for validating BAST documents against the BAST meta-schema. + +### Error payload (D-BAST-009) + +`AlkTypeError::Validation(jsonschema::ValidationError<'static>)` is +**retained** as the error variant for both paths. The bytes path no +longer uses `jsonschema` for validation, so its error payload is +constructed via `jsonschema::ValidationError::custom` purely to keep +the variant's type unchanged. The rationale is consumer ergonomics on +the *combined* path: consumers like alkcall use both `validate_json` +and `validate_bytes` and handle `AlkTypeError::Validation` in one +place. A single uniform payload type means one match arm covers both +sources. + +The alternative (`Validation(String)`) was rejected — it would force +`validate_json` to flatten its structured errors (instance path, schema +path, keyword) to a `String` via `Display`. The more information-rich +path would lose data to accommodate the less rich one. That is the +wrong direction. + +The `no_std`/minimal-build angle (OQ-002) that the alternative was +meant to enable is moot: `validate_json` requires `jsonschema` +regardless, so a bytes-only `no_std` build already has to give up +`validate_json` as a separate, larger decision. The right place to +revisit is when/if OQ-002 is actually pursued. + +## Consequences + +### Positive + +- **Right validator for each input.** Bytes are validated by the BAST + document that describes them; JSON values are validated by a JSON + Schema that describes them. No forced isomorphism between two + different input types. +- **No external JSON Schema needed for `validate_bytes`.** The BAST + document is both the layout spec and the validation spec for bytes. + This is the "schema is the format" principle from ADR-001, now fully + realized. +- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed + — the BAST-native validator checks the materialized index against + `values.len()` directly. +- **Per-variant constraint enforcement (OQ-008) without custom + keywords.** The validator recurses into the selected variant's BAST + definition on `__discriminator` lookup, enforcing every field + constraint the variant declares (e.g., `maxLength` on a `bytes` field + inside a variant struct). +- **Wasm binary-size win.** The `validate_bytes` path no longer touches + `jsonschema` for validation (it still uses `ValidationError::custom` + for the error payload type, per D-BAST-009 — but no validator + compilation, no keyword registration, no sub-validators). +- **Architecture simplification.** ~200 lines of custom-keyword + factories become ~250 lines of straightforward pattern matching. No + `with_keyword` registration; no `inline_union_variant_refs` compile + step. +- **Uniform error payload.** Consumers handle one + `AlkTypeError::Validation` match arm for both paths (D-BAST-009). + +### Negative + +- **Two validators, not one.** The engine struct carries an + `Option` (for `validate_json`) and re-parses + the BAST typed tree on each `validate_bytes` call (the BAST-native + validator is not pre-built — it's a recursive walker over the + on-demand `BastDoc`). This is a small cost; the validators serve + different inputs and don't share structure. +- **`validate_json` requires a consumer-provided JSON Schema.** The + engine no longer builds a validator from the alktype schema; the + consumer must supply a JSON Schema at `compile` time (or accept that + `validate_json` returns `AlkTypeError::Schema`). This is a behavioral + break from v0.1.0, intentional under the pivot. +- **`AlkTypeError::Validation` payload is `jsonschema`'s type even on + the bytes path.** The bytes path constructs it via + `ValidationError::custom`, which is slightly awkward but keeps the + variant uniform. The `no_std` revisit (OQ-002) is the place to + reconsider if a bytes-only minimal build ever materializes. + +## Scope Boundaries (What This Is Not) + +- **Not a `Validator` trait abstraction.** Two methods on one struct, + not a trait with impls for JSON-only and BAST-binary schemas. The two + impls share little internally (`validate_json` is a single + `jsonschema` call; `validate_bytes` is materialize + BAST-native + walk), so a trait would add a layer without unifying behavior. See + [ADR-010](010-generalized-validation-validate-bytes.md) §"Not a + `Validator` trait abstraction". +- **Not a binary-aware validator that skips the `Value` tree.** The + `Value`-materialization path is the validation path. A future + "validate bytes without materializing" path is a two-way door but + explicitly out of scope for v1 (would re-introduce a hand-rolled + validator, ADR-001). +- **Not framing-aware.** `validate_bytes` validates the bytes of *one* + schema instance. Framing stays in the consumer (ADR-010). + +## References + +- [`bast-format.md` §Validation Model](../bast-format.md#validation-model) + — the normative validation model +- [ADR-BAST](bast-bast-format.md) — the BAST format decision +- [ADR-004](004-error-handling-validation-strategy.md) — error handling + and validation strategy (load-time build, access-time check, + `AlkTypeError` enum — retained; validation-strategy section refined + by this ADR) +- [ADR-010](010-generalized-validation-validate-bytes.md) — + `validate_bytes` (the two-step concept retained; the validation step + amended to the BAST-native validator by this ADR) +- [`validation.md`](../validation.md) — the validation layer + documentation +- `src/bast_validation.rs` — the BAST-native validator implementation +- `src/validation.rs` — the `build_validator` helper +- [BAST pivot research record](../../research/bast-pivot.md) — + D-BAST-006, D-BAST-007, D-BAST-009 \ No newline at end of file diff --git a/docs/architecture/layout-engine.md b/docs/architecture/layout-engine.md index 2ee988b..15077c6 100644 --- a/docs/architecture/layout-engine.md +++ b/docs/architecture/layout-engine.md @@ -7,8 +7,8 @@ last_updated: 2026-07-22 The layout engine: offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, and variable-length -field handling. This is the novel code — the recursive walk of the schema -JSON that computes byte positions for each field. +field handling. This is the novel code — the recursive walk of the BAST +typed tree that computes byte positions for each field. ## The Two Layout Modes @@ -24,8 +24,8 @@ protocols. **Components:** -- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `AlkType:Struct` at the top level), then `builder.build(&var_sizes) -> Result` where `var_sizes: &HashMap` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions. -- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame. +- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(bast_doc, root_name)` (requires a `struct` at the root), then `builder.build(&var_sizes) -> Result` where `var_sizes: &HashMap` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions. +- **`SequentialReader`** — constructed via `SequentialReader::new(bast_doc, root_name)`, then driven by `reader.read_next(&buffer) -> Result, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame. **How it works:** @@ -69,7 +69,7 @@ and safetensors. **Component:** -- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result` (requires `AlkType:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets. +- **`OffsetMap`** — constructed via `OffsetMap::compute(&doc) -> Result` (requires a `struct` at the root). Walks the BAST typed tree once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets. **How it works:** @@ -118,14 +118,15 @@ fields must use `maxLength` (fixed-size reservation) or ## Offset Computation Algorithm -The offset computation is a recursive walk of the schema JSON. The +The offset computation is a recursive walk of the BAST typed tree +([`BastDoc`](schema-layer.md#the-bast-parser-bast-module)). The algorithm is the same for both modes; the difference is whether alignment padding is inserted between fields. ### Fixed-size types For each fixed-size type, the algorithm: -1. Determines the type's byte size from the `AlkType:*` kind. +1. Determines the type's byte size from the `AlkTypeKind`. 2. In aligned mode: inserts padding to satisfy the type's alignment (or the field's `align` annotation, or the struct's `align` default). 3. Records the field's `(start, end)` range. @@ -133,14 +134,14 @@ For each fixed-size type, the algorithm: ### Composite types -**`TStruct`:** Recurse into the struct's `properties`. The inner fields +**`struct`:** Recurse into the struct's `fields` array. The inner fields are computed relative to the struct's start offset. The struct's total size is the sum of its fields' sizes (plus alignment padding in aligned mode). The struct itself may have an `align` annotation that rounds up its total size. -**`TUnion`:** TUnion is supported in packed sequential mode only. In -aligned static mode, `OffsetMap::compute` rejects `TUnion` fields with +**`union`:** TUnion is supported in packed sequential mode only. In +aligned static mode, `OffsetMap::compute` rejects `union` fields with `AlkTypeError::Offset` — see [ADR-008](decisions/008-reject-tunion-in-aligned-mode.md). Unions are the protocol dispatch pattern (SFTP type bytes, call protocol event @@ -161,12 +162,12 @@ total size. The `SequentialReader` reads the discriminator first, looks up the variant schema, then reads the variant struct sequentially — it doesn't need to know the union's total size upfront. -**`TArray` of fixed-size elements:** Element stride = element size (plus +**`array` of fixed-size elements:** Element stride = element size (plus alignment padding in aligned mode). Element `i` starts at `array_offset + i × stride`. The array's total size is `count × stride`. -**`TArray` of variable-length-element structs:** Deferred for v1 -(OQ-001). +**`array` of variable-length-element structs:** Deferred for v1 +(OQ-001, D-BAST-004 — BAST arrays require `count` in v1). ### Variable-length types @@ -289,7 +290,7 @@ pub struct FieldPosition { A field's computed position in a packed layout, produced by `LayoutBuilder::build`. For variable-length fields, `size` is `4` (the length prefix); for fixed-size fields, `size` is the type's byte size. -`kind` records the field's `AlkType:*` kind so the consumer can dispatch +`kind` records the field's `AlkTypeKind` so the consumer can dispatch to the correct `data_access` read/write function. ### `PackedLayout` (packed mode) @@ -317,16 +318,16 @@ A flat table of `(field_path, byte_range)` pairs computed from a schema. ```rust impl OffsetMap { - pub fn compute(schema: &Value) -> Result; + pub fn compute<'a>(doc: &'a BastDoc<'a>) -> Result; pub fn get(&self, field_path: &str) -> Option<&ByteRange>; pub fn total_size(&self) -> usize; pub fn iter(&self) -> impl Iterator; } ``` -`compute` requires a `AlkType:Struct` at the top level. `total_size` -includes trailing alignment padding. `iter` yields fields in insertion -order (schema `properties` order, nested struct fields appearing inline). +`compute` requires a `struct` at the root. `total_size` +includes trailing alignment padding. `iter` yields fields in the BAST +`fields` array order (nested struct fields appearing inline). ## Design Decisions diff --git a/docs/architecture/open-questions.md b/docs/architecture/open-questions.md index 28302ca..029a5ed 100644 --- a/docs/architecture/open-questions.md +++ b/docs/architecture/open-questions.md @@ -106,31 +106,35 @@ architect's desk" is answerable at a glance. ### OQ-006: Builder spec Example 3 — wrap `Union` in a `Struct` — RESOLVED - **Status**: resolved. [builder.md](builder.md) Example 3 now wraps - the `Union` in a `Schema::struct_().field("payload", ...)` and merges - `$defs` via `Definitions::merge_into`. Matches the realistic SFTP - wire shape and the engine's `AlkType:Struct`-at-root constraint. + the `Union` in a `Schema::struct_().field("payload", ...)` and builds + a complete BAST document via `Definitions::build_doc`. Matches the + realistic SFTP wire shape and the engine's struct-at-root constraint + (the root `$defs` entry must be a `struct`). - **Full file**: [OQ-006](questions/006-builder-spec-example-3-wrap-union.md) ### OQ-007: `Bytes` materialization — lossy UTF-8 conversion — RESOLVED - **Status**: resolved. Array of u8: the materializer produces `Value::Array` of `Value::Number` (one entry per byte, 0..=255) for - `AlkType:Bytes` fields. The `BytesValidator` accepts both - `Value::String` (for `validate_json`) and `Value::Array` (for - `validate_bytes`). `maxLength` = max byte count. Implemented in - `src/materialize.rs` and `src/validation.rs`. + `bytes` fields. The BAST-native validator (`bast_validation::check_bytes`) + accepts both `Value::String` (for `validate_json`-style inputs) and + `Value::Array` (for `validate_bytes`). `maxLength` = max byte count. + Implemented in `src/materialize.rs` and `src/bast_validation.rs`. - **Full file**: [OQ-007](questions/007-bytes-materialization-lossy-utf8.md) ### OQ-008: `UnionValidator` variant dispatch — RESOLVED -- **Status**: resolved. `UnionValidator` now builds a sub-validator for - each variant at factory time and dispatches on `__discriminator` at - validation time. `AlkTypeEngine::compile` calls - `schema::inline_union_variant_refs` before `build_validator` to inline - `$ref`s in union `mapping` entries (necessary because the - `union_factory` receives the union node, but `$defs` live at the - schema root). Implemented in `src/validation.rs`, `src/schema.rs`, - and `src/engine.rs`. +- **Status**: resolved. Under the BAST pivot, the BAST-native validator + (`bast_validation::validate_union`) reads `__discriminator`, looks up + the variant `BastType` in the union's `mapping`, and recurses into the + variant's BAST definition via `validate_typeref` — enforcing every + field constraint the variant declares (e.g. `maxLength` on a `bytes` + field inside a variant struct). Variant `$ref`s resolve lazily via + `BastDoc::resolve_typeref` — no `inline_union_variant_refs` compile + step (removed under BAST). No custom keywords, no `jsonschema` + involvement on the bytes path. Implemented in `src/bast_validation.rs` + and `src/bast.rs`. See + [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md). - **Full file**: [OQ-008](questions/008-unionvalidator-variant-dispatch.md) ## Deferred / Blocked diff --git a/docs/architecture/overview.md b/docs/architecture/overview.md index 1f5e47c..a9a6fab 100644 --- a/docs/architecture/overview.md +++ b/docs/architecture/overview.md @@ -1,12 +1,12 @@ --- -status: draft -last_updated: 2026-08-11 +status: accepted +last_updated: 2026-08-15 --- # alktype — Overview -The binary struct engine: a small Rust crate that takes a JSON Schema -with `AlkType:*` custom keywords and produces an offset map, read/write +The binary struct engine: a small Rust crate that takes a BAST (Binary +Abstract Syntax Tree) document and produces an offset map, read/write functions, and validation — all driven by the schema. The schema is the format definition; the engine is generic. @@ -16,29 +16,43 @@ Component details are in the sibling documents. ## What -`alktype` is a library crate that consumes JSON Schemas annotated -with `AlkType:*` custom keywords (the same kinds defined in TypeBox's -`typedef.ts`, plus `AlkType:Bytes`, `AlkType:Int64`, and `AlkType:Uint64` -as alktype additions) and produces three capabilities: +`alktype` is a library crate that consumes BAST documents and produces +three capabilities: -1. **An offset map** — walks the schema, computes byte offsets for each - field based on type sizes, field order, and alignment. +1. **An offset map** — walks the BAST typed tree, computes byte offsets + for each field based on type sizes, field order, and alignment. 2. **Read/write functions** — given a `&[u8]` buffer and a field path, read the field's bytes at its offset (zero-copy for fixed-size types). Given a `&mut [u8]` buffer, write a value at its offset. -3. **Validation** — via `jsonschema` custom keywords, validates that a - buffer's bytes match the schema's type constraints. +3. **Validation** — two validators for two input types: + - `validate_bytes(&[u8])` uses the BAST-native validator (a recursive + walker over the BAST type tree) to check the value-domain + constraints the materializer doesn't (integer ranges, `maxLength`, + timestamp shape, enum index bounds, union variant constraints). + - `validate_json(&Value)` uses a standard `jsonschema::Validator` + compiled from a consumer-provided JSON Schema (BAST is not involved + — BAST describes bytes, not JSON shape). -The heavy lifting is done by the `jsonschema` crate (validation) and -`serde_json` (schema parsing). The novel code is the offset computation -— a recursive walk of the schema JSON that computes byte positions for -each field. The custom keyword implementations are small (a few lines -each, generated from shared macros — see [validation.md](validation.md)). +BAST is a JSON document that describes binary data layouts using a +`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is +itself a valid JSON Schema instance (it has a meta-schema), making it +self-validating, editor-friendly, and trivially consumable from any +language with a JSON parser. See [ADR-BAST](decisions/bast-bast-format.md) +and [`bast-format.md`](bast-format.md). + +The heavy lifting is done by the `jsonschema` crate (the +`validate_json` path and BAST document meta-schema validation) and +`serde_json` (BAST document parsing). The novel code is the offset +computation — a recursive walk of the BAST typed tree that computes +byte positions for each field — and the BAST-native validator — a flat +recursive match over the same tree. See +[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) +(purpose/scope) and [ADR-BAST](decisions/bast-bast-format.md) (format). The crate replaces two prior attempts that built their own jsonschema engines — typebox-rs (~8,400 lines) and the @alkdev/alktype prototype -(~5,600 lines) — with `jsonschema` + an offset map + small custom keyword -implementations. See +(~5,600 lines) — with a BAST parser + an offset map + a BAST-native +validator + the `jsonschema` crate for the JSON-validation path. See [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md). ## Why @@ -48,21 +62,20 @@ read or write binary data at computed offsets. Instead of per-protocol serde structs (russh-sftp's 29 packet types), per-handler wire format code (TTY's 5-byte format parser), or per-format offset computation (metatensor's tensor access), all of these become instances of the same -engine with different schemas. +engine with different BAST documents. The guiding insight: -> **The schema is the format.** A JSON Schema with `AlkType:Float32`, -> `AlkType:Struct`, `AlkType:Union` etc. is both the validation spec and -> the layout spec. No separate format definition, no separate parser, no -> separate validator. One schema, three uses: validate, compute offsets, -> access data. +> **The schema is the format.** A BAST document is both the layout spec +> and the validation spec for bytes. No separate format definition, no +> separate parser, no separate validator. One schema, three uses: +> validate, compute offsets, access data. This is the convergence of three threads identified in the call-channels-unification research: the `typedef.ts` schema kinds from TypeBox, the russh-sftp protocol packets, and the metatensor format. The -common pattern: a JSON Schema describes the shape of binary data, and -the binary data is the struct's bytes at computed offsets. +common pattern: a schema describes the shape of binary data, and the +binary data is the struct's bytes at computed offsets. The crate was bumped up in the timeline when the call-channels-unification research surfaced that channels, TTY, and the binary call protocol are @@ -75,41 +88,45 @@ read/write the binary payload." ## The "Schema Is the Format" Principle -A JSON Schema with `AlkType:*` custom keywords serves three roles -simultaneously: +A BAST document serves three roles simultaneously: | Role | Mechanism | When | |------|-----------|------| -| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) | +| **Validation spec (bytes)** | BAST-native validator (recursive walker over the BAST type tree) | Load time (parse typed tree), access time (`validate_bytes`) | +| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) | | **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) | | **Data access** | Read/write at computed offsets | Access time (read field, write field) | -No separate format definition, no separate parser, no separate validator. -The schema is the single source of truth for the binary format. Adding a -new field to a protocol is adding a property to the schema JSON — the -engine computes the new offsets automatically. +No separate format definition, no separate parser, no separate +validator. The BAST document is the single source of truth for the +binary format. Adding a new field to a protocol is adding an entry to +the BAST `fields` array — the engine computes the new offsets +automatically. This is the same principle as `#[repr(C)]` struct field access, but at -runtime from a portable JSON Schema instead of at compile-time from -language-specific annotations. The schema is the ABI contract. +runtime from a portable JSON document instead of at compile-time from +language-specific annotations. The BAST document is the ABI contract. ## Dependencies ``` alktype -├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support -├── serde_json (with preserve_order) — schema parsing; field order is load-bearing -└── (no tokio, no platform deps) — WASM-clean by construction +├── jsonschema (v0.46, Draft 2020-12, default-features=false) — validate_json path + BAST meta-schema validation +├── serde_json (with preserve_order) — BAST document parsing; mapping iteration order is load-bearing +└── (no tokio, no platform deps) — WASM-clean by construction ``` `alktype` is dependency-light: `jsonschema` + `serde_json` only. No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for -browser use. The `jsonschema` crate is already in the workspace at -`@alkdev/alknet: jsonschema/` — alktype is its first consumer. +browser use. The `validate_bytes` path does not touch `jsonschema` for +validation (it uses `jsonschema::ValidationError::custom` only for the +error payload type, D-BAST-009) — a small wasm binary-size win. -`serde_json` requires the `preserve_order` feature because field order -is load-bearing for binary layouts. The order of properties in the -schema JSON determines the order of fields in the binary struct. +`serde_json`'s `preserve_order` feature remains a dependency. Under +BAST, struct field order is explicit (the `fields` array), so layout +correctness no longer depends on it; but `mapping` iteration order and +the `Definitions` `$defs` block order are still load-bearing for the +parser's lazy resolution and the builder's output. ## Consumers @@ -133,14 +150,17 @@ schema roles (binary layout + JSON payloads) from one library. See The russh-sftp case is the most instructive and the highest-value POC target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written dispatch on a type byte followed by serde deserialization. Under alktype, -the dispatch is `TUnion` with a byte-offset discriminator — the schema -says "byte 0 is the discriminator, bytes 1..N are the variant struct." -The engine reads the discriminator, looks up the variant schema, computes -offsets, reads fields. Same result, no per-packet-type code. +the dispatch is a BAST `union` with a byte-offset discriminator — the +schema says "byte 0 is the discriminator, bytes 1..N are the variant +struct." The engine reads the discriminator, looks up the variant +schema, computes offsets, reads fields. Same result, no per-packet-type +code. ## Scope Boundaries (What This Is Not) -These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md). +These boundaries are decided in +[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) +and [ADR-BAST](decisions/bast-bast-format.md). - **Not metatensor.** alktype is the binary struct *engine*. Metatensor is a *format* (8-byte header + JSON header + binary data) that uses the @@ -150,60 +170,73 @@ These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-js should not do anything that explicitly blocks adding a Value system later. - **Not a code generator.** typebox-rs's `codegen/` module is a separate - concern. The alktype engine consumes schemas; it does not generate them. + concern. The alktype engine consumes BAST documents; it does not + generate them. - **Schema builder is in scope as of v0.1.0.** A fluent Rust API for - constructing schemas at runtime, producing `serde_json::Value`, is - shipped in v0.1.0 ([ADR-009](decisions/009-builder-api.md), resolves - OQ-003). The builder covers AlkType kinds and standard JSON Schema; - see [builder.md](builder.md). Schemas may still be authored in - TypeBox, generated by ujsx components, or hand-written — the builder - is an additional construction path, not a replacement. + constructing BAST documents and standard JSON Schemas at runtime, + producing `serde_json::Value`, is shipped in v0.1.0 + ([ADR-009](decisions/009-builder-api.md), resolves OQ-003). The + builder covers BAST kinds (`struct_()`) and standard JSON Schema + (`object()`); see [builder.md](builder.md). BAST documents may still + be authored in TypeBox, generated by ujsx components, or hand-written + — the builder is an additional construction path, not a replacement. - **Not a serialization framework.** The alktype engine is not a general-purpose serde replacement. It operates on raw byte buffers at - computed offsets — no intermediate `Value` tree, no reflection, no - dynamic dispatch per field. For JSON data, use serde. For binary data - with a known schema, use alktype. + computed offsets — no intermediate `Value` tree (except for the + `validate_bytes` materialization step), no reflection, no dynamic + dispatch per field. For JSON data, use serde. For binary data with a + known schema, use alktype. +- **Not a JSON-payload validator.** BAST describes bytes, not JSON + shape. `validate_json` validates a JSON `Value` against a + consumer-provided standard JSON Schema, not against the BAST document. + See [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md). ## Architecture (component pointers) -- **[schema-layer.md](schema-layer.md)** — the 19 `AlkType:*` kinds, - jsonschema custom keyword integration, TypeBox interop, schema - annotations (endianness, alignment, encoding, TUnion discriminators). +- **[schema-layer.md](schema-layer.md)** — the BAST parser (the typed + surface every engine module walks), the 19 BAST kinds, the + `AlkTypeKind` enum, and the foundational annotation types. +- **[`bast-format.md`](bast-format.md)** — the normative BAST format + specification (meta-schema, TypeRef, examples, validation model). - **[layout-engine.md](layout-engine.md)** — offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length field handling. - **[data-access.md](data-access.md)** — read/write functions, TUnion dispatch, field paths, zero-copy access for fixed-size types, length-prefix reading for variable-length types. -- **[validation.md](validation.md)** — custom keyword validators for all - 19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time - validation, `AlkTypeEngine` as the compiled form of a schema. - `validate_json` for JSON values; `validate_bytes` for binary buffers - (ADR-010). -- **[builder.md](builder.md)** — fluent Rust API for constructing - alktype JSON Schemas at runtime, producing `serde_json::Value`. - Covers AlkType kinds and standard JSON Schema (ADR-009). +- **[validation.md](validation.md)** — the two-validator model + (BAST-native for `validate_bytes`, standard `jsonschema` for + `validate_json`), `AlkTypeError`, load-time vs access-time validation, + `AlkTypeEngine` as the compiled form of a BAST document. +- **[builder.md](builder.md)** — fluent Rust API for constructing BAST + documents and standard JSON Schemas at runtime, producing + `serde_json::Value`. Covers BAST kinds and standard JSON Schema + (ADR-009, D-BAST-008). ## Design Decisions | Decision | ADR | Summary | |----------|-----|---------| -| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries | +| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries (format-specific content superseded by ADR-BAST) | +| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | BAST as the schema format; meta-schema, `$defs`/`$ref`/`kind` vocabulary; supersedes ADR-001's format-specific content | | Two layout modes | [ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats | -| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) | -| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping | +| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) — semantics carry forward; location moved to BAST type-level properties | +| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping (validation-strategy section refined by ADR-VAL-SPLIT) | +| Two-validator model | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | BAST-native validator for `validate_bytes`; standard `jsonschema::Validator` for `validate_json`; D-BAST-006/007/009 | | Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (SFTP offsets, metatensor data_offsets) | | Non-final inline variable fields | [ADR-006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields) | | Packed-mode read factory | [ADR-007](decisions/007-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader | | TUnion in aligned mode | [ADR-008](decisions/008-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) | -| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 | -| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate | +| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 (output format amended to BAST / standard JSON Schema by ADR-BAST) | +| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate (validation step amended to the BAST-native validator by ADR-VAL-SPLIT) | ## Open Questions See [open-questions.md](open-questions.md) for full details. -- **OQ-001** (deferred(scope)): Arrays of variable-length-element structs. +- **OQ-001** (deferred(scope)): Arrays of variable-length-element + structs — BAST arrays require `count` in v1 (D-BAST-004), aligning + with this deferral. - **OQ-002** (deferred(scope)): `no_std` + `alloc` support. - **OQ-003** (resolved by [ADR-009](decisions/009-builder-api.md)): Builder API for schema construction. Shipped in v0.1.0; see @@ -223,8 +256,12 @@ See [open-questions.md](open-questions.md) for full details. - `@alkdev/alknet: alknet-typedef-poc/` — the POC code (disposable) - `@alkdev/alknet: typebox-rs/` — prior attempt, replaced by alktype - `@alkdev/alknet: alktype-prototype/` — prior attempt (the @alkdev/alktype prototype; not to be confused with this crate, which reuses the name but is backed by the `jsonschema` crate) +- [BAST pivot research record](../research/bast-pivot.md) — motivation, + POC scope and result, decisions D-BAST-001..009, risks +- [BAST pivot implementation plan](../plans/bast-implementation.md) — + ordered implementation steps, semver contract, ADR-sync checklist > **Note**: The research findings, POC code, and prior-attempt paths above > refer to the parent `@alkdev/alknet` workspace where this crate originated. > They are preserved here as historical context for the architectural -> decisions; the artifacts themselves are not part of this standalone repo. +> decisions; the artifacts themselves are not part of this standalone repo. \ No newline at end of file diff --git a/docs/architecture/questions/008-unionvalidator-variant-dispatch.md b/docs/architecture/questions/008-unionvalidator-variant-dispatch.md index b27dfdc..de9f8a0 100644 --- a/docs/architecture/questions/008-unionvalidator-variant-dispatch.md +++ b/docs/architecture/questions/008-unionvalidator-variant-dispatch.md @@ -1,5 +1,19 @@ # OQ-008: `UnionValidator` variant dispatch — validate variant fields against variant schema +> **Note (post-BAST-pivot):** The v0.1.0 resolution below — +> `UnionValidator` + `inline_union_variant_refs` + custom-keyword +> `jsonschema` sub-validators — was superseded by the BAST pivot. The +> current implementation is the BAST-native validator +> (`bast_validation::validate_union`), which reads `__discriminator`, +> looks up the variant `BastType` in the union's `mapping`, and +> recurses via `validate_typeref` into the variant's BAST definition. +> Variant `$ref`s resolve lazily via `BastDoc::resolve_typeref` — no +> `inline_union_variant_refs` compile step (removed under BAST). No +> custom keywords, no `jsonschema` involvement on the bytes path. See +> [ADR-VAL-SPLIT](../decisions/val-split-two-validator-model.md). The +> v0.1.0 resolution text is preserved below as the historical record +> of how the question was originally resolved. + - **Origin**: Raised during the v0.1.0 POC round 2 (SFTP Packet `validate_bytes` POC). Surfaced when the over-`maxLength` `Bytes` test failed: the materializer read the bytes correctly, but the diff --git a/docs/architecture/schema-layer.md b/docs/architecture/schema-layer.md index f79fc1b..2042149 100644 --- a/docs/architecture/schema-layer.md +++ b/docs/architecture/schema-layer.md @@ -1,58 +1,61 @@ --- -status: draft -last_updated: 2026-07-22 +status: accepted +last_updated: 2026-08-15 --- # alktype — Schema Layer -The schema layer: the 19 `AlkType:*` custom type kinds, their mapping to -Rust types and byte sizes, the `jsonschema` custom keyword integration, -TypeBox interop, and the concrete JSON shapes for schema-level -annotations. +The schema layer: the BAST (Binary Abstract Syntax Tree) format and the +typed parser that the layout engines, materializer, and BAST-native +validator walk. BAST replaces the v0.1.0 `AlkType:*` custom-keyword JSON +Schema format decided in ADR-001; the pivot is recorded in +[ADR-BAST](decisions/bast-bast-format.md) and grounded in +[D-BAST-001..009](../research/bast-pivot.md#decisions). -## The 19 AlkType Kinds +The **normative format specification** is +[`bast-format.md`](bast-format.md) (meta-schema, TypeRef, examples, +validation model). This document describes the *implementation* — the +typed parser in `src/bast.rs` and the foundational `AlkTypeKind` enum +in `src/schema.rs` — and points at the format spec for shape details. -These are the custom schema kinds defined in TypeBox's `typedef.ts` -(`@alkdev/alknet: typebox/example/typedef/typedef.ts`, 619 lines) and -ported to Rust via `jsonschema` custom keywords. Each kind carries binary -layout semantics — a known byte size (for fixed-size types) or a known -encoding strategy (for variable-length types). +## The 19 BAST Kinds -| Kind | TypeBox key | Rust type | Size | Category | -|------|-------------|-----------|------|----------| -| `TFloat32` | `AlkType:Float32` | `f32` | 4 | fixed | -| `TFloat64` | `AlkType:Float64` | `f64` | 8 | fixed | -| `TInt8` | `AlkType:Int8` | `i8` | 1 | fixed | -| `TInt16` | `AlkType:Int16` | `i16` | 2 | fixed | -| `TInt32` | `AlkType:Int32` | `i32` | 4 | fixed | -| `TInt64` | `AlkType:Int64` | `i64` | 8 | fixed | -| `TUint8` | `AlkType:Uint8` | `u8` | 1 | fixed | -| `TUint16` | `AlkType:Uint16` | `u16` | 2 | fixed | -| `TUint32` | `AlkType:Uint32` | `u32` | 4 | fixed | -| `TUint64` | `AlkType:Uint64` | `u64` | 8 | fixed | -| `TBoolean` | `AlkType:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed | -| `TString` | `AlkType:String` | length-prefixed UTF-8 | variable | variable | -| `TBytes` | `AlkType:Bytes` | length-prefixed raw bytes | variable | variable | -| `TStruct` | `AlkType:Struct` | record of fields | sum of field sizes | composite | -| `TUnion` | `AlkType:Union` | tagged union | discriminator + variant | composite | -| `TArray` | `AlkType:Array` | repeated element | count × element size | composite | -| `TEnum` | `AlkType:Enum` | u32 index into enum values | 4 (fixed) | fixed | -| `TRecord` | `AlkType:Record` | count-prefixed sequence of (key, value) pairs | variable | variable | -| `TTimestamp` | `AlkType:Timestamp` | length-prefixed RFC 3339 string | variable | variable | +BAST uses lowercase `kind` strings (`"uint32"`, `"struct"`, `"union"`, +etc.). The engine represents them as the `AlkTypeKind` Rust enum — one +variant per kind — providing compile-time exhaustiveness checking and +integer-discriminant dispatch (a jump table) instead of string +comparison at every field access. -`AlkType:Int64` and `AlkType:Uint64` are alktype additions — -TypeBox's `typedef.ts` tops out at 32-bit integers. They are required by -the primary POC targets: SFTP `Read`/`Write` packets have `offset: u64`, -and metatensor `data_offsets` are `u64`. See +| BAST kind | `AlkTypeKind` | Rust type | Size | Category | +|-----------|---------------|-----------|------|----------| +| `int8` | `Int8` | `i8` | 1 | fixed | +| `int16` | `Int16` | `i16` | 2 | fixed | +| `int32` | `Int32` | `i32` | 4 | fixed | +| `int64` | `Int64` | `i64` | 8 | fixed | +| `uint8` | `Uint8` | `u8` | 1 | fixed | +| `uint16` | `Uint16` | `u16` | 2 | fixed | +| `uint32` | `Uint32` | `u32` | 4 | fixed | +| `uint64` | `Uint64` | `u64` | 8 | fixed | +| `float32` | `Float32` | `f32` | 4 | fixed | +| `float64` | `Float64` | `f64` | 8 | fixed | +| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed | +| `string` | `String` | length-prefixed UTF-8 | variable | variable | +| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable | +| `timestamp` | `Timestamp` | length-prefixed RFC 3339 string | variable | variable | +| `struct` | `Struct` | record of fields | sum of field sizes | composite | +| `union` | `Union` | tagged union | discriminator + variant | composite | +| `array` | `Array` | repeated element | count × element size | composite | +| `enum` | `Enum` | u32 index into enum values | 4 (fixed) | fixed | +| `record` | `Record` | count-prefixed (key, value) pairs | variable | variable | + +`int64`/`uint64` are alktype additions — TypeBox's `typedef.ts` tops +out at 32-bit integers. Required by SFTP `Read`/`Write` `offset: u64` +and metatensor `data_offsets`. See [ADR-005](decisions/005-int64-uint64-first-class-kinds.md). ### The `AlkTypeKind` enum -The engine represents the 19 kinds as a Rust enum — `AlkTypeKind` — with -one variant per kind (`AlkTypeKind::Float32`, `AlkTypeKind::Struct`, etc.). -The enum provides compile-time exhaustiveness checking and integer -discriminant dispatch (a jump table) instead of string comparison at -every field access. It is `pub` and re-exported from the crate root. +`src/schema.rs` defines the enum: ```rust pub enum AlkTypeKind { @@ -69,7 +72,8 @@ The enum carries the kind's binary-layout metadata as inherent methods: | Method | Returns | Notes | |--------|---------|-------| -| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"AlkType:Uint8"` | +| `to_bast_str(self)` | `&'static str` | The lowercase BAST kind string (`"uint32"`) | +| `from_bast_str(s)` | `Result` | Parses a lowercase BAST kind string; `AlkTypeError::Schema` for unknowns | | `type_size(self)` | `Option` | `Some(N)` for fixed-size kinds; `None` for variable/composite | | `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array | | `is_fixed_size(self)` | `bool` | True for the 12 fixed-size primitive kinds | @@ -77,431 +81,214 @@ The enum carries the kind's binary-layout metadata as inherent methods: | `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record | | `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter | -`AlkTypeKind` implements `Display` (renders the keyword string) and -`FromStr` (parses the keyword string back into the variant, returning -`AlkTypeError::Schema` for unknown kinds). The layout engines and the -validator dispatch on the enum, not on strings. +`AlkTypeKind` implements `Display`, backed by `to_bast_str` so the +layout engines, materializer, validator, and parser surface the +canonical BAST name in error messages. `from_bast_str` is the inverse +and is the dispatch point the BAST parser uses to map a `kind` string +to the enum variant (D-BAST-002). -### Fixed-size types +### Foundational annotation types -`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`, -`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset -computation uses these sizes directly. Read/write is zero-copy pointer -cast for these types. - -**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other -values are invalid and produce a `AlkTypeError::Access` on read. - -**`TEnum` binary representation:** A `u32` index into the enum's declared -values, in declaration order. The first declared value is index 0, the -second is index 1, etc. The enum's values are declared via the standard -JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`). -The `AlkType:Enum` custom keyword signals that the type is an enum for -layout purposes; the built-in `enum` keyword provides the value list. - -**Design note:** TypeBox's `TEnum` is a string enum (variable-length). The -alktype engine uses a `u32` index instead — a deliberate deviation from -TypeBox fidelity in favor of binary efficiency. Most enums have a small -number of variants (e.g., the call protocol's 5 event types); a `u32` -index is compact, fixed-size, and sufficient for any realistic enum. The -JSON representation (for validation) remains a string; the binary -representation is the `u32` index. -The `u32` index follows the schema's endianness annotation (ADR-003), like -all other fixed-size types. In little-endian mode the index is -`u32::from_le_bytes`; in big-endian mode it is `u32::from_be_bytes`. - -### Variable-length types - -`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes. -The alktype engine supports three strategies for handling variable-length -types in binary layouts, selected by the `encoding` annotation and the -standard JSON Schema `maxLength` keyword: - -| Strategy | Encoding annotation | Layout behavior | Use case | -|----------|-------------------|-----------------|----------| -| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) | -| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) | -| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) | - -**Strategy 1: Inline length-prefixing (default).** The field's fixed -portion is a 4-byte length prefix at a computed offset. The variable data -follows immediately after. In packed sequential mode, the length prefix -determines the position of subsequent fields. In aligned static mode, the -length prefix is at a known offset; the variable data is not included in -the static layout. This is the universal pattern used by channels, SFTP, -TTY, and most binary protocols. - -**Strategy 2: Fixed-size reservation.** When a variable-length field -declares `maxLength` (a standard JSON Schema keyword), the engine reserves -`maxLength` bytes at a fixed offset in aligned static mode. Data shorter -than `maxLength` is zero-padded; data longer than `maxLength` is a -validation error. This makes the field fixed-size from the layout -perspective — subsequent fields have known, unchanging offsets. This is -the database `VARCHAR(N)` pattern and the metatensor struct-tensor -pattern for fields with known maximum sizes. - -In packed sequential mode, `maxLength` is a validation constraint only — -the engine still uses inline length-prefixing (strategy 1) because -protocols don't benefit from fixed-size reservation. - -**Strategy 3: Offset indirection.** The field is a struct -`{offset: u32, length: u32}` at a known position. The consumer provides -the data region separately; the engine reads the offset and length, then -slices the data region. This is the metatensor blob tensor pattern — the -index struct lives in one region, the blob data lives in another. Enables -mmap-friendly random access to variable-length data without parsing -length prefixes and without reserving worst-case space. - -**Default strategy selection:** -- In packed sequential mode: always strategy 1 (inline length-prefixing). - `maxLength` is a validation constraint only. -- In aligned static mode: strategy 2 (fixed-size reservation) if - `maxLength` is declared; strategy 3 (offset indirection) if - `"encoding": "offset-indirect"` is declared; strategy 1 (inline - length-prefixing) otherwise. - -**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3) -respects the schema's `"endian"` annotation (ADR-003). In little-endian -mode, the length is `u32::from_le_bytes`. In big-endian mode, the length -is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have -consistent byte order for both field values and length prefixes. - -**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`. -Otherwise identical to `TString` in layout (same three strategies). - -**Design note:** `AlkType:Bytes` is an alktype addition — it does -not exist in TypeBox's `typedef.ts` (which defines 16 kinds). It is -included because raw byte arrays are a common binary protocol primitive -(SFTP data payloads, channels payloads, tensor data) and are semantically -distinct from UTF-8 strings. In the binary representation, TBytes is raw -bytes with no encoding (not base64, not hex). In the JSON representation -(for validation), TBytes is a string (JSON has no native byte type). - -**`TRecord`:** A string-keyed map. The value type is declared via the -schema's `"values"` property (e.g., `"values": { "AlkType:Float32": true }`). -Binary layout is a count-prefixed sequence of `(key, value)` pairs: -`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times. -The count is the number of entries. Each key is a length-prefixed UTF-8 -string. Each value is encoded according to its declared `AlkType:*` kind -— a `Record` value is 4 raw bytes; a `Record` value is -itself a length-prefixed string; a `Record` value is the struct's -fields laid out inline. There is **no separate `value_len` prefix** — -the value's size is determined by its kind (fixed-size kinds have a -known size; variable-length kinds carry their own length prefix). The -count and key-length prefixes respect the schema's endianness. In -aligned static mode with `maxLength`, the entire record is reserved at -`maxLength` bytes (zero-padded). - -**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of -ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or -fixed-size reservation (strategy 2 with `maxLength`). The data-access -layer treats timestamps as opaque length-prefixed strings — it does not -parse or validate the timestamp format. The jsonschema custom keyword -validator checks RFC 3339 conformance at the JSON level (see -[validation.md](validation.md)). - -`TArray` is variable-length when the element type is variable-length or -when the count is not known at schema time. For fixed-size element arrays -with a known count, the size is `element_size × count`. - -**`TArray` count declaration:** The array count is declared via the -standard JSON Schema `"minItems"` and `"maxItems"` keywords. When -`minItems == maxItems`, the array has a fixed count known at schema time. -When they differ or are absent, the count is variable and the array uses -a length-prefixed encoding: `[count: u32][element_0]...[element_N]`. -The count prefix respects the schema's endianness. - -### Composite types - -`TStruct` and `TUnion` are composite — their size is the sum of their -fields' sizes (plus alignment padding in aligned static mode). The offset -computation recurses into their properties. - -## Schema-Layer Public API - -The `schema` module exposes the foundational types and functions every -other module depends on. These are re-exported from the crate root. - -### `get_alktype_kind` vs `get_alktype_kind_loose` - -The engine recognizes a `AlkType:*` kind on a schema node two ways, -because the keyword value may be either a boolean (`true`) or an -annotation object (`{ "encoding": "..." }`): - -| Function | Recognizes | Returns | -|----------|------------|---------| -| `get_alktype_kind(node) -> Option<&str>` | Boolean form only (`{ "AlkType:String": true }`) | The keyword string, e.g. `"AlkType:String"` | -| `get_alktype_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string | -| `get_alktype_kind_enum(node) -> Option` | Boolean form only | The parsed enum variant | -| `get_alktype_kind_loose_enum(node) -> Option` | Boolean form **and** object form | The parsed enum variant | - -The boolean-form-only functions are used by the validator factories -(which reject the object form as a schema error) and the top-level -kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new` -(which require `AlkType:Struct` at the root). The "loose" variants are -used by the layout engines during field traversal, so that a variable- -length field with an `encoding` annotation (`{ "AlkType:String": -{ "encoding": "offset-indirect" } }`) is still recognized as a `String`. - -### Annotation parsers - -Each schema-level annotation has a dedicated parser that reads it from a -`serde_json::Value` node and returns a sensible default when absent: - -| Function | Annotation | Default | -|----------|------------|---------| -| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` | -| `parse_align(node) -> Option` | `"align"` | `None` | -| `parse_max_length(node) -> Option` | `"maxLength"` | `None` | -| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` | -| `parse_discriminator(node) -> Result` | `"discriminator"` | (required — returns `AlkTypeError::Schema` if absent) | - -### Public enums +`src/schema.rs` also defines the two annotation enums (semantics +unchanged from ADR-003; only their *location* in the document moved — +see [ADR-BAST](decisions/bast-bast-format.md) and +[bast-format.md §Variable-Length Encoding](bast-format.md#variable-length-encoding)): ```rust pub enum Endian { Little, Big } pub enum VariableEncoding { LengthPrefixed, OffsetIndirect } -pub enum DiscriminatorKind { - Byte { offset: usize, disc_type: AlkTypeKind }, - Field { name: String }, -} ``` -`DiscriminatorKind::Byte` carries the byte position (`offset`) and the -discriminator's `AlkType:*` kind (`disc_type`, restricted to `Uint8`/ -`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator -field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for -how these drive dispatch. +The `Discriminator` builder enum lives in +[`src/builder.rs`](../../src/builder.rs) (the builder's domain); the BAST +parser's typed discriminator view is +[`BastDiscriminator`](#the-bast-parser-bast-module). -### `$ref` resolution and normalization +## The BAST Parser (`bast` module) -| Function | Purpose | -|----------|---------| -| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/"`. Idempotent. Runs once at `AlkTypeEngine::compile` time. | -| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. | -| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). | +`src/bast.rs` is the typed surface over a BAST document. Three +consumers walk the same tree — the layout engines +([`offset_map`](layout-engine.md), [`layout_builder`](layout-engine.md), +[`sequential_reader`](layout-engine.md)), the +[`materialize`](data-access.md) layer, and the +[`bast_validation`](validation.md) validator — so a typed view pays for +itself: each walks matched arms over `BastType` instead of re-parsing +raw JSON at every node. Borrowing (not cloning) the source +`serde_json::Value` keeps the parse allocation-free beyond the small +typed nodes themselves. -`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s -JSON Pointer requirement. The layout engines call `resolve_ref_or_inline` -on every `$ref`-bearing node they encounter during traversal. +### Document shape -## jsonschema Custom Keyword Integration +Every BAST document has the same top-level shape: -The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords -via the `with_keyword` API. Each `AlkType:*` kind is registered as a -custom keyword: - -```rust -let validator = jsonschema::options() - .with_keyword("AlkType:Float32", factory) - .with_keyword("AlkType:Int32", factory) - .with_keyword("AlkType:Struct", factory) - // ... all 19 kinds - .build(&schema)?; +```json +{ "$defs": { "": { ...TypeDef... }, ... } } ``` -The factory closure receives the parent schema object, the keyword's -value, and the schema path — enabling cross-keyword awareness. The -`AlkType:Struct` validator, for example, inspects the parent's -`properties` to validate each field against its declared `AlkType:*` kind. +- The `$defs` block is **required** (D-BAST-003). Single-type documents + are a special case with one entry. +- The **root type name** is a required parameter to + `AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001). + It selects which `$defs` entry is the top-level type; convention + (first entry) is fragile and depends on JSON key order, so an + explicit parameter is used instead. -Each custom keyword implementation is ~10 lines. The `jsonschema` crate -handles all structural validation (object properties, required fields, -array items, enum values) — the custom keywords only need to validate -the leaf type constraints. See [validation.md](validation.md) for the -validator implementations. +See [`bast-format.md`](bast-format.md) for the normative TypeDef shapes +(Struct, Union, Enum, FieldDef, TypeRef) and the meta-schema. -This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side. -Same semantics, different language, same JSON Schema wire format. A -TypeBox schema serialized to JSON feeds into the alktype engine after a -single pre-processing step: normalizing `$ref` values (see below). +### Typed tree -## TypeBox Interop +The parser produces a borrowed typed tree: -TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox -schema like: +| Type | Role | +|------|------| +| `BastDoc<'a>` | The parsed document: the root `Value`, the chosen root name, and the parsed root `BastDef`. Entry point via `BastDoc::new(root, root_name)`. | +| `BastDef<'a>` | A named `$defs` entry — `{ name, kind: BastDefKind, source }`. Only `struct`/`union`/`enum` can live at the top level. | +| `BastDefKind<'a>` | `Struct(BastStruct)` / `Union(BastUnion)` / `Enum(BastEnum)`. | +| `BastStruct<'a>` | `{ endian, align, fields: Vec, source }`. Field order is the `fields` array order (BAST design principle #4 — no reliance on `serde_json`'s `preserve_order`). | +| `BastField<'a>` | `{ name, ty: BastType, endian, align, encoding, max_length, source }`. Annotations are field-level properties (ADR-003 semantics, BAST location). | +| `BastUnion<'a>` | `{ endian, discriminator, fields, mapping: Vec<(key, BastType)>, source }`. Variant `$ref`s resolve **lazily** — no compile-time inlining. | +| `BastDiscriminator<'a>` | `Byte { offset, disc_type }` / `Field { name }`. The typed view of the `discriminator` object. | +| `BastEnum<'a>` | `{ values: Vec<&'a str>, source }`. Non-empty (enforced). | +| `BastType<'a>` | A TypeRef — `Primitive(AlkTypeKind)` / `Ref(BastRef)` / `Array(BastArray)` / `Record(BastRecord)` / `Struct(...)` / `Union(...)` / `Enum(...)`. The central mechanism for typing fields, array elements, record values, and union variants. | +| `BastRef<'a>` | A `$ref` restricted to `#/$defs/`. Carries just the name. | +| `BastArray<'a>` | `{ element: Box, count, source }`. `count` is required in v1 (D-BAST-004). | +| `BastRecord<'a>` | `{ values: Box, source }`. | -```typescript -const TensorRef = Type.Object({ - dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]), - shape: Type.Array(Type.Number()), - data_offsets: Type.Tuple([Type.Number(), Type.Number()]) -}); -``` +All of these are re-exported from the crate root (`pub use bast::{...}` +in `src/lib.rs`). -serialized to JSON is a standard JSON Schema with `type: "object"`, -`properties`, and `required`. That JSON feeds into the alktype engine -after `$ref` normalization. The `AlkType:*` custom keywords are added by -TypeBox's `TypeRegistry.Set` — they appear in the serialized JSON as -additional properties on the schema object. +### `$ref` resolution -### `$ref` normalization +BAST `$ref`s are always full JSON Pointers restricted to +`#/$defs/` — no external references, no fragment-only pointers, +no bare names (rejected by the parser). The restriction keeps +resolution a single hash lookup and eliminates the v0.1.0 +`normalize_refs` pass that rewrote TypeBox's bare-name refs. -TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`), -referencing sibling definitions within the same `$defs` block. The -`jsonschema` crate requires full JSON Pointer paths (e.g., -`"$ref": "#/$defs/Read"`). The alktype engine normalizes TypeBox-style -refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization) -— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to -`"#/$defs/"`. The normalization is idempotent — full JSON Pointer -refs pass through unchanged. It runs once at `AlkTypeEngine::compile` -time, before the schema is passed to `jsonschema` or the offset -computation. +`BastDoc` exposes three resolution helpers: -**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs -with `Resource 'Read' is not present in a registry`. Full JSON Pointer -refs (`#/$defs/Read`) resolve correctly. The normalization step bridges -the gap between TypeBox's output and jsonschema's input. +| Method | Purpose | +|--------|---------| +| `lookup_def(name) -> Result<&'a Value, AlkTypeError>` | The single hash lookup into `$defs`. | +| `resolve_ref(r: &BastRef) -> Result` | Resolve a `BastRef` to its `BastDef`. | +| `resolve_typeref(ty: &BastType) -> Result` | Deref one `$ref` level, or return the inline type unchanged. The composite-walkers call this. | +| `resolve_typeref_as_def(ty, path) -> Result` | Resolve a `BastType` to a `BastDef`, wrapping inline composites in a synthetic def. Convenient for the validator/materializer. | -The alktype engine does not depend on TypeBox or any JS toolchain. It -consumes JSON — whether that JSON was authored in TypeBox, generated by -a ujsx component, or hand-written. The schema is the interface. +Variant `$ref`s (union `mapping` entries) are resolved **lazily** by +the materializer and validator via these helpers — no +`inline_union_variant_refs` compile step (removed under BAST). The +parser only records the `BastRef` target name. + +### Untrusted input + +Every path that walks a BAST document returns +`Err(AlkTypeError::Schema)` on a malformed document, never +`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream +`alkcall` consumer accepts schemas from arbitrary internet peers in its +hub/spoke topology). Overflow-safe arithmetic (`checked_add`, +`usize::try_from`) is used for any offset/count cast (AGENTS.md §4). + +A malformed document (missing `$defs`, missing `kind`, unknown kind +string, dangling `$ref`, empty `mapping`, non-struct/union/enum at the +top level, a field-name union without a `fields` array, etc.) surfaces +as `AlkTypeError::Schema` with a path-annotated message. + +### What the parser does *not* do + +- **No meta-schema validation.** `BastDoc::new` parses structurally + (every reachable def parses to a typed `BastDef`) but does not run + the BAST meta-schema. Consumers that want full structural validation + can run the meta-schema via `jsonschema` directly + ([`BAST_META_SCHEMA`](bast-format.md#the-meta-schema) is re-exported + from the crate root). The parser's structural checks catch the cases + that matter for layout/materialize/validate; the meta-schema is the + authoritative well-formedness check. +- **No eager full-document parse.** Only the root definition and the + definitions it (transitively) references are parsed eagerly; orphan + `$defs` entries are not checked. Lazy `$ref` resolution reaches the + rest at access time. +- **No annotation interpretation.** The parser *records* `endian`/ + `align`/`encoding`/`maxLength` on `BastField`/`BastStruct`; the + layout engines and validator *interpret* them (ADR-003 semantics). ## Schema Annotations -Schema-level annotations control binary layout behavior. These are -decided in [ADR-003](decisions/003-schema-annotations.md). +Annotation *semantics* carry forward unchanged from ADR-003; only their +*location* moved from v0.1.0's custom-keyword objects to BAST +type-level properties. The concrete BAST shapes are in +[`bast-format.md`](bast-format.md): -### Endianness +- [Endianness](bast-format.md#endianness) — struct/union-level `endian` + with field-level override. +- [Alignment](bast-format.md#alignment) — struct/field-level `align` + (aligned mode only). +- [Variable-length encoding](bast-format.md#variable-length-encoding) — + field-level `encoding` and `maxLength`. +- [Union discriminators](bast-format.md#union) — `discriminator` object + on the union def (`byte` or `field`). -Schema-level annotation with a default of little-endian: +The `maxLength` keyword is *not* a BAST invention — it is the standard +JSON Schema `maxLength`, repurposed as a byte-length cap. In aligned +mode it reserves a fixed-size slot; in packed mode it is a validation +constraint only. See [bast-format.md §Variable-Length +Encoding](bast-format.md#variable-length-encoding) and +[ADR-003](decisions/003-schema-annotations.md). -```json -{ "AlkType:Struct": true, "endian": "big", "properties": { ... } } -``` +## What Was Removed -- `"endian": "little"` (default) — read/write in little-endian byte order. -- `"endian": "big"` — read/write in big-endian byte order. -- Applies to the entire schema and all nested types. +The v0.1.0 custom-keyword accessor layer was removed in step 8 of the +BAST pivot. The `schema` module retains only the foundational types +(`AlkTypeKind`, `Endian`, `VariableEncoding`, shared constants); the +BAST parser is the typed surface every engine module walks. Removed: -### Alignment +- `get_alktype_kind` / `get_alktype_kind_enum` / `get_alktype_kind_loose` + / `get_alktype_kind_loose_enum` — superseded by the parser's + `kind`-string dispatch. +- `normalize_refs` / `inline_union_variant_refs` (+ recursive helpers) + — BAST refs are always `#/$defs/`; one hash lookup, variant + refs resolve lazily. +- `parse_encoding` / `parse_align` / `parse_max_length` / `parse_endian` + — `bast.rs` has its own BAST-property-form copies (internal to the + parser). +- `parse_discriminator` + `DiscriminatorKind` — replaced by + `bast::BastDiscriminator`; the builder has its own `Discriminator` + enum. +- `resolve_ref` / `resolve_ref_or_inline` — replaced by + `BastDoc::lookup_def` / `resolve_typeref`. +- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`, + `BYTE_DISCRIMINATOR_TYPES`, and their unit tests. -Both struct-level and field-level, with field-level overriding: - -```json -{ - "AlkType:Struct": true, - "align": 256, - "properties": { - "weight": { "AlkType:Float32": true, "align": 16 } - } -} -``` - -- Struct-level `"align"` sets the default for all fields. -- Field-level `"align"` overrides the struct default. -- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/ - enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix), - 1 for struct/union/array. -- Only meaningful in aligned static mode (ADR-002). Ignored in packed - sequential mode. - -### Variable-length encoding - -The alktype engine supports three strategies for variable-length types -(see §Variable-length types above for full details). The strategy is -selected by the `encoding` annotation and the standard JSON Schema -`maxLength` keyword: - -```json -// Strategy 1: Inline length-prefixing (default, shorthand) -{ "AlkType:String": true } - -// Strategy 1: Explicit inline length-prefixing -{ "AlkType:String": { "encoding": "length-prefixed" } } - -// Strategy 2: Fixed-size reservation (uses standard maxLength) -{ "AlkType:String": true, "maxLength": 256 } - -// Strategy 3: Offset indirection (opt-in) -{ "AlkType:String": { "encoding": "offset-indirect" } } -``` - -- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at - computed offset, variable data follows immediately. Used by protocol - wire formats. -- `maxLength` (standard JSON Schema keyword) — in aligned static mode, - reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the - field fixed-size from the layout perspective. In packed sequential - mode, `maxLength` is a validation constraint only. -- `"encoding": "offset-indirect"` — field is a struct - `{offset: u32, length: u32}` pointing into a separate data region. - The consumer provides the data region separately. Used by metatensor - blob tensors. -- Applies to all variable-length types: `AlkType:String`, `AlkType:Bytes`, - `AlkType:Array`, `AlkType:Record`, `AlkType:Timestamp`. - -### TUnion discriminators - -Two discriminator kinds: byte-offset (protocol dispatch) and field-name -(typedef.ts pattern). - -**Byte-offset discriminator** (SFTP type bytes, call protocol event types): - -```json -{ - "AlkType:Union": true, - "discriminator": { - "kind": "byte", - "offset": 0, - "type": "AlkType:Uint8" - }, - "mapping": { - "5": { "$ref": "#/$defs/Read" }, - "6": { "$ref": "#/$defs/Write" }, - "101": { "$ref": "#/$defs/Status" } - } -} -``` - -- `"offset"` — byte position of the discriminator. -- `"type"` — the `AlkType:*` kind of the discriminator (typically - `AlkType:Uint8`). -- Mapping keys are stringified integers. The variant struct starts at - `offset + discriminator_size`. - -**Field-name discriminator** (typedef.ts pattern): - -```json -{ - "AlkType:Union": true, - "discriminator": { - "kind": "field", - "name": "type" - }, - "mapping": { - "read": { "$ref": "#/$defs/Read" }, - "write": { "$ref": "#/$defs/Write" } - } -} -``` - -- `"name"` — the field name holding the discriminator value. -- Mapping keys are string values matching the discriminator field's value. -- The discriminator field is just another field in the struct. - -Mapping values may be either inline schemas or `$ref` pointers. Both work. +See [ADR-BAST](decisions/bast-bast-format.md) §"What is removed" and +[`bast-format.md` §What is removed](bast-format.md#what-is-removed). ## Design Decisions | Decision | ADR | Summary | |----------|-----|---------| -| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators | +| BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary | [ADR-BAST](decisions/bast-bast-format.md) | Supersedes ADR-001's format-specific content; records D-BAST-001..009 | +| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Annotation semantics (carry forward unchanged; only location moves) | | Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (required by SFTP offsets and metatensor data_offsets) | -| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle | +| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle (format-specific content superseded by ADR-BAST) | ## Open Questions See [open-questions.md](open-questions.md) for full details. -- **OQ-003** (deferred(scope)): Builder API for schema construction. +- **OQ-001** (deferred(scope)): Arrays of variable-length-element + structs — BAST arrays require `count` in v1 (D-BAST-004), aligning + with this deferral. ## References -- `@alkdev/alknet: typebox/example/typedef/typedef.ts` — the TypeBox - schema kinds (619 lines) -- `@alkdev/alknet: jsonschema/` — the jsonschema crate (v0.46.5, Draft - 2020-12) -- [ADR-003](decisions/003-schema-annotations.md) — schema - annotation shapes -- [validation.md](validation.md) — custom keyword validator implementations +- [`bast-format.md`](bast-format.md) — the normative BAST format + specification (meta-schema, TypeRef, examples, validation model) +- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format decision +- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics +- [BAST pivot research record](../research/bast-pivot.md) — motivation, + POC scope and result, decisions D-BAST-001..009 +- [validation.md](validation.md) — the BAST-native validator and the + `validate_json` JSON-Schema path +- [`src/bast.rs`](../../src/bast.rs) — the parser implementation +- [`src/schema.rs`](../../src/schema.rs) — the `AlkTypeKind` enum and + foundational annotation types \ No newline at end of file diff --git a/docs/architecture/validation.md b/docs/architecture/validation.md index d1a35d0..0b1a023 100644 --- a/docs/architecture/validation.md +++ b/docs/architecture/validation.md @@ -1,63 +1,147 @@ --- -status: draft -last_updated: 2026-08-11 +status: accepted +last_updated: 2026-08-15 --- # alktype — Validation -The validation layer: custom keyword validators for all 19 `AlkType:*` -kinds, the `AlkTypeError` enum, load-time vs access-time validation -strategy, and the `AlkTypeEngine` as the compiled form of a schema. +The validation layer: two validators for two input types, the +`AlkTypeError` enum, the load-time-build / access-time-check strategy, +and the `AlkTypeEngine` as the compiled form of a BAST document. + +## The Validator Split + +BAST separates two concerns that the v0.1.0 format conflated, and in +doing so reveals that the engine has **two distinct validation paths** +with different inputs and guarantees. This is the validator split, +decided in [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model), +[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model), +and +[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape), +and recorded in [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md). + +| Path | Input | Validator | Schema source | +|------|-------|-----------|---------------| +| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator (`bast_validation`) | The BAST document (binary layout + value constraints) | +| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema | + +### `validate_bytes` — bytes in, BAST is the validator + +The materializer produces a `serde_json::Value` tree from bytes +(walking the layout engine). By construction, this `Value` is +*structurally correct*: all declared fields are present (the +materializer iterates the field list), types are correct (`read_u32` +produces `Value::Number`), bounds are checked (via +`data_access::check_bounds`), UTF-8 is valid (via `from_utf8`), the +discriminator is in the mapping, and the boolean byte is 0 or 1. + +What the materializer does NOT check — and what the BAST-native +validator checks afterward — are **value-domain constraints expressed +in the BAST document**. The BAST-native validator +(`src/bast_validation.rs`) is a recursive walker over the BAST typed +tree ([`crate::bast::BastDoc`]/[`BastType`]) that checks exactly these: + +| Constraint | Validator arm | +|------------|---------------| +| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` with `as_i64`/`as_u64` + range check | +| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` (JSON precision caveat per ADR-005) | +| Float finiteness (Float32/64) | `validate_float` with `as_f64().is_finite()` | +| String `maxLength` (byte length) | `check_string` reads the field-level `maxLength` | +| Bytes `maxLength` (array length) | `check_bytes` accepts `Value::String` and `Value::Array` forms | +| RFC 3339 timestamp shape | `validate_timestamp` — non-strict check (matching v0.1.0) | +| Enum index bounds | `validate_enum` checks `idx < values.len()` — **fixes the v0.1.0 dead constraint** | +| Union variant dispatch | `validate_union` reads `__discriminator`, looks up the variant, recurses via `validate_typeref` | +| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses | +| Array count | `validate_array` checks `arr.len() == count` and recurses per element | +| Record values | `validate_record` recurses into each value's `values` type | +| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) | + +No external JSON Schema is required for `validate_bytes`. The BAST +document is the complete specification of the binary format — it +describes both the layout (how to read) and the constraints (what +values are valid). This is the "schema is the format" principle from +ADR-001, now fully realized. + +An optional external JSON Schema can be layered on top for constraints +BAST doesn't express (cross-field consistency, regex patterns on string +content). This is additive, not load-bearing. + +### `validate_json` — JSON in, JSON Schema is the validator + +The consumer provides a JSON `Value` (e.g., an incoming JSON-RPC +request). The BAST document is irrelevant — BAST describes bytes, not +JSON shape. The right validator for a JSON value is a standard +`jsonschema::Validator` built from a standard JSON Schema document the +consumer provides at `AlkTypeEngine::compile` time. This is the path +alkcall uses for its `OperationSpec` JSON validation. No custom +keywords; BAST is not involved. + +If no JSON Schema was supplied to `compile`, the JSON-validation +methods return `AlkTypeError::Schema` (`validate_json`) or `false` +(`is_valid_json`). + +### `AlkTypeError::Validation` payload shape (D-BAST-009) + +The `validate_bytes` path no longer uses `jsonschema`, so its error +payload is constructed via `jsonschema::ValidationError::custom` purely +to keep the `Validation` variant's type unchanged. The rationale is +consumer ergonomics on the *combined* path: consumers like alkcall use +both `validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary +channels) and handle `AlkTypeError::Validation` in one place. A single +uniform payload type means one match arm covers both sources. + +The alternative (`Validation(String)`) would force `validate_json` to +flatten its structured errors (instance path, schema path, keyword) to +a `String` via `Display` — the more information-rich path loses data to +accommodate the less rich one. That is the wrong direction. + +### What is removed + +Under the BAST pivot, the v0.1.0 validation machinery is removed from +the `validate_bytes` path: + +- All 19 `jsonschema::Keyword` implementations (~200 lines of validator + factories) — replaced by the BAST-native validator (~250 lines, a + flat match with no factories, no trait objects, no sub-validator + pre-computation). +- `inline_union_variant_refs()` — union variant refs are resolved lazily + by the validator and materializer. +- The custom-keyword `build_validator` path — `build_validator` is + repurposed to build a *standard* `jsonschema::Validator` from a + consumer-provided JSON Schema (no custom keywords). See + [`build_validator`](#build_validator). + +The `jsonschema` crate **remains a direct dependency** for +`validate_json` and for validating BAST documents against the BAST +meta-schema. The only thing removed is the custom keyword integration +path. The `validate_bytes` path no longer touches `jsonschema` — a +small wasm binary-size win in addition to the architecture +simplification. ## Validation Strategy -Validation is delegated to the `jsonschema` crate (v0.46.5, Draft -2020-12). The alktype engine does not implement its own validation — -it registers custom keyword validators for each `AlkType:*` kind and -lets `jsonschema` handle the structural validation (object properties, -required fields, array items, enum values). +The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md) +and refined by [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md): -The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md): - -1. **Load time:** Parse the schema JSON, build the layout engine, build the - jsonschema validator. This is the `AlkTypeEngine::compile(schema)` constructor. +1. **Load time:** Parse the BAST document into the typed tree, compute + the layout engine, and (optionally) build the standard + `jsonschema::Validator` for the JSON-validation path. This is the + `AlkTypeEngine::compile` constructor. 2. **Access time:** Use the compiled engine for repeated read/write operations. Validation is opt-in per operation. -### What validation validates - -The jsonschema validator operates on `serde_json::Value` instances — it -validates JSON representations of data, not raw byte buffers. This is -the correct separation of concerns: - -- **JSON validation** (jsonschema): validates that a JSON document - conforms to the schema. Used for validating hand-written schemas, - TypeBox output, JSON payloads, or the JSON representation of a binary - struct after deserialization. -- **Binary access validation** (data access layer): the read/write - functions perform type-level validation at access time — range checks - for integers, UTF-8 validity for strings, buffer bounds checking. - These return `AlkTypeError::Access` with field paths. - -The "schema is the format" principle means the same schema describes -both the JSON shape and the binary layout. The jsonschema validator -checks the JSON shape; the data access layer checks the binary layout. -A consumer that wants to validate a binary buffer end-to-end reads the -buffer into a `Value` tree via the data access layer, then validates -that `Value` against the jsonschema validator. This is a two-step -process, not a single `validate(buffer)` call. - ### The `AlkTypeEngine` struct -The `AlkTypeEngine` is the compiled form of a schema. It supports both -layout modes (ADR-002) via an internal `Layout` enum: +The `AlkTypeEngine` is the compiled form of a BAST document. It +supports both layout modes (ADR-002) via an internal `Layout` enum: ```rust pub struct AlkTypeEngine { - layout: Layout, // packed or aligned (private enum) - validator: jsonschema::Validator, // compiled once at load time - endian: Endian, // parsed from the schema's "endian" annotation - schema: Value, // the normalized schema (refs resolved) + layout: Layout, // packed or aligned (private enum) + json_validator: Option, // None when no JSON Schema supplied + endian: Endian, // parsed from the root struct's "endian" + bast_doc: Value, // retained for sequential_reader/read_field + root_name: String, // the selected $defs entry } // Private — the consumer selects via LayoutMode at compile time. @@ -68,205 +152,105 @@ enum Layout { ``` The consumer selects the mode at construction time via `LayoutMode` -(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout` -enum is private — the engine exposes mode-appropriate accessors instead: +(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The +`Layout` enum is private — the engine exposes mode-appropriate +accessors instead: ```rust impl AlkTypeEngine { - pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result; + pub fn compile( + bast_doc: &Value, + root_name: &str, + mode: LayoutMode, + json_schema: Option<&Value>, + ) -> Result; pub fn mode(&self) -> LayoutMode; pub fn endian(&self) -> Endian; pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode pub fn sequential_reader(&self) -> Option; // owned fresh reader (ADR-007) - pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // ADR-004 - pub fn is_valid_json(&self, instance: &Value) -> bool; // ADR-004 - pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // ADR-010 + pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str) + -> Result, AlkTypeError>; // aligned mode + pub fn write_field(&self, buffer: &mut [u8], field_path: &str, + value: &FieldValue<'_>) -> Result<(), AlkTypeError>; // aligned mode + pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // D-BAST-007 + pub fn is_valid_json(&self, instance: &Value) -> bool; // D-BAST-007 + pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // D-BAST-006 } ``` -`compile` takes `&mut Value` because it normalizes `$ref` values in place -(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization)) -before computing the layout and building the validator. The `schema` -field retains the normalized schema for `read_field`'s kind lookup and -for `sequential_reader()`'s factory construction. The validator is -mode-agnostic (it operates on `Value`, not raw bytes). +`compile` takes `&Value` (not `&mut Value`) — BAST needs no in-place +`normalize_refs`. `root_name` selects which `$defs` entry is the +top-level type (D-BAST-001). `json_schema` is the optional +consumer-provided standard JSON Schema for the `validate_json` path +(D-BAST-007); pass `None` when JSON validation is not needed. The +engine retains a clone of the BAST `Value` so `sequential_reader` and +`read_field` can re-parse the typed tree on demand without lifetime +entanglement with the caller's `Value`. -The `Layout::Packed` variant stores only the `LayoutBuilder` (write-side). -The `SequentialReader` (read-side) is not stored — it has mutable cursor -state that the consumer owns, so `sequential_reader()` constructs a fresh -reader on each call (ADR-007). +The `Layout::Packed` variant stores only the `LayoutBuilder` +(write-side). The `SequentialReader` (read-side) is not stored — it has +mutable cursor state that the consumer owns, so `sequential_reader()` +constructs a fresh reader on each call (ADR-007). The `read_field`/`write_field` methods on `AlkTypeEngine` are the aligned-mode data-access API — see [data-access.md](data-access.md) §"Higher-level read/write". -## Custom Keyword Validators +## `build_validator` -Each `AlkType:*` kind gets a `Keyword` implementation registered via -`jsonschema::options().with_keyword(...)`. The validators check leaf -type constraints; `jsonschema` handles all structural validation. - -### Numeric type validators - -**`AlkType:Float32` / `AlkType:Float64`:** -- Value must be a finite number. -- For `Float32`: value must be representable as `f32` (no precision loss - beyond `f32`'s mantissa). - -**`AlkType:Int8` / `AlkType:Int16` / `AlkType:Int32`:** -- Value must be an integer within the type's range. -- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647. - -**`AlkType:Uint8` / `AlkType:Uint16` / `AlkType:Uint32`:** -- Value must be a non-negative integer within the type's range. -- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295. - -### String and binary validators - -**`AlkType:String`:** -- Value must be a valid UTF-8 string. -- If `maxLength` is specified in the schema, the string's byte length - must not exceed it. - -**`AlkType:Bytes`:** -- Value must be a string (the JSON form for `validate_json` consumers) - or an array of integers 0..=255 (the materialized form for - `validate_bytes`). JSON has no native byte type; the string form is - the JSON convention, the array form is the round-trippable form for - non-UTF-8 bytes (see [OQ-007](questions/007-bytes-materialization-lossy-utf8.md)). -- If `maxLength` is specified, the byte length must not exceed it. For - the string form, this is the string's byte length; for the array - form, this is the array length (one entry per byte). -- **Binary representation:** In the binary layout, `TBytes` is raw bytes - with no encoding (not base64, not hex). The JSON representation (for - validation) uses a string or array; the binary representation (for - data access) uses `&[u8]` directly. - -**`AlkType:Enum`:** -- The `AlkType:Enum` custom keyword signals that the type is an enum for - *layout* purposes (the engine needs to know it's a fixed-size u32 index, - not a variable-length string). The built-in `enum` keyword provides the - value list and handles value-membership validation. The custom keyword - validator is a no-op beyond the built-in check — it exists solely for - the layout engine to recognize the type. - -**`AlkType:Timestamp`:** -- Value must be a valid RFC 3339 timestamp string (the internet profile - of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`). - -### Composite type validators - -**`AlkType:Struct`:** -- Value must be an object. -- Each property must match its declared `AlkType:*` kind. -- Required fields must be present. -- The `jsonschema` crate's built-in `properties` and `required` keywords - handle the structural checks — the custom keyword only needs to - validate that each field's value matches its `AlkType:*` kind. - -**`AlkType:Union`:** -- The instance must be an object with a `__discriminator` field - carrying the mapping key (stringified discriminator value for - byte-offset discriminators, string value for field-name - discriminators). This is the shape the materializer produces for - `validate_bytes`; `validate_json` consumers produce the same shape - when validating a union instance. -- The `UnionValidator` builds a sub-validator for each variant at - factory time (when the parent validator tree is constructed) and - dispatches on `__discriminator` at validation time, validating the - full instance (including the variant fields) against the selected - variant's schema. This closes the OQ-008 gap: variant field - constraints (e.g. `maxLength` on a `Bytes` field inside a variant) - are checked. -- `$ref`s in the union's `mapping` are inlined by - `schema::inline_union_variant_refs` during `AlkTypeEngine::compile` - (before `build_validator`), so the `union_factory` sees full inline - variant schemas. See [OQ-008](questions/008-unionvalidator-variant-dispatch.md). - -**`AlkType:Array`:** -- Value must be an array. -- Each element must match the array's declared element type. -- If `minItems`/`maxItems` is specified, the array length must be within - bounds. - -### Other validators - -**`AlkType:Boolean`:** -- Value must be `true` or `false`. - -**`AlkType:Record`:** -- Value must be an object. -- All values must match the record's declared value type (specified via - the `"values"` property in the schema, e.g., - `"values": { "AlkType:Float32": true }`). - -### Validator implementation pattern - -Each custom keyword implementation is ~10 lines. Example for -`AlkType:Float32`: +`src/validation.rs` exposes one function: ```rust -struct Float32Validator; - -impl Keyword for Float32Validator { - fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> { - match instance { - Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()), - _ => Err(ValidationError::custom("expected finite f32-compatible number")), - } - } - fn is_valid(&self, instance: &Value) -> bool { - instance.as_f64().map_or(false, |f| f.is_finite()) - } -} +pub fn build_validator(schema: &Value) -> Result; ``` -Registration: +Under the pivot this is **repurposed** (D-BAST-007): it builds a +*standard* `jsonschema::Validator` from a consumer-provided plain JSON +Schema — no custom keywords, no BAST involvement. The engine calls it +internally during `compile` when `json_schema` is `Some`. Consumers +that only need a one-off validator may call `jsonschema::options().build(schema)` +directly; `build_validator` exists so the engine's error mapping +(`jsonschema` build error → `AlkTypeError::Schema`) is reused. -```rust -let validator = jsonschema::options() - .with_keyword("AlkType:Float32", |parent, value, path| { - Ok(Box::new(Float32Validator)) - }) - .build(&schema)?; -``` - -The factory closure receives the parent schema object, the keyword's -value, and the schema path. This enables cross-keyword awareness — for -example, a `AlkType:Struct` validator can inspect the parent's -`properties` to validate each field against its declared `AlkType:*` kind. +The v0.1.0 custom-keyword `build_validator` (registered 19 +`with_keyword(...)` factories) is removed. ## AlkTypeError A single `AlkTypeError` enum covers all error conditions across the -engine's three phases (schema parsing, offset computation, read/write) -plus validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md). +engine's phases (schema parsing, offset computation, read/write) plus +validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md); +the variant shapes are unchanged under the pivot (D-BAST-009). ```rust pub enum AlkTypeError { - /// Schema parsing errors (invalid JSON, missing keywords, unknown AlkType kinds). + /// Schema parsing errors (malformed BAST, dangling $ref, unknown kind). Schema(String), /// Offset computation errors (field not found, unsupported type). Offset { field_path: String, reason: String }, /// Read/write errors (buffer too short, invalid UTF-8, value out of range). Access { field_path: String, reason: String }, - /// Validation errors (delegated to jsonschema). - Validation(ValidationError<'static>), + /// Validation errors (both paths — D-BAST-009 uniform payload). + Validation(jsonschema::ValidationError<'static>), } ``` -- **`Schema`** — for errors during `AlkTypeEngine::compile()`. Invalid - JSON, missing required keywords, unknown `AlkType:*` kinds. +- **`Schema`** — for errors during `AlkTypeEngine::compile()` or any + BAST-walking path. Malformed BAST, missing `$defs`, unknown `kind` + string, dangling `$ref`, empty `mapping`, etc. - **`Offset`** — for errors during offset computation. Field not found - in the schema, type not supported for offset computation, recursive - depth exceeded. Carries the field path. -- **`Access`** — for errors during read/write. Buffer too short, invalid - UTF-8 in a string field, value out of range for the target type. - Carries the field path. -- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The - `'static` lifetime is correct — the validator owns its schema reference - and lives for the lifetime of the `AlkTypeEngine`. + in the BAST tree, type not supported for offset computation. Carries + the field path. +- **`Access`** — for errors during read/write. Buffer too short, + invalid UTF-8 in a string field, value out of range for the target + type. Carries the field path. +- **`Validation`** — wraps a `jsonschema::ValidationError<'static>`. + On the `validate_json` path, this is the `jsonschema` crate's own + structured error. On the `validate_bytes` path, it is constructed via + `jsonschema::ValidationError::custom` from the BAST-native + validator's path + reason string. The `'static` lifetime is correct — + the payload owns its data. ### Field-path-carrying errors @@ -287,20 +271,54 @@ you exactly which field failed and why. ### Load time: `AlkTypeEngine::compile()` The expensive work happens once at schema load time: -1. Normalize `$ref` values in the schema (`normalize_refs`). -2. Parse the schema's `"endian"` annotation. -3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned). -4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`). -The result is a `AlkTypeEngine` that can be used for repeated operations. +1. Parse the BAST document into the typed tree (`BastDoc::new`). +2. Parse the root struct's `"endian"` annotation. +3. Compute the layout (`LayoutBuilder` for packed, `OffsetMap` for + aligned). +4. If `json_schema` is `Some`, build the standard + `jsonschema::Validator` via `validation::build_validator`. + +The result is an `AlkTypeEngine` that can be used for repeated +operations. The BAST-native validator is not pre-built — it is a +recursive walker that runs on the materialized `Value` at access time, +re-using the `BastDoc` (re-parsed on demand from the retained +`bast_doc`). + +### Access time: `engine.validate_bytes(&[u8])` + +For binary-layout schemas, `validate_bytes` runs the two phases in +sequence (D-BAST-006): + +1. **Materialize `Value` from bytes.** `materialize::materialize_packed` + or `materialize::materialize_aligned` walks the buffer against the + BAST typed tree and the engine's `Endian`, producing a + `serde_json::Value` tree. Composites are recursed into (`Struct` → + object of field values; `Array` → array of element values; `Union` + → dispatch then recurse; `Record` → object of key/value entries). + The read phase reuses the existing data-access functions and returns + `AlkTypeError::Access` (with field paths) on read failures. +2. **Validate the `Value`.** The materialized `Value` is passed to + `bast_validation::validate_value(&doc, &value)`, producing + `AlkTypeError::Validation` on the first violated value-domain + constraint. + +Mode dispatch: + +- **Packed mode** — materializes fields in declaration order. +- **Aligned mode** — uses the `OffsetMap` to read fields at their + computed offsets. + +Both modes produce the same `Value` form; the BAST-native validator is +mode-agnostic. ### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)` Validation is opt-in per operation. The consumer calls `engine.validate_json(instance)` when validation is desired, or -`engine.is_valid_json(instance)` for a boolean check. The jsonschema -validator is already compiled — these are fast checks against the -compiled validator. +`engine.is_valid_json(instance)` for a boolean check. The +`jsonschema::Validator` is already compiled — these are fast checks +against the compiled validator. ```rust pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; @@ -308,79 +326,37 @@ pub fn is_valid_json(&self, instance: &Value) -> bool; ``` The argument is a `serde_json::Value` (the JSON representation of the -data), not a raw byte buffer — see §"What validation validates" above. -To validate a binary buffer end-to-end, the consumer reads it into a -`Value` tree via the data access layer, then validates that `Value`. +data), not a raw byte buffer. `validate_json` validates against the +consumer-provided JSON Schema supplied at `compile` time (D-BAST-007); +the BAST document is not involved. If no JSON Schema was supplied, +`validate_json` returns `AlkTypeError::Schema` and `is_valid_json` +returns `false`. High-throughput paths can skip validation. Security-sensitive paths -(parsing incoming frames from untrusted peers) can validate every frame. -The choice is the consumer's. +(parsing incoming frames from untrusted peers) can validate every +frame. The choice is the consumer's. -### Access time: `engine.validate_bytes(&[u8])` — binary buffer validation - -For binary-layout schemas (schemas declaring `AlkType:*` kinds), the -engine offers a single-call form of the two-step dance: walk the bytes -against the layout to materialize a `Value` tree, then validate that -`Value` against the compiled jsonschema validator. Decided in -[ADR-010](decisions/010-generalized-validation-validate-bytes.md). - -```rust -pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; -``` - -`validate_bytes` runs the existing machinery in sequence: - -1. **Materialize `Value` from bytes.** A new internal helper - (`materialize_value`, alongside `SequentialReader::read_field_value` - in `src/sequential_reader.rs`) walks the buffer against the schema - and the engine's `Endian`, producing a `serde_json::Value` tree. - Composites are recursed into (`Struct` → object of field values; - `Array` → array of element values; `Union` → dispatch then recurse; - `Record` → object of key/value entries). The read phase reuses the - existing data-access functions and returns `AlkTypeError::Access` - (with field paths) on read failures. -2. **Validate the `Value`.** The materialized `Value` is passed to the - existing `self.validator.validate(&value)`, producing - `AlkTypeError::Validation` on failure. - -Mode dispatch: - -- **Packed mode** — walks with a fresh `SequentialReader` (the engine - is already a reader factory per ADR-007), materializing fields in - declaration order. -- **Aligned mode** — uses the `OffsetMap` to read fields at their - computed offsets, then materializes composites by recursing into the - offset map's nested entries. - -Both modes produce the same `Value` form; the validator is -mode-agnostic (it operates on `Value`, not bytes — ADR-004). - -#### When to use which entry point +### When to use which entry point | Entry point | Schema form | Input form | When | |-------------|--------------|------------|------| -| `validate_json(&Value)` | Any (AlkType or plain JSON Schema) | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); TypeBox output; anything off `serde_json::from_slice` / `from_str` | -| `validate_bytes(&[u8])` | AlkType binary-layout schema | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs | +| `validate_json(&Value)` | Consumer-provided standard JSON Schema | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); anything off `serde_json::from_slice` / `from_str` | +| `validate_bytes(&[u8])` | BAST document (binary layout) | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs | -`validate_bytes` requires the engine's schema to declare `AlkType:*` -kinds — it materializes `Value` via the layout engine, which needs -binary-layout semantics. A pure JSON Schema (call's `input_schema`, -no AlkType kinds) compiled via `AlkTypeEngine::compile` would fail at -the materialize step (no `AlkType:Struct` at the root). For pure JSON -payloads, the consumer uses `serde_json::from_slice` then -`validate_json`. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md) -§"Not a binary-payload validator for JSON-only schemas". +`validate_bytes` requires the engine's root type to be a struct (the +layout engine enforces this) — it materializes `Value` via the layout +engine, which needs binary-layout semantics. For pure JSON payloads, +the consumer uses `serde_json::from_slice` then `validate_json`. See +[ADR-010](decisions/010-generalized-validation-validate-bytes.md) and +[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md). #### What `validate_bytes` is not -- **Not a new validation engine.** It runs the existing `jsonschema` - validator against the existing materialized `Value`. No new - validator code, no parallel validation path (ADR-001). - **Not framing-aware.** It validates the bytes of *one* schema instance. It does not strip length prefixes, parse `[length: u32][payload]` framing, or handle multiple frames in a - buffer. That's the consumer's job. alktype validates what one - schema describes; it does not parse the wire envelope around it. + buffer. That's the consumer's job. alktype validates what one schema + describes; it does not parse the wire envelope around it. - **Not a `Validator` trait.** Two methods on one struct, not a trait abstraction. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md) §"Not a `Validator` trait abstraction". @@ -390,24 +366,24 @@ payloads, the consumer uses `serde_json::from_slice` then Validation and data access are independent operations on the same data. The consumer can: -1. Validate the JSON representation of a buffer to ensure it conforms to - the schema. +1. Validate the bytes of a buffer to ensure it conforms to the BAST + document's value constraints. 2. Read fields from the binary buffer at computed offsets. -3. Both — validate the JSON representation first, then read the binary - buffer (defense in depth). +3. Both — validate first, then read (defense in depth). -The engine does not couple validation and access. A consumer that trusts -its data source can skip validation and go straight to read/write. A -consumer that parses untrusted input can validate the JSON -representation first, then access the binary buffer. +The engine does not couple validation and access. A consumer that +trusts its data source can skip validation and go straight to +read/write. A consumer that parses untrusted input can validate first, +then access the binary buffer. ## Design Decisions | Decision | ADR | Summary | |----------|-----|---------| -| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping | +| Two-validator model (BAST-native + standard jsonschema) | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | `validate_bytes` uses the BAST-native validator; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema; D-BAST-006/007/009 | +| Error handling and validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping | | Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation (materialize `Value` from bytes, then validate); two methods on one struct, not a trait | -| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine | +| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | The BAST document is the complete binary-format spec (layout + value constraints) | ## Open Questions @@ -418,11 +394,19 @@ see [builder.md](builder.md). ## References -- `@alkdev/alknet: docs/research/alknet-typedef/findings.md` - §"Validation" — the POC's custom keyword validators for all 17 kinds +- [`bast-format.md` §Validation Model](bast-format.md#validation-model) + — the normative validation model +- [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) — the + two-validator decision - [ADR-004](decisions/004-error-handling-validation-strategy.md) — error handling and validation strategy -- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds that the - validators check -- [data-access.md](data-access.md) — read/write functions that operate - on the same buffers +- [ADR-010](decisions/010-generalized-validation-validate-bytes.md) — + `validate_bytes` (the collapsed two-step dance) +- [schema-layer.md](schema-layer.md) — the BAST parser that the + BAST-native validator walks +- [data-access.md](data-access.md) — read/write functions and the + materializer that produce the `Value` the validator checks +- [`src/bast_validation.rs`](../../src/bast_validation.rs) — the + BAST-native validator implementation +- [`src/validation.rs`](../../src/validation.rs) — the `build_validator` + helper \ No newline at end of file diff --git a/docs/plans/bast-implementation.md b/docs/plans/bast-implementation.md index dec7cc8..f06b621 100644 --- a/docs/plans/bast-implementation.md +++ b/docs/plans/bast-implementation.md @@ -1,10 +1,27 @@ --- -status: draft +status: complete created: 2026-08-15 +last_updated: 2026-08-15 --- # BAST Pivot — Implementation Plan +**Status: complete.** All 10 steps are implemented and pushed to +`origin/main` (steps 1–8 in commits `66ab9d7` → `54fd112`; step 9 was +a no-op — steps 4–8 converted the tests as they went, leaving only the +intentional `from_bast_str` rejection test referencing the +`"AlkType:Uint32"` string; step 10 synced the architecture docs and +ADRs in this commit). The two new ADRs +([ADR-BAST](../architecture/decisions/bast-bast-format.md), +[ADR-VAL-SPLIT](../architecture/decisions/val-split-two-validator-model.md)) +record the decisions; the amended ADRs (001, 002, 003, 004, 009, 010) +carry supersession/amendment notes. The research record +([`bast-pivot.md`](../research/bast-pivot.md)) is flipped to +`implemented`. What follows is the original plan, preserved as the +historical execution record. + +--- + This is the execution plan for the BAST pivot: replacing alktype's v0.1.0 `AlkType:*` custom-keyword JSON Schema format with the BAST (Binary Abstract Syntax Tree) format. It is the **entry point** an diff --git a/docs/research/bast-pivot.md b/docs/research/bast-pivot.md index aeb34b7..b827ba8 100644 --- a/docs/research/bast-pivot.md +++ b/docs/research/bast-pivot.md @@ -1,11 +1,32 @@ --- -status: draft +status: implemented created: 2026-08-14 last_updated: 2026-08-15 --- # BAST Pivot — Research Record +**Status: implemented.** The BAST pivot landed in steps 1–10 of the +[implementation plan](../plans/bast-implementation.md) (commits +`66ab9d7` → `54fd112` on `origin/main`). The decisions D-BAST-001..009 +are recorded in two new ADRs — +[ADR-BAST](../architecture/decisions/bast-bast-format.md) (the format) +and [ADR-VAL-SPLIT](../architecture/decisions/val-split-two-validator-model.md) +(the two-validator model) — which supersede the format-specific content +of [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) +and refine the validation strategy of +[ADR-004](../architecture/decisions/004-error-handling-validation-strategy.md) +and [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md). +The normative format spec is +[`docs/architecture/bast-format.md`](../architecture/bast-format.md); +the parser is documented in +[`docs/architecture/schema-layer.md`](../architecture/schema-layer.md). + +What follows is the original research record — the *why* and *what was +proved*, preserved as the historical grounding for the decisions. + +--- + Replace alktype's custom JSON Schema keywords (`AlkType:Uint32`, `AlkType:Struct`, etc.) with a standalone JSON format — BAST (Binary Abstract Syntax Tree) — that describes binary data layouts using a