Decompose BAST pivot doc into normative spec + implementation plan

The bast-pivot.md research doc had grown to 1477 lines (~58KB) through
iterative editing, pushing its most actionable content (D-BAST
decisions, POC result, migration steps) past the 50KB Read tool cap.
Agents peeking at the truncated file landed in duplicated/out-of-order
sections. Decompose into three readable-sized files with distinct roles:

- docs/architecture/bast-format.md (28KB, new): the normative BAST
  format spec -- meta-schema, TypeRef, examples, validation model.
  Grounded in the POC and D-BAST-001..009. Stable and safe to write
  now; schema-layer.md/validation.md stay describing current code and
  are rewritten post-implementation (per AGENTS.md ADR-grounding rule).
- docs/plans/bast-implementation.md (31KB, new): the execution entry
  point -- ordered 10-step plan with per-step goal/files/spec-ref/
  verification, the public-API semver contract table up front as a
  scope-creep guardrail, and the ADR-sync checklist at the end. Each
  step links to the specific bast-format.md section and D-BAST anchor.
- docs/research/bast-pivot.md (28KB, trimmed): now the research record
  only -- Summary, Motivation, POC scope/result, Decisions, Risks,
  References. The normative format spec, what-changes tables,
  validator-split details, and migration steps moved to the two new
  docs; pointers added. 1155 lines removed, 216 added.
- docs/architecture/README.md: index updated to list bast-format.md
  and the two in-progress pivot docs, with notes on schema-layer.md
  and validation.md being rewritten when the pivot lands.

All three files are under the 50KB Read cap, so an implementing agent
gets the whole document in one call. Cross-reference anchors verified
to resolve. No code changes; cargo test --release (396 tests) green.

Verification: cargo test --release (310 crate + 86 integration, all pass).
This commit is contained in:
glm-5.2 committed 2026-08-15 10:59:14 +00:00
1 parent 5796d1c22f
commit f5f52c61e8
4 files changed
+1423 -1155

No files matched your search

+10 -2
View File
@@ -15,12 +15,20 @@ format definition; the engine is generic.
| Document | Status | Description |
|----------|--------|-------------|
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [schema-layer.md](schema-layer.md) | draft | The 19 `AlkType:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [schema-layer.md](schema-layer.md) | draft | The 19 `AlkType:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations. *(Current v0.1.0 format; will be rewritten for BAST — see [bast-format.md](bast-format.md).)* |
| [bast-format.md](bast-format.md) | draft | **Target schema format.** BAST (Binary Abstract Syntax Tree): meta-schema, TypeRef, examples, the two-validator model. Supersedes the format-spec content of schema-layer.md when the [BAST pivot](../plans/bast-implementation.md) lands. |
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine`; `validate_bytes` for binary buffers (ADR-010) |
| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine`; `validate_bytes` for binary buffers (ADR-010). *(Current v0.1.0 validation; will be rewritten for the validator split — see [bast-format.md §Validation Model](bast-format.md#validation-model).)* |
| [builder.md](builder.md) | draft | Fluent Rust API for constructing alktype JSON Schemas at runtime, producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema (ADR-009) |
### In-progress work
| Document | Status | Description |
|----------|--------|-------------|
| [BAST pivot — research record](../research/bast-pivot.md) | draft | Motivation, POC scope and result, decisions D-BAST-001..009, risks for the BAST format pivot |
| [BAST pivot — implementation plan](../plans/bast-implementation.md) | draft | Ordered implementation steps, the public-API semver contract, and the ADR-sync checklist for the BAST pivot |
## Applicable ADRs
| ADR | Title | Relevance |
+675
View File
@@ -0,0 +1,675 @@
---
status: draft
last_updated: 2026-08-15
---
# alktype — BAST Format
**BAST** (Binary Abstract Syntax Tree) is alktype's schema format: a
JSON document that describes binary data layouts using a `kind`-based
vocabulary with `$defs`/`$ref` for composition. BAST replaces the
v0.1.0 `AlkType:*` custom-keyword JSON Schema format.
This document is the **normative format specification**. It is grounded
in the POC on branch `bast-validator-poc` (commit `f371fe4`) and the
decisions D-BAST-001 through D-BAST-009 in
[the pivot research record](../research/bast-pivot.md#decisions). The
implementation plan is
[`docs/plans/bast-implementation.md`](../plans/bast-implementation.md).
Until the BAST pivot lands in code, [`schema-layer.md`](schema-layer.md)
describes the *current* (custom-keyword) schema layer. This document
describes the *target* (BAST) schema layer. They coexist temporarily;
the implementation plan's final step retires `schema-layer.md`'s
custom-keyword content.
## Design Principles
1. **BAST is a JSON Schema instance.** A BAST document is valid JSON
that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON
Schema). Any JSON Schema validator can check whether a BAST document
is well-formed; editors with JSON Schema support provide autocomplete
and inline validation for free.
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
top-level `$defs` block. `$ref` handles cross-references and union
variant references — the same pattern as JSON Schema's own `$defs`
and TypeBox's `Type.Module`. No custom reference resolution mechanism.
3. **`kind`-based vocabulary.** Every type has a `kind` field whose
value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.).
This replaces the `AlkType:*` custom-keyword pattern with a flat,
easily-matched string. The 19 `AlkTypeKind` enum variants are
unchanged; `AlkTypeKind::from_str`/`to_str` map between the enum and
the lowercase BAST strings (D-BAST-002).
4. **Order is explicit.** Struct fields are an ordered array, not an
object with `properties`. Field order is unambiguous — no reliance on
`serde_json`'s `preserve_order` for correctness — and matches the
mental model of binary layouts.
5. **Annotations are type-level properties.** Endianness, alignment,
encoding, and discriminators are properties of the type definition
or field, not custom keywords on a separate schema object. Their
*semantics* carry forward unchanged from ADR-003; only their
*location* moves.
## Document Shape
Every BAST document has the same top-level shape:
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
- The `$defs` block is **required** (D-BAST-003). Single-type documents
are a special case with one entry. A bare struct at the top level
would be a different shape with different parsing logic and no home
for additional definitions — rejected.
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode)` (D-BAST-001). It
selects which `$defs` entry is the top-level type. Convention (first
entry) is fragile and depends on JSON key order; a `$root` marker is
redundant with an explicit parameter.
## The Meta-Schema
The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that
validates the *structure* of BAST documents (is it well-formed?). It
lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded
in the crate for offline use. A different validator — the BAST-native
validator (see [Validation Model](#validation-model) below) — validates
*binary data* against a BAST document (are the bytes a valid instance?).
These are different validators for different inputs.
```json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://alk.dev/bast/v1/schema",
"title": "Binary Abstract Syntax Tree (BAST) v1",
"description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
"type": "object",
"properties": {
"$defs": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeDef" }
}
},
"required": ["$defs"],
"$defs": {
"TypeDef": {
"oneOf": [
{ "$ref": "#/$defs/StructDef" },
{ "$ref": "#/$defs/UnionDef" },
{ "$ref": "#/$defs/EnumDef" }
]
},
"StructDef": {
"type": "object",
"properties": {
"kind": { "const": "struct" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
}
},
"required": ["kind", "fields"],
"additionalProperties": false
},
"FieldDef": {
"type": "object",
"properties": {
"name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
"kind": { "$ref": "#/$defs/TypeRef" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"encoding": { "enum": ["length-prefixed", "offset-indirect"] },
"maxLength": { "type": "integer", "minimum": 0 }
},
"required": ["name", "kind"],
"additionalProperties": false
},
"TypeRef": {
"oneOf": [
{
"description": "Primitive type",
"type": "string",
"enum": [
"int8", "int16", "int32", "int64",
"uint8", "uint16", "uint32", "uint64",
"float32", "float64",
"bool", "string", "bytes", "timestamp"
]
},
{
"description": "Reference to a named $defs entry",
"type": "object",
"properties": {
"$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
},
"required": ["$ref"],
"additionalProperties": false
},
{
"description": "Array type (fixed-size only in v1 — count is required)",
"type": "object",
"properties": {
"kind": { "const": "array" },
"element": { "$ref": "#/$defs/TypeRef" },
"count": { "type": "integer", "minimum": 0 }
},
"required": ["kind", "element", "count"],
"additionalProperties": false
},
{
"description": "Record (string-keyed map) type",
"type": "object",
"properties": {
"kind": { "const": "record" },
"values": { "$ref": "#/$defs/TypeRef" }
},
"required": ["kind", "values"],
"additionalProperties": false
}
]
},
"UnionDef": {
"type": "object",
"properties": {
"kind": { "const": "union" },
"endian": { "enum": ["little", "big"] },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
},
"discriminator": {
"oneOf": [
{
"type": "object",
"properties": {
"kind": { "const": "byte" },
"offset": { "type": "integer", "minimum": 0 },
"type": { "enum": ["uint8", "uint16", "uint32"] }
},
"required": ["kind", "offset", "type"],
"additionalProperties": false
},
{
"type": "object",
"properties": {
"kind": { "const": "field" },
"name": { "type": "string" }
},
"required": ["kind", "name"],
"additionalProperties": false
}
]
},
"mapping": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeRef" }
}
},
"required": ["kind", "discriminator", "mapping"],
"additionalProperties": false
},
"EnumDef": {
"type": "object",
"properties": {
"kind": { "const": "enum" },
"values": {
"type": "array",
"items": { "type": "string" },
"minItems": 1
}
},
"required": ["kind", "values"],
"additionalProperties": false
}
}
}
```
## Type Definitions
### Struct
```json
{
"kind": "struct",
"endian": "big",
"align": 256,
"fields": [
{ "name": "channel_id", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
}
```
- `kind` (required): `"struct"`.
- `endian` (optional): `"little"` (default) or `"big"`. Sets the default
for all fields; field-level `endian` overrides.
- `align` (optional): struct-level alignment, only meaningful in aligned
static mode (ADR-002/003).
- `fields` (required): ordered array of [FieldDef](#fielddef). Array
position is field order — no reliance on JSON key order.
### FieldDef
```json
{ "name": "handle", "kind": "string", "encoding": "offset-indirect", "maxLength": 256 }
```
- `name` (required): identifier, `^[a-zA-Z_][a-zA-Z0-9_]*$`.
- `kind` (required): a [TypeRef](#typeref) — primitive string, `$ref`
object, array object, or record object.
- `endian` (optional): overrides the struct/union default for this field.
- `align` (optional): field-level alignment (aligned mode only).
- `encoding` (optional): `"length-prefixed"` (default) or
`"offset-indirect"`. See [Variable-length encoding](#variable-length-encoding).
- `maxLength` (optional): byte-length cap. See
[Variable-length encoding](#variable-length-encoding).
### TypeRef
`TypeRef` is the central mechanism for referencing types. Four forms:
| Form | Example | Meaning |
|------|---------|---------|
| Primitive string | `"uint32"` | A built-in primitive (see [Primitives](#primitives)) |
| `$ref` object | `{ "$ref": "#/$defs/Read" }` | Reference to a named `$defs` entry |
| Array object | `{ "kind": "array", "element": "uint32", "count": 3 }` | Fixed-size array |
| Record object | `{ "kind": "record", "values": "string" }` | String-keyed map |
The `$ref` form uses standard JSON Pointer syntax **restricted to
`#/$defs/<name>`** — no external references, no fragment-only pointers,
no bare names. The restriction keeps resolution a single hash lookup
and eliminates the `normalize_refs` step the v0.1.0 engine needed for
TypeBox's bare-name refs.
Arrays and records are inline type constructors, not top-level `$defs`
entries. Complex element types use nested `$ref`:
```json
{ "kind": "array", "element": { "$ref": "#/$defs/ComplexElement" }, "count": 4 }
```
### Primitives
The 14 primitive `kind` strings map to the unchanged `AlkTypeKind`
variants (D-BAST-002 — lowercase strings, PascalCase enum variants):
| BAST kind | `AlkTypeKind` | Rust type | Size | Category |
|-----------|---------------|-----------|------|----------|
| `int8` | `Int8` | `i8` | 1 | fixed |
| `int16` | `Int16` | `i16` | 2 | fixed |
| `int32` | `Int32` | `i32` | 4 | fixed |
| `int64` | `Int64` | `i64` | 8 | fixed |
| `uint8` | `Uint8` | `u8` | 1 | fixed |
| `uint16` | `Uint16` | `u16` | 2 | fixed |
| `uint32` | `Uint32` | `u32` | 4 | fixed |
| `uint64` | `Uint64` | `u64` | 8 | fixed |
| `float32` | `Float32` | `f32` | 4 | fixed |
| `float64` | `Float64` | `f64` | 8 | fixed |
| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed |
| `string` | `String` | length-prefixed UTF-8 | variable | variable |
| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable |
| `timestamp` | `Timestamp` | length-prefixed RFC 3339 string | variable | variable |
`int64`/`uint64` are alktype additions (not in TypeBox's `typedef.ts`),
required by SFTP `offset: u64` and metatensor `data_offsets`. JSON
precision caveat per ADR-005 applies: integers beyond `2^53` lose
precision in `serde_json::Value::Number`; the binary path is exact.
### Enum
```json
{
"kind": "enum",
"values": ["Ok", "PermissionDenied", "NoSuchFile", "Failure"]
}
```
- `kind` (required): `"enum"`.
- `values` (required): non-empty array of strings, in declaration order.
- Binary representation: a `u32` index into `values` (0-based), encoded
per the struct's endianness. This is a deliberate deviation from
TypeBox's string enum in favor of binary efficiency — a `u32` index is
compact, fixed-size, and sufficient for any realistic enum.
**Bug fix vs v0.1.0:** The v0.1.0 engine has a dead constraint on the
bytes path — the built-in `enum` keyword checks string membership, but
the materializer emits `Value::Number(index)`, which never matches. The
BAST-native validator (see [Validation Model](#validation-model)) checks
the materialized index against `values.len()` bounds, fixing this.
### Union
```json
{
"kind": "union",
"endian": "big",
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "uint8"
},
"mapping": {
"1": { "$ref": "#/$defs/Init" },
"3": { "$ref": "#/$defs/Open" },
"5": { "$ref": "#/$defs/Read" }
}
}
```
- `kind` (required): `"union"`.
- `endian` (optional): default endianness for variant fields.
- `discriminator` (required): one of:
- **Byte-offset**: `{ "kind": "byte", "offset": <N>, "type": "uint8"|"uint16"|"uint32" }`.
The discriminator byte is at `offset`; the variant struct starts at
`offset + discriminator_size`. Mapping keys are stringified integers.
- **Field-name**: `{ "kind": "field", "name": "<field>" }`. The
discriminator is a length-prefixed string field; mapping keys are
string values matching the field's value. The union's `fields` array
(optional, only valid with field-name discriminators per D-BAST-005)
provides the field definitions including the discriminator field.
- `mapping` (required): object mapping discriminator values to
[TypeRef](#typeref) entries (typically `$ref` to `$defs` variants).
Variant `$ref`s are resolved **lazily** by the materializer and
validator — no `inline_union_variant_refs` compile step (removed under
BAST). The validator recurses into the selected variant's BAST
definition on `__discriminator` lookup, recovering the OQ-008
per-variant constraint enforcement (e.g., `maxLength` on a `bytes`
field inside a variant) without custom keywords.
### Array
```json
{ "kind": "array", "element": "float32", "count": 3 }
```
- `kind` (required): `"array"`.
- `element` (required): a [TypeRef](#typeref).
- `count` (required in v1): the fixed element count. **Arrays of
variable-length elements without a `count` are not supported in v1**
(D-BAST-004, aligning with OQ-001). The meta-schema enforces this:
`count` is in `required`. Variable-length collections are available
via `record` instead.
For fixed-size elements with a known count, the array size is
`element_size × count`. For variable-length elements (e.g.,
`"element": "string"`) with a known count, each element carries its own
length prefix — the array is count-prefixed in the sense that the count
is known at schema time, but the total byte size is not.
### Record
```json
{ "kind": "record", "values": "string" }
```
- `kind` (required): `"record"`.
- `values` (required): a [TypeRef](#typeref) for the value type.
Binary layout: a count-prefixed sequence of `(key, value)` pairs —
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
Each key is a length-prefixed UTF-8 string. Each value is encoded per
its `values` type. There is no separate `value_len` prefix — the value's
size is determined by its kind (fixed-size kinds have a known size;
variable-length kinds carry their own length prefix). The count and
key-length prefixes respect the struct's endianness.
## Variable-Length Encoding
The three strategies from ADR-003 carry forward with the same semantics,
expressed as field-level properties instead of keyword-value objects:
| Strategy | BAST syntax | Behavior |
|----------|------------|----------|
| Inline length-prefixed (default) | `{ "name": "handle", "kind": "string" }` | `[u32 length][data]` |
| Fixed-size reservation | `{ "name": "name", "kind": "string", "maxLength": 256 }` | Reserve `maxLength` bytes (aligned mode); validation constraint (packed mode) |
| Offset indirection | `{ "name": "blob", "kind": "bytes", "encoding": "offset-indirect" }` | `{offset: u32, length: u32}` pointing to separate data region |
**Default strategy selection (unchanged from v0.1.0):**
- **Packed sequential mode:** always inline length-prefixing.
`maxLength` is a validation constraint only.
- **Aligned static mode:** fixed-size reservation if `maxLength` is
declared; offset indirection if `"encoding": "offset-indirect"` is
declared; inline length-prefixing otherwise.
**Length prefix endianness:** The 4-byte length prefix (strategies 1
and 3) respects the effective endianness (struct default or field
override). In little-endian mode, `u32::from_le_bytes`; in big-endian
mode, `u32::from_be_bytes`. Ensures SFTP consumers (big-endian) have
consistent byte order for field values and length prefixes.
Applies to all variable-length types: `string`, `bytes`, `timestamp`,
`record`, and arrays of variable-length elements.
## Endianness
Struct-level or union-level property with per-field override (same
semantics as ADR-003):
- Struct/union-level `"endian"` sets the default for all fields.
- Field-level `"endian"` overrides the struct/union default.
- Default is `"little"` when neither is specified.
- The length prefix for variable-length fields respects the effective
endianness.
```json
{
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "id", "kind": "uint32" },
{ "name": "handle", "kind": "string" },
{ "name": "crc", "kind": "uint32", "endian": "little" }
]
}
```
## Alignment
Struct-level or field-level property, only meaningful in aligned static
mode (same as ADR-003):
```json
{
"kind": "struct",
"align": 256,
"fields": [
{ "name": "header", "kind": { "$ref": "#/$defs/Header" } },
{ "name": "weight", "kind": "float32", "align": 16 }
]
}
```
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/
f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length
prefix), 1 for struct/union/array. Unchanged from v0.1.0.
- Ignored in packed sequential mode.
## Validation Model
BAST separates two concerns that the v0.1.0 format conflates, and in
doing so reveals that the engine has **two distinct validation paths**
with different inputs and guarantees. This is the validator split,
decided in D-BAST-006, D-BAST-007, and D-BAST-009. See
[`validation.md`](validation.md) for the current (pre-pivot) validation
layer; this section specifies the target model.
### Two validators, two inputs
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
**`validate_bytes` — bytes in, BAST is the validator.** The materializer
produces a `Value` tree from bytes. By construction, this `Value` is
*structurally correct*: all declared fields are present (the
materializer iterates the field list), types are correct (`read_u32`
produces `Value::Number`), bounds are checked (via `data_access::
check_bounds`), UTF-8 is valid (via `from_utf8`), the discriminator is
in the mapping, and the boolean byte is 0 or 1. What the materializer
does NOT check — and what the 19 v0.1.0 custom keyword validators check
afterward — are **value-domain constraints expressed in the BAST
document**. The BAST-native validator is a recursive walker over the
BAST type tree that checks exactly these:
| Constraint | Validator arm |
|------------|---------------|
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` with `as_i64`/`as_u64` + range check |
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `validate_float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `check_string` reads the field-level `maxLength` |
| Bytes `maxLength` (array length) | `check_bytes` accepts `Value::String` and `Value::Array` forms |
| RFC 3339 timestamp shape | `validate_timestamp` — same non-strict check as v0.1.0 |
| Enum index bounds | `validate_enum` checks `idx < values.len()` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `validate_union` reads `__discriminator`, looks up the variant, recurses via `validate_typeref` |
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
| Record values | `validate_record` recurses into each value's `values` type |
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
No external JSON Schema is required for `validate_bytes`. The BAST
document is the complete specification of the binary format — it
describes both the layout (how to read) and the constraints (what
values are valid). This is the "schema is the format" principle from
ADR-001, now fully realized.
An optional external JSON Schema can be layered on top for constraints
BAST doesn't express (cross-field consistency, regex patterns on string
content). This is additive, not load-bearing.
**`validate_json` — JSON in, JSON Schema is the validator.** The
consumer provides a JSON `Value` (e.g., an incoming JSON-RPC request).
The BAST document is irrelevant — BAST describes bytes, not JSON shape.
The right validator for a JSON value is a standard
`jsonschema::Validator` built from a standard JSON Schema document the
consumer provides. This is the path alkcall uses for its `OperationSpec`
JSON validation. No custom keywords; BAST is not involved.
### What is removed
Under the BAST pivot, the v0.1.0 validation machinery is removed from
the `validate_bytes` path:
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
factories) — replaced by the BAST-native validator (~250 lines, a
flat match with no factories, no trait objects, no sub-validator
pre-computation).
- `inline_union_variant_refs()` — union variant refs are resolved lazily
by the validator and materializer.
- `build_validator()`'s custom-keyword path — repurposed or removed (see
the implementation plan's step 6 for the decision on its fate).
The `jsonschema` crate **remains a direct dependency** for
`validate_json` and for validating BAST documents against the BAST
meta-schema. The only thing removed is the custom keyword integration
path. The `validate_bytes` path no longer touches `jsonschema` — a
small wasm binary-size win in addition to the architecture
simplification.
### `AlkTypeError::Validation` payload shape
**Decided (D-BAST-009):** Keep
`Validation(jsonschema::ValidationError<'static>)`.
The `validate_bytes` path no longer uses `jsonschema`, so its error
payload is constructed via `jsonschema::ValidationError::custom` purely
to keep the variant's type unchanged. The rationale is consumer
ergonomics on the *combined* path: consumers like alkcall use both
`validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary
channels) and handle `AlkTypeError::Validation` in one place. A single
uniform payload type means one match arm covers both sources. The
alternative (`Validation(String)`) would force `validate_json` to
flatten its structured errors (instance path, schema path, keyword) to
a `String` via `Display` — the more information-rich path loses data to
accommodate the less rich one. That is the wrong direction.
The `no_std`/minimal-build angle (OQ-002) that the alternative was
meant to enable is moot: `validate_json` requires `jsonschema`
regardless, so a bytes-only `no_std` build already has to give up
`validate_json` as a separate, larger decision. The right place to
revisit is when/if OQ-002 is actually pursued.
## Relationship to JSON Schema and TypeBox
### BAST is a JSON Schema dialect
BAST is a specific JSON Schema instance format — like how JSON Schema
itself is a JSON document conforming to the JSON Schema meta-schema.
BAST documents conform to the BAST meta-schema. The entire JSON Schema
tooling ecosystem works with BAST:
- **Validation:** `jsonschema::options().build(&bast_meta_schema)?.validate(&bast_doc)`
- **Editors:** VSCode with `$schema` pointing to the BAST meta-schema URL
- **Documentation:** JSON Schema generators produce human-readable docs
from the meta-schema
### TypeBox interop
TypeBox's `Type.Module({...})` pattern maps naturally to BAST's `$defs`
structure. A TypeBox module defining binary types can serialize to BAST
JSON. The relationship:
- TypeBox → BAST JSON → alktype engine (binary layout)
- TypeBox → standard JSON Schema → jsonschema (JSON validation)
Same TypeBox source, two output formats, two validators.
### Not a replacement for JSON Schema
BAST does not replace JSON Schema for JSON data validation. A BAST
document cannot validate a JSON payload — it describes binary data
layouts and value-domain constraints for bytes. For JSON validation,
consumers use standard JSON Schema documents (which may be derived from
BAST via future codegen, or authored separately). The `validate_json`
path accepts a consumer-provided JSON Schema and uses a standard
`jsonschema::Validator` — BAST is not involved.
This is the split: `validate_bytes` is BAST-native (the BAST document
is both the layout spec and the validation spec for bytes);
`validate_json` is JSON-Schema-native (a standard JSON Schema is the
validation spec for JSON values). One crate, two validators, two input
types.
## Decisions
The BAST format is grounded in decisions D-BAST-001 through D-BAST-009,
recorded in [the pivot research record](../research/bast-pivot.md#decisions).
The implementation-relevant summary:
| Decision | Summary |
|----------|---------|
| [D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention |
| [D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase |
| [D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape |
| [D-BAST-004](../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) |
| [D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` |
| [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed |
| [D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema |
| [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema |
| [D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths |
## References
- [Pivot research record](../research/bast-pivot.md) — motivation, POC
scope and result, decisions D-BAST-001..009, risks
- [Implementation plan](../plans/bast-implementation.md) — ordered
steps, public-API semver contract, ADR-sync checklist
- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics
(carry forward unchanged; only location moves)
- [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) — Int64/
Uint64 as first-class kinds; JSON precision caveat
- [`schema-layer.md`](schema-layer.md) — the current (v0.1.0) schema
layer; superseded by this document when the pivot lands
- [`validation.md`](validation.md) — the current (v0.1.0) validation
layer; rewritten for the validator split when the pivot lands
+532
View File
@@ -0,0 +1,532 @@
---
status: draft
created: 2026-08-15
---
# BAST Pivot — Implementation Plan
This is the execution plan for the BAST pivot: replacing alktype's
v0.1.0 `AlkType:*` custom-keyword JSON Schema format with the BAST
(Binary Abstract Syntax Tree) format. It is the **entry point** an
implementing agent reads first.
Companion documents:
- [`docs/architecture/bast-format.md`](../architecture/bast-format.md) —
the normative BAST format spec (meta-schema, TypeRef, examples,
validation model). Read this for *what* the format is.
- [`docs/research/bast-pivot.md`](../research/bast-pivot.md) — the
research record: motivation, POC scope and result, decisions
D-BAST-001..009, risks. Read this for *why* and *what was proved*.
The POC lives on branch `bast-validator-poc` (commit `f371fe4`) as
`src/bast_poc.rs` — reference scaffolding, deliberately not merged.
**Working order:** read this plan top-to-bottom. The Semver Contract
section is the scope-creep guardrail — consult it before each step.
Each step links to the specific spec section it implements and the
relevant D-BAST-* decision anchor. Implement steps in order; each step
lists its verification gate.
## Semver Contract
The crate is on crates.io at 0.1.0 with zero real consumers, so a
breaking bump is free — but the contract is explicit so the
implementation doesn't drift. Per AGENTS.md, the 0.1.0 public surface
is the items re-exported from `src/lib.rs`. This table is the
authoritative scope-creep guardrail for the pivot.
| Public item (from `lib.rs` re-exports) | Class | Change |
|---|---|---|
| `AlkTypeKind` (enum + variants + methods) | **Additive** | Unchanged. 19 variants, same methods. New `from_str()`/`to_str()` mapping for lowercase BAST kind strings (`"uint32"` ↔ `AlkTypeKind::Uint32`) — additive methods. |
| `Endian`, `VariableEncoding`, `DiscriminatorKind` | **Unchanged** | — |
| `AlkTypeEngine::compile` | **Breaking** | Signature: `compile(schema: &mut Value, mode)` → `compile(bast_doc: &Value, root_name: &str, mode)`. Adds required `root_name` param (D-BAST-001); drops `&mut` (BAST needs no in-place `normalize_refs`); input is a BAST document, not a custom-keyword JSON Schema. |
| `AlkTypeEngine::validate_json` | **Breaking (behavioral)** | Signature unchanged `(instance: &Value) -> Result<...>`, but the validator it runs is now a standard `jsonschema::Validator` from a consumer-provided JSON Schema, not a custom-keyword validator built from the alktype schema. The *contract* of what schema validates the instance changes. |
| `AlkTypeEngine::validate_bytes` | **Unchanged (contract)** | Same signature. Internally the validation step switches from `jsonschema::Validator` to the BAST-native validator. Error type unchanged (D-BAST-009). |
| `AlkTypeEngine::is_valid_json` | **Breaking (behavioral)** | Same caveat as `validate_json` — validates against the consumer JSON Schema, not the alktype schema. |
| `AlkTypeEngine` accessors (`endian`, `mode`, `offset_map`, `layout_builder`, `sequential_reader`, `read_field`, `write_field`, etc.) | **Unchanged** | Layout-layer accessors are format-agnostic. |
| `LayoutMode`, `OffsetMap`, `ByteRange` | **Unchanged** | — |
| `LayoutBuilder`, `PackedLayout`, `FieldPosition` | **Unchanged** | — |
| `SequentialReader`, `FieldValue` | **Unchanged** | — |
| `UnionDispatch` | **Unchanged** | — |
| `data_access::*` functions | **Unchanged** | — |
| `AlkTypeError` (all 4 variants) | **Unchanged** | D-BAST-009 keeps `Validation(jsonschema::ValidationError<'static>)`. |
| `Schema` builder (`struct_`, `object`, `field`, `build`, all setters) | **Breaking (output format)** | Public method signatures unchanged. `build()` output changes from custom-keyword JSON to BAST JSON (for `struct_`) / standard JSON Schema (for `object`). Callers that introspect the built `Value` break; callers that pass it straight to `compile` are source-compatible once `compile` takes BAST. |
| `Definitions` builder (`new`, `define`, `define_value`, `build`, `merge_into`) | **Breaking (output format)** | Same as `Schema` — signatures unchanged, `build()`/`merge_into()` output shape changes to BAST `$defs`. |
| `Discriminator` builder enum | **Unchanged** | — |
| `build_validator` (from `validation`) | **Breaking (signature or removal)** | Currently `build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError>` builds a custom-keyword validator. Under the pivot it either (a) is removed (consumers call `jsonschema` directly for standard JSON Schema) or (b) is repurposed to build a standard `jsonschema::Validator` from a consumer-provided standard JSON Schema (no custom keywords). Decision belongs to step 6. Either way the current signature's contract breaks. |
| `get_alktype_kind`, `get_alktype_kind_enum`, `get_alktype_kind_loose`, `get_alktype_kind_loose_enum`, `normalize_refs`, `inline_union_variant_refs`, `resolve_ref`, `resolve_ref_or_inline`, `parse_align`, `parse_discriminator`, `parse_encoding`, `parse_endian`, `parse_max_length` | **Breaking (removal or rework)** | All currently re-exported from `lib.rs`. `normalize_refs` and `inline_union_variant_refs` are removed (BAST needs neither). The `get_alktype_kind*` family is removed (replaced by direct `kind` parsing). The `parse_*` and `resolve_*` functions are reworked to read BAST properties instead of keyword-value objects, or removed if subsumed by the BAST parser. **Open: which of these stay public vs become internal.** Current leaning — drop all from `lib.rs` re-exports (they're engine-internal accessors, not consumer API); the BAST parser exposes a new typed surface instead. Confirmed during step 3. |
**Net breaking surface:** `compile`, `validate_json`/`is_valid_json`
(contract), `Schema::build`/`Definitions::build` (output format),
`build_validator` (signature/removal), and the ~13 `schema::*` helper
re-exports. **Net additive:** BAST parser, BAST-native validator,
`AlkTypeKind::from_str`/`to_str`. **Net unchanged:** the entire layout
+ data-access + materialize + tunion layer, `AlkTypeError`, the
`Discriminator` builder, `AlkTypeKind` variants.
### Decisions deferred to their implementation steps
These are small enough to decide when the step is reached, but are
flagged here so they don't become drive-by semver changes:
1. **`validate_json` JSON Schema source** (step 6): does the consumer
pass the JSON Schema to `compile` (engine carries a second
validator) or to `validate_json` at call time? The former preserves
the current single-call ergonomics; the latter is more flexible. Not
semver-relevant either way if `validate_json`'s signature can absorb
a new param or stay as-is — needs the call-site analysis.
2. **`build_validator` fate** (step 6): removed vs repurposed. If
repurposed, its signature stays but its contract (no custom
keywords) changes — a behavioral break, not a type break.
3. **`schema::*` helper re-exports** (step 3): drop from `lib.rs`
(engine-internal) vs keep public for consumers that walk schemas.
Leaning: drop — they're accessors for the old format, and the BAST
parser exposes a cleaner typed surface. Confirmed during step 3.
## Steps
### Step 1 — Add `AlkTypeKind::from_str`/`to_str` for BAST kind strings
**Goal:** Add the lowercase-string mapping (`"uint32"` ↔
`AlkTypeKind::Uint32`) that the BAST parser and validator dispatch on.
This is the additive-only, zero-risk foundation — no existing code
changes.
**Spec reference:** [bast-format.md §Primitives](../architecture/bast-format.md#primitives),
[D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set).
**Files:** `src/schema.rs` (the `AlkTypeKind` impl block). No `lib.rs`
change needed — the methods are inherent on the already-re-exported
enum.
**Implementation notes:**
- `to_str(self) -> &'static str` returns the lowercase BAST string.
- `from_str(s: &str) -> Result<AlkTypeKind, AlkTypeError>` returns
`AlkTypeError::Schema` for unknown strings. This is a new inherent
method, distinct from the existing `FromStr` impl that parses the
v0.1.0 `"AlkType:Uint32"` keyword form. Do not remove the existing
`FromStr` yet — step 8 removes the v0.1.0 accessors.
- Cover all 14 primitive kinds plus `struct`, `union`, `array`,
`record`, `enum` (19 total, matching the enum variants). The
lowercase strings are in the [primitives table](../architecture/bast-format.md#primitives);
composite kinds are `"struct"`, `"union"`, `"array"`, `"record"`,
`"enum"`.
**Verification:** `cargo test --release` (new unit tests for the
mapping, both directions; existing tests unaffected). `cargo clippy
--all-targets -- -D warnings`.
---
### Step 2 — Embed the BAST meta-schema
**Goal:** Embed the BAST meta-schema as a `serde_json::Value` constant
in the crate, available for validating BAST documents at compile time
and for publishing at `https://alk.dev/bast/v1/schema`.
**Spec reference:** [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
**Files:** New `src/bast_meta.rs` (or a `const` in `src/schema.rs` —
match existing module conventions). Re-export the meta-schema `Value`
from `lib.rs` if consumers should be able to validate BAST documents
themselves (likely yes — additive, not semver-relevant).
**Implementation notes:**
- The meta-schema JSON is in [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
Copy it verbatim into a `serde_json::json! {...}` macro invocation or
parse it from an embedded string via `serde_json::from_str`.
- No feature flags (AGENTS.md §6). The meta-schema is a compile-time
constant, no I/O.
- WASM-clean: no `include_str!` of an external file is needed if the
`json!` macro is used; either way is wasm-safe.
**Verification:** `cargo test --release`. `cargo build --target
wasm32-unknown-unknown --release` (meta-schema is a `Value` constant —
wasm-relevant). `cargo clippy --all-targets -- -D warnings`.
---
### Step 3 — BAST document parser
**Goal:** Implement the BAST document parser that the layout engines
and materializer use instead of the `get_alktype_kind*` custom-keyword
accessors. This is the natural entry point for the pivot — the largest
step, and the one the rest of the steps build on.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(the whole document — the parser implements the format spec).
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection),
[D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement),
[D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions).
**Files:** New `src/bast.rs` (the parser). The existing `src/schema.rs`
stays for now — steps 4–8 migrate callers off it. Update `src/lib.rs`
to add `pub mod bast;` and re-export the parser's public surface.
**Implementation notes:**
- The parser reads `kind`/`fields`/annotation properties from BAST
nodes. It produces a typed surface (a small `BastNode` enum or
equivalent) that the layout engines, materializer, and validator can
walk without re-parsing the raw JSON at every node. The POC parsed
lazily from raw JSON in both passes to keep the model honest; a typed
tree is a straightforward follow-on optimization (POC observation 5).
Either is acceptable for the production version; the typed tree is
recommended since three consumers (layout, materialize, validate)
walk the same tree.
- `$ref` resolution: `#/$defs/<name>` only — a single hash lookup. No
`normalize_refs` (BAST refs are always full JSON Pointers), no
`inline_union_variant_refs` (union variant refs resolved lazily by
the validator and materializer). See [bast-format.md §TypeRef](../architecture/bast-format.md#typeref).
- Untrusted input: every path that walks a BAST document must return
`Err(AlkTypeError::Schema)` on a malformed document, never `panic!`/
`unreachable!` (AGENTS.md §3). The POC's
`malformed_document_produces_schema_error_not_panic` test is the
template.
- **Decide deferred decision #3 here:** drop the `schema::*` helper
re-exports from `lib.rs`, or keep them public. Leaning: drop. The
BAST parser exposes a cleaner typed surface; the v0.1.0 accessors
are engine-internal and not consumer API.
**Verification:** `cargo test --release` (port the POC's parser tests
— the malformed-document test, the type-ref resolution tests).
`cargo clippy --all-targets -- -D warnings`. The layout engines don't
use the parser yet (step 4 wires it in), so the existing suite still
passes on the old path.
---
### Step 4 — Wire `compile()` to accept a BAST document + root name
**Goal:** Change `AlkTypeEngine::compile` to the new signature and
have it use the BAST parser instead of the custom-keyword accessors.
The layout engines (`offset_map`, `layout_builder`,
`sequential_reader`) consume the BAST parser's typed output instead of
walking raw JSON with `get_alktype_kind*`.
**Spec reference:** [bast-format.md §Document Shape](../architecture/bast-format.md#document-shape),
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection).
Semver contract: `compile` is **Breaking**.
**Files:** `src/engine.rs` (the `compile` signature and body). The
layout modules (`src/offset_map.rs`, `src/layout_builder.rs`,
`src/sequential_reader.rs`) — their schema-walking code changes from
`get_alktype_kind*` calls to BAST parser calls. `src/lib.rs` if the
parser's public surface needs re-exporting (step 3 may have done this).
**Implementation notes:**
- New signature: `pub fn compile(bast_doc: &Value, root_name: &str,
mode: LayoutMode) -> Result<Self, AlkTypeError>`. Note `&Value` (not
`&mut Value`) — BAST needs no in-place `normalize_refs`.
- The engine stores the BAST document (or the parsed typed tree) for
`sequential_reader()`'s factory construction and `read_field`'s kind
lookup. The `Layout` enum and mode dispatch are unchanged.
- `parse_endian`, `parse_align`, `parse_encoding`, `parse_discriminator`
are reworked to read BAST properties (struct/field-level) instead of
keyword-value objects. Their *semantics* are unchanged (ADR-003);
only their *input location* moves. Whether they stay as free
functions or become methods on the typed `BastNode` is an
implementation choice — the POC read properties inline.
- The layout engines are format-agnostic beneath the accessors
(checked offset arithmetic, the two modes, union dispatch). This
step is an accessor swap, not a layout-engine rewrite.
**Verification:** `cargo test --release` (test inputs must be converted
to BAST format — see step 9 for the full test conversion; this step
converts the layout tests as a sanity check). `cargo clippy
--all-targets -- -D warnings`. `cargo build --target
wasm32-unknown-unknown --release` (layout/wasm-relevant).
---
### Step 5 — BAST-native validator (production version)
**Goal:** Port the POC's BAST-native validator into a production module
and wire it into `validate_bytes` as the validation step, replacing the
`jsonschema::Validator` call on the bytes path.
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
[D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model),
[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape).
POC reference: `src/bast_poc.rs` on branch `bast-validator-poc`.
**Files:** New `src/bast_validation.rs`. `src/engine.rs`
(`validate_bytes` body — swap the `self.validator.validate(&value)` call
for the BAST-native validator). `src/lib.rs` — add `pub mod
bast_validation;` (the validator is engine-internal; whether it's
re-exported is an implementation choice, leaning no).
**Implementation notes:**
- The POC is the reference. The validator is a single recursive
function (`validate_typeref`) that dispatches on the BAST `kind`. The
constraint table is in [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model).
- Construct `AlkTypeError::Validation` via
`jsonschema::ValidationError::custom` — the variant's payload type is
unchanged (D-BAST-009). The bytes path no longer touches `jsonschema`
for validation, but the error type retains the `jsonschema` type for
uniformity with the `validate_json` path.
- The validator and materializer share the BAST-walking code structure.
If step 3 produced a typed `BastNode` tree, both consume it. If step
3 parses lazily, the validator parses lazily too (POC approach).
- Enum index bounds: check the materialized index against
`values.len()` — this **fixes the v0.1.0 dead constraint** (the
built-in `enum` keyword checked string membership, but the
materializer emits `Value::Number(index)`, which never matched). Net
improvement.
- Union variant dispatch: read `__discriminator`, look up the variant's
BAST definition, recurse. Recovers OQ-008 per-variant constraint
enforcement without custom keywords.
**Verification:** `cargo test --release` — the existing `validate_bytes`
tests are the regression target (test *inputs* change to BAST format
in step 9; expected validation outcomes must be identical). The POC's
20 tests are the reference. `cargo clippy --all-targets -- -D warnings`.
`cargo build --target wasm32-unknown-unknown --release`.
---
### Step 6 — `validate_json` against a consumer-provided JSON Schema
**Goal:** Update `validate_json`/`is_valid_json` to validate against a
standard `jsonschema::Validator` compiled from a consumer-provided JSON
Schema, not a custom-keyword validator built from the alktype schema.
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model).
Semver contract: `validate_json`/`is_valid_json` are **Breaking
(behavioral)**; `build_validator` is **Breaking (signature or
removal)**.
**Files:** `src/engine.rs` (`validate_json`/`is_valid_json` bodies, and
the engine's stored validator field if the JSON Schema is supplied at
compile time). `src/validation.rs` (`build_validator` — repurposed or
removed). `src/lib.rs` (the `build_validator` re-export if removed).
**Implementation notes:**
- **Decide deferred decision #1 here:** does the consumer pass the JSON
Schema to `compile` (engine carries a second validator) or to
`validate_json` at call time? The former preserves single-call
ergonomics; the latter is more flexible. Needs the alkcall call-site
analysis. Not semver-relevant either way if the signature can absorb
the change.
- **Decide deferred decision #2 here:** `build_validator` removed vs
repurposed. If repurposed, its signature stays but its contract
changes (no custom keywords) — a behavioral break. If removed, drop
the `lib.rs` re-export.
- The `jsonschema` crate remains a direct dependency (for `validate_json`
and for validating BAST documents against the meta-schema). Only the
custom keyword integration is removed.
- The engine may carry two validators: the BAST-native validator (for
`validate_bytes`, from step 5) and the standard `jsonschema::Validator`
(for `validate_json`, from this step). Or `validate_json` takes the
JSON Schema at call time and builds a transient validator. The
decision shapes the engine struct's fields.
**Verification:** `cargo test --release` (new tests for the
consumer-provided JSON Schema path; existing `validate_json` tests
converted — their schemas were custom-keyword, now standard). `cargo
clippy --all-targets -- -D warnings`.
---
### Step 7 — Builder API produces BAST JSON
**Goal:** Update the builder's `build()` methods to produce BAST JSON
(for `struct_()`) and standard JSON Schema (for `object()`). Public
method signatures are unchanged; only the output `Value` shape changes.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(the output format), [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats).
Semver contract: `Schema::build`/`Definitions::build` are **Breaking
(output format)**.
**Files:** `src/builder.rs`. `src/lib.rs` if the builder's public
surface changes (it shouldn't — method signatures are unchanged).
**Implementation notes:**
- `Schema::struct_().field(...).build()` → BAST JSON (a `$defs` entry
with `kind: "struct"`, ordered `fields` array, type-level
annotations).
- `Schema::object().field(...).build()` → standard JSON Schema (no
`AlkType:*` keywords, no BAST `kind` — just `type`/`properties`/
`required`).
- `Definitions::build()`/`merge_into()` → a BAST `$defs` block.
- The builder already distinguishes AlkType kinds from JSON Schema types
via naming conventions (`string()` vs `string_()`). The construction
API is the same; only the serialization differs.
- The `Discriminator` builder is unchanged (semver contract:
**Unchanged**).
**Verification:** `cargo test --release` (builder tests assert on the
output `Value` — update the expected shapes). `cargo clippy
--all-targets -- -D warnings`.
---
### Step 8 — Remove v0.1.0 custom-keyword machinery
**Goal:** Remove the dead code now that all callers use the BAST parser
and BAST-native validator.
**Spec reference:** [bast-format.md §What is removed](../architecture/bast-format.md#what-is-removed).
Semver contract: the ~13 `schema::*` helper re-exports are **Breaking
(removal or rework)** (decision #3, confirmed in step 3).
**Files:** `src/schema.rs` (remove `get_alktype_kind*`,
`normalize_refs`, `inline_union_variant_refs`; rework or remove
`parse_*`/`resolve_*`). `src/validation.rs` (remove the 19
`jsonschema::Keyword` implementations if not already removed in step 5/6).
`src/lib.rs` (drop the removed items from the `pub use` block).
**Implementation notes:**
- Remove: all 19 `jsonschema::Keyword` implementations (~200 lines),
`normalize_refs()`, `inline_union_variant_refs()`, the
`get_alktype_kind*` family.
- Rework or remove: `parse_align`, `parse_discriminator`,
`parse_encoding`, `parse_endian`, `parse_max_length`, `resolve_ref`,
`resolve_ref_or_inline`. If the BAST parser subsumes them (likely),
remove them. If any remain useful as free functions over the typed
`BastNode`, keep them internal (not re-exported from `lib.rs`).
- The `jsonschema` crate's `with_keyword(...)` registration calls are
removed from `compile`/`build_validator`. The crate itself stays.
- Drop the removed items from `lib.rs`'s `pub use schema::{ ... }`
block. The BAST parser's public surface replaces them.
**Verification:** `cargo test --release`. `cargo clippy --all-targets
-- -D warnings`. `cargo doc --no-deps` (the public API surface
changed — doc comments must build). `cargo build --target
wasm32-unknown-unknown --release` (removing code shouldn't add
platform deps).
---
### Step 9 — Convert all tests to BAST format
**Goal:** Update the full test suite to use BAST format for inputs.
Test assertions (expected validation outcomes, expected offsets,
expected materialized values) must be identical — only the input
schema shape changes.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(input format).
**Files:** `tests/*.rs` (integration tests), `src/*.rs` inline `#[cfg(test)]`
modules (unit tests).
**Implementation notes:**
- This may be partially done by steps 4–8 (each step converts the tests
it touches as a sanity check). This step is the sweep: every test
using `AlkType:*` keywords converts to BAST `kind`/`fields`.
- The POC's 20 tests are the reference for BAST-shaped test inputs.
- Expected validation outcomes are the regression target. The
enum-index-bounds test is new behavior (the v0.1.0 dead constraint
is now enforced) — that test's expectation *changes* (was: silently
passed; now: `AlkTypeError::Validation`). This is the intended fix,
not a regression.
- Coverage: 310 crate + 86 integration tests (~396 total). All must
pass.
**Verification:** `cargo test --release` (the full suite — this is the
gate). `cargo clippy --all-targets -- -D warnings`.
---
### Step 10 — Sync architecture docs and ADRs
**Goal:** Sync the descriptive docs and ADRs to the shipped code. This
is the final step — per AGENTS.md, ADRs are written post-implementation,
grounded in shipped code.
**Spec reference:** [Semver Contract §ADR impact](#adr-impact-checklist)
below.
**Files:** `docs/architecture/README.md`, `docs/architecture/overview.md`,
`docs/architecture/schema-layer.md` (rewrite for the BAST parser),
`docs/architecture/validation.md` (rewrite for the validator split),
`docs/architecture/builder.md` (update `build()` output examples),
`src/lib.rs` (module doc comment). New ADRs: ADR-BAST, ADR-VAL-SPLIT.
Amended ADRs: 001 (superseded), 003, 004, 009, 010.
**Implementation notes:**
- Rewrite `schema-layer.md` to describe the BAST parser (replaces the
custom-keyword accessor walk-through). The current `schema-layer.md`
content is the v0.1.0 reference; `bast-format.md` already contains
the target spec. Either fold `bast-format.md` into `schema-layer.md`
or keep both with `schema-layer.md` pointing at `bast-format.md` for
the format and describing the parser module.
- Rewrite `validation.md` for the validator split (the [bast-format.md
§Validation Model](../architecture/bast-format.md#validation-model)
content moves here, expanded with the production validator's
details).
- Update `builder.md` output examples to BAST JSON.
- Update `src/lib.rs` module doc comment: "Takes a JSON Schema with
`AlkType:*` custom keywords" → "Takes a BAST document".
- Update `docs/architecture/README.md` index — the document table, the
ADR table (new ADRs, superseded ADR-001), the key design principles
(#1, #2, #7, #10 change wording).
- Remove stale TODOs referencing custom-keyword normalization,
`inline_union_variant_refs`, or the rejected bare-name-ref design
(AGENTS.md §"Architecture Context").
- `docs/research/bast-pivot.md` is the research record — its status
flips from `draft` to `accepted`/`implemented` and it gains a pointer
to the ADRs that superseded its decisions.
**Verification:** `cargo doc --no-deps` (doc comments build).
Cross-reference check: every link in this plan, `bast-format.md`, and
the new/updated ADRs resolves. `cargo test --release` (no code change,
but the doc sweep shouldn't break anything).
## ADR Impact Checklist
Sync these ADRs when step 10 lands. Per AGENTS.md, ADRs are written
post-implementation, grounded in shipped code.
| ADR | Action | Reason |
|---|---|---|
| [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) (purpose, scope, "schema is the format") | **Supersede** | The "schema is the format" principle is retained and strengthened (BAST *is* the format), but the concrete format changes from custom-keyword JSON Schema to BAST. A new ADR (ADR-BAST) records the BAST format as the realization of the principle. ADR-001 Status → Superseded by ADR-BAST. |
| [ADR-002](../architecture/decisions/002-two-layout-modes-packed-vs-aligned.md) (two layout modes) | **Unchanged** | Layout modes are format-agnostic. One-line note that the input format changed but the modes didn't. |
| [ADR-003](../architecture/decisions/003-schema-annotations.md) (annotations) | **Amend** | Annotation *semantics* carry forward unchanged; annotation *location* moves from custom-keyword objects to BAST type-level properties. Amend the "where annotations live" sections, keep the semantics. |
| [ADR-004](../architecture/decisions/004-error-handling-validation-strategy.md) (error handling, validation strategy) | **Amend** | Error enum shape unchanged (D-BAST-009). The "validation strategy" section updates: bytes path uses BAST-native validator, JSON path uses standard `jsonschema`. The load-time/access-time split is retained. |
| [ADR-005](../architecture/decisions/005-int64-uint64-first-class-kinds.md) (Int64/Uint64) | **Unchanged** | Kinds carry forward; JSON precision caveat unchanged. |
| [ADR-006](../architecture/decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) (reject non-final inline in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
| [ADR-007](../architecture/decisions/007-packed-mode-read-factory.md) (packed-mode read factory) | **Unchanged** | Reader factory semantics are format-agnostic. |
| [ADR-008](../architecture/decisions/008-reject-tunion-in-aligned-mode.md) (reject TUnion in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
| [ADR-009](../architecture/decisions/009-builder-api.md) (builder API) | **Amend** | Public method surface unchanged; `build()` output format changes (BAST for `struct_`, standard JSON Schema for `object`). Amend the "output format" section; keep the method catalog. |
| [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md) (`validate_bytes`) | **Amend** | The two-step concept (materialize → validate) is retained. The validation step's *implementation* changes from `jsonschema` custom keywords to the BAST-native validator. Amend the "validation step" section; add a pointer to D-BAST-006/D-BAST-009 and ADR-VAL-SPLIT. |
**New ADRs to write (post-implementation, grounded in shipped code):**
- **ADR-BAST** — the BAST format, meta-schema, and `$defs`/`$ref`/
`kind` vocabulary. Supersedes ADR-001's format-specific content.
- **ADR-VAL-SPLIT** (or fold into ADR-004's amend) — the two-validator
model: BAST-native for `validate_bytes`, standard `jsonschema` for
`validate_json`. Records D-BAST-006, D-BAST-007, D-BAST-009.
**Descriptive docs to sync (post-implementation):**
- `docs/architecture/schema-layer.md` — rewrite for the BAST parser
(replaces the custom-keyword accessor walk-through).
- `docs/architecture/validation.md` — rewrite for the validator split.
- `docs/architecture/builder.md` — update the `build()` output examples
to BAST JSON.
- `src/lib.rs` module doc comment — update the "Takes a JSON Schema
with `AlkType:*` custom keywords" preamble to BAST.
- `docs/architecture/README.md` — update the document table, ADR table,
and key design principles for the pivot.
- `docs/architecture/overview.md` — update the "what" and "why" for
BAST (the crate now takes a BAST document, not a custom-keyword JSON
Schema).
**Stale TODOs to remove:** any TODO referencing custom-keyword
normalization, `inline_union_variant_refs`, or the rejected
bare-name-ref design — align with the ADRs as AGENTS.md §"Architecture
Context" requires.
## Verification Commands
Run these before committing each step. All must pass. Per AGENTS.md:
```bash
cargo test --release # full suite (~396 tests: 310 crate + 86 integration)
cargo clippy --all-targets -- -D warnings
cargo doc --no-deps # if docs changed (step 8, step 10)
cargo build --target wasm32-unknown-unknown --release # if layout/wasm-relevant code changed (step 2, 4, 5, 8)
cargo publish --dry-run --allow-dirty # before a release (post-step 10)
```
+206 -1153
View File
File diff suppressed because it is too large. Load diff