--- status: draft created: 2026-08-14 --- # BAST Pivot — Binary Abstract Syntax Tree as the Schema Format ## Summary Replace alktype's custom JSON Schema keywords (`AlkType:Uint32`, `AlkType:Struct`, etc.) with a standalone JSON format — BAST (Binary Abstract Syntax Tree) — that describes binary data layouts using a `kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is itself a valid JSON Schema instance (it has a meta-schema), making it self-validating, editor-friendly, and trivially consumable from any language with a JSON parser. The engine's core logic (layout computation, data access, union dispatch, two layout modes) is unchanged. Only the JSON parsing layer changes: instead of detecting `AlkType:*` keywords scattered through a JSON Schema tree, the engine parses a purpose-built `kind`-based format. The builder API's public surface stays the same; only the JSON output format changes internally. ## Motivation ### Current state alktype v0.1.0 embeds binary layout information inside standard JSON Schema documents via custom keywords: ```json { "AlkType:Struct": true, "type": "object", "properties": { "channel_id": { "AlkType:Uint32": true, "type": "integer" }, "length": { "AlkType:Uint32": true, "type": "integer" } }, "endian": "big" } ``` This works for the Rust engine — it walks the tree, detects keywords, computes offsets. But it creates friction for everything outside Rust: 1. **Cross-language consumption.** A Python, Go, or TypeScript consumer that wants to parse an alktype schema must re-implement custom keyword detection. The format is not self-describing — you need to know that `AlkType:Uint32` means "4-byte unsigned integer" and that it can appear as either `true` or `{ "encoding": "..." }`. 2. **Code generation.** Generating Rust/TypeScript/Python readers and writers from a schema requires walking an arbitrary JSON Schema tree looking for custom keywords. A `kind`-based format with known keys makes this a straightforward structural walk. 3. **Tooling.** Editors, linters, and schema validators don't understand `AlkType:*` keywords. A BAST document with a published meta-schema gets autocomplete, validation, and documentation in any JSON Schema- aware editor for free. 4. **Two concerns in one document.** The current format conflates binary layout (what the engine needs) with JSON validation (what jsonschema needs). A `type: "object"` with `properties` and `required` is a JSON validation concern; `AlkType:Uint32` is a binary layout concern. They live in the same JSON object but serve different masters. ### The downstream pain is real The alkcall agent's review identified that the channels 8-byte chunk header is hand-rolled with manual bit shifts — alktype's binary layout capability is unused because the custom-keyword format is awkward to integrate for a simple 2-field struct. alktty plans to hand-roll its 5-byte TTY chunk format for the same reason. SFTP's 29 packet types were proven byte-identical with alktype in the POC, but the production path requires defining 29 schemas in the custom-keyword format. All three cases are the same pattern: a small binary struct that needs a schema-driven reader/writer. BAST makes this trivial — a 10-line JSON file replaces hand-rolled bit shifts. ### Timing v0.1.0 was published but has zero real consumers (only bots/scanners have downloaded it). A breaking change now is free. Waiting until adoption creates migration cost. ## The BAST Format ### Design principles 1. **BAST is a JSON Schema instance.** A BAST document is valid JSON that conforms to the BAST meta-schema. Any standard JSON Schema validator can validate a BAST document's structure. 2. **`$defs`/`$ref` for composition.** Named type definitions live in a top-level `$defs` block. `$ref` handles recursion, cross-references, and union variant references. This is the same pattern as TypeBox's `Type.Module` and JSON Schema's own `$defs` — no custom reference resolution mechanism needed. 3. **`kind`-based vocabulary.** Every type has a `kind` field whose value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.). This replaces the `AlkType:*` custom keyword pattern with a flat, easily-matched string. 4. **Order is explicit.** Struct fields are an ordered array, not an object with `properties`. This makes field order unambiguous (no reliance on `serde_json`'s `preserve_order` for correctness) and matches the mental model of binary layouts. 5. **Annotations are type-level properties.** Endianness, alignment, encoding, and discriminators are properties of the type definition, not custom keywords on a separate schema object. ### Examples #### Channels chunk header (2-field struct, big-endian) ```json { "$defs": { "ChunkHeader": { "kind": "struct", "endian": "big", "fields": [ { "name": "channel_id", "kind": "uint32" }, { "name": "length", "kind": "uint32" } ] } } } ``` #### TTY chunk (3-field struct, big-endian) ```json { "$defs": { "TtyChunk": { "kind": "struct", "endian": "big", "fields": [ { "name": "stream_id", "kind": "uint8" }, { "name": "length", "kind": "uint32" } ] } } } ``` #### SFTP Read packet (struct with mixed fixed/variable fields) ```json { "$defs": { "Read": { "kind": "struct", "endian": "big", "fields": [ { "name": "id", "kind": "uint32" }, { "name": "handle", "kind": "string" }, { "name": "offset", "kind": "uint64" }, { "name": "len", "kind": "uint32" } ] } } } ``` #### SFTP Packet union (byte-offset discriminator) ```json { "$defs": { "SftpPacket": { "kind": "union", "endian": "big", "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" }, "mapping": { "1": { "$ref": "#/$defs/Init" }, "3": { "$ref": "#/$defs/Open" }, "5": { "$ref": "#/$defs/Read" }, "6": { "$ref": "#/$defs/Write" }, "101": { "$ref": "#/$defs/Status" } } }, "Read": { "kind": "struct", "endian": "big", "fields": [ { "name": "id", "kind": "uint32" }, { "name": "handle", "kind": "string" }, { "name": "offset", "kind": "uint64" }, { "name": "len", "kind": "uint32" } ] }, "Write": { "kind": "struct", "endian": "big", "fields": [ { "name": "id", "kind": "uint32" }, { "name": "handle", "kind": "string" }, { "name": "offset", "kind": "uint64" }, { "name": "data", "kind": "bytes" } ] }, "Status": { "kind": "struct", "endian": "big", "fields": [ { "name": "id", "kind": "uint32" }, { "name": "status_code", "kind": "uint32" }, { "name": "error_message", "kind": "string" }, { "name": "language_tag", "kind": "string" } ] } } } ``` #### Metatensor header (aligned mode, little-endian, custom alignment) ```json { "$defs": { "TensorHeader": { "kind": "struct", "endian": "little", "align": 256, "fields": [ { "name": "magic", "kind": "uint64" }, { "name": "json_length", "kind": "uint64" }, { "name": "data_offset", "kind": "uint64" } ] } } } ``` #### Recursive type (e.g., a tree node) ```json { "$defs": { "TreeNode": { "kind": "struct", "fields": [ { "name": "value", "kind": "uint32" }, { "name": "left", "kind": { "$ref": "#/$defs/TreeNode" } }, { "name": "right", "kind": { "$ref": "#/$defs/TreeNode" } } ] } } } ``` #### Array of fixed-size elements with known count ```json { "$defs": { "Vector3": { "kind": "struct", "fields": [ { "name": "components", "kind": { "kind": "array", "element": "float32", "count": 3 } } ] } } } ``` #### Array of variable-length elements (count-prefixed) ```json { "$defs": { "StringList": { "kind": "struct", "fields": [ { "name": "items", "kind": { "kind": "array", "element": "string" } } ] } } } ``` #### Record (string-keyed map) ```json { "$defs": { "Headers": { "kind": "struct", "fields": [ { "name": "entries", "kind": { "kind": "record", "values": "string" } } ] } } } ``` #### Enum ```json { "$defs": { "StatusCode": { "kind": "enum", "values": ["Ok", "PermissionDenied", "NoSuchFile", "Failure"] } } } ``` ### The BAST meta-schema A BAST document is valid JSON that conforms to the BAST meta-schema. The meta-schema is a standard JSON Schema (Draft 2020-12) that validates the structure of BAST documents. This means: - Any JSON Schema validator can check whether a BAST document is well-formed before the engine compiles it. - Editors with JSON Schema support (VSCode, JetBrains) provide autocomplete and inline validation for BAST documents. - The format is self-describing — a consumer can inspect the meta-schema to understand the vocabulary without reading Rust source code. The meta-schema lives at a stable URL (e.g., `https://alk.dev/bast/v1/schema`) and is embedded in the crate for offline use. #### Meta-schema sketch ```json { "$schema": "https://json-schema.org/draft/2020-12/schema", "$id": "https://alk.dev/bast/v1/schema", "title": "Binary Abstract Syntax Tree (BAST) v1", "description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.", "type": "object", "properties": { "$defs": { "type": "object", "additionalProperties": { "$ref": "#/$defs/TypeDef" } } }, "required": ["$defs"], "$defs": { "TypeDef": { "oneOf": [ { "$ref": "#/$defs/StructDef" }, { "$ref": "#/$defs/UnionDef" }, { "$ref": "#/$defs/EnumDef" } ] }, "StructDef": { "type": "object", "properties": { "kind": { "const": "struct" }, "endian": { "enum": ["little", "big"] }, "align": { "type": "integer", "minimum": 1 }, "fields": { "type": "array", "items": { "$ref": "#/$defs/FieldDef" } } }, "required": ["kind", "fields"], "additionalProperties": false }, "FieldDef": { "type": "object", "properties": { "name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" }, "kind": { "$ref": "#/$defs/TypeRef" }, "endian": { "enum": ["little", "big"] }, "align": { "type": "integer", "minimum": 1 }, "encoding": { "enum": ["length-prefixed", "offset-indirect"] }, "maxLength": { "type": "integer", "minimum": 0 } }, "required": ["name", "kind"], "additionalProperties": false }, "TypeRef": { "oneOf": [ { "description": "Primitive type", "type": "string", "enum": [ "int8", "int16", "int32", "int64", "uint8", "uint16", "uint32", "uint64", "float32", "float64", "bool", "string", "bytes", "timestamp" ] }, { "description": "Reference to a named $defs entry", "type": "object", "properties": { "$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" } }, "required": ["$ref"], "additionalProperties": false }, { "description": "Array type", "type": "object", "properties": { "kind": { "const": "array" }, "element": { "$ref": "#/$defs/TypeRef" }, "count": { "type": "integer", "minimum": 0 } }, "required": ["kind", "element"], "additionalProperties": false }, { "description": "Record (string-keyed map) type", "type": "object", "properties": { "kind": { "const": "record" }, "values": { "$ref": "#/$defs/TypeRef" } }, "required": ["kind", "values"], "additionalProperties": false } ] }, "UnionDef": { "type": "object", "properties": { "kind": { "const": "union" }, "endian": { "enum": ["little", "big"] }, "discriminator": { "oneOf": [ { "type": "object", "properties": { "kind": { "const": "byte" }, "offset": { "type": "integer", "minimum": 0 }, "type": { "enum": ["uint8", "uint16", "uint32"] } }, "required": ["kind", "offset", "type"], "additionalProperties": false }, { "type": "object", "properties": { "kind": { "const": "field" }, "name": { "type": "string" } }, "required": ["kind", "name"], "additionalProperties": false } ] }, "mapping": { "type": "object", "additionalProperties": { "$ref": "#/$defs/TypeRef" } } }, "required": ["kind", "discriminator", "mapping"], "additionalProperties": false }, "EnumDef": { "type": "object", "properties": { "kind": { "const": "enum" }, "values": { "type": "array", "items": { "type": "string" }, "minItems": 1 } }, "required": ["kind", "values"], "additionalProperties": false } } } ``` ### Type reference resolution `TypeRef` is the central mechanism for referencing types. It has four forms: | Form | Example | Meaning | |------|---------|---------| | Primitive string | `"uint32"` | A built-in primitive type | | `$ref` object | `{ "$ref": "#/$defs/Read" }` | Reference to a named definition | | Array object | `{ "kind": "array", "element": "uint32" }` | Array of elements | | Record object | `{ "kind": "record", "values": "string" }` | String-keyed map | The `$ref` form uses standard JSON Pointer syntax restricted to `#/$defs/`. This is a subset of JSON Schema's `$ref` — no external references, no fragment-only pointers, no bare names. The restriction keeps resolution simple (single hash lookup) and avoids the normalization step that the current engine needs for TypeBox's bare-name refs. Arrays and records are inline type constructors, not top-level `$defs` entries. This keeps the common cases concise while allowing complex element types via nested `$ref`: ```json { "kind": "array", "element": { "$ref": "#/$defs/ComplexElement" } } ``` ### Variable-length encoding The three strategies from ADR-003 carry forward with the same semantics, expressed as field-level properties instead of keyword-value objects: | Strategy | BAST syntax | Behavior | |----------|------------|----------| | Inline length-prefixed (default) | `{ "name": "handle", "kind": "string" }` | `[u32 length][data]` | | Fixed-size reservation | `{ "name": "name", "kind": "string", "maxLength": 256 }` | Reserve `maxLength` bytes (aligned mode); validation constraint (packed mode) | | Offset indirection | `{ "name": "blob", "kind": "bytes", "encoding": "offset-indirect" }` | `{offset: u32, length: u32}` pointing to separate data region | ### Endianness Endianness is a struct-level or union-level property with per-field override, same as ADR-003: - Struct-level `"endian"` sets the default for all fields. - Field-level `"endian"` overrides the struct default. - Default is `"little"` when neither is specified. - The length prefix for variable-length fields respects the effective endianness (struct default or field override). ```json { "kind": "struct", "endian": "big", "fields": [ { "name": "id", "kind": "uint32" }, { "name": "handle", "kind": "string" }, { "name": "offset", "kind": "uint64" }, { "name": "crc", "kind": "uint32", "endian": "little" } ] } ``` ### Alignment Alignment is a struct-level or field-level property, only meaningful in aligned static mode (same as ADR-003): ```json { "kind": "struct", "align": 256, "fields": [ { "name": "header", "kind": { "$ref": "#/$defs/Header" } }, { "name": "weight", "kind": "float32", "align": 16 } ] } ``` ## What Changes ### JSON format | Aspect | Current (custom keywords) | BAST | |--------|--------------------------|------| | Type declaration | `"AlkType:Uint32": true` on a property | `"kind": "uint32"` in a field definition | | Struct fields | `"properties": { "x": {...}, "y": {...} }` | `"fields": [{ "name": "x", ... }, { "name": "y", ... }]` | | Field order | Implicit via `serde_json` `preserve_order` | Explicit via array position | | Endianness | `"endian": "big"` on the schema object | `"endian": "big"` on the struct/union definition | | Variable encoding | `"AlkType:String": { "encoding": "offset-indirect" }` | `{ "name": "x", "kind": "string", "encoding": "offset-indirect" }` | | Discriminator | `"discriminator": { "kind": "byte", "offset": 0, "type": "AlkType:Uint8" }` | `"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" }` | | `$ref` | Bare names (`"Read"`) normalized to `#/$defs/Read` | Always `#/$defs/Read` — no normalization needed | | Top-level container | Schema object with `AlkType:Struct` at root | `{ "$defs": { ... } }` with a root type name | | JSON validation info | `"type": "object"`, `"required"`, etc. on the same object | Separate concern — not in BAST | ### Engine internals | Component | Change | |-----------|--------| | `get_alktype_kind()` / `get_alktype_kind_loose()` | Replaced by parsing `"kind"` field directly from BAST nodes | | `normalize_refs()` | Removed — BAST `$ref` values are always full JSON Pointers | | `inline_union_variant_refs()` | Removed — union mapping values are resolved lazily via `$ref` | | `parse_endian()`, `parse_align()`, `parse_encoding()`, `parse_discriminator()` | Adapted to read from BAST field/struct properties instead of keyword-value objects | | Custom keyword validators (19 `jsonschema::Keyword` impls) | Removed from alktype. BAST validation uses the BAST meta-schema + standard JSON Schema validator. Binary data validation uses `validate_bytes` (materialize + validate against a derived JSON Schema or the BAST structure directly). | | `build_validator()` | Repurposed or removed. The engine no longer registers custom keywords with `jsonschema`. | | `validate_json()` | May be removed or changed to validate against a separately-provided JSON Schema. | | `validate_bytes()` | Unchanged in concept — materialize `Value` from bytes, then validate. The materialization step walks BAST instead of custom-keyword JSON. | ### Public API | Item | Change | |------|--------| | `AlkTypeKind` enum | Unchanged — same 19 variants, same methods | | `Endian`, `VariableEncoding`, `DiscriminatorKind` | Unchanged | | `AlkTypeEngine` | `compile()` takes a BAST document + root type name instead of a custom-keyword JSON Schema. `validate_json()` may change. `validate_bytes()` unchanged. | | `LayoutMode`, `OffsetMap`, `ByteRange` | Unchanged | | `LayoutBuilder`, `PackedLayout`, `FieldPosition` | Unchanged | | `SequentialReader`, `FieldValue` | Unchanged | | `UnionDispatch` | Unchanged | | `data_access` functions | Unchanged | | `AlkTypeError` | Unchanged (Schema/Offset/Access/Validation variants) | | `Schema` builder | Public methods unchanged. `build()` produces BAST JSON instead of custom-keyword JSON. | | `Definitions` builder | Public methods unchanged. `build()` produces a BAST `$defs` block. | | `Discriminator` builder | Unchanged | ### What is removed - All 19 `jsonschema::Keyword` implementations (~200 lines of validator factories) - `normalize_refs()` — BAST `$ref` values are always full JSON Pointers - `inline_union_variant_refs()` — union mapping values are resolved lazily - `get_alktype_kind()` / `get_alktype_kind_loose()` and their `_enum` variants — replaced by direct `kind` field parsing - The `jsonschema` crate dependency for custom keyword registration (the crate may remain as a transitive dependency for BAST meta-schema validation, but the engine no longer registers custom keywords with it) ### What is added - BAST meta-schema (embedded in the crate, published at a stable URL) - BAST document parser — validates a BAST document against the meta-schema, then extracts type definitions - `AlkTypeEngine::compile()` takes `(bast_document: &Value, root_name: &str, mode: LayoutMode)` — the root name selects which `$defs` entry is the top-level type - BAST meta-schema validation at compile time (optional but recommended — the engine can skip it and trust the caller, or validate as a guard) ## The Validator Split A key architectural clarification: BAST separates two concerns that the current format conflates. ### Current model (conflated) ``` ┌─────────────────────────────────────────────┐ │ JSON Schema with AlkType:* custom keywords │ │ ┌───────────────────────────────────────┐ │ │ │ Binary layout info (AlkType:Uint32) │ │ │ │ JSON validation info (type, enum) │ │ │ └───────────────────────────────────────┘ │ └─────────────────────────────────────────────┘ │ ▼ AlkTypeEngine::compile() │ ├──► Layout (offset map / layout builder) └──► Validator (jsonschema with custom keywords) ``` One document serves two roles. The engine extracts both layout and validation from the same JSON tree. ### BAST model (separated) ``` ┌──────────────────────┐ ┌──────────────────────────┐ │ BAST document │ │ JSON Schema document │ │ (binary layout) │ │ (JSON validation) │ │ │ │ │ │ kind: "struct" │ │ type: "object" │ │ fields: [ │ │ properties: { │ │ { name, kind } │ │ id: { type: "integer"} │ │ ] │ │ } │ │ endian: "big" │ │ required: ["id"] │ └──────────┬───────────┘ └────────────┬─────────────┘ │ │ ▼ ▼ AlkTypeEngine::compile() jsonschema::options() │ │ ▼ ▼ Layout (offset map / JSON Schema validator layout builder) (standard, no custom keywords) │ │ ▼ ▼ validate_bytes(&[u8]) validate_json(&Value) ``` Two documents, two validators, two concerns. The BAST document describes binary layout. The JSON Schema document describes JSON data shape. They can be linked (a BAST struct can reference a JSON Schema by `$id` for validation) but they are separate documents. ### What this means for consumers **alkcall** today uses alktype for two roles: 1. Binary layout (channels chunk header) — `AlkTypeEngine::compile()` in packed mode 2. JSON validation (call's `OperationSpec` schemas) — `build_validator()` or `jsonschema` directly Under BAST: 1. Binary layout — `AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed)` — same flow, different input format 2. JSON validation — standard `jsonschema::options().build(&json_schema)` — unchanged, but no longer goes through alktype's `build_validator()` The builder API can produce both formats: - `Schema::struct_().field(...).build()` → BAST JSON - `Schema::object().field(...).build()` → standard JSON Schema ### Validation of binary data `validate_bytes()` still works: materialize a `Value` tree from binary bytes using the BAST layout, then validate that `Value` against a JSON Schema. The JSON Schema can be: - Derived from the BAST definition (a codegen step, future) - Provided separately by the consumer - A standard JSON Schema that the consumer already has for JSON validation of the same logical type The materialization step walks the BAST layout (unchanged from current behavior — it walks the schema to compute offsets and read fields). The validation step uses a standard `jsonschema::Validator` (no custom keywords needed — the `Value` tree is already typed by the materialization). ## Relationship to JSON Schema and TypeBox ### BAST is a JSON Schema dialect BAST is a specific JSON Schema instance format — like how JSON Schema itself is a JSON document that conforms to the JSON Schema meta-schema. BAST documents conform to the BAST meta-schema. The meta-schema is a standard JSON Schema (Draft 2020-12). This means the entire JSON Schema tooling ecosystem works with BAST: - Validation: `jsonschema::options().build(&bast_meta_schema)?.validate(&bast_doc)` - Editors: VSCode with `$schema` pointing to the BAST meta-schema URL - Documentation: JSON Schema generators can produce human-readable docs from the meta-schema ### TypeBox interop TypeBox's `Type.Module({...})` pattern maps naturally to BAST's `$defs` structure. A TypeBox module that defines binary types can serialize to BAST JSON instead of custom-keyword JSON. The codegen (`ts-to-module.ts`) could target BAST as an output format. The relationship is: - TypeBox → BAST JSON → alktype engine (binary layout) - TypeBox → standard JSON Schema → jsonschema (JSON validation) Same TypeBox source, two output formats, two validators. ### Not a replacement for JSON Schema BAST does not replace JSON Schema for JSON data validation. A BAST document cannot validate a JSON payload. It describes binary data layouts. For JSON validation, consumers use standard JSON Schema documents (which may be derived from BAST definitions via codegen, or authored separately). ## Codegen (Future) BAST enables code generation that the custom-keyword format makes awkward. A codegen module (feature-gated behind `codegen`) would: 1. **Input**: A BAST document (or `SchemaRegistry` equivalent) 2. **Walk**: Iterate `$defs` entries, inspect `kind` values 3. **Map**: `"uint32"` → `u32` (Rust), `number` (TypeScript), `int` (Python) 4. **Emit**: Handlebars templates for struct/enum/union definitions ### Generated artifacts | Target | Type definitions | Binary reader/writer | |--------|-----------------|---------------------| | Rust | `struct ChunkHeader { channel_id: u32, length: u32 }` | `fn read_header(buf: &[u8]) -> Result` | | TypeScript | `interface ChunkHeader { channelId: number; length: number }` | `function readHeader(buf: Uint8Array): ChunkHeader` | | Python | `@dataclass class ChunkHeader: ...` | `def read_header(buf: bytes) -> ChunkHeader` | ### Relationship to typebox-rs codegen The typebox-rs `codegen/` module (in `/workspace/@alkimiadev/typebox-rs`) is the reference architecture: - `SchemaRegistry` for named types with `$ref` resolution - `RustGenerator` / `TypeScriptGenerator` wrapping `Handlebars` templates - `schema_to_rust_type()` / `schema_to_ts_type()` mapping functions - Feature-gated behind `codegen = ["handlebars"]` alktype's codegen would follow the same pattern but walk BAST `kind` values instead of `SchemaKind` enum variants. The handlebars-rs dependency is WASM-compatible. ### Scope boundary Codegen is out of scope for the BAST pivot itself. The pivot changes the schema format; codegen builds on top of the new format. It is described here to show that BAST enables it, not to commit to a specific implementation timeline. ## ABI Adapter (Future) A BAST document describes the binary interface of a protocol — it is essentially an ABI specification in JSON. This enables: - **Version negotiation**: Two peers exchange BAST documents to agree on a protocol version. The engine can detect mismatches (field added, type changed, endianness differs) and either reject or adapt. - **Schema migration**: A consumer with schema v1 can read data written by schema v2 if the changes are compatible (fields added at the end, types widened). The engine can compute a migration plan from the diff of two BAST documents. - **WASM interop**: A WASM component can export its BAST schema as part of its WIT interface, enabling host languages to generate readers/writers for the component's binary protocol without manual bindings. This is a future capability, not part of the pivot. It is mentioned because BAST makes it possible in a way that custom keywords scattered through JSON Schema trees do not. ## Migration Path ### Phase 1: BAST format and meta-schema (this pivot) 1. Define the BAST meta-schema (the JSON Schema that validates BAST documents) 2. Implement BAST document parsing in the engine (replace custom keyword detection with `kind` field parsing) 3. Update `AlkTypeEngine::compile()` to accept a BAST document + root type name 4. Update the builder API to produce BAST JSON (public methods unchanged) 5. Remove custom keyword validators, `normalize_refs()`, `inline_union_variant_refs()`, and the `get_alktype_kind*` functions 6. Update all tests to use BAST format 7. Update architecture docs (ADRs, schema-layer.md, etc.) ### Phase 2: Downstream adoption 1. Port alkcall's chunk header to BAST (replace hand-rolled `wire.rs` with `AlkTypeEngine` + `SequentialReader`/`LayoutBuilder`) 2. Port alktty's TTY chunk format to BAST 3. Port the SFTP POC schemas to BAST format ### Phase 3: Codegen (future) 1. Add `codegen` feature flag with `handlebars` dependency 2. Implement `RustGenerator` and `TypeScriptGenerator` walking BAST `kind` values 3. External Handlebars templates in `src/codegen/templates/` ## Open Questions ### OQ-BAST-001: Root type selection How does the engine know which `$defs` entry is the root type? Options: - Explicit: `AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed)` - Convention: the first entry in `$defs` (fragile — depends on JSON key order) - Marker: a `"$root": "ChunkHeader"` property on the BAST document **Leaning**: Explicit. The root type name is a required parameter to `compile()`. This is unambiguous and matches how consumers think about it ("compile the ChunkHeader schema"). ### OQ-BAST-002: JSON validation schema linkage How does a BAST struct reference a JSON Schema for `validate_bytes`? Options: - Separate parameter: `engine.validate_bytes(buf, Some(&json_schema))` - Embedded reference: `{ "kind": "struct", "validation": { "$ref": "https://..." } }` - Convention: same `$id` base, different fragment **Leaning**: Separate parameter for v1. The JSON Schema is a separate document; the engine doesn't need to know about it at compile time. `validate_bytes()` accepts an optional `&jsonschema::Validator` that the consumer provides. This keeps BAST focused on binary layout. ### OQ-BAST-003: Primitive type string set The current proposal uses lowercase strings for primitives: `"uint32"`, `"int8"`, `"float64"`, `"bool"`, `"string"`, `"bytes"`, `"timestamp"`. Alternatives: - PascalCase: `"Uint32"`, `"Int8"` (matches `AlkTypeKind` variant names) - UPPER_CASE: `"UINT32"`, `"INT8"` - Prefixed: `"bast:uint32"` (namespaced, but verbose) **Leaning**: Lowercase. Matches JSON Schema's own convention (`"string"`, `"integer"`, `"boolean"`), is easier to type, and is the convention in the TypeBox research examples. The `AlkTypeKind` enum variants remain PascalCase in Rust — the mapping is a simple `from_str()` impl. ### OQ-BAST-004: Array count — fixed vs variable When is an array fixed-size (no count prefix) vs variable-size (count prefix in binary)? Options: - Explicit `count` field: `{ "kind": "array", "element": "uint32", "count": 3 }` → fixed - Absent `count`: `{ "kind": "array", "element": "uint32" }` → variable (count-prefixed) - Separate `minItems`/`maxItems` like current JSON Schema convention **Leaning**: Explicit `count` for fixed, absent for variable. This is clearer than `minItems == maxItems` and matches the BAST principle of explicit layout information. ### OQ-BAST-005: Top-level `$defs` requirement Should a BAST document always have a top-level `$defs` block, or can a single struct be the root? Options: - Always `$defs`: `{ "$defs": { "ChunkHeader": { "kind": "struct", ... } } }` - Bare struct: `{ "kind": "struct", "fields": [...] }` (no `$defs`) **Leaning**: Always `$defs`. Consistency — every BAST document has the same top-level shape. Single-type documents are a special case of the general form. The `$defs` block is the namespace; the root type name selects the entry point. ### OQ-BAST-006: `validate_json` on `AlkTypeEngine` With custom keyword validators removed, what does `AlkTypeEngine::validate_json()` do? Options: - Remove it — the engine is for binary layout; JSON validation is a separate concern - Keep it with a separately-provided JSON Schema — the engine holds a `jsonschema::Validator` compiled from a JSON Schema the consumer provides at compile time - Keep it as a convenience that validates the materialized `Value` against the BAST structure itself (type checks only, no range constraints) **Leaning**: Keep it with a separately-provided JSON Schema. The engine already compiles a validator at load time (ADR-004). The validator just comes from a standard JSON Schema instead of custom keywords. This preserves the `validate_json` / `validate_bytes` symmetry from ADR-010. ### OQ-BAST-007: Builder API — two output formats The builder API currently produces one JSON format (custom keywords). Under BAST, it needs to produce two: 1. BAST JSON (for `Schema::struct_()`, `Schema::uint32()`, etc.) 2. Standard JSON Schema (for `Schema::object()`, `Schema::string_()`, etc.) Should these be two separate builder types, or one builder with a mode flag? Options: - Two builders: `BastBuilder` and `JsonSchemaBuilder` (or `Schema::bast` and `Schema::json`) - One builder with mode: `Schema::new(Mode::Bast)` / `Schema::new(Mode::JsonSchema)` - One builder, two build methods: `schema.build_bast()` / `schema.build_json_schema()` **Leaning**: One builder, two build methods. The construction API is the same (field names, types, annotations); only the output format differs. `Schema::struct_().field(...).build()` → BAST. `Schema::object().field(...).build()` → standard JSON Schema. The builder already distinguishes AlkType kinds from JSON Schema types via naming conventions (`string()` vs `string_()`). ## Risks and Mitigations | Risk | Mitigation | |------|-----------| | BAST format doesn't cover all 19 type kinds | The format is designed to cover all 19. The meta-schema is the spec — if a kind can't be expressed, the meta-schema is wrong. | | `$ref` resolution complexity moves from engine to schema authoring | BAST `$ref` values are always full JSON Pointers (`#/$defs/Name`). No normalization, no bare names. Resolution is a single hash lookup. | | Losing `jsonschema`'s structural validation (required fields, etc.) | BAST is for binary layout, not JSON validation. Structural constraints belong in the JSON Schema document, not the BAST document. | | Builder API output format change breaks consumers | No real consumers exist yet (v0.1.0 has zero adoption). The builder's public methods are unchanged; only the JSON output format changes. | | Meta-schema maintenance burden | The meta-schema is small (~150 lines) and changes rarely. It's embedded in the crate and published at a stable URL. | ## Proposed POCs Before committing to the full pivot, these focused POCs would de-risk the key decision points: ### POC 1: BAST meta-schema validation Prove that a standard JSON Schema validator can validate BAST documents. - Write the BAST meta-schema as a JSON Schema document - Validate the example BAST documents (chunk header, SFTP packets, recursive types) against it using `jsonschema` - Verify that malformed BAST documents (wrong `kind`, missing `fields`, invalid `$ref`) are rejected with clear errors ### POC 2: BAST → layout compilation Prove the engine can compile BAST documents into layouts. - Implement a minimal BAST parser that extracts type definitions from a BAST document - Feed a BAST ChunkHeader schema through the existing `LayoutBuilder` and `SequentialReader` - Verify byte-identical output with the current custom-keyword path ### POC 3: BAST round-trip against russh-sftp Prove byte-identical SFTP packet serialization using BAST schemas. - Port the SFTP Read/Write/Status schemas from the alknet-typedef-poc to BAST format - Run the existing round-trip tests (`sftp_roundtrip_test.rs`) against BAST-compiled layouts - Verify byte-identical output with `russh_sftp::protocol` serialization ### POC 4: Builder API output switch Prove the builder API can produce BAST JSON without changing its public methods. - Add a `build_bast()` method (or change `build()` output) to the existing `Schema` builder - Construct the ChunkHeader schema via the builder - Verify the output is valid BAST JSON that passes meta-schema validation ## References - [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) — current "schema is the format" principle (to be updated) - [ADR-003](../architecture/decisions/003-schema-annotations.md) — annotation shapes (endianness, alignment, encoding, discriminators — carry forward to BAST) - [ADR-009](../architecture/decisions/009-builder-api.md) — builder API (public surface unchanged, output format changes) - [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md) — `validate_bytes` (unchanged in concept) - `/workspace/research/typebox_research/ujsx/jpath.gen.ts` — TypeBox `Type.Module` pattern (the `$defs`/`$ref` model BAST follows) - `/workspace/research/typebox_research/ujsx/mdast.gen.ts` — TypeBox cross-module references and composite types - `/workspace/research/typebox_research/codegen/ts-to-module.ts` — TypeScript-to-TypeBox codegen (reference for future BAST codegen) - `/workspace/@alkimiadev/typebox-rs/src/codegen/` — Rust/TypeScript codegen from schemas (reference architecture) - `/workspace/alknet-typedef-poc/tests/sftp_roundtrip_test.rs` — SFTP POC proving byte-identical output (to be replicated with BAST) - `/workspace/@alkdev/alkcall/src/channels/wire.rs` — hand-rolled chunk header (target for BAST replacement)