Files
alktype/docs/research/bast-pivot.md
T
deepseek-v4-pro 82bc8f29c0 Clean up BAST pivot: drop POCs, remove recursion, add spec gaps
- Remove the four proposed POCs: they were implementation smoke tests,
  not de-risking probes. The pivot is a backend swap on a proven layout
  engine; byte-identity is already proven and the layout code is
  unchanged, so there is nothing empirical left to de-risk.
- Remove the recursive TreeNode example and the recursion mention in
  design principle 2: recursion is not a binary-layout concern and the
  engine has no cycle detection.
- Add a Spec Gaps section: validate_bytes semantics after keyword
  validator removal (UnionValidator variant dispatch regression),
  arrays of variable-length elements (engine rejects them), and
  field-name discriminator unions (meta-schema cannot express them).
- Reframe the engine change as an accessor-layer refactor: the
  walkers' (kind, field list, annotations) reads change; everything
  beneath them carries over unchanged.
2026-08-15 07:46:03 +00:00

40 KiB

status, created
status created
draft 2026-08-14

BAST Pivot — Binary Abstract Syntax Tree as the Schema Format

Summary

Replace alktype's custom JSON Schema keywords (AlkType:Uint32, AlkType:Struct, etc.) with a standalone JSON format — BAST (Binary Abstract Syntax Tree) — that describes binary data layouts using a kind-based vocabulary with $defs/$ref for composition. BAST is itself a valid JSON Schema instance (it has a meta-schema), making it self-validating, editor-friendly, and trivially consumable from any language with a JSON parser.

The engine's core logic (layout computation, data access, union dispatch, two layout modes) is unchanged. Only the schema-walking accessor layer changes: instead of detecting AlkType:* keywords scattered through a JSON Schema tree, the walkers read kind/fields/ annotation properties from a purpose-built format.

The builder API's public surface stays the same; only the JSON output format changes internally.

Motivation

Current state

alktype v0.1.0 embeds binary layout information inside standard JSON Schema documents via custom keywords:

{
  "AlkType:Struct": true,
  "type": "object",
  "properties": {
    "channel_id": { "AlkType:Uint32": true, "type": "integer" },
    "length":     { "AlkType:Uint32": true, "type": "integer" }
  },
  "endian": "big"
}

This works for the Rust engine — it walks the tree, detects keywords, computes offsets. But it creates friction for everything outside Rust:

  1. Cross-language consumption. A Python, Go, or TypeScript consumer that wants to parse an alktype schema must re-implement custom keyword detection. The format is not self-describing — you need to know that AlkType:Uint32 means "4-byte unsigned integer" and that it can appear as either true or { "encoding": "..." }.

  2. Code generation. Generating Rust/TypeScript/Python readers and writers from a schema requires walking an arbitrary JSON Schema tree looking for custom keywords. A kind-based format with known keys makes this a straightforward structural walk.

  3. Tooling. Editors, linters, and schema validators don't understand AlkType:* keywords. A BAST document with a published meta-schema gets autocomplete, validation, and documentation in any JSON Schema- aware editor for free.

  4. Two concerns in one document. The current format conflates binary layout (what the engine needs) with JSON validation (what jsonschema needs). A type: "object" with properties and required is a JSON validation concern; AlkType:Uint32 is a binary layout concern. They live in the same JSON object but serve different masters.

The downstream pain is real

The alkcall agent's review identified that the channels 8-byte chunk header is hand-rolled with manual bit shifts — alktype's binary layout capability is unused because the custom-keyword format is awkward to integrate for a simple 2-field struct. alktty plans to hand-roll its 5-byte TTY chunk format for the same reason. SFTP's 29 packet types were proven byte-identical with alktype in the POC, but the production path requires defining 29 schemas in the custom-keyword format.

All three cases are the same pattern: a small binary struct that needs a schema-driven reader/writer. BAST makes this trivial — a 10-line JSON file replaces hand-rolled bit shifts.

Timing

v0.1.0 was published but has zero real consumers (only bots/scanners have downloaded it). A breaking change now is free. Waiting until adoption creates migration cost.

The BAST Format

Design principles

  1. BAST is a JSON Schema instance. A BAST document is valid JSON that conforms to the BAST meta-schema. Any standard JSON Schema validator can validate a BAST document's structure.

  2. $defs/$ref for composition. Named type definitions live in a top-level $defs block. $ref handles cross-references and union variant references. This is the same pattern as TypeBox's Type.Module and JSON Schema's own $defs — no custom reference resolution mechanism needed.

  3. kind-based vocabulary. Every type has a kind field whose value is a known string ("uint32", "struct", "union", etc.). This replaces the AlkType:* custom keyword pattern with a flat, easily-matched string.

  4. Order is explicit. Struct fields are an ordered array, not an object with properties. This makes field order unambiguous (no reliance on serde_json's preserve_order for correctness) and matches the mental model of binary layouts.

  5. Annotations are type-level properties. Endianness, alignment, encoding, and discriminators are properties of the type definition, not custom keywords on a separate schema object.

Examples

Channels chunk header (2-field struct, big-endian)

{
  "$defs": {
    "ChunkHeader": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "channel_id", "kind": "uint32" },
        { "name": "length",     "kind": "uint32" }
      ]
    }
  }
}

TTY chunk (3-field struct, big-endian)

{
  "$defs": {
    "TtyChunk": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "stream_id", "kind": "uint8" },
        { "name": "length",    "kind": "uint32" }
      ]
    }
  }
}

SFTP Read packet (struct with mixed fixed/variable fields)

{
  "$defs": {
    "Read": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",     "kind": "uint32" },
        { "name": "handle", "kind": "string" },
        { "name": "offset", "kind": "uint64" },
        { "name": "len",    "kind": "uint32" }
      ]
    }
  }
}

SFTP Packet union (byte-offset discriminator)

{
  "$defs": {
    "SftpPacket": {
      "kind": "union",
      "endian": "big",
      "discriminator": {
        "kind": "byte",
        "offset": 0,
        "type": "uint8"
      },
      "mapping": {
        "1":   { "$ref": "#/$defs/Init" },
        "3":   { "$ref": "#/$defs/Open" },
        "5":   { "$ref": "#/$defs/Read" },
        "6":   { "$ref": "#/$defs/Write" },
        "101": { "$ref": "#/$defs/Status" }
      }
    },
    "Read": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",     "kind": "uint32" },
        { "name": "handle", "kind": "string" },
        { "name": "offset", "kind": "uint64" },
        { "name": "len",    "kind": "uint32" }
      ]
    },
    "Write": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",     "kind": "uint32" },
        { "name": "handle", "kind": "string" },
        { "name": "offset", "kind": "uint64" },
        { "name": "data",   "kind": "bytes" }
      ]
    },
    "Status": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",            "kind": "uint32" },
        { "name": "status_code",   "kind": "uint32" },
        { "name": "error_message", "kind": "string" },
        { "name": "language_tag",  "kind": "string" }
      ]
    }
  }
}

Metatensor header (aligned mode, little-endian, custom alignment)

{
  "$defs": {
    "TensorHeader": {
      "kind": "struct",
      "endian": "little",
      "align": 256,
      "fields": [
        { "name": "magic",       "kind": "uint64" },
        { "name": "json_length", "kind": "uint64" },
        { "name": "data_offset", "kind": "uint64" }
      ]
    }
  }
}

Array of fixed-size elements with known count

{
  "$defs": {
    "Vector3": {
      "kind": "struct",
      "fields": [
        { "name": "components", "kind": { "kind": "array", "element": "float32", "count": 3 } }
      ]
    }
  }
}

Array of variable-length elements (count-prefixed)

{
  "$defs": {
    "StringList": {
      "kind": "struct",
      "fields": [
        { "name": "items", "kind": { "kind": "array", "element": "string" } }
      ]
    }
  }
}

Record (string-keyed map)

{
  "$defs": {
    "Headers": {
      "kind": "struct",
      "fields": [
        { "name": "entries", "kind": { "kind": "record", "values": "string" } }
      ]
    }
  }
}

Enum

{
  "$defs": {
    "StatusCode": {
      "kind": "enum",
      "values": ["Ok", "PermissionDenied", "NoSuchFile", "Failure"]
    }
  }
}

The BAST meta-schema

A BAST document is valid JSON that conforms to the BAST meta-schema. The meta-schema is a standard JSON Schema (Draft 2020-12) that validates the structure of BAST documents. This means:

  • Any JSON Schema validator can check whether a BAST document is well-formed before the engine compiles it.
  • Editors with JSON Schema support (VSCode, JetBrains) provide autocomplete and inline validation for BAST documents.
  • The format is self-describing — a consumer can inspect the meta-schema to understand the vocabulary without reading Rust source code.

The meta-schema lives at a stable URL (e.g., https://alk.dev/bast/v1/schema) and is embedded in the crate for offline use.

Meta-schema sketch

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://alk.dev/bast/v1/schema",
  "title": "Binary Abstract Syntax Tree (BAST) v1",
  "description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
  "type": "object",
  "properties": {
    "$defs": {
      "type": "object",
      "additionalProperties": { "$ref": "#/$defs/TypeDef" }
    }
  },
  "required": ["$defs"],
  "$defs": {
    "TypeDef": {
      "oneOf": [
        { "$ref": "#/$defs/StructDef" },
        { "$ref": "#/$defs/UnionDef" },
        { "$ref": "#/$defs/EnumDef" }
      ]
    },
    "StructDef": {
      "type": "object",
      "properties": {
        "kind": { "const": "struct" },
        "endian": { "enum": ["little", "big"] },
        "align": { "type": "integer", "minimum": 1 },
        "fields": {
          "type": "array",
          "items": { "$ref": "#/$defs/FieldDef" }
        }
      },
      "required": ["kind", "fields"],
      "additionalProperties": false
    },
    "FieldDef": {
      "type": "object",
      "properties": {
        "name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
        "kind": { "$ref": "#/$defs/TypeRef" },
        "endian": { "enum": ["little", "big"] },
        "align": { "type": "integer", "minimum": 1 },
        "encoding": { "enum": ["length-prefixed", "offset-indirect"] },
        "maxLength": { "type": "integer", "minimum": 0 }
      },
      "required": ["name", "kind"],
      "additionalProperties": false
    },
    "TypeRef": {
      "oneOf": [
        {
          "description": "Primitive type",
          "type": "string",
          "enum": [
            "int8", "int16", "int32", "int64",
            "uint8", "uint16", "uint32", "uint64",
            "float32", "float64",
            "bool", "string", "bytes", "timestamp"
          ]
        },
        {
          "description": "Reference to a named $defs entry",
          "type": "object",
          "properties": {
            "$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
          },
          "required": ["$ref"],
          "additionalProperties": false
        },
        {
          "description": "Array type",
          "type": "object",
          "properties": {
            "kind": { "const": "array" },
            "element": { "$ref": "#/$defs/TypeRef" },
            "count": { "type": "integer", "minimum": 0 }
          },
          "required": ["kind", "element"],
          "additionalProperties": false
        },
        {
          "description": "Record (string-keyed map) type",
          "type": "object",
          "properties": {
            "kind": { "const": "record" },
            "values": { "$ref": "#/$defs/TypeRef" }
          },
          "required": ["kind", "values"],
          "additionalProperties": false
        }
      ]
    },
    "UnionDef": {
      "type": "object",
      "properties": {
        "kind": { "const": "union" },
        "endian": { "enum": ["little", "big"] },
        "discriminator": {
          "oneOf": [
            {
              "type": "object",
              "properties": {
                "kind": { "const": "byte" },
                "offset": { "type": "integer", "minimum": 0 },
                "type": { "enum": ["uint8", "uint16", "uint32"] }
              },
              "required": ["kind", "offset", "type"],
              "additionalProperties": false
            },
            {
              "type": "object",
              "properties": {
                "kind": { "const": "field" },
                "name": { "type": "string" }
              },
              "required": ["kind", "name"],
              "additionalProperties": false
            }
          ]
        },
        "mapping": {
          "type": "object",
          "additionalProperties": { "$ref": "#/$defs/TypeRef" }
        }
      },
      "required": ["kind", "discriminator", "mapping"],
      "additionalProperties": false
    },
    "EnumDef": {
      "type": "object",
      "properties": {
        "kind": { "const": "enum" },
        "values": {
          "type": "array",
          "items": { "type": "string" },
          "minItems": 1
        }
      },
      "required": ["kind", "values"],
      "additionalProperties": false
    }
  }
}

Type reference resolution

TypeRef is the central mechanism for referencing types. It has four forms:

Form Example Meaning
Primitive string "uint32" A built-in primitive type
$ref object { "$ref": "#/$defs/Read" } Reference to a named definition
Array object { "kind": "array", "element": "uint32" } Array of elements
Record object { "kind": "record", "values": "string" } String-keyed map

The $ref form uses standard JSON Pointer syntax restricted to #/$defs/<name>. This is a subset of JSON Schema's $ref — no external references, no fragment-only pointers, no bare names. The restriction keeps resolution simple (single hash lookup) and avoids the normalization step that the current engine needs for TypeBox's bare-name refs.

Arrays and records are inline type constructors, not top-level $defs entries. This keeps the common cases concise while allowing complex element types via nested $ref:

{ "kind": "array", "element": { "$ref": "#/$defs/ComplexElement" } }

Variable-length encoding

The three strategies from ADR-003 carry forward with the same semantics, expressed as field-level properties instead of keyword-value objects:

Strategy BAST syntax Behavior
Inline length-prefixed (default) { "name": "handle", "kind": "string" } [u32 length][data]
Fixed-size reservation { "name": "name", "kind": "string", "maxLength": 256 } Reserve maxLength bytes (aligned mode); validation constraint (packed mode)
Offset indirection { "name": "blob", "kind": "bytes", "encoding": "offset-indirect" } {offset: u32, length: u32} pointing to separate data region

Endianness

Endianness is a struct-level or union-level property with per-field override, same as ADR-003:

  • Struct-level "endian" sets the default for all fields.
  • Field-level "endian" overrides the struct default.
  • Default is "little" when neither is specified.
  • The length prefix for variable-length fields respects the effective endianness (struct default or field override).
{
  "kind": "struct",
  "endian": "big",
  "fields": [
    { "name": "id",     "kind": "uint32" },
    { "name": "handle", "kind": "string" },
    { "name": "offset", "kind": "uint64" },
    { "name": "crc",    "kind": "uint32", "endian": "little" }
  ]
}

Alignment

Alignment is a struct-level or field-level property, only meaningful in aligned static mode (same as ADR-003):

{
  "kind": "struct",
  "align": 256,
  "fields": [
    { "name": "header", "kind": { "$ref": "#/$defs/Header" } },
    { "name": "weight", "kind": "float32", "align": 16 }
  ]
}

What Changes

JSON format

Aspect Current (custom keywords) BAST
Type declaration "AlkType:Uint32": true on a property "kind": "uint32" in a field definition
Struct fields "properties": { "x": {...}, "y": {...} } "fields": [{ "name": "x", ... }, { "name": "y", ... }]
Field order Implicit via serde_json preserve_order Explicit via array position
Endianness "endian": "big" on the schema object "endian": "big" on the struct/union definition
Variable encoding "AlkType:String": { "encoding": "offset-indirect" } { "name": "x", "kind": "string", "encoding": "offset-indirect" }
Discriminator "discriminator": { "kind": "byte", "offset": 0, "type": "AlkType:Uint8" } "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" }
$ref Bare names ("Read") normalized to #/$defs/Read Always #/$defs/Read — no normalization needed
Top-level container Schema object with AlkType:Struct at root { "$defs": { ... } } with a root type name
JSON validation info "type": "object", "required", etc. on the same object Separate concern — not in BAST

The engine change is an accessor-layer refactor, not a rewrite. The schema walkers currently read (kind, field list, annotations) out of JSON nodes via custom-keyword accessors; under BAST they read the same information from kind/fields/annotation properties. Everything beneath the accessors — checked offset arithmetic, union dispatch, materialization, the two layout modes — is format-agnostic and carries over unchanged.

Engine internals

Component Change
get_alktype_kind() / get_alktype_kind_loose() Replaced by parsing "kind" field directly from BAST nodes
normalize_refs() Removed — BAST $ref values are always full JSON Pointers
inline_union_variant_refs() Removed — union mapping values are resolved lazily via $ref
parse_endian(), parse_align(), parse_encoding(), parse_discriminator() Adapted to read from BAST field/struct properties instead of keyword-value objects
Custom keyword validators (19 jsonschema::Keyword impls) Removed. These were the only thing using jsonschema's custom keyword API.
build_validator() Repurposed. Still builds a jsonschema::Validator, but from a standard JSON Schema (no custom keywords). Used for JSON payload validation (call's OperationSpec schemas, etc.) and for validating BAST documents against the BAST meta-schema.
validate_json() Unchanged in signature. Validates a Value against the engine's compiled jsonschema::Validator (now built from a standard JSON Schema instead of one with custom keywords).
validate_bytes() Unchanged in concept — materialize Value from bytes, then validate. The materialization step walks BAST instead of custom-keyword JSON.

Public API

Item Change
AlkTypeKind enum Unchanged — same 19 variants, same methods
Endian, VariableEncoding, DiscriminatorKind Unchanged
AlkTypeEngine compile() takes a BAST document + root type name instead of a custom-keyword JSON Schema. validate_json() may change. validate_bytes() unchanged.
LayoutMode, OffsetMap, ByteRange Unchanged
LayoutBuilder, PackedLayout, FieldPosition Unchanged
SequentialReader, FieldValue Unchanged
UnionDispatch Unchanged
data_access functions Unchanged
AlkTypeError Unchanged (Schema/Offset/Access/Validation variants)
Schema builder Public methods unchanged. build() produces BAST JSON instead of custom-keyword JSON.
Definitions builder Public methods unchanged. build() produces a BAST $defs block.
Discriminator builder Unchanged

What is removed

  • All 19 jsonschema::Keyword implementations (~200 lines of validator factories)
  • normalize_refs() — BAST $ref values are always full JSON Pointers
  • inline_union_variant_refs() — union mapping values are resolved lazily
  • get_alktype_kind() / get_alktype_kind_loose() and their _enum variants — replaced by direct kind field parsing
  • Custom keyword registration with jsonschema — the engine no longer calls jsonschema::options().with_keyword(...). The jsonschema crate remains a direct dependency for standard JSON Schema validation (validating JSON payloads like call's OperationSpec schemas, and validating BAST documents against the BAST meta-schema). The only thing removed is the custom keyword integration path.

What is added

  • BAST meta-schema (embedded in the crate, published at a stable URL)
  • BAST document parser — validates a BAST document against the meta-schema, then extracts type definitions
  • AlkTypeEngine::compile() takes (bast_document: &Value, root_name: &str, mode: LayoutMode) — the root name selects which $defs entry is the top-level type
  • BAST meta-schema validation at compile time (optional but recommended — the engine can skip it and trust the caller, or validate as a guard)

The Validator Split

A key architectural clarification: BAST separates two concerns that the current format conflates.

Current model (conflated)

┌─────────────────────────────────────────────┐
│  JSON Schema with AlkType:* custom keywords  │
│  ┌───────────────────────────────────────┐  │
│  │  Binary layout info (AlkType:Uint32)  │  │
│  │  JSON validation info (type, enum)    │  │
│  └───────────────────────────────────────┘  │
└─────────────────────────────────────────────┘
         │
         ▼
  AlkTypeEngine::compile()
         │
         ├──► Layout (offset map / layout builder)
         └──► Validator (jsonschema with custom keywords)

One document serves two roles. The engine extracts both layout and validation from the same JSON tree.

BAST model (separated)

┌──────────────────────┐     ┌──────────────────────────┐
│   BAST document       │     │  JSON Schema document     │
│   (binary layout)     │     │  (JSON validation)        │
│                       │     │                          │
│  kind: "struct"       │     │  type: "object"          │
│  fields: [            │     │  properties: {           │
│    { name, kind }     │     │    id: { type: "integer"} │
│  ]                    │     │  }                       │
│  endian: "big"        │     │  required: ["id"]         │
└──────────┬───────────┘     └────────────┬─────────────┘
           │                              │
           ▼                              ▼
  AlkTypeEngine::compile()      build_validator(&json_schema)
           │                              │
           ▼                              ▼
  Layout (offset map /          jsonschema::Validator
  layout builder)               (standard JSON Schema)
           │                              │
           ▼                              ▼
  validate_bytes(&[u8])         validate_json(&Value)

Two documents, two validators, two concerns. The BAST document describes binary layout. The JSON Schema document describes JSON data shape. They can be linked (a BAST struct can reference a JSON Schema by $id for validation) but they are separate documents.

Both paths use jsonschema under the hood — the BAST path uses it to validate BAST documents against the BAST meta-schema at compile time; the JSON path uses it to validate JSON payloads against standard JSON Schema documents. The only thing removed is the custom keyword registration path (jsonschema::options().with_keyword(...)).

What this means for consumers

alkcall today uses alktype for two roles:

  1. Binary layout (channels chunk header) — AlkTypeEngine::compile() in packed mode
  2. JSON validation (call's OperationSpec schemas) — build_validator() or jsonschema directly

Under BAST:

  1. Binary layout — AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed) — same flow, different input format
  2. JSON validation — build_validator(&json_schema) — same API, now builds a standard jsonschema::Validator (no custom keywords). Consumers that don't use binary at all (e.g., adapters around remote JSON Schema APIs) use this path exclusively.

The builder API can produce both formats:

  • Schema::struct_().field(...).build() → BAST JSON
  • Schema::object().field(...).build() → standard JSON Schema

Validation of binary data

validate_bytes() still works: materialize a Value tree from binary bytes using the BAST layout, then validate that Value against a JSON Schema. The JSON Schema can be:

  • Derived from the BAST definition (a codegen step, future)
  • Provided separately by the consumer
  • A standard JSON Schema that the consumer already has for JSON validation of the same logical type

The materialization step walks the BAST layout (unchanged from current behavior — it walks the schema to compute offsets and read fields). The validation step uses a standard jsonschema::Validator (no custom keywords needed — the Value tree is already typed by the materialization).

Relationship to JSON Schema and TypeBox

BAST is a JSON Schema dialect

BAST is a specific JSON Schema instance format — like how JSON Schema itself is a JSON document that conforms to the JSON Schema meta-schema. BAST documents conform to the BAST meta-schema. The meta-schema is a standard JSON Schema (Draft 2020-12).

This means the entire JSON Schema tooling ecosystem works with BAST:

  • Validation: jsonschema::options().build(&bast_meta_schema)?.validate(&bast_doc)
  • Editors: VSCode with $schema pointing to the BAST meta-schema URL
  • Documentation: JSON Schema generators can produce human-readable docs from the meta-schema

TypeBox interop

TypeBox's Type.Module({...}) pattern maps naturally to BAST's $defs structure. A TypeBox module that defines binary types can serialize to BAST JSON instead of custom-keyword JSON. The codegen (ts-to-module.ts) could target BAST as an output format.

The relationship is:

  • TypeBox → BAST JSON → alktype engine (binary layout)
  • TypeBox → standard JSON Schema → jsonschema (JSON validation)

Same TypeBox source, two output formats, two validators.

Not a replacement for JSON Schema

BAST does not replace JSON Schema for JSON data validation. A BAST document cannot validate a JSON payload. It describes binary data layouts. For JSON validation, consumers use standard JSON Schema documents (which may be derived from BAST definitions via codegen, or authored separately).

Codegen (Future)

BAST enables code generation that the custom-keyword format makes awkward. A codegen module (feature-gated behind codegen) would:

  1. Input: A BAST document (or SchemaRegistry equivalent)
  2. Walk: Iterate $defs entries, inspect kind values
  3. Map: "uint32" → u32 (Rust), number (TypeScript), int (Python)
  4. Emit: Handlebars templates for struct/enum/union definitions

Generated artifacts

Target Type definitions Binary reader/writer
Rust struct ChunkHeader { channel_id: u32, length: u32 } fn read_header(buf: &[u8]) -> Result<ChunkHeader, Error>
TypeScript interface ChunkHeader { channelId: number; length: number } function readHeader(buf: Uint8Array): ChunkHeader
Python @dataclass class ChunkHeader: ... def read_header(buf: bytes) -> ChunkHeader

Relationship to typebox-rs codegen

The typebox-rs codegen/ module (in /workspace/@alkimiadev/typebox-rs) is the reference architecture:

  • SchemaRegistry for named types with $ref resolution
  • RustGenerator / TypeScriptGenerator wrapping Handlebars templates
  • schema_to_rust_type() / schema_to_ts_type() mapping functions
  • Feature-gated behind codegen = ["handlebars"]

alktype's codegen would follow the same pattern but walk BAST kind values instead of SchemaKind enum variants. The handlebars-rs dependency is WASM-compatible.

Scope boundary

Codegen is out of scope for the BAST pivot itself. The pivot changes the schema format; codegen builds on top of the new format. It is described here to show that BAST enables it, not to commit to a specific implementation timeline.

ABI Adapter (Future)

A BAST document describes the binary interface of a protocol — it is essentially an ABI specification in JSON. This enables:

  • Version negotiation: Two peers exchange BAST documents to agree on a protocol version. The engine can detect mismatches (field added, type changed, endianness differs) and either reject or adapt.
  • Schema migration: A consumer with schema v1 can read data written by schema v2 if the changes are compatible (fields added at the end, types widened). The engine can compute a migration plan from the diff of two BAST documents.
  • WASM interop: A WASM component can export its BAST schema as part of its WIT interface, enabling host languages to generate readers/writers for the component's binary protocol without manual bindings.

This is a future capability, not part of the pivot. It is mentioned because BAST makes it possible in a way that custom keywords scattered through JSON Schema trees do not.

Migration Path

Phase 1: BAST format and meta-schema (this pivot)

  1. Define the BAST meta-schema (the JSON Schema that validates BAST documents)
  2. Implement BAST document parsing in the engine (replace custom keyword detection with kind field parsing)
  3. Update AlkTypeEngine::compile() to accept a BAST document + root type name
  4. Update the builder API to produce BAST JSON (public methods unchanged)
  5. Remove custom keyword validators, normalize_refs(), inline_union_variant_refs(), and the get_alktype_kind* functions
  6. Update all tests to use BAST format
  7. Update architecture docs (ADRs, schema-layer.md, etc.)

Phase 2: Downstream adoption

  1. Port alkcall's chunk header to BAST (replace hand-rolled wire.rs with AlkTypeEngine + SequentialReader/LayoutBuilder)
  2. Port alktty's TTY chunk format to BAST
  3. Port the SFTP POC schemas to BAST format

Phase 3: Codegen (future)

  1. Add codegen feature flag with handlebars dependency
  2. Implement RustGenerator and TypeScriptGenerator walking BAST kind values
  3. External Handlebars templates in src/codegen/templates/

Spec Gaps

No POCs are needed for this pivot. It is a backend swap (custom keywords → kind-based format) on top of a proven layout engine, not new protocol invention. The layout engine's byte-identity is already proven (alknet-typedef-poc, alktype-builder-poc) and the layout code is unchanged, so there is nothing empirical left to de-risk. What remains are three spec gaps where the draft promises more than the engine delivers, or expresses less than the engine supports. Each is a decision, not an unknown.

Gap 1: validate_bytes semantics after keyword validator removal

The current engine's UnionValidator dispatches to variant schemas at validation time (OQ-008), so validate_bytes on a union checks variant field constraints (e.g. maxLength on a Bytes field inside a variant). Removing the 19 custom keyword validators deletes this dispatch. The spec must state what validate_bytes validates:

  • Structural readability only (bounds, discriminator lookup, UTF-8) — constraint checks move to a separately-provided JSON Schema
  • Full constraint validation via a JSON Schema the consumer provides (OQ-BAST-006's leaning) — but variant constraints then require the consumer to hand-author __discriminator dispatch, a regression from today's behavior

Gap 2: Arrays of variable-length elements

The example in §"The BAST Format" shows { "kind": "array", "element": "string" } (count-prefixed). The engine explicitly rejects arrays of variable-length elements ("TArray of variable-length element kind ... is not supported"). Either the example is out of scope for v1 (remove it and note the restriction in the meta-schema), or variable-element arrays are a new feature with their own layout semantics.

Gap 3: Field-name discriminator unions

The meta-schema's UnionDef has no fields array, but the engine's field-name discriminator reads the discriminator field from the union's properties object. As sketched, BAST cannot express a feature the engine already supports. The meta-schema needs a field-name- discriminator shape (e.g. a fields array on UnionDef), or field-name discriminators are dropped for v1.

Open Questions

OQ-BAST-001: Root type selection

How does the engine know which $defs entry is the root type? Options:

  • Explicit: AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed)
  • Convention: the first entry in $defs (fragile — depends on JSON key order)
  • Marker: a "$root": "ChunkHeader" property on the BAST document

Leaning: Explicit. The root type name is a required parameter to compile(). This is unambiguous and matches how consumers think about it ("compile the ChunkHeader schema").

OQ-BAST-002: JSON validation schema linkage

How does a BAST struct reference a JSON Schema for validate_bytes? Options:

  • Separate parameter: engine.validate_bytes(buf, Some(&json_schema))
  • Embedded reference: { "kind": "struct", "validation": { "$ref": "https://..." } }
  • Convention: same $id base, different fragment

Leaning: Separate parameter for v1. The JSON Schema is a separate document; the engine doesn't need to know about it at compile time. validate_bytes() accepts an optional &jsonschema::Validator that the consumer provides. This keeps BAST focused on binary layout.

OQ-BAST-003: Primitive type string set

The current proposal uses lowercase strings for primitives: "uint32", "int8", "float64", "bool", "string", "bytes", "timestamp". Alternatives:

  • PascalCase: "Uint32", "Int8" (matches AlkTypeKind variant names)
  • UPPER_CASE: "UINT32", "INT8"
  • Prefixed: "bast:uint32" (namespaced, but verbose)

Leaning: Lowercase. Matches JSON Schema's own convention ("string", "integer", "boolean"), is easier to type, and is the convention in the TypeBox research examples. The AlkTypeKind enum variants remain PascalCase in Rust — the mapping is a simple from_str() impl.

OQ-BAST-004: Array count — fixed vs variable

When is an array fixed-size (no count prefix) vs variable-size (count prefix in binary)? Options:

  • Explicit count field: { "kind": "array", "element": "uint32", "count": 3 } → fixed
  • Absent count: { "kind": "array", "element": "uint32" } → variable (count-prefixed)
  • Separate minItems/maxItems like current JSON Schema convention

Leaning: Explicit count for fixed, absent for variable. This is clearer than minItems == maxItems and matches the BAST principle of explicit layout information.

OQ-BAST-005: Top-level $defs requirement

Should a BAST document always have a top-level $defs block, or can a single struct be the root? Options:

  • Always $defs: { "$defs": { "ChunkHeader": { "kind": "struct", ... } } }
  • Bare struct: { "kind": "struct", "fields": [...] } (no $defs)

Leaning: Always $defs. Consistency — every BAST document has the same top-level shape. Single-type documents are a special case of the general form. The $defs block is the namespace; the root type name selects the entry point.

OQ-BAST-006: validate_json on AlkTypeEngine

With custom keyword validators removed, what does AlkTypeEngine::validate_json() do? Options:

  • Remove it — the engine is for binary layout; JSON validation is a separate concern
  • Keep it with a separately-provided JSON Schema — the engine holds a jsonschema::Validator compiled from a JSON Schema the consumer provides at compile time
  • Keep it as a convenience that validates the materialized Value against the BAST structure itself (type checks only, no range constraints)

Leaning: Keep it with a separately-provided JSON Schema. The engine already compiles a validator at load time (ADR-004). The validator just comes from a standard JSON Schema instead of custom keywords. This preserves the validate_json / validate_bytes symmetry from ADR-010.

OQ-BAST-007: Builder API — two output formats

The builder API currently produces one JSON format (custom keywords). Under BAST, it needs to produce two:

  1. BAST JSON (for Schema::struct_(), Schema::uint32(), etc.)
  2. Standard JSON Schema (for Schema::object(), Schema::string_(), etc.)

Should these be two separate builder types, or one builder with a mode flag? Options:

  • Two builders: BastBuilder and JsonSchemaBuilder (or Schema::bast and Schema::json)
  • One builder with mode: Schema::new(Mode::Bast) / Schema::new(Mode::JsonSchema)
  • One builder, two build methods: schema.build_bast() / schema.build_json_schema()

Leaning: One builder, two build methods. The construction API is the same (field names, types, annotations); only the output format differs. Schema::struct_().field(...).build() → BAST. Schema::object().field(...).build() → standard JSON Schema. The builder already distinguishes AlkType kinds from JSON Schema types via naming conventions (string() vs string_()).

Risks and Mitigations

Risk Mitigation
BAST format doesn't cover all 19 type kinds The format is designed to cover all 19. The meta-schema is the spec — if a kind can't be expressed, the meta-schema is wrong.
$ref resolution complexity moves from engine to schema authoring BAST $ref values are always full JSON Pointers (#/$defs/Name). No normalization, no bare names. Resolution is a single hash lookup.
Losing jsonschema's structural validation (required fields, etc.) BAST is for binary layout, not JSON validation. Structural constraints belong in the JSON Schema document, not the BAST document.
Builder API output format change breaks consumers No real consumers exist yet (v0.1.0 has zero adoption). The builder's public methods are unchanged; only the JSON output format changes.
Meta-schema maintenance burden The meta-schema is small (~150 lines) and changes rarely. It's embedded in the crate and published at a stable URL.

References

  • ADR-001 — current "schema is the format" principle (to be updated)
  • ADR-003 — annotation shapes (endianness, alignment, encoding, discriminators — carry forward to BAST)
  • ADR-009 — builder API (public surface unchanged, output format changes)
  • ADR-010 — validate_bytes (unchanged in concept)
  • /workspace/research/typebox_research/ujsx/jpath.gen.ts — TypeBox Type.Module pattern (the $defs/$ref model BAST follows)
  • /workspace/research/typebox_research/ujsx/mdast.gen.ts — TypeBox cross-module references and composite types
  • /workspace/research/typebox_research/codegen/ts-to-module.ts — TypeScript-to-TypeBox codegen (reference for future BAST codegen)
  • /workspace/@alkimiadev/typebox-rs/src/codegen/ — Rust/TypeScript codegen from schemas (reference architecture)
  • /workspace/alknet-typedef-poc/tests/sftp_roundtrip_test.rs — SFTP POC proving byte-identical output (to be replicated with BAST)
  • /workspace/@alkdev/alkcall/src/channels/wire.rs — hand-rolled chunk header (target for BAST replacement)