Files
alktype/docs/research/bast-pivot.md
T
glm-5.2 30b01c1696 Sync BAST pivot doc with main (semver/ADR impact mapping)
Brings the 'Semver and ADR Impact' section from main onto this branch
so the two doc copies don't drift. No code change.
2026-08-15 10:29:27 +00:00

72 KiB

status, created
status created
draft 2026-08-14

BAST Pivot — Binary Abstract Syntax Tree as the Schema Format

Summary

Replace alktype's custom JSON Schema keywords (AlkType:Uint32, AlkType:Struct, etc.) with a standalone JSON format — BAST (Binary Abstract Syntax Tree) — that describes binary data layouts using a kind-based vocabulary with $defs/$ref for composition. BAST is itself a valid JSON Schema instance (it has a meta-schema), making it self-validating, editor-friendly, and trivially consumable from any language with a JSON parser.

The engine's core logic (layout computation, data access, union dispatch, two layout modes) is unchanged. Only the schema-walking accessor layer changes: instead of detecting AlkType:* keywords scattered through a JSON Schema tree, the walkers read kind/fields/ annotation properties from a purpose-built format.

The builder API's public surface stays the same; only the JSON output format changes internally.

Motivation

Current state

alktype v0.1.0 embeds binary layout information inside standard JSON Schema documents via custom keywords:

{
  "AlkType:Struct": true,
  "type": "object",
  "properties": {
    "channel_id": { "AlkType:Uint32": true, "type": "integer" },
    "length":     { "AlkType:Uint32": true, "type": "integer" }
  },
  "endian": "big"
}

This works for the Rust engine — it walks the tree, detects keywords, computes offsets. But it creates friction for everything outside Rust:

  1. Cross-language consumption. A Python, Go, or TypeScript consumer that wants to parse an alktype schema must re-implement custom keyword detection. The format is not self-describing — you need to know that AlkType:Uint32 means "4-byte unsigned integer" and that it can appear as either true or { "encoding": "..." }.

  2. Code generation. Generating Rust/TypeScript/Python readers and writers from a schema requires walking an arbitrary JSON Schema tree looking for custom keywords. A kind-based format with known keys makes this a straightforward structural walk.

  3. Tooling. Editors, linters, and schema validators don't understand AlkType:* keywords. A BAST document with a published meta-schema gets autocomplete, validation, and documentation in any JSON Schema- aware editor for free.

  4. Two concerns in one document. The current format conflates binary layout (what the engine needs) with JSON validation (what jsonschema needs). A type: "object" with properties and required is a JSON validation concern; AlkType:Uint32 is a binary layout concern. They live in the same JSON object but serve different masters.

The downstream pain is real

The alkcall agent's review identified that the channels 8-byte chunk header is hand-rolled with manual bit shifts — alktype's binary layout capability is unused because the custom-keyword format is awkward to integrate for a simple 2-field struct. alktty plans to hand-roll its 5-byte TTY chunk format for the same reason. SFTP's 29 packet types were proven byte-identical with alktype in the POC, but the production path requires defining 29 schemas in the custom-keyword format.

All three cases are the same pattern: a small binary struct that needs a schema-driven reader/writer. BAST makes this trivial — a 10-line JSON file replaces hand-rolled bit shifts.

Timing

v0.1.0 was published but has zero real consumers (only bots/scanners have downloaded it). A breaking change now is free. Waiting until adoption creates migration cost.

The BAST Format

Design principles

  1. BAST is a JSON Schema instance. A BAST document is valid JSON that conforms to the BAST meta-schema. Any standard JSON Schema validator can validate a BAST document's structure.

  2. $defs/$ref for composition. Named type definitions live in a top-level $defs block. $ref handles cross-references and union variant references. This is the same pattern as TypeBox's Type.Module and JSON Schema's own $defs — no custom reference resolution mechanism needed.

  3. kind-based vocabulary. Every type has a kind field whose value is a known string ("uint32", "struct", "union", etc.). This replaces the AlkType:* custom keyword pattern with a flat, easily-matched string.

  4. Order is explicit. Struct fields are an ordered array, not an object with properties. This makes field order unambiguous (no reliance on serde_json's preserve_order for correctness) and matches the mental model of binary layouts.

  5. Annotations are type-level properties. Endianness, alignment, encoding, and discriminators are properties of the type definition, not custom keywords on a separate schema object.

Examples

Channels chunk header (2-field struct, big-endian)

{
  "$defs": {
    "ChunkHeader": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "channel_id", "kind": "uint32" },
        { "name": "length",     "kind": "uint32" }
      ]
    }
  }
}

TTY chunk (3-field struct, big-endian)

{
  "$defs": {
    "TtyChunk": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "stream_id", "kind": "uint8" },
        { "name": "length",    "kind": "uint32" }
      ]
    }
  }
}

SFTP Read packet (struct with mixed fixed/variable fields)

{
  "$defs": {
    "Read": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",     "kind": "uint32" },
        { "name": "handle", "kind": "string" },
        { "name": "offset", "kind": "uint64" },
        { "name": "len",    "kind": "uint32" }
      ]
    }
  }
}

SFTP Packet union (byte-offset discriminator)

{
  "$defs": {
    "SftpPacket": {
      "kind": "union",
      "endian": "big",
      "discriminator": {
        "kind": "byte",
        "offset": 0,
        "type": "uint8"
      },
      "mapping": {
        "1":   { "$ref": "#/$defs/Init" },
        "3":   { "$ref": "#/$defs/Open" },
        "5":   { "$ref": "#/$defs/Read" },
        "6":   { "$ref": "#/$defs/Write" },
        "101": { "$ref": "#/$defs/Status" }
      }
    },
    "Read": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",     "kind": "uint32" },
        { "name": "handle", "kind": "string" },
        { "name": "offset", "kind": "uint64" },
        { "name": "len",    "kind": "uint32" }
      ]
    },
    "Write": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",     "kind": "uint32" },
        { "name": "handle", "kind": "string" },
        { "name": "offset", "kind": "uint64" },
        { "name": "data",   "kind": "bytes" }
      ]
    },
    "Status": {
      "kind": "struct",
      "endian": "big",
      "fields": [
        { "name": "id",            "kind": "uint32" },
        { "name": "status_code",   "kind": "uint32" },
        { "name": "error_message", "kind": "string" },
        { "name": "language_tag",  "kind": "string" }
      ]
    }
  }
}

Metatensor header (aligned mode, little-endian, custom alignment)

{
  "$defs": {
    "TensorHeader": {
      "kind": "struct",
      "endian": "little",
      "align": 256,
      "fields": [
        { "name": "magic",       "kind": "uint64" },
        { "name": "json_length", "kind": "uint64" },
        { "name": "data_offset", "kind": "uint64" }
      ]
    }
  }
}

Array of fixed-size elements with known count

{
  "$defs": {
    "Vector3": {
      "kind": "struct",
      "fields": [
        { "name": "components", "kind": { "kind": "array", "element": "float32", "count": 3 } }
      ]
    }
  }
}

Array of variable-length elements (deferred — see Decisions)

Arrays of variable-length elements (e.g., { "kind": "array", "element": "string" } without a count) are not supported in v1. The engine rejects them today (OQ-001), and the meta-schema requires count on all array types. This is a known limitation, not a gap — see D-BAST-004.

Record (string-keyed map)

{
  "$defs": {
    "Headers": {
      "kind": "struct",
      "fields": [
        { "name": "entries", "kind": { "kind": "record", "values": "string" } }
      ]
    }
  }
}

Enum

{
  "$defs": {
    "StatusCode": {
      "kind": "enum",
      "values": ["Ok", "PermissionDenied", "NoSuchFile", "Failure"]
    }
  }
}

The BAST meta-schema

A BAST document is valid JSON that conforms to the BAST meta-schema. The meta-schema is a standard JSON Schema (Draft 2020-12) that validates the structure of BAST documents. This means:

  • Any JSON Schema validator can check whether a BAST document is well-formed before the engine compiles it.
  • Editors with JSON Schema support (VSCode, JetBrains) provide autocomplete and inline validation for BAST documents.
  • The format is self-describing — a consumer can inspect the meta-schema to understand the vocabulary without reading Rust source code.

The meta-schema lives at a stable URL (e.g., https://alk.dev/bast/v1/schema) and is embedded in the crate for offline use.

Meta-schema sketch

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://alk.dev/bast/v1/schema",
  "title": "Binary Abstract Syntax Tree (BAST) v1",
  "description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
  "type": "object",
  "properties": {
    "$defs": {
      "type": "object",
      "additionalProperties": { "$ref": "#/$defs/TypeDef" }
    }
  },
  "required": ["$defs"],
  "$defs": {
    "TypeDef": {
      "oneOf": [
        { "$ref": "#/$defs/StructDef" },
        { "$ref": "#/$defs/UnionDef" },
        { "$ref": "#/$defs/EnumDef" }
      ]
    },
    "StructDef": {
      "type": "object",
      "properties": {
        "kind": { "const": "struct" },
        "endian": { "enum": ["little", "big"] },
        "align": { "type": "integer", "minimum": 1 },
        "fields": {
          "type": "array",
          "items": { "$ref": "#/$defs/FieldDef" }
        }
      },
      "required": ["kind", "fields"],
      "additionalProperties": false
    },
    "FieldDef": {
      "type": "object",
      "properties": {
        "name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
        "kind": { "$ref": "#/$defs/TypeRef" },
        "endian": { "enum": ["little", "big"] },
        "align": { "type": "integer", "minimum": 1 },
        "encoding": { "enum": ["length-prefixed", "offset-indirect"] },
        "maxLength": { "type": "integer", "minimum": 0 }
      },
      "required": ["name", "kind"],
      "additionalProperties": false
    },
    "TypeRef": {
      "oneOf": [
        {
          "description": "Primitive type",
          "type": "string",
          "enum": [
            "int8", "int16", "int32", "int64",
            "uint8", "uint16", "uint32", "uint64",
            "float32", "float64",
            "bool", "string", "bytes", "timestamp"
          ]
        },
        {
          "description": "Reference to a named $defs entry",
          "type": "object",
          "properties": {
            "$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
          },
          "required": ["$ref"],
          "additionalProperties": false
        },
        {
          "description": "Array type (fixed-size only in v1 — count is required)",
          "type": "object",
          "properties": {
            "kind": { "const": "array" },
            "element": { "$ref": "#/$defs/TypeRef" },
            "count": { "type": "integer", "minimum": 0 }
          },
          "required": ["kind", "element", "count"],
          "additionalProperties": false
        },
        {
          "description": "Record (string-keyed map) type",
          "type": "object",
          "properties": {
            "kind": { "const": "record" },
            "values": { "$ref": "#/$defs/TypeRef" }
          },
          "required": ["kind", "values"],
          "additionalProperties": false
        }
      ]
    },
    "UnionDef": {
      "type": "object",
      "properties": {
        "kind": { "const": "union" },
        "endian": { "enum": ["little", "big"] },
        "fields": {
          "type": "array",
          "items": { "$ref": "#/$defs/FieldDef" }
        },
        "discriminator": {
          "oneOf": [
            {
              "type": "object",
              "properties": {
                "kind": { "const": "byte" },
                "offset": { "type": "integer", "minimum": 0 },
                "type": { "enum": ["uint8", "uint16", "uint32"] }
              },
              "required": ["kind", "offset", "type"],
              "additionalProperties": false
            },
            {
              "type": "object",
              "properties": {
                "kind": { "const": "field" },
                "name": { "type": "string" }
              },
              "required": ["kind", "name"],
              "additionalProperties": false
            }
          ]
        },
        "mapping": {
          "type": "object",
          "additionalProperties": { "$ref": "#/$defs/TypeRef" }
        }
      },
      "required": ["kind", "discriminator", "mapping"],
      "additionalProperties": false
    },
    "EnumDef": {
      "type": "object",
      "properties": {
        "kind": { "const": "enum" },
        "values": {
          "type": "array",
          "items": { "type": "string" },
          "minItems": 1
        }
      },
      "required": ["kind", "values"],
      "additionalProperties": false
    }
  }
}

Type reference resolution

TypeRef is the central mechanism for referencing types. It has four forms:

Form Example Meaning
Primitive string "uint32" A built-in primitive type
$ref object { "$ref": "#/$defs/Read" } Reference to a named definition
Array object { "kind": "array", "element": "uint32" } Array of elements
Record object { "kind": "record", "values": "string" } String-keyed map

The $ref form uses standard JSON Pointer syntax restricted to #/$defs/<name>. This is a subset of JSON Schema's $ref — no external references, no fragment-only pointers, no bare names. The restriction keeps resolution simple (single hash lookup) and avoids the normalization step that the current engine needs for TypeBox's bare-name refs.

Arrays and records are inline type constructors, not top-level $defs entries. This keeps the common cases concise while allowing complex element types via nested $ref:

{ "kind": "array", "element": { "$ref": "#/$defs/ComplexElement" } }

Variable-length encoding

The three strategies from ADR-003 carry forward with the same semantics, expressed as field-level properties instead of keyword-value objects:

Strategy BAST syntax Behavior
Inline length-prefixed (default) { "name": "handle", "kind": "string" } [u32 length][data]
Fixed-size reservation { "name": "name", "kind": "string", "maxLength": 256 } Reserve maxLength bytes (aligned mode); validation constraint (packed mode)
Offset indirection { "name": "blob", "kind": "bytes", "encoding": "offset-indirect" } {offset: u32, length: u32} pointing to separate data region

Endianness

Endianness is a struct-level or union-level property with per-field override, same as ADR-003:

  • Struct-level "endian" sets the default for all fields.
  • Field-level "endian" overrides the struct default.
  • Default is "little" when neither is specified.
  • The length prefix for variable-length fields respects the effective endianness (struct default or field override).
{
  "kind": "struct",
  "endian": "big",
  "fields": [
    { "name": "id",     "kind": "uint32" },
    { "name": "handle", "kind": "string" },
    { "name": "offset", "kind": "uint64" },
    { "name": "crc",    "kind": "uint32", "endian": "little" }
  ]
}

Alignment

Alignment is a struct-level or field-level property, only meaningful in aligned static mode (same as ADR-003):

{
  "kind": "struct",
  "align": 256,
  "fields": [
    { "name": "header", "kind": { "$ref": "#/$defs/Header" } },
    { "name": "weight", "kind": "float32", "align": 16 }
  ]
}

What Changes

JSON format

Aspect Current (custom keywords) BAST
Type declaration "AlkType:Uint32": true on a property "kind": "uint32" in a field definition
Struct fields "properties": { "x": {...}, "y": {...} } "fields": [{ "name": "x", ... }, { "name": "y", ... }]
Field order Implicit via serde_json preserve_order Explicit via array position
Endianness "endian": "big" on the schema object "endian": "big" on the struct/union definition
Variable encoding "AlkType:String": { "encoding": "offset-indirect" } { "name": "x", "kind": "string", "encoding": "offset-indirect" }
Discriminator "discriminator": { "kind": "byte", "offset": 0, "type": "AlkType:Uint8" } "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" }
$ref Bare names ("Read") normalized to #/$defs/Read Always #/$defs/Read — no normalization needed
Top-level container Schema object with AlkType:Struct at root { "$defs": { ... } } with a root type name
JSON validation info "type": "object", "required", etc. on the same object Separate concern — not in BAST

The engine change is an accessor-layer refactor, not a rewrite. The schema walkers currently read (kind, field list, annotations) out of JSON nodes via custom-keyword accessors; under BAST they read the same information from kind/fields/annotation properties. Everything beneath the accessors — checked offset arithmetic, union dispatch, materialization, the two layout modes — is format-agnostic and carries over unchanged.

Engine internals

Component Change
get_alktype_kind() / get_alktype_kind_loose() Replaced by parsing "kind" field directly from BAST nodes
normalize_refs() Removed — BAST $ref values are always full JSON Pointers
inline_union_variant_refs() Removed — union mapping values are resolved lazily via $ref
parse_endian(), parse_align(), parse_encoding(), parse_discriminator() Adapted to read from BAST field/struct properties instead of keyword-value objects
Custom keyword validators (19 jsonschema::Keyword impls) Removed. Replaced by the BAST-native validator (see below).
build_validator() Repurposed. Still builds a jsonschema::Validator, but from a standard JSON Schema (no custom keywords). Used for JSON payload validation (call's OperationSpec schemas, etc.) and for validating BAST documents against the BAST meta-schema.
BAST-native validator New. A recursive walker over the BAST type tree that checks value-domain constraints on a materialized Value: integer ranges, float finiteness, maxLength, timestamp shape, enum index bounds, union variant dispatch. Replaces the 19 custom keyword validators for the validate_bytes path. See The Validator Split.
validate_json() Unchanged in signature. Validates a Value against a jsonschema::Validator compiled from a standard JSON Schema the consumer provides. No longer uses custom keywords.
validate_bytes() Unchanged in concept — materialize Value from bytes, then validate. The validation step uses the BAST-native validator (not jsonschema) to check value-domain constraints from the BAST document.

Public API

Item Change
AlkTypeKind enum Unchanged — same 19 variants, same methods
Endian, VariableEncoding, DiscriminatorKind Unchanged
AlkTypeEngine compile() takes a BAST document + root type name instead of a custom-keyword JSON Schema. validate_json() validates against a consumer-provided JSON Schema. validate_bytes() validates against the BAST document via the BAST-native validator.
LayoutMode, OffsetMap, ByteRange Unchanged
LayoutBuilder, PackedLayout, FieldPosition Unchanged
SequentialReader, FieldValue Unchanged
UnionDispatch Unchanged
data_access functions Unchanged
AlkTypeError Unchanged (Schema/Offset/Access/Validation variants)
Schema builder Public methods unchanged. build() produces BAST JSON instead of custom-keyword JSON.
Definitions builder Public methods unchanged. build() produces a BAST $defs block.
Discriminator builder Unchanged

What is removed

  • All 19 jsonschema::Keyword implementations (~200 lines of validator factories)
  • normalize_refs() — BAST $ref values are always full JSON Pointers
  • inline_union_variant_refs() — union mapping values are resolved lazily
  • get_alktype_kind() / get_alktype_kind_loose() and their _enum variants — replaced by direct kind field parsing
  • Custom keyword registration with jsonschema — the engine no longer calls jsonschema::options().with_keyword(...). The jsonschema crate remains a direct dependency for standard JSON Schema validation (validating JSON payloads like call's OperationSpec schemas, and validating BAST documents against the BAST meta-schema). The only thing removed is the custom keyword integration path.

What is added

  • BAST meta-schema (embedded in the crate, published at a stable URL)
  • BAST document parser — validates a BAST document against the meta-schema, then extracts type definitions
  • AlkTypeEngine::compile() takes (bast_document: &Value, root_name: &str, mode: LayoutMode) — the root name selects which $defs entry is the top-level type
  • BAST-native validator — a recursive walker over the BAST type tree that checks value-domain constraints on a materialized Value. Replaces the 19 custom keyword validators for the validate_bytes path. Enforces: integer ranges, float finiteness, string/bytes maxLength, RFC 3339 timestamp shape, enum index bounds, union variant dispatch (recursing into variants). See The Validator Split.
  • BAST meta-schema validation at compile time (optional but recommended — the engine can skip it and trust the caller, or validate as a guard)

Semver and ADR Impact

This section is the scope-creep guardrail for the public API during implementation. The crate is on crates.io at 0.1.0 with zero real consumers, so a breaking bump is free — but the contract still needs to be explicit so the implementation doesn't drift. Per AGENTS.md, the 0.1.0 public surface is the items re-exported from src/lib.rs.

Public API delta

Public item (from lib.rs re-exports) Class Change
AlkTypeKind (enum + variants + methods) Additive Unchanged. 19 variants, same methods. New from_str()/to_str() mapping for lowercase BAST kind strings ("uint32" ↔ AlkTypeKind::Uint32) — additive methods.
Endian, VariableEncoding, DiscriminatorKind Unchanged —
AlkTypeEngine::compile Breaking Signature: compile(schema: &mut Value, mode) → compile(bast_doc: &Value, root_name: &str, mode). Adds required root_name param (D-BAST-001); drops &mut (BAST needs no in-place normalize_refs); input is a BAST document, not a custom-keyword JSON Schema.
AlkTypeEngine::validate_json Breaking (behavioral) Signature unchanged (instance: &Value) -> Result<...>, but the validator it runs is now a standard jsonschema::Validator from a consumer-provided JSON Schema supplied at compile, not a custom-keyword validator built from the alktype schema. The compiled engine must carry a separate JSON-Schema validator (or validate_json takes the JSON Schema at call time — to be decided in step 6). Either way the contract of what schema validates the instance changes.
AlkTypeEngine::validate_bytes Unchanged (contract) Same signature. Internally the validation step switches from jsonschema::Validator to the BAST-native validator. Error type unchanged (D-BAST-009).
AlkTypeEngine::is_valid_json Breaking (behavioral) Same caveat as validate_json — validates against the consumer JSON Schema, not the alktype schema.
AlkTypeEngine accessors (endian, mode, offset_map, layout_builder, sequential_reader, read_field, write_field, etc.) Unchanged Layout-layer accessors are format-agnostic.
LayoutMode, OffsetMap, ByteRange Unchanged —
LayoutBuilder, PackedLayout, FieldPosition Unchanged —
SequentialReader, FieldValue Unchanged —
UnionDispatch Unchanged —
data_access::* functions Unchanged —
AlkTypeError (all 4 variants) Unchanged D-BAST-009 keeps Validation(jsonschema::ValidationError<'static>).
Schema builder (struct_, object, field, build, all setters) Breaking (output format) Public method signatures unchanged. build() output changes from custom-keyword JSON to BAST JSON (for struct_) / standard JSON Schema (for object). Callers that introspect the built Value break; callers that pass it straight to compile are source-compatible once compile takes BAST.
Definitions builder (new, define, define_value, build, merge_into) Breaking (output format) Same as Schema — signatures unchanged, build()/merge_into() output shape changes to BAST $defs.
Discriminator builder enum Unchanged —
build_validator (from validation) Breaking (signature or removal) Currently build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError> builds a custom-keyword validator. Under the pivot it either (a) is removed (consumers call jsonschema directly for standard JSON Schema) or (b) is repurposed to build a standard jsonschema::Validator from a consumer-provided standard JSON Schema (no custom keywords). Decision belongs to step 6. Either way the current signature's contract breaks.
get_alktype_kind, get_alktype_kind_enum, get_alktype_kind_loose, get_alktype_kind_loose_enum, normalize_refs, inline_union_variant_refs, resolve_ref, resolve_ref_or_inline, parse_align, parse_discriminator, parse_encoding, parse_endian, parse_max_length Breaking (removal or rework) All currently re-exported from lib.rs. normalize_refs and inline_union_variant_refs are removed (BAST needs neither). The get_alktype_kind* family is removed (replaced by direct kind parsing). The parse_* and resolve_* functions are reworked to read BAST properties instead of keyword-value objects, or removed if subsumed by the BAST parser. Open: which of these stay public vs become internal. Current leaning — drop all from lib.rs re-exports (they're engine-internal accessors, not consumer API); the BAST parser exposes a new typed surface instead.

Net breaking surface: compile, validate_json/is_valid_json (contract), Schema::build/Definitions::build (output format), build_validator (signature/removal), and the ~13 schema::* helper re-exports. Net additive: BAST parser, BAST-native validator, AlkTypeKind::from_str/to_str. Net unchanged: the entire layout

  • data-access + materialize + tunion layer, AlkTypeError, the Discriminator builder, AlkTypeKind variants.

ADR impact

ADR Action Reason
ADR-001 (purpose, scope, "schema is the format") Supersede The "schema is the format" principle is retained and strengthened (BAST is the format), but the concrete format changes from custom-keyword JSON Schema to BAST. A new ADR (ADR-BAST or renumbered) records the BAST format as the realization of the principle. ADR-001's Status → Superseded by ADR-BAST.
ADR-002 (two layout modes) Unchanged Layout modes are format-agnostic. No content change; maybe a one-line note that the input format changed but the modes didn't.
ADR-003 (endianness, alignment, encoding, discriminators) Amend Annotation semantics carry forward unchanged; annotation location moves from custom-keyword objects to BAST type-level properties. Amend the "where annotations live" sections, keep the semantics.
ADR-004 (AlkTypeError, load-time build, access-time check) Amend Error enum shape unchanged (D-BAST-009). The "validation strategy" section updates: bytes path uses BAST-native validator, JSON path uses standard jsonschema. The load-time/access-time split is retained.
ADR-005 (Int64/Uint64, JSON precision caveat) Unchanged Kinds carry forward; JSON precision caveat is unchanged.
ADR-006 (reject non-final inline length-prefixed in aligned mode) Unchanged Layout rule, format-agnostic.
ADR-007 (packed-mode read factory, sequential_reader() returns owned reader) Unchanged Reader factory semantics are format-agnostic.
ADR-008 (reject TUnion in aligned mode for v1) Unchanged Layout rule, format-agnostic.
ADR-009 (builder API) Amend Public method surface unchanged; build() output format changes (BAST for struct_, standard JSON Schema for object). Amend the "output format" section; keep the method catalog.
ADR-010 (validate_bytes — materialize then validate) Amend The two-step concept (materialize → validate) is retained. The validation step's implementation changes from jsonschema custom keywords to the BAST-native validator. Amend the "validation step" section; add a pointer to D-BAST-006/D-BAST-009 and the Validator Split section.

New ADRs to write (post-implementation, grounded in shipped code):

  • ADR-BAST — the BAST format, meta-schema, and $defs/$ref/ kind vocabulary. Supersedes ADR-001's format-specific content.
  • ADR-VAL-SPLIT (or fold into ADR-004's amend) — the two-validator model: BAST-native for validate_bytes, standard jsonschema for validate_json. Records D-BAST-006, D-BAST-007, D-BAST-009.

Descriptive docs to sync (post-implementation):

  • docs/architecture/schema-layer.md — rewrite for the BAST parser (replaces the custom-keyword accessor walk-through).
  • docs/architecture/validation.md — rewrite for the validator split.
  • docs/architecture/builder.md — update the build() output examples to BAST JSON.
  • src/lib.rs module doc comment — update the "Takes a JSON Schema with AlkType:* custom keywords" preamble to BAST.

Stale TODOs to remove: any TODO referencing custom-keyword normalization, inline_union_variant_refs, or the rejected bare-name-ref design — align with the ADRs as AGENTS.md §"Architecture Context" requires.

Open decisions deferred to implementation steps

These are small enough to decide when the step is reached, but are flagged here so they don't become drive-by semver changes:

  1. validate_json JSON Schema source (step 6): does the consumer pass the JSON Schema to compile (engine carries a second validator) or to validate_json at call time? The former preserves the current single-call ergonomics; the latter is more flexible. Not semver-relevant either way if validate_json's signature can absorb a new param or stay as-is — needs the call-site analysis.
  2. build_validator fate (step 6): removed vs repurposed. If repurposed, its signature stays but its contract (no custom keywords) changes — a behavioral break, not a type break.
  3. schema::* helper re-exports (step 3): drop from lib.rs (engine-internal) vs keep public for consumers that walk schemas. Leaning: drop — they're accessors for the old format, and the BAST parser exposes a cleaner typed surface. Confirmed during step 3.

The Validator Split

A key architectural clarification: BAST separates two concerns that the current format conflates, and in doing so reveals that the current engine has two distinct validation paths that custom keywords paper over with a single mechanism.

Current model (conflated)

┌─────────────────────────────────────────────┐
│  JSON Schema with AlkType:* custom keywords  │
│  ┌───────────────────────────────┐  │
│  │  Binary layout info (AlkType:Uint32)  │  │
│  │  JSON validation info (type, enum)    │  │
│  └───────────────────────────────┘  │
└─────────────────────────────────────────────┘
         │
         ▼
  AlkTypeEngine::compile()
         │
         ├──► Layout (offset map / layout builder)
         └──► Validator (jsonschema with 19 custom keywords)
                 │
                 ├──► validate_json(&Value)  — JSON value validation
                 └──► validate_bytes(&[u8]) — materialize, then validate

One document serves two roles. The engine extracts both layout and validation from the same JSON tree. A single jsonschema::Validator (with custom keywords) serves both validate_json and validate_bytes.

BAST model (three layers, two validators)

The current engine's two validation paths have different needs, and BAST makes the split explicit:

                          ┌──────────────────────┐
                          │   BAST document       │
                          │   (binary layout +   │
                          │    value constraints) │
                          └──────────┬───────────┘
                                     │
                                     ▼
                          AlkTypeEngine::compile()
                                     │
                    ┌────────────────┼────────────────┐
                    ▼                                  ▼
          ┌─────────────────┐                ┌──────────────────┐
          │  Layout          │                │  BAST-native      │
          │  (offset map /  │                │  validator         │
          │   layout builder)│                │  (walks BAST,     │
          └────────┬────────┘                │   checks value    │
                   │                          │   constraints)    │
                   ▼                          └────────┬──────────┘
          ┌─────────────────┐                          │
          │  Materializer    │                          │
          │  (bytes → Value, │                          │
          │   guarantees     │                          │
          │   structure)     │                          │
          └────────┬────────┘                          │
                   │                                    │
                   ▔──────────────┬─────────────────────┘
                                  ▼
                       validate_bytes(&[u8])
                       (materialize, then check
                        value constraints against
                        the BAST schema — no external
                        JSON Schema needed)

  ┌──────────────────────────┐
  │  JSON Schema document    │     (separate concern)
  │  (JSON validation)       │
  │  type: "object"          │
  │  properties: {           │
  │    id: { type: "integer"}│
  │  }                       │
  │  required: ["id"]         │
  └────────────┬─────────────┘
               │
               ▼
      build_validator(&json_schema)
               │
               ▼
      jsonschema::Validator
      (standard JSON Schema,
       no custom keywords)
               │
               ▼
      validate_json(&Value)
      (validates consumer-supplied
       JSON against a standard
       JSON Schema — BAST not
       involved)

Why two validators, not one

The two paths have fundamentally different inputs and guarantees:

validate_bytes(&[u8]) — bytes in, BAST is the validator. The materializer produces a Value tree from bytes. By construction, this Value is structurally correct: all declared fields are present (because the materializer iterates the field list), types are correct (because read_u32 produces a Value::Number), bounds are checked (via data_access::check_bounds), UTF-8 is valid (via from_utf8), the discriminator is in the mapping, and the boolean byte is 0 or 1. What the materializer does NOT check — and what the 19 custom keyword validators currently check afterward — are value-domain constraints expressed in the schema:

Constraint Current enforcer BAST-native validator
Integer range (Int8..Uint64) Custom keyword BAST kind + range check
Float finiteness (Float32/64) Custom keyword BAST kind + is_finite()
String maxLength Custom keyword (reads parent) BAST field maxLength property
Bytes maxLength Custom keyword (reads parent) BAST field maxLength property
RFC 3339 timestamp shape Custom keyword BAST kind: "timestamp" + shape check
Union variant dispatch UnionValidator (OQ-008) BAST union — recurse into variant
Enum value membership Built-in enum keyword (broken on bytes path — see below) BAST kind: "enum" + index-in-range

All of these are expressible directly from the BAST document — no external JSON Schema needed. The BAST-native validator is a recursive walker over the BAST type tree, checking the materialized Value against each type's constraints. Union dispatch is just recursion: read __discriminator, look up the variant's BAST definition, recurse.

This recovers the OQ-008 union dispatch behavior without custom keywords and without requiring the consumer to hand-author discriminator dispatch in a separate JSON Schema. The BAST document already knows the union's variants and their fields.

validate_json(&Value) — JSON in, JSON Schema is the validator. The consumer provides a JSON Value (e.g., an incoming JSON-RPC request on alkcall's channel 0). The BAST document is irrelevant — BAST describes bytes, not JSON shape. The right validator for a JSON value is a standard jsonschema::Validator built from a standard JSON Schema document (which the consumer provides, or which is generated from BAST via future codegen). This is the path alkcall uses today for its OperationSpec JSON validation.

Bug fix: enum validation on the bytes path

The current engine has a dead constraint: when validate_bytes materializes an AlkType:Enum field, it reads a raw u32 index and emits Value::Number(index). The built-in enum keyword lists string members (e.g., ["Ok", "Error"]). A numeric index never matches a string-membered enum, so enum membership is silently unenforced on the bytes path. The BAST-native validator fixes this: for kind: "enum", it checks that the materialized index is within the values array bounds (0..len-1), which is the correct validation for a binary enum encoded as an index.

What this means for consumers

alkcall today uses alktype for two roles:

  1. Binary layout (channels chunk header) — AlkTypeEngine::compile() in packed mode
  2. JSON validation (call's OperationSpec schemas) — build_validator() or jsonschema directly

Under BAST:

  1. Binary layout + validation — AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed) then engine.validate_bytes(&buf). The engine uses the BAST-native validator — no external JSON Schema needed for binary data. All value constraints (maxLength, ranges, enum bounds, union dispatch) are enforced from the BAST document.
  2. JSON validation — build_validator(&json_schema) — builds a standard jsonschema::Validator from a standard JSON Schema document (no custom keywords). Consumers that don't use binary at all (e.g., adapters around remote JSON Schema APIs) use this path exclusively.

The builder API produces both formats:

  • Schema::struct_().field(...).build() → BAST JSON (binary layout)
  • Schema::object().field(...).build() → standard JSON Schema (JSON validation)

Both live in the same crate, same wasm target, same builder surface. This is the point of alktype: one small wasm-compatible codebase that handles both binary layout and JSON validation for protocol crates like alkcall, alktty, and future channel implementations.

Validation of binary data

validate_bytes() works as follows:

  1. Materialize a Value tree from binary bytes using the BAST layout. The materializer guarantees structural correctness (bounds, UTF-8, bool byte, discriminator lookup, all fields present).
  2. Validate the materialized Value against value-domain constraints from the BAST document, using the BAST-native validator. This checks integer ranges, float finiteness, maxLength, timestamp shape, enum index bounds, and recurses into union variants.

No external JSON Schema is required for validate_bytes. The BAST document is the complete specification of the binary format — it describes both the layout (how to read) and the constraints (what values are valid). This is the "schema is the format" principle from ADR-001, now fully realized: the BAST document is the binary format, and it is also the validation spec.

An optional external JSON Schema can be layered on top for constraints BAST doesn't express (e.g., cross-field consistency, regex patterns on string content). This is additive, not load-bearing — the binary path works without it.

Relationship to JSON Schema and TypeBox

BAST is a JSON Schema dialect

BAST is a specific JSON Schema instance format — like how JSON Schema itself is a JSON document that conforms to the JSON Schema meta-schema. BAST documents conform to the BAST meta-schema. The meta-schema is a standard JSON Schema (Draft 2020-12).

This means the entire JSON Schema tooling ecosystem works with BAST:

  • Validation: jsonschema::options().build(&bast_meta_schema)?.validate(&bast_doc)
  • Editors: VSCode with $schema pointing to the BAST meta-schema URL
  • Documentation: JSON Schema generators can produce human-readable docs from the meta-schema

Note: the BAST meta-schema validates the structure of a BAST document (is it well-formed?). The BAST-native validator validates binary data against the BAST document (are the bytes a valid instance of this format?). These are different validators for different inputs.

TypeBox interop

TypeBox's Type.Module({...}) pattern maps naturally to BAST's $defs structure. A TypeBox module that defines binary types can serialize to BAST JSON instead of custom-keyword JSON. The codegen (ts-to-module.ts) could target BAST as an output format.

The relationship is:

  • TypeBox → BAST JSON → alktype engine (binary layout)
  • TypeBox → standard JSON Schema → jsonschema (JSON validation)

Same TypeBox source, two output formats, two validators.

Not a replacement for JSON Schema

BAST does not replace JSON Schema for JSON data validation. A BAST document cannot validate a JSON payload — it describes binary data layouts and value-domain constraints for bytes. For JSON validation, consumers use standard JSON Schema documents (which may be derived from BAST definitions via codegen, or authored separately). The validate_json path on AlkTypeEngine accepts a consumer-provided JSON Schema and uses a standard jsonschema::Validator — BAST is not involved.

This is the split: validate_bytes is BAST-native (the BAST document is both the layout spec and the validation spec for bytes); validate_json is JSON-Schema-native (a standard JSON Schema is the validation spec for JSON values). One crate, two validators, two input types.

Codegen (Future)

BAST enables code generation that the custom-keyword format makes awkward. A codegen module (feature-gated behind codegen) would:

  1. Input: A BAST document (or SchemaRegistry equivalent)
  2. Walk: Iterate $defs entries, inspect kind values
  3. Map: "uint32" → u32 (Rust), number (TypeScript), int (Python)
  4. Emit: Handlebars templates for struct/enum/union definitions

Generated artifacts

Target Type definitions Binary reader/writer
Rust struct ChunkHeader { channel_id: u32, length: u32 } fn read_header(buf: &[u8]) -> Result<ChunkHeader, Error>
TypeScript interface ChunkHeader { channelId: number; length: number } function readHeader(buf: Uint8Array): ChunkHeader
Python @dataclass class ChunkHeader: ... def read_header(buf: bytes) -> ChunkHeader

Relationship to typebox-rs codegen

The typebox-rs codegen/ module (in /workspace/@alkimiadev/typebox-rs) is the reference architecture:

  • SchemaRegistry for named types with $ref resolution
  • RustGenerator / TypeScriptGenerator wrapping Handlebars templates
  • schema_to_rust_type() / schema_to_ts_type() mapping functions
  • Feature-gated behind codegen = ["handlebars"]

alktype's codegen would follow the same pattern but walk BAST kind values instead of SchemaKind enum variants. The handlebars-rs dependency is WASM-compatible.

Scope boundary

Codegen is out of scope for the BAST pivot itself. The pivot changes the schema format; codegen builds on top of the new format. It is described here to show that BAST enables it, not to commit to a specific implementation timeline.

ABI Adapter (Future)

A BAST document describes the binary interface of a protocol — it is essentially an ABI specification in JSON. This enables:

  • Version negotiation: Two peers exchange BAST documents to agree on a protocol version. The engine can detect mismatches (field added, type changed, endianness differs) and either reject or adapt.
  • Schema migration: A consumer with schema v1 can read data written by schema v2 if the changes are compatible (fields added at the end, types widened). The engine can compute a migration plan from the diff of two BAST documents.
  • WASM interop: A WASM component can export its BAST schema as part of its WIT interface, enabling host languages to generate readers/writers for the component's binary protocol without manual bindings.

This is a future capability, not part of the pivot. It is mentioned because BAST makes it possible in a way that custom keywords scattered through JSON Schema trees do not.

Migration Path

Phase 1: BAST format and meta-schema (this pivot)

  1. Define the BAST meta-schema (the JSON Schema that validates BAST documents)
  2. Run the BAST-native validator POC — implement the validator, wire it into validate_bytes, confirm the existing test suite passes. This de-risks the validation model before the full pivot. Done. See POC Result — BAST-native validator. The hypothesis is confirmed; the production refactor (step 5 below) is now a straightforward port. POC code lives on branch bast-validator-poc (commit f371fe4) as a self-contained src/bast_poc.rs reference — deliberately not merged to main (it is throwaway scaffolding that gets replaced by the production module in step 5).
  3. Implement BAST document parsing in the engine (replace custom keyword detection with kind field parsing)
  4. Update AlkTypeEngine::compile() to accept a BAST document + root type name
  5. Implement the BAST-native validator (production version, informed by the POC). Constructs AlkTypeError::Validation via jsonschema::ValidationError::custom per D-BAST-009 — the Validation variant's payload type is unchanged.
  6. Update validate_json() to accept a consumer-provided JSON Schema and build a standard jsonschema::Validator (no custom keywords)
  7. Update the builder API to produce BAST JSON (public methods unchanged)
  8. Remove custom keyword validators, normalize_refs(), inline_union_variant_refs(), and the get_alktype_kind* functions
  9. Update all tests to use BAST format
  10. Update architecture docs (ADRs, schema-layer.md, etc.)

Phase 2: Downstream adoption

  1. Port alkcall's chunk header to BAST (replace hand-rolled wire.rs with AlkTypeEngine + SequentialReader/LayoutBuilder)
  2. Port alktty's TTY chunk format to BAST
  3. Port the SFTP POC schemas to BAST format

Phase 3: Codegen (future)

  1. Add codegen feature flag with handlebars dependency
  2. Implement RustGenerator and TypeScriptGenerator walking BAST kind values
  3. External Handlebars templates in src/codegen/templates/

POC Scope

The layout swap needs no POC — it is a backend swap (custom keywords → kind-based format) on top of a proven layout engine. The layout engine's byte-identity is already proven (alknet-typedef-poc, alktype-builder-poc) and the layout code is unchanged, so there is nothing empirical to de-risk there.

One targeted POC was needed to de-risk the validation model. The risk was specific and falsifiable: can a BAST-native validator — a recursive walker over the BAST type tree — fully replace the 19 custom keyword validators on the validate_bytes path, including the OQ-008 union variant dispatch, without regression?

The POC has been run and succeeded. The remainder of this section records the original scope and success criteria (kept for the record); the outcome is in POC Result — BAST-native validator. The POC code is on branch bast-validator-poc (commit f371fe4), not merged to main — it is reference scaffolding superseded by Phase 1 step 5.

POC: BAST-native validator for validate_bytes

Hypothesis: A recursive walker over the BAST type tree can enforce all value-domain constraints that the 19 custom keyword validators currently enforce, recovering the OQ-008 union variant dispatch behavior, and fixing the enum-membership dead constraint on the bytes path — all without jsonschema custom keywords and without requiring the consumer to provide an external JSON Schema.

Scope:

  1. Implement the BAST-native validator as a new module (src/bast_validation.rs or similar)
  2. The validator walks a materialized Value tree against the BAST type definitions, checking:
    • Integer ranges (Int8..Uint64)
    • Float finiteness (Float32/64)
    • String maxLength (UTF-8 byte length)
    • Bytes maxLength (byte length)
    • RFC 3339 timestamp shape (non-strict, matching current behavior)
    • Enum index bounds (0..values.len()-1 — fixes the dead constraint)
    • Union variant dispatch (read __discriminator, look up variant BAST definition, recurse)
    • Boolean validity (materializer already checks, but the validator should confirm)
  3. Wire it into validate_bytes() as the validation step (replacing the jsonschema::Validator call)
  4. Run the existing test suite — the tests encode all current expected validation behavior. If they pass, the POC succeeds.

Success criteria:

  • All existing validate_bytes tests pass without modification to their assertions (test inputs will change to BAST format, but the expected validation outcomes must be identical)
  • The union variant dispatch tests (OQ-008) pass — maxLength on a Bytes field inside a union variant is enforced
  • The enum index-bounds validation works (new behavior — currently broken, so this is a fix, not a regression)

Failure path: If the POC reveals that the BAST-native validator cannot cleanly express some constraint (e.g., a constraint that relies on JSON Schema's structural keywords in a way that's hard to reimplement), the fallback is the "structural-only + external JSON Schema" model from the original Gap 1 — but this is unlikely given that the materializer already guarantees structure, leaving only value-domain checks.

Out of scope for this POC:

  • validate_json — this path uses a standard jsonschema::Validator from a consumer-provided JSON Schema, not the BAST-native validator. No POC needed; it's a standard jsonschema usage.
  • BAST document parsing / meta-schema validation — the parser is straightforward JSON walking; no empirical risk.
  • Layout computation — unchanged, already proven.

Decisions

The following were open questions in earlier drafts. Each is now resolved. They are recorded here as decisions, not re-litigated.

D-BAST-001: Root type selection

Decision: Explicit. The root type name is a required parameter to compile(): AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed). This is unambiguous and matches how consumers think about it ("compile the ChunkHeader schema"). Convention (first entry in $defs) is fragile and depends on JSON key order; a $root marker is redundant with an explicit parameter.

D-BAST-002: Primitive type string set

Decision: Lowercase ("uint32", "int8", "float64", "bool", "string", "bytes", "timestamp"). Matches JSON Schema's own convention ("string", "integer", "boolean"), is easier to type, and is the convention in the TypeBox research examples. The AlkTypeKind enum variants remain PascalCase in Rust — the mapping is a simple from_str() impl.

D-BAST-003: Top-level $defs requirement

Decision: Always $defs. Every BAST document has the same top-level shape: { "$defs": { ... } }. Single-type documents are a special case with one entry. The $defs block is the namespace; the root type name (D-BAST-001) selects the entry point. A bare struct at the top level would be a special case with different parsing logic and no home for additional definitions.

D-BAST-004: Arrays of variable-length elements (deferred)

Decision: Arrays of variable-length elements are not supported in v1. The meta-schema requires count on all array types, making arrays fixed-size only. The no-count example has been removed. This matches the engine's current behavior (it rejects arrays of variable-length elements) and aligns with OQ-001 (deferred, blocked on a concrete consumer needing interleaved variable-stride arrays).

Variable-length collections are still available via record (a count-prefixed string-keyed map), which the engine supports. If a consumer needs a variable-length array of fixed-size elements, they can use a record with integer-stringified keys as a workaround, or wait for OQ-001 to be addressed.

D-BAST-005: Field-name discriminator unions

Decision: Supported. The meta-schema includes an optional fields array on UnionDef. When discriminator.kind == "field", the fields array provides the union's field definitions (including the discriminator field). When discriminator.kind == "byte", fields is absent — the union's layout is purely the variant layout. This preserves a feature the engine already supports. The meta-schema update is in the meta-schema sketch above.

D-BAST-006: validate_bytes validation model

Decision: BAST-native validator. The validate_bytes path uses a recursive walker over the BAST type tree to check value-domain constraints on the materialized Value — no external JSON Schema needed. This recovers the OQ-008 union variant dispatch behavior (the validator recurses into the variant's BAST definition) and fixes the enum-membership dead constraint (the validator checks the materialized index against the values array bounds). See The Validator Split and the POC.

An optional external JSON Schema can be layered on top for constraints BAST doesn't express (cross-field consistency, regex patterns on string content). This is additive, not load-bearing.

D-BAST-007: validate_json validation model

Decision: Standard JSON Schema validator. validate_json on AlkTypeEngine validates a consumer-provided JSON Value against a jsonschema::Validator compiled from a standard JSON Schema document the consumer provides at compile time. No custom keywords. The BAST document is not involved in this path — BAST describes bytes, not JSON shape. This preserves the validate_json / validate_bytes symmetry from ADR-010, but the two paths now use different validators (standard jsonschema for JSON, BAST-native for bytes), reflecting their different inputs and guarantees.

The JSON Schema may be authored separately or derived from BAST via future codegen. For alkcall's channel 0 (JSON-RPC), the JSON Schema is the OperationSpec schema, authored independently of any BAST document.

D-BAST-008: Builder API — two output formats

Decision: One builder, two build methods. The construction API is the same (field names, types, annotations); only the output format differs. Schema::struct_().field(...).build() → BAST JSON (binary layout). Schema::object().field(...).build() → standard JSON Schema (JSON validation). The builder already distinguishes AlkType kinds from JSON Schema types via naming conventions (string() vs string_()).

Both output formats live in the same crate. This is the point of alktype: one small wasm-compatible codebase that handles both binary layout and JSON validation for protocol crates. alkcall uses both — channel 0 is JSON (standard JSON Schema), binary channels use BAST. Future crates (alktty, tunnels, sftp, git) will use BAST for their binary formats. The codegen feature (future) will generate readers/writers from BAST documents for these crates.

D-BAST-009: AlkTypeError::Validation payload shape

Status: decided. Keep Validation(jsonschema::ValidationError<'static>).

AlkTypeError::Validation currently wraps jsonschema::ValidationError<'static>. Under the BAST pivot the validate_bytes path no longer uses jsonschema at all (confirmed by the POC — observation 1), so the error payload on that path is constructed via jsonschema::ValidationError::custom purely to keep the variant's type unchanged. The two options were:

  1. Keep Validation(jsonschema::ValidationError<'static>). Simplest — ValidationError::custom is public and 'static, so the bytes path can construct it without a real jsonschema validator. Cost: the error type retains its jsonschema dependency even though the bytes path no longer drives it. validate_json still uses jsonschema, so the dependency isn't removable either way — but the error type carries jsonschema only for one of its two callers.
  2. Introduce Validation(String) (or a small structured payload). Drops the jsonschema type from the public error enum. This is a semver-relevant public-API change (the Validation variant's payload type changes), so per AGENTS.md it requires an explicit decision, not a drive-by. Benefit: the error type is jsonschema-free, which matters if a future no_std/minimal build wants to drop jsonschema from the bytes-only path (relates to OQ-002).

Rationale for option 1: The deciding factor is consumer ergonomics on the combined path. Consumers like alkcall use both validate_json (channel 0, JSON-RPC) and validate_bytes (binary channels) and handle AlkTypeError::Validation in one place. A single uniform payload type means one match arm covers both sources — no Validation(jsonschema_err) vs Validation(string) branching downstream. Option 2 would force validate_json to flatten its structured errors (instance path, schema path, keyword) to a String via Display just to match a bytes-path shape — the more information-rich path loses data to accommodate the less rich one. That is the wrong direction.

The no_std/minimal-build angle (OQ-002) that option 2 was meant to enable is moot in practice: validate_json requires jsonschema regardless, so a bytes-only no_std build already has to give up validate_json as a separate, larger decision. Dropping the type from one error variant does not unlock that build — the dependency is load- bearing on the other validation path. The right place to revisit this is when/if OQ-002 is actually pursued, not preemptively.

The POC already used option 1 (via ValidationError::custom); the production refactor (Phase 1 step 5) follows the same construction pattern. No semver-relevant change to the Validation variant.

Risks and Mitigations

Risk Mitigation
BAST format doesn't cover all 19 type kinds The format is designed to cover all 19. The meta-schema is the spec — if a kind can't be expressed, the meta-schema is wrong.
$ref resolution complexity moves from engine to schema authoring BAST $ref values are always full JSON Pointers (#/$defs/Name). No normalization, no bare names. Resolution is a single hash lookup.
BAST-native validator misses a constraint the custom keywords enforced The POC runs the existing test suite, which encodes all current expected validation behavior. If a constraint is missed, a test fails before the pivot lands.
Losing OQ-008 union variant dispatch The BAST-native validator recurses into the variant's BAST definition on __discriminator lookup — same behavior, no custom keywords. Covered by the POC.
Enum membership broken on bytes path Already broken today (dead constraint). The BAST-native validator fixes it by checking the materialized index against the values array bounds. Net improvement.
validate_json loses custom keyword validation validate_json uses a standard jsonschema::Validator from a consumer-provided JSON Schema. Consumers that relied on custom keywords for JSON validation need to provide equivalent standard JSON Schema keywords. No real consumers exist yet.
Builder API output format change breaks consumers No real consumers exist yet (v0.1.0 has zero adoption). The builder's public methods are unchanged; only the JSON output format changes.
Meta-schema maintenance burden The meta-schema is small and changes rarely. It's embedded in the crate and published at a stable URL.

References

  • ADR-001 — current "schema is the format" principle (to be updated)
  • ADR-003 — annotation shapes (endianness, alignment, encoding, discriminators — carry forward to BAST)
  • ADR-009 — builder API (public surface unchanged, output format changes)
  • ADR-010 — validate_bytes (unchanged in concept)
  • /workspace/research/typebox_research/ujsx/jpath.gen.ts — TypeBox Type.Module pattern (the $defs/$ref model BAST follows)
  • /workspace/research/typebox_research/ujsx/mdast.gen.ts — TypeBox cross-module references and composite types
  • /workspace/research/typebox_research/codegen/ts-to-module.ts — TypeScript-to-TypeBox codegen (reference for future BAST codegen)
  • /workspace/@alkimiadev/typebox-rs/src/codegen/ — Rust/TypeScript codegen from schemas (reference architecture)
  • /workspace/alknet-typedef-poc/tests/sftp_roundtrip_test.rs — SFTP POC proving byte-identical output (to be replicated with BAST)
  • /workspace/@alkdev/alkcall/src/channels/wire.rs — hand-rolled chunk header (target for BAST replacement)

POC Result — BAST-native validator

Status: succeeded. The POC is on branch bast-validator-poc in src/bast_poc.rs (20 tests, all passing; full crate suite — 416 tests — green; cargo clippy --all-targets -- -D warnings clean; cargo build --target wasm32-unknown-unknown --release clean).

The POC implements the validate_bytes validation model from D-BAST-006 as a self-contained module that does not touch the production schema / materializer / validator paths. It reuses only data_access (read primitives), AlkTypeError (error type), and Endian. The BAST document parser, a packed-mode materializer, and the BAST-native validator are all implemented from scratch — that is the point: prove the model works end-to-end before refactoring the production code.

What the POC proves

The hypothesis from the POC section is confirmed: a recursive walker over the BAST type tree fully replaces the 19 custom keyword validators on the validate_bytes path, including the OQ-008 union variant dispatch, and fixes the enum-membership dead constraint — all without jsonschema custom keywords and without an external JSON Schema.

The hard cases that were the actual de-risking targets all pass:

  • Union byte-offset discriminator + maxLength inside a variant (OQ-008). union_byte_disc_max_length_inside_variant_enforced materializes a union with two $ref variants, dispatches on a byte-offset uint8 discriminator, and enforces maxLength on a bytes field inside the selected variant. The validator reads __discriminator, looks up the variant's BAST definition, and recurses — same behavior as the current UnionValidator's per-variant sub-validators, but with no jsonschema involvement.
  • Union field-name discriminator + maxLength inside a variant. union_field_disc_max_length_inside_variant_enforced covers the typedef.ts-style discriminator (a length-prefixed string field selects the variant). Same recursion model.
  • Enum index-bounds fix. enum_index_out_of_bounds_rejected exercises the constraint that is broken in the current engine (the built-in enum keyword checks string membership; the materializer emits Value::Number(index), which never matches — a dead constraint). The BAST-native validator checks the materialized index against the values array bounds (0..len-1), which is the correct validation for a binary enum encoded as an index. Net improvement, not a regression.
  • Nested struct wrapping a union wrapping a struct. nested_struct_with_union_variant confirms the recursion composes through multiple type layers.
  • Arrays of fixed-size structs with count. array_of_structs_with_count covers the Vector3-style array (D-BAST-004).
  • Records (count-prefixed string-keyed maps). record_of_uint32 covers the TRecord shape.
  • Untrusted schema input. malformed_document_produces_schema_error_not_panic confirms a malformed BAST document surfaces as AlkTypeError::Schema, not a panic (AGENTS.md §3).
  • Basic cases (chunk header, int8/uint32 ranges, string/bytes maxLength, timestamp, bool, short buffer) all pass — if a couple of basic examples work, all of them do, since the validator is a flat per-kind dispatch with no per-kind special-casing beyond the range bounds.

How the validator works

The validator is a single recursive function (validate_typeref) that dispatches on the BAST kind. Each arm checks the value-domain constraint for that kind and, for composites, recurses into the child type definitions. The materializer (also implemented in the POC) guarantees structural correctness — bounds, UTF-8, bool byte, discriminator lookup, all fields present — so the validator only enforces what the materializer cannot:

Constraint Validator arm
Integer range (Int8..Uint64) validate_int / validate_uint with as_i64/as_u64 + range check
Int64/Uint64 (full range) validate_int64 / validate_uint64 (JSON precision caveat per ADR-005)
Float finiteness (Float32/64) validate_float with as_f64().is_finite()
String maxLength (byte length) check_string reads the field-level maxLength annotation
Bytes maxLength (array length) check_bytes accepts both Value::String and Value::Array forms
RFC 3339 timestamp shape validate_timestamp reuses the same non-strict check as the current engine
Enum index bounds validate_enum checks idx < values.len() — the dead-constraint fix
Union variant dispatch validate_union reads __discriminator, looks up the variant, recurses via validate_typeref
Struct fields validate_struct walks fields, requires each declared field present, recurses
Array count validate_array checks arr.len() == count and recurses per element
Record values validate_record recurses into each value's values type
Boolean validate_bool (materializer already rejects non-0/1 bytes)

Observations for the production implementation

  1. No jsonschema dependency for validate_bytes. The validator only needs serde_json (for Value) and the BAST document. The jsonschema crate is still a direct dependency for validate_json and for validating BAST documents against the BAST meta-schema, but the validate_bytes path no longer touches it. This is a small wasm binary-size win in addition to the architecture simplification.

  2. $ref resolution is a single hash lookup. The POC's resolve_ref_or_inline handles only #/$defs/Name pointers — the only form BAST allows. The current engine's normalize_refs / inline_union_variant_refs / resolve_ref_or_inline machinery for bare-name refs and inlined union variants is no longer needed: BAST $refs are always full JSON Pointers, and union variant refs are resolved lazily by the validator (the materializer already does this for the read path). The inline_union_variant_refs compile step can be removed entirely.

  3. The validator is ~250 lines. The 19 custom keyword validators (src/validation.rs) plus the macro definitions are ~500 lines and require the jsonschema::Keyword trait plumbing (factory closures, Box<dyn Keyword>, sub-validator construction at factory time). The BAST-native validator is a flat match — no factories, no trait objects, no sub-validator pre-computation. The recursion is direct.

  4. The AlkTypeError::Validation variant still wraps jsonschema::ValidationError<'static>. The POC uses jsonschema::ValidationError::custom to construct these so the error type is unchanged. This is now the decided shape for the production refactor — see D-BAST-009. The rationale is consumer ergonomics: a single uniform payload type means one match arm covers both validate_json and validate_bytes errors downstream, and validate_json's structured errors are worth preserving rather than flattening to a String.

  5. The materializer and validator share the BAST-walking code structure. Both walk the same kind/fields/mapping tree. The production refactor could share a typed BAST tree (a small BastNode enum) between them so the walk is parsed once. The POC parses lazily from the raw JSON in both passes to keep the model honest; a typed tree is a straightforward follow-on optimization, not a risk.

Verdict

The "how do we reproduce the same behavior?" question is answered: walk the BAST tree the same way the materializer does, checking the same value-domain constraints the custom keyword validators check today. The model is a strict simplification — fewer moving parts, no jsonschema integration on the bytes path, no compile-time inline_union_variant_refs step, no factory closures or trait objects, and the enum dead-constraint is fixed as a side effect.

The POC does not wire into AlkTypeEngine::validate_bytes — that is the production refactor (Phase 1 step 5 in the Migration Path), which replaces validation::build_validator usage on the bytes path with the BAST-native validator. The POC's job was to de-risk the model before that refactor; that job is done.