Port alknet-typedef crate from alknet

Copy the binary struct engine (src/, tests/) verbatim from
alknet/crates/alknet-typedef and create a standalone Cargo.toml
(workspace-inherited fields inlined). Port the architecture docs
(specs, ADRs 095-102, OQs 069-071) from alknet's nested multi-crate
layout to a flat single-crate layout, fixing relative link paths.

Build, 295 tests, and clippy all pass clean.
This commit is contained in:
glm-5.2 committed 2026-08-02 05:59:12 +00:00
1 parent eac7ad88b3
commit 2c4a4994dc
36 files changed
+13805

No files matched your search

+110
View File
@@ -0,0 +1,110 @@
---
status: draft
last_updated: 2026-07-22
---
# alknet-typedef
The binary struct engine: a small Rust crate that takes a JSON Schema
with `TypeDef:*` custom keywords and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
## Documents
| Document | Status | Description |
|----------|--------|-------------|
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [schema-layer.md](schema-layer.md) | draft | The 19 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation, `TypedefEngine` |
## Applicable ADRs
| ADR | Title | Relevance |
|-----|-------|-----------|
| [095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| [096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` |
| [097](decisions/097-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations |
| [098](decisions/098-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors |
| [099](decisions/099-int64-uint64-first-class-kinds.md) | Int64/Uint64 as First-Class Kinds | 64-bit integers (SFTP offsets, metatensor data_offsets); JSON precision caveat |
| [100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode | Prevents silent data corruption (inline variable data clobbering subsequent fields) |
| [101](decisions/101-packed-mode-read-factory.md) | Packed-Mode Read API — Engine as SequentialReader Factory | `engine.sequential_reader()` returns an owned reader, not a reference |
| [102](decisions/102-reject-tunion-in-aligned-mode.md) | Reject TUnion in Aligned Mode for v1 | Unions are the protocol pattern; aligned-mode union semantics were broken |
## Relevant Open Questions
| OQ | Title | Status | Relevance |
|----|-------|--------|-----------|
| OQ-069 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it |
| OQ-070 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case |
| OQ-071 | Builder API for schema construction | deferred(scope) | Schemas are authored in TypeBox or hand-written JSON for v1; blocked on a concrete need |
## Key Design Principles
1. **The schema is the format.** A JSON Schema with `TypeDef:*` custom
keywords is both the validation spec and the layout spec. No separate
format definition, no separate parser, no separate validator. One
schema, three uses: validate, compute offsets, access data. See
[overview.md](overview.md) and [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
2. **jsonschema is the validation engine, not a custom engine.** The
`jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with
custom keyword support. The novel code is the offset computation, not
the validation. This eliminates ~14,000 lines of hand-rolled schema
engines (typebox-rs, alktype). See [schema-layer.md](schema-layer.md)
and [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
3. **Two layout modes for two use cases.** Packed sequential
(`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP,
channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats
(metatensor). The consumer selects the mode; the schema is the same.
See [layout-engine.md](layout-engine.md) and
[ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md).
4. **Variable-length types default to inline length-prefixing.**
`[length: u32][data]` is the universal pattern used by channels, SFTP,
TTY, and most binary protocols. Offset indirection (the metatensor
blob tensor pattern) is opt-in via the `encoding` annotation. See
[layout-engine.md](layout-engine.md) and
[ADR-097](decisions/097-schema-annotations.md).
5. **TUnion supports both byte-offset and field-name discriminators.**
Byte-offset for protocol dispatch (SFTP type bytes, call protocol
event types). Field-name for the typedef.ts string pattern. See
[data-access.md](data-access.md) and
[ADR-097](decisions/097-schema-annotations.md).
6. **Endianness is per-schema, default little-endian.** The engine reads
the `"endian"` annotation and byte-swaps accordingly. SFTP consumers
specify `"endian": "big"`. See [layout-engine.md](layout-engine.md)
and [ADR-097](decisions/097-schema-annotations.md).
7. **Validation is opt-in, built once at load time.** The jsonschema
validator is compiled once at schema load time. Access-time validation
is a fast `is_valid()` check. High-throughput paths can skip
validation; security-sensitive paths can validate every frame. See
[validation.md](validation.md) and
[ADR-098](decisions/098-error-handling-validation-strategy.md).
8. **Not a serialization framework.** The typedef engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use typedef. See [overview.md](overview.md) and
[ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
## References
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
passing, two layout modes, TUnion dispatch, endianness)
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine" — the origin of this research
thread
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
+385
View File
@@ -0,0 +1,385 @@
---
status: draft
last_updated: 2026-07-22
---
# alknet-typedef — Data Access
The data access layer: read/write functions, TUnion dispatch, field paths,
zero-copy access for fixed-size types, and length-prefix reading for
variable-length types. This is the consumer-facing API — given a compiled
`TypedefEngine` and a byte buffer, read and write fields at
schema-computed offsets.
This document covers two layers:
- **Primitive read/write functions** in the `data_access` module —
typed reads/writes at a caller-provided offset. These are the building
blocks used by the layout types (`OffsetMap`, `LayoutBuilder`,
`SequentialReader`) and the `TypedefEngine`. Each operates on a raw
byte buffer at a known offset and returns a `TypedefError::Access`
carrying the field path on bounds or encoding failures.
- **The `FieldValue` enum and the higher-level APIs** —
`TypedefEngine::read_field`/`write_field` (aligned mode) and
`SequentialReader::read_next`/`read_field` (packed mode) — which look
up a field's offset via the layout and dispatch to the primitive
functions, returning a unified `FieldValue<'a>`.
## The `FieldValue` enum
The higher-level read APIs return a single unified type — `FieldValue<'a>`
— so one method can read any field kind without the caller dispatching on
schema kind first. The variant carries the typed value; the lifetime
borrows from the input buffer for variable-length kinds (zero-copy).
```rust
pub enum FieldValue<'a> {
I8(i8), I16(i16), I32(i32), I64(i64),
U8(u8), U16(u16), U32(u32), U64(u64),
F32(f32), F64(f64),
Bool(bool),
Enum(u32), // u32 index into the schema's "enum" array
String(&'a str), // borrows from the buffer
Bytes(&'a [u8]), // borrows from the buffer
Struct { start: usize, end: usize }, // consumer recurses with a fresh reader
Union { discriminator: String, variant_start: usize },
Array { count: u32, element_start: usize, element_stride: usize },
}
```
For composite kinds (`Struct`, `Union`, `Array`), `FieldValue` returns a
layout descriptor, not the decoded contents — the consumer recurses with
a fresh `SequentialReader` (or a sub-range read) scoped to the reported
byte range. `Array`'s `element_stride` is `0` for variable-length element
types, signalling the consumer must walk each element sequentially.
## Read/Write Model
The typedef engine operates on raw byte buffers (`&[u8]` for reading,
`&mut [u8]` for writing). There is no intermediate `Value` tree, no
reflection, no dynamic dispatch per field. The engine uses the offset map
(or `LayoutBuilder`/`SequentialReader`) to locate fields, then performs
typed access at the computed positions.
### Higher-level read/write
The `TypedefEngine` and `SequentialReader` provide the primary
consumer-facing read/write APIs. They look up a field's offset via the
layout and dispatch to the primitive `data_access` functions, returning
`FieldValue` (read) or accepting `&FieldValue` (write).
```rust
impl TypedefEngine {
// Aligned mode: looks up the field's ByteRange in the OffsetMap,
// dispatches to the right data_access function by TypeDefKind.
// Returns TypedefError::Access if compiled in packed mode
// (use sequential_reader() for packed mode).
pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str)
-> Result<FieldValue<'a>, TypedefError>;
pub fn write_field(&self, buffer: &mut [u8], field_path: &str,
value: &FieldValue<'_>) -> Result<(), TypedefError>;
// Packed mode: returns an owned fresh SequentialReader (ADR-101).
// Each call returns a new reader with the cursor at position 0.
// The consumer owns the reader and drives read_next/read_field/reset.
pub fn sequential_reader(&self) -> Option<SequentialReader>;
}
impl SequentialReader {
// Packed mode: walks the buffer field-by-field, reading length
// prefixes to find each field's position. read_field walks all
// preceding fields to reach the target.
pub fn read_next<'a>(&mut self, buffer: &'a [u8])
-> Result<Option<(String, FieldValue<'a>)>, TypedefError>;
pub fn read_field<'a>(&mut self, buffer: &'a [u8], field_path: &str)
-> Result<FieldValue<'a>, TypedefError>;
pub fn reset(&mut self);
pub fn position(&self) -> usize;
pub fn endian(&self) -> Endian;
}
```
`read_field`/`write_field` on `TypedefEngine` work for the fixed-size
primitive kinds and the length-prefixed `String`/`Bytes`/`Timestamp`
fields. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a
`FieldValue` carrying a layout descriptor (byte range, variant start,
or array stride) for the consumer to recurse on — see §"FieldValue" above.
For writing in packed mode, the consumer uses `LayoutBuilder::build` to
compute positions, then calls the primitive `data_access::write_*`
functions at the computed offsets. There is no packed-mode
`engine.write_field` — the layout depends on the actual data sizes,
which the builder consumes at `build` time.
### Primitive read/write functions
The `data_access` module exposes typed read/write functions for each
primitive kind. Each takes `field_path: &str` for error attribution
(produces a `TypedefError::Access` carrying the path on bounds or
encoding failures) and, for multi-byte types, an `Endian` parameter.
### Fixed-size types
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, `TEnum`, etc.) are
accessed via zero-copy reads of N bytes at the offset:
```rust
// Read a u32 at a known offset, applying endianness. Bounds-checked.
fn read_u32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
-> Result<u32, TypedefError> {
let bytes: [u8; 4] = read_array(buffer, offset, field_path)?;
Ok(match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
})
}
// Write a u32 at a known offset, applying endianness. Bounds-checked.
fn write_u32(buffer: &mut [u8], offset: usize, value: u32,
field_path: &str, endian: Endian) -> Result<(), TypedefError> {
let bytes = match endian {
Endian::Little => value.to_le_bytes(),
Endian::Big => value.to_be_bytes(),
};
write_array(buffer, offset, bytes, field_path)
}
```
The engine applies endianness at access time based on the schema's
`"endian"` annotation (ADR-097). The offset computation is
endian-agnostic. The `read_array`/`write_array` helpers perform the
bounds check and produce `TypedefError::Access` with the field path on
failure.
### TEnum access
`TEnum` is a fixed-size type (4 bytes, `u32` index). Read/write delegates
to the `u32` primitives, applying the schema's endianness:
```rust
pub fn read_enum(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
-> Result<u32, TypedefError> {
read_u32(buffer, offset, field_path, endian)
}
```
The consumer maps the `u32` index back to the enum's string values using
the schema's `"enum"` array (index 0 → first value, index 1 → second
value, etc.). The engine does not perform this mapping — it operates on
the raw `u32` index. The jsonschema validator checks that the index
corresponds to a valid enum value at the JSON level.
### Variable-length types (inline length-prefixing)
For variable-length types with inline length-prefixing (the default),
the `data_access` module provides `read_string`/`write_string`/
`read_bytes`/`write_bytes`. Each takes `field_path: &str` for error
attribution and `endian` for the length prefix:
```rust
// Read a length-prefixed string, borrowing from the buffer.
fn read_string<'a>(buffer: &'a [u8], offset: usize,
field_path: &str, endian: Endian) -> Result<&'a str, TypedefError>;
// Write a length-prefixed string. Returns total bytes written (4 + data.len()).
fn write_string(buffer: &mut [u8], offset: usize, value: &str,
field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
// read_bytes / write_bytes have the same shape — raw bytes, no UTF-8 check.
```
The engine reads the 4-byte length prefix at the field's offset, then
slices the data that follows. For writing, the engine writes the length
prefix + data. `read_string` validates UTF-8 and returns a `&str`
borrowing from the input buffer (zero-copy); `read_bytes` returns a
`&[u8]` slice with no encoding check.
In packed sequential mode, the `SequentialReader` uses the length prefix
to determine the position of the next field. In aligned static mode, the
`OffsetMap` records the position of the length prefix; the variable data
is accessed separately.
### Variable-length types (offset indirection)
For variable-length types with offset indirection (opt-in), the
`data_access` module provides `read_string_indirect`/`read_bytes_indirect`.
The 8-byte struct at `buffer[offset..offset+8]` is
`{ data_offset: u32, data_length: u32 }` (endian-aware); the actual
bytes live in a separate `data_region`:
```rust
fn read_string_indirect<'a>(buffer: &'a [u8], offset: usize,
data_region: &'a [u8], field_path: &str,
endian: Endian) -> Result<&'a str, TypedefError>;
fn read_bytes_indirect<'a>(buffer: &'a [u8], offset: usize,
data_region: &'a [u8], field_path: &str,
endian: Endian) -> Result<&'a [u8], TypedefError>;
```
The field is a struct `{offset: u32, length: u32}` at a known position
in the `OffsetMap`. The consumer provides the data region separately; the
engine reads the offset and length, then slices the data region.
## TUnion Dispatch
The `tunion` module provides TUnion discriminator dispatch — reading the
discriminator value from a byte buffer, looking up the variant schema in
the union's `mapping`, and reporting the offset where the variant struct
begins. All reads go through the `data_access` primitives so bounds checks
and endianness handling are uniform with the rest of the engine.
The result of dispatch is a `UnionDispatch` struct:
```rust
pub struct UnionDispatch {
pub key: String, // mapping key (stringified disc value)
pub variant_offset: usize, // byte offset where the variant struct starts
pub discriminator_size: usize, // discriminator's byte size
}
```
After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)`
to get the variant schema, then reads the variant's fields at
`dispatch.variant_offset` using the normal `data_access` functions (or a
fresh `SequentialReader` scoped to the variant).
### Byte-offset discriminator
```rust
/// Read the discriminator value from a byte-offset TUnion. The discriminator
/// is a fixed-size integer (TypeDef:Uint8/Uint16/Uint32) at a known byte
/// offset. Returns the mapping key (stringified integer) and the variant
/// struct offset.
pub fn read_byte_discriminator(
buffer: &[u8],
union_schema: &Value,
endian: Endian,
) -> Result<UnionDispatch, TypedefError>;
```
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
1..N are the variant struct. The call protocol's 5 event types
(`call.requested` → 0x01, etc.) use the same pattern. The variant struct
starts at `offset + discriminator_size`.
### Field-name discriminator
```rust
/// Read the discriminator value from a field-name TUnion. The
/// discriminator is a named field within the struct — the consumer
/// provides the field's computed offset (from the OffsetMap or
/// LayoutBuilder). Supports TypeDef:String, Uint8, and Enum discriminator
/// fields.
pub fn read_field_discriminator(
buffer: &[u8],
union_schema: &Value,
disc_field_offset: usize,
endian: Endian,
) -> Result<UnionDispatch, TypedefError>;
```
The discriminator is a named field within the struct. Its offset is
computed like any other field (the consumer passes it in as
`disc_field_offset`). The mapping keys are string values. After reading
the discriminator, the consumer looks up the variant schema and reads
the variant's fields starting at the end of the discriminator field.
### Variant resolution
```rust
/// Look up a variant schema from the union's mapping. Inline schemas
/// are returned directly. $ref pointers of the form "#/$defs/<name>"
/// are resolved against the union schema's own $defs block.
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str)
-> Result<&'a Value, TypedefError>;
/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a
/// byte-offset TUnion. Field-name discriminators have no fixed size
/// and produce a TypedefError::Schema.
pub fn discriminator_size(union_schema: &Value) -> Result<usize, TypedefError>;
```
### TUnion in the layout engines
The `LayoutBuilder` and `SequentialReader` also handle TUnion fields
inline during traversal (the consumer does not need to call the `tunion`
functions for a union field reached during a sequential walk). For
`LayoutBuilder`, the consumer supplies the discriminator value (byte-offset)
or variant index (field-name) in `var_sizes` under the synthetic key
`"<union_path>.__discriminator"` or `"<union_path>.__variant"`. For
`SequentialReader`, a union field yields
`FieldValue::Union { discriminator, variant_start }`. The standalone
`tunion` functions are for dispatch outside the layout walk — e.g., a
consumer that receives a bare union buffer and needs to identify the
variant before recursing.
## Field Paths
Fields are addressed by dotted paths: `"header.version"`, `"payload.data"`.
Both `OffsetMap` and `PackedLayout` store fully-qualified paths (nested
struct fields appear under their parent's path prefix). The higher-level
APIs (`TypedefEngine::read_field`/`write_field`, `SequentialReader::read_field`)
accept a field path, look up the byte range/position in the layout, and
dispatch to the primitive `data_access` function for the field's kind.
For aligned-mode access, `TypedefEngine::read_field(&buffer, "header.version")`
returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds
the field's `TypeDef:*` kind in the schema, and calls the matching
`data_access::read_*` function. `write_field` is the mirror. Composite
kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue`
carrying a layout descriptor; the consumer recurses with a fresh reader
or sub-range read.
For packed-mode access, `SequentialReader::read_field(&buffer, "c")` walks
all preceding fields to reach the target (sequential access is inherent
to packed layouts). `read_next` walks fields in declaration order.
Nested structs produce nested field paths. The offset computation
propagates the field path prefix during recursion, so the `OffsetMap`
and `PackedLayout` contain entries like `"header.version"` and
`"header.magic"`.
## Zero-Copy Access
For fixed-size types, the engine provides zero-copy access — the consumer
gets a reference to the bytes in the buffer, not a copy. This is
important for performance-sensitive paths (metatensor tensor access,
high-throughput protocol parsing).
For variable-length types with inline length-prefixing, the engine
returns a slice of the buffer — the string or byte array data is not
copied. The consumer gets a `&str` or `&[u8]` that borrows from the
input buffer.
For offset-indirect types, the consumer provides the data region; the
engine returns a slice of that region.
## Error Handling
Read/write errors carry the field path for debugging. See
[ADR-098](decisions/098-error-handling-validation-strategy.md) and
[validation.md](validation.md) for the full error model.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Two layout modes | [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Determines whether offsets are fixed (OffsetMap) or sequential (SequentialReader) |
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Endianness, encoding, and TUnion discriminator shapes that control data access |
| Error handling | [ADR-098](decisions/098-error-handling-validation-strategy.md) | Field-path-carrying errors for read/write operations |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
— affects the sequential walking logic for array access.
## References
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
(read/write round-trip) and POC 2 (SFTP byte-identical round-trip)
- [layout-engine.md](layout-engine.md) — offset computation that produces
the positions this layer reads/writes at
- [validation.md](validation.md) — validation that runs on the same
buffers
@@ -0,0 +1,173 @@
# ADR-095: alknet-typedef — Purpose, Scope, and the jsonschema Engine
## Status
Accepted
## Context
Three threads in the codebase converge on the same pattern: a JSON Schema
describes the shape of binary data, and the binary data is the struct's
bytes at computed offsets.
1. **typedef.ts** (`/workspace/@alkdev/typebox/example/typedef/typedef.ts`,
619 lines) defines custom TypeBox schema kinds (`TFloat32`, `TStruct`,
`TUnion`, etc.) that carry binary layout semantics. These are registered
via `TypeRegistry.Set` with custom validators.
2. **russh-sftp** has 29 packet types, each a struct with typed fields
(`Read { id: u32, handle: String, offset: u64, len: u32 }`). The wire
format is `[length: u32][type: u8][payload]` where payload is the
struct's serde bytes. The `Packet` enum dispatches on the type byte —
a tagged union of structs. Under the typedef lens, each packet is a
`TStruct`; the `Packet` enum is a `TUnion` with a byte-offset
discriminator.
3. **metatensor** needs an offset map for mmap-friendly tensor access —
given a schema describing a model layout (ConvNet struct, tensor refs),
compute byte offsets for each field so the consumer can read tensor
data at known positions without parsing.
The common pattern: **a JSON Schema with `TypeDef:*` custom keywords
describes the shape of binary data; the binary data is the struct's bytes
at computed offsets.** The schema is the format definition; the engine is
generic.
Two prior attempts built their own jsonschema engines — the fatal flaw:
- **typebox-rs** (`/workspace/@alkimiadev/typebox-rs/`, ~8,400 lines):
a full 26-variant `SchemaKind` enum, a custom `Value` type with typed
arrays, and a 912-line hand-written validator.
- **alktype** (`/workspace/@alkimiadev/alktype/`, ~5,600 lines): a
handler-registry pattern that also implements its own validation for
each type.
The `jsonschema` crate (v0.46.5, Draft 2020-12) is already in the
workspace at `/workspace/jsonschema/`. It handles validation with custom
keyword support — the novel code is the offset computation, not the
validation.
The call-channels-unification research
(`docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine") identified the convergence and
bumped typedef up in the timeline. The POC
(`docs/research/alknet-typedef/findings.md`, 26 tests passing) validated
the approach: a ~1,900-line Rust crate that takes a JSON Schema with
`TypeDef:*` custom keywords and produces an offset map, read/write
functions, and validation — all driven by the schema.
## Decision
**alknet-typedef is a small Rust crate that takes a JSON Schema with
`TypeDef:*` custom keywords and produces three capabilities:**
1. **An offset map** — walks the schema, computes byte offsets for each
field based on type sizes, field order, and alignment.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that data
conforms to the schema's type constraints. The jsonschema validator
operates on `serde_json::Value` instances (JSON representations), not
raw byte buffers directly. A consumer that wants to validate a binary
buffer reads it into a `Value` tree via the data access layer, then
validates that `Value` against the jsonschema validator.
**The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing).** The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are ~10 lines each.
**The schema is the format.** A JSON Schema with `TypeDef:Float32`,
`TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
the layout spec. No separate format definition, no separate parser, no
separate validator. One schema, three uses: validate, compute offsets,
access data.
**The crate depends on `jsonschema` and `serde_json` (with
`preserve_order`).** No tokio, no platform deps. Compiles to
`wasm32-unknown-unknown` for browser use. The `jsonschema` crate's
`with_keyword("TypeDef:Float32", factory)` API is the integration point
for custom type kinds — each `TypeDef:*` kind maps to a custom keyword
validator in Rust. Same semantics as TypeBox's `TypeRegistry.Set`, same
JSON Schema wire format.
**The crate targets `std` for v1.** The WASM target has `std` available
via `wasm-bindgen`. If embedded use cases emerge, `no_std` + `alloc` can
be added as a feature gate later — the engine's core (offset computation,
read/write) is already allocation-free. See OQ-070.
## Consequences
### Positive
- **Eliminates ~14,000 lines of hand-rolled schema engines.** typebox-rs
and alktype are replaced by `jsonschema` + an offset map + ~50 lines of
custom keyword implementations. The codebase drops from "a port of
TypeBox" to "jsonschema + an offset map."
- **One schema, three uses.** The same JSON Schema validates, computes
offsets, and drives data access. No separate format definition, parser,
or validator per protocol.
- **Schema-driven, not code-driven.** Adding a new SFTP packet type is
adding a variant to the schema JSON, not writing a new Rust struct +
serde impl. The engine is generic; the schema is the configuration.
- **WASM-clean.** `serde_json` + `jsonschema` + byte manipulation. No
tokio, no platform deps. The same typedef schemas work in browser,
Node, Python (via `wasmtime-py`), Go (via `wazero`), and any other
WASM host.
- **TypeBox interop.** TypeBox modules render to standard JSON Schema
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
on the Rust side. Zero translation. The same schema validates in both
ecosystems.
- **Defense in depth.** Schema validation via jsonschema custom keywords —
a malformed binary payload can be read into a `Value` tree via the data
access layer and validated against the schema before any consumer
touches it. The `jsonschema` crate's compiled validators are fast enough
to run on every incoming frame.
### Negative
- **New dependency on `jsonschema`.** The crate is already in the
workspace but not yet used by any alknet crate. This is the first
consumer.
- **`serde_json` with `preserve_order` is required.** Field order is
load-bearing for binary layouts. The `preserve_order` feature adds a
small compile-time cost.
- **Schema authoring is external.** Schemas are authored in TypeBox (JS)
or hand-written JSON. The typedef engine consumes schemas; it does not
generate them. A builder API is deferred (OQ-071).
## Scope Boundaries (What This Is Not)
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
is a *format* (8-byte header + JSON header + binary data) that uses the
typedef engine for its offset computation and tensor access.
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
`Value.Convert` — schema evolution — is out of scope for v1. The engine
should not do anything that explicitly blocks adding a Value system
later.
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
concern. The typedef engine consumes schemas; it does not generate them.
- **Not a schema builder.** The typedef engine does not provide a fluent
API for constructing schemas. Schemas are plain JSON.
- **Not a serialization framework.** The typedef engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use typedef.
## References
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
passing, two layout modes, TUnion dispatch, endianness)
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine" — the origin of this research
thread
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes decision
- [ADR-097](097-schema-annotations.md) — schema annotation shapes
- [ADR-098](098-error-handling-validation-strategy.md) — error handling
and validation strategy
@@ -0,0 +1,137 @@
# ADR-096: Two Layout Modes — Packed Sequential vs Aligned Static
## Status
Accepted
## Context
The POCs surfaced that protocols and mmap-friendly formats need different
layout strategies. POC 1 built an aligned `OffsetMap` with natural
alignment padding — correct for mmap-friendly formats (metatensor) but
wrong for protocol wire formats (SFTP, channels, TTY). POC 2 built a
`LayoutBuilder` and `SequentialReader` for packed sequential layouts —
correct for protocol wire formats but wrong for mmap-friendly formats.
This is the most important architectural finding from the POCs. The
engine must support both modes; a single layout strategy cannot serve
both use cases.
### Packed sequential layout (protocol wire formats)
Protocols pack fields sequentially with no alignment padding.
Variable-length fields shift all subsequent fields. Writing requires
knowing actual data sizes upfront; reading walks the buffer sequentially,
reading length prefixes to determine positions.
This is the layout used by SFTP (all strings and byte arrays are
length-prefixed inline), channels (`[channel_id: u32][size: u32][payload]`),
TTY (`[stream_type: u8][length: u32][payload]`), and most binary protocols.
### Aligned static layout (mmap-friendly formats)
Fields have fixed positions with natural alignment padding.
Variable-length fields get a 4-byte length prefix at a known offset; the
variable data is not included in the static layout. This enables
mmap-friendly random access — the consumer can read field N at a known
offset without parsing the fields before it.
This is the layout used by metatensor (blob tensor pattern: index struct
in one region, blob data in another) and safetensors (header + aligned
tensor data).
## Decision
**The typedef engine supports two layout modes, selected by the consumer
at engine construction time:**
### Mode 1: Packed sequential (`LayoutBuilder` / `SequentialReader`)
For protocol wire formats. Fields are packed with no alignment padding.
Variable-length fields shift all subsequent fields.
- **LayoutBuilder** — takes a schema and actual data sizes for
variable-length fields, computes byte positions for each field in a
packed layout. Used at write time when the consumer knows the data
sizes upfront.
- **SequentialReader** — walks a buffer field-by-field according to the
schema, reading length prefixes to determine variable-length data
positions. Used at read time when the consumer is parsing an incoming
frame.
The `LayoutBuilder` and `SequentialReader` are the primary interface for
protocol consumers (SFTP, binary call frames, TTY negotiation).
### Mode 2: Aligned static (`OffsetMap`)
For mmap-friendly formats. Fields have fixed positions with natural
alignment padding. Variable-length fields get a 4-byte length prefix at
a known offset; the variable data is not included in the static layout.
- **OffsetMap** — walks the schema once, computes fixed byte positions
for each field based on type sizes and alignment. The output is a flat
table of `(field_path, byte_range)` pairs. Used for both read and write
at known offsets.
The `OffsetMap` is the primary interface for mmap consumers (metatensor).
### Variable-length handling in each mode
**Packed sequential mode:** Variable-length fields are inline
length-prefixed by default (`[length: u32][data]`). The `LayoutBuilder`
takes the actual data size to compute the length prefix value and the
position of subsequent fields. The `SequentialReader` reads the length
prefix to determine the data extent and the position of the next field.
**Aligned static mode:** Variable-length fields get a 4-byte length
prefix at a known offset. The variable data lives outside the static
layout — either immediately after the fixed fields (inline
length-prefixing) or in a separate data region (offset indirection, the
metatensor blob tensor pattern). The `OffsetMap` records the position of
the length prefix (or the `{offset, length}` pair for offset-indirect
fields).
### Default for variable-length types
Inline length-prefixing (`[length: u32][data]`) is the default for all
variable-length types in both modes. This is the universal pattern used
by channels, SFTP, TTY, and most binary protocols. Offset indirection is
opt-in via the `encoding` annotation (see ADR-097).
## Consequences
### Positive
- **One engine, two modes.** The same schema can be used in either mode.
A schema describing an SFTP packet can be consumed by a `SequentialReader`
(for parsing incoming frames) and a `LayoutBuilder` (for constructing
outgoing frames). A schema describing a metatensor layout can be
consumed by an `OffsetMap` (for mmap access).
- **Correct for both use cases.** Packed sequential mode produces
byte-identical output to hand-written protocol serialization (validated
by POC 2's russh-sftp round-trip tests). Aligned static mode produces
correct offsets for mmap-friendly access (validated by POC 1's
alignment tests).
- **No mode confusion.** The consumer explicitly selects the mode at
engine construction time. A protocol consumer never accidentally gets
alignment padding; an mmap consumer never accidentally gets
variable-length field shifting.
### Negative
- **Two APIs to learn.** Consumers must choose between
`LayoutBuilder`/`SequentialReader` and `OffsetMap`. The choice is
determined by the use case (protocol vs mmap), not by the schema.
- **Variable-length fields in packed mode require size foreknowledge.**
The `LayoutBuilder` needs actual data sizes for variable-length fields
to compute correct positions for subsequent fields. This is inherent
to packed layouts — the consumer must know the data sizes before
writing.
## References
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-097](097-schema-annotations.md) — schema annotations including
the `encoding` field for variable-length types
@@ -0,0 +1,259 @@
# ADR-097: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
## Status
Accepted
## Context
The typedef engine needs concrete JSON shapes for schema-level
annotations that control binary layout behavior. The POCs validated the
semantics; this ADR pins the shapes.
Four annotation categories need concrete shapes:
1. **Endianness** — safetensors is little-endian, SFTP is big-endian.
The engine needs to know which to use.
2. **Alignment** — different backends have different alignment
requirements (wgpu: 256-byte, protocols: natural, mmap: page).
3. **Variable-length encoding** — inline length-prefixing vs offset
indirection for strings, byte arrays, and other variable-length types.
4. **TUnion discriminators** — byte-offset (protocol dispatch) vs
field-name (typedef.ts pattern).
## Decision
### 1. Endianness
**Schema-level annotation with a default of little-endian.**
```json
{
"TypeDef:Struct": true,
"endian": "big",
"properties": { ... }
}
```
- `"endian": "little"` (default) — read/write in little-endian byte order.
- `"endian": "big"` — read/write in big-endian byte order.
- The annotation applies to the entire schema and all nested types.
- Mixed endianness within one schema is not supported (pathological; no
known protocol requires it).
- The default is little-endian, matching safetensors, wgpu, and most
modern formats. SFTP consumers specify `"endian": "big"`.
### 2. Alignment
**Both struct-level and field-level, with field-level overriding
struct-level.**
```json
{
"TypeDef:Struct": true,
"align": 256,
"properties": {
"header": { "TypeDef:Struct": true, "properties": { ... } },
"weight": { "TypeDef:Float32": true, "align": 16 }
}
}
```
- Struct-level `"align"` sets the default alignment for all fields in
that struct. The struct's total size is rounded up to this alignment.
- Field-level `"align"` overrides the struct default for that specific
field.
- Default alignment (when no annotation is present): 1 for u8/bool, 2
for u16/i16, 4 for u32/i32/f32, 8 for u64/i64/f64, max field alignment
for structs.
- Alignment is only meaningful in aligned static mode (ADR-096). In
packed sequential mode, alignment annotations are ignored — fields are
packed with no padding.
### 3. Variable-length encoding
**Three strategies for variable-length types, selected by the `encoding`
annotation and the standard JSON Schema `maxLength` keyword.**
```json
// Strategy 1: Inline length-prefixing (default, shorthand)
{ "TypeDef:String": true }
// Strategy 1: Explicit inline length-prefixing
{ "TypeDef:String": { "encoding": "length-prefixed" } }
// Strategy 2: Fixed-size reservation (uses standard maxLength)
{ "TypeDef:String": true, "maxLength": 256 }
// Strategy 3: Offset indirection (opt-in)
{ "TypeDef:String": { "encoding": "offset-indirect" } }
```
**Strategy 1: Inline length-prefixing (default).** The field's fixed
portion is a 4-byte length prefix at a computed offset. The variable data
follows immediately after. In packed sequential mode, the length prefix
determines the position of subsequent fields. In aligned static mode, the
length prefix is at a known offset; the variable data is not included in
the static layout. This is the universal pattern used by channels, SFTP,
TTY, and most binary protocols.
**Strategy 2: Fixed-size reservation.** When a variable-length field
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
than `maxLength` is zero-padded; data longer than `maxLength` is a
validation error. This makes the field fixed-size from the layout
perspective — subsequent fields have known, unchanging offsets. This is
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
pattern for fields with known maximum sizes. In packed sequential mode,
`maxLength` is a validation constraint only — the engine still uses
inline length-prefixing (strategy 1).
**Strategy 3: Offset indirection.** The field is a struct
`{offset: u32, length: u32}` that points into a separate data region.
This is the metatensor blob tensor pattern — the index struct lives in
one region, the blob data lives in another. The consumer provides the
data region separately. Enables mmap-friendly random access to
variable-length data without parsing length prefixes and without
reserving worst-case space.
**Default strategy selection:**
- In packed sequential mode: always strategy 1 (inline length-prefixing).
`maxLength` is a validation constraint only.
- In aligned static mode: strategy 2 (fixed-size reservation) if
`maxLength` is declared; strategy 3 (offset indirection) if
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
length-prefixing) otherwise.
- `true` is a shorthand for the default (length-prefixed). This keeps
the common case concise and the override explicit.
- The `encoding` annotation and `maxLength` apply to all variable-length
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
`TypeDef:Record`, `TypeDef:Timestamp`.
### 3a. TRecord value type
`TypeDef:Record` is a string-keyed map. The value type is declared via
the `"values"` property in the schema:
```json
{
"TypeDef:Record": true,
"values": { "TypeDef:Float32": true }
}
```
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
values in the record. All values share the same type.
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count`
times. Each key is a length-prefixed UTF-8 string. Each value is
encoded according to its declared `TypeDef:*` kind — a `Record<Uint32>`
value is 4 raw bytes; a `Record<String>` value is itself a
length-prefixed string; a `Record<Struct>` value is the struct's
fields laid out inline. There is **no separate `value_len` prefix** —
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix).
- The count and key-length prefixes respect the schema's endianness.
- In aligned static mode with `maxLength`, the entire record is reserved
at `maxLength` bytes (zero-padded).
### 4. TUnion discriminators
**Two discriminator kinds: byte-offset (protocol dispatch) and
field-name (typedef.ts pattern).**
#### Kind A: Byte-offset discriminator
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "TypeDef:Uint8"
},
"mapping": {
"1": { "$ref": "#/$defs/Init" },
"3": { "$ref": "#/$defs/Open" },
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" },
"101": { "$ref": "#/$defs/Status" }
}
}
```
- The discriminator is a fixed-size integer at a known byte offset.
- `"offset"` is the byte position of the discriminator within the union's
buffer.
- `"type"` is the `TypeDef:*` kind of the discriminator (typically
`TypeDef:Uint8` for protocol type bytes).
- The mapping keys are stringified integers (`"1"`, `"5"`, `"101"`).
The engine parses the key to match the discriminator value.
- The variant struct starts at `offset + discriminator_size`.
- This is the SFTP `Packet` enum pattern and the call protocol's event
type dispatch.
#### Kind B: Field-name discriminator
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "field",
"name": "type"
},
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
}
```
- The discriminator is a named field within the struct.
- `"name"` is the field name that holds the discriminator value.
- The mapping keys are string values matching the discriminator field's
value.
- The discriminator field is just another field in the struct — its
offset is computed like any other field.
- This is the typedef.ts `TUnion` pattern.
#### Mapping values
Mapping values may be either inline schemas or `$ref` pointers. `$ref`
is cleaner for large unions (29 SFTP variants) but requires a `$defs`
section. Inline schemas are simpler for small unions (5 call protocol
event types). Both work.
## Consequences
### Positive
- **Concrete, validated shapes.** All four annotation categories have
concrete JSON shapes that were validated by the POCs.
- **Sensible defaults.** Little-endian, natural alignment, inline
length-prefixing — the common case requires no annotations.
- **Explicit overrides.** Big-endian, custom alignment, offset
indirection — the uncommon case is explicit and self-documenting.
- **TUnion covers both protocol and typedef.ts patterns.** The
byte-offset discriminator handles SFTP type bytes and call protocol
event types. The field-name discriminator handles the typedef.ts string
pattern. No separate union type needed.
### Negative
- **Keyword value shape change.** `"TypeDef:String": true` (boolean) and
`"TypeDef:String": { "encoding": "length-prefixed" }` (object) are both
valid. The engine must handle both shapes. This is a minor parsing
concern — the POC already handles it.
- **Alignment annotations are mode-specific.** Alignment is only
meaningful in aligned static mode. In packed sequential mode, alignment
annotations are ignored. This is documented, not enforced — a consumer
that specifies alignment in packed mode gets no error, just no effect.
## References
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — the
annotation shape questions this ADR resolves
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes (alignment only meaningful in aligned static mode)
@@ -0,0 +1,157 @@
# ADR-098: Error Handling and Validation Strategy
## Status
Accepted
## Context
The typedef engine operates in three phases, each with distinct error
conditions:
1. **Schema parsing** — invalid JSON, missing required keywords, unknown
`TypeDef:*` kinds, malformed annotations.
2. **Offset computation** — field not found, type not supported for
offset computation, recursive schema depth exceeded.
3. **Read/write** — buffer too short, invalid UTF-8, value out of range
for the target type.
4. **Validation** — type constraint violations (range, UTF-8, field
presence, discriminator membership).
The engine also needs a clear strategy for *when* validation happens:
once at schema load time (build the validator) vs repeatedly at access
time (validate each buffer).
## Decision
### Error type: `TypedefError`
A single `TypedefError` enum with variants for each error category:
```rust
pub enum TypedefError {
/// Schema parsing errors.
Schema(String),
/// Offset computation errors.
Offset { field_path: String, reason: String },
/// Read/write errors.
Access { field_path: String, reason: String },
/// Validation errors (delegated to jsonschema).
Validation(ValidationError<'static>),
}
```
- `Schema` — for invalid JSON, missing required keywords, unknown
`TypeDef:*` kinds. The error message describes the problem.
- `Offset` — for field-not-found, unsupported type for offset
computation, etc. Carries the field path for debugging.
- `Access` — for buffer-too-short, invalid UTF-8, value out of range.
Carries the field path for debugging.
- `Validation` — wraps `jsonschema`'s `ValidationError`. The
`jsonschema` crate already provides rich error messages with schema
paths; the typedef engine does not re-wrap or re-interpret them.
The `Validation` variant uses `ValidationError<'static>` because the
validator is built once at schema load time and lives for the lifetime
of the `TypedefEngine`. The `'static` lifetime is correct — the validator
owns its schema reference.
### Validation timing: load-time build, access-time check
The jsonschema validator is built once at schema load time
(`validator_for(&schema)?`) and then called repeatedly
(`validator.is_valid(&instance)`). The typedef engine follows the same
pattern:
1. **Load time:** Parse the schema JSON, build the offset map (or
`LayoutBuilder`/`SequentialReader`), build the jsonschema validator.
This is the `TypedefEngine::compile(schema: &Value) -> Result<Self,
TypedefError>` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation — the consumer calls
`engine.validate(buffer)` when validation is desired.
The `TypedefEngine` struct is the compiled form of a schema:
```rust
pub struct TypedefEngine {
offset_map: OffsetMap, // or LayoutBuilder/SequentialReader
validator: jsonschema::Validator, // compiled once at load time
}
```
### Custom keyword validators
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
`jsonschema::options().with_keyword(...)`. The validators check:
- **Numeric types** (`TypeDef:Float32`, `TypeDef:Int8`, etc.): range
constraints (Int8: -128..127, Uint8: 0..255, etc.), finiteness for
floats.
- **`TypeDef:String`**: UTF-8 validity.
- **`TypeDef:Struct`**: field presence and types (delegated to
jsonschema's structural validation — the custom keyword only needs to
validate that the struct's fields match their declared `TypeDef:*`
kinds).
- **`TypeDef:Union`**: discriminator value membership in the mapping.
- **`TypeDef:Array`**: element type conformance.
- **`TypeDef:Boolean`**: value is `true` or `false`.
- **`TypeDef:Timestamp`**: RFC 3339 string format (the internet profile of ISO 8601).
The `jsonschema` crate handles all the structural validation (object
properties, required fields, array items, enum values) — the custom
keywords only need to validate the leaf type constraints. Each custom
keyword implementation is ~10 lines.
### Read/write errors carry field paths
Read/write errors include the field path for debugging:
```rust
// Example: reading a u32 from a buffer that's too short
Err(TypedefError::Access {
field_path: "header.version".to_string(),
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
})
```
This makes debugging binary format issues tractable — the error tells
you exactly which field failed and why.
## Consequences
### Positive
- **Single error type.** Consumers handle one `TypedefError` enum, not
multiple error types from different engine phases.
- **Field-path-carrying errors.** Read/write errors include the field
path, making binary format debugging tractable.
- **Validation is opt-in.** The consumer decides when to validate.
High-throughput paths can skip validation; security-sensitive paths
can validate every frame.
- **jsonschema integration is clean.** The `ValidationError` is wrapped
as-is — no re-interpretation, no information loss.
- **Load-time build, access-time use.** The expensive work (schema
parsing, validator compilation, offset computation) happens once at
load time. Access-time operations are cheap (pointer casts, slice
operations, length-prefix reads).
### Negative
- **`ValidationError<'static>` lifetime.** The `'static` lifetime on the
`Validation` variant means the error cannot borrow from the buffer
being validated. This is correct (the validator owns its schema
reference) but may surprise readers who expect a shorter lifetime.
- **No error recovery.** The engine does not attempt to recover from
partial reads or writes. A buffer-too-short error on field N means
fields N+1.. are also unreadable. This is inherent to binary formats
— there is no "skip to next field" without a schema-driven parser.
## References
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — error
handling strategy question (OQ 8)
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes
- [ADR-097](097-schema-annotations.md) — schema annotations
@@ -0,0 +1,114 @@
# ADR-099: Int64/Uint64 as First-Class Kinds
## Status
Accepted
## Context
The typedef engine's kind set (ADR-095, ADR-097) tops out at 32-bit
integers. The POC included `u64` read/write primitives, and the
call-channels-unification research's own SFTP schema example uses
`"TypeDef:Uint64"` for the `offset` field (`Read`/`Write` packets have
`offset: u64`). Metatensor/safetensors `data_offsets` are also `u64`.
A `TypeDef:Uint64` variant was added to the `TypeDefKind` enum during
implementation (the task decomposition correctly identified the gap),
but without an ADR the addition was half-finished: `type_size()`
returned `None`, the layout engines couldn't compute offsets for it, and
the validator didn't register a `TypeDef:Uint64` keyword. The variant
was then removed (commit `14d9cf2`) on the grounds that it was
unintended and a latent panic — but the underlying gap is real: SFTP and
metatensor, the two primary POC targets, both require 64-bit integers.
The presumed reason 64-bit integers were left out of the original
specification is a JSON-level concern: `serde_json::Number` loses
precision past 2^53 when parsing from JSON text. This is a
*validation-layer* caveat, not a *layout-layer* one — the binary layout
is 8 raw bytes, and `from_le_bytes`/`from_be_bytes` work correctly for
the full `u64`/`i64` range. The validation concern is handled by
accepting integer-form JSON values (the `jsonschema` crate's
`as_i64`/`as_u64` methods handle the common range; values past 2^53 are
a JSON representation limitation, not a typedef limitation).
## Decision
**Add `TypeDef:Int64` and `TypeDef:Uint64` as first-class kinds.**
Both are fixed-size (8 bytes), with natural alignment 8. They follow
the schema's endianness annotation like all other fixed-size types.
Read/write is via `data_access::read_i64`/`write_i64`/`read_u64`/
`write_u64` (endian-aware, 8 bytes).
### Kind table additions
| Kind | TypeBox key | Rust type | Size | Alignment |
|------|-------------|-----------|------|-----------|
| `TInt64` | `TypeDef:Int64` | `i64` | 8 | 8 |
| `TUint64` | `TypeDef:Uint64` | `u64` | 8 | 8 |
### Validation
The custom keyword validators check:
- `TypeDef:Int64`: value must be an integer in `i64::MIN..=i64::MAX`
(`-9223372036854775808` to `9223372036854775807`).
- `TypeDef:Uint64`: value must be a non-negative integer in
`0..=u64::MAX` (`0` to `18446744073709551615`).
The `jsonschema` crate's `as_i64`/`as_u64` handle the common range.
JSON numbers past 2^53 lose precision in the JSON representation —
this is a JSON limitation, not a typedef limitation. The binary
representation (8 raw bytes) is always exact. A consumer that needs
to validate the full 64-bit range from JSON should provide the value
as a JSON integer (which `serde_json` preserves for values up to
`u64::MAX`/`i64::MIN` when the `arbitrary_precision` feature is
enabled, or when the value fits in `i64`/`u64` without the feature).
### `FieldValue` additions
`FieldValue::I64(i64)` and `FieldValue::U64(u64)` are added to the
unified return type. The `SequentialReader`, `TypedefEngine::read_field`,
and `TypedefEngine::write_field` dispatch on the new kinds.
### Kind count
The engine now has **19** first-class kinds (17 + Int64 + Uint64).
`TypeDefKind::is_fixed_size()` returns `true` for both new kinds.
`type_size()` returns `Some(8)`. `natural_alignment()` returns `8`.
`needs_endian()` returns `true`.
## Consequences
### Positive
- **Unblocks the two primary POC targets.** SFTP `Read`/`Write` packets
(`offset: u64`) and metatensor `data_offsets` (`u64`) are now
expressible in typedef schemas.
- **Completes the half-finished addition.** The `TypeDefKind` enum,
`data_access` primitives, and `FieldValue` variants for 64-bit
integers now have matching layout, validator, and engine support.
- **No new design surface.** Int64/Uint64 are fixed-size types that
follow all existing patterns (endianness, alignment, zero-copy
read/write). They are mechanical additions.
### Negative
- **JSON precision caveat.** Values past 2^53 lose precision in the
JSON representation (not in the binary representation). This is a
JSON limitation, not a typedef limitation, but it means the
validation layer cannot perfectly round-trip the full 64-bit range
through JSON `Number` without `arbitrary_precision`. In practice,
SFTP offsets and tensor data offsets are well within 2^53.
- **Two more kinds to maintain.** The kind table, validator
registration, `FieldValue` enum, and dispatch arms all grow by two
variants. This is the cost of completeness.
## References
- `docs/research/call-channels-unification/findings.md` §"russh-sftp" —
the SFTP schema with `"offset": { "TypeDef:Uint64": true }`
- `docs/research/alknet-typedef/findings.md` §"POC 1" — the POC included
u64 read/write
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope (the kind set)
- [ADR-097](097-schema-annotations.md) — schema annotations
(endianness applies to the new kinds)
@@ -0,0 +1,112 @@
# ADR-100: Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode
## Status
Accepted
## Context
The aligned static layout mode (ADR-096) is designed for mmap-friendly
formats: fields have fixed positions with natural alignment padding,
enabling random access by field path without parsing preceding fields.
The spec (layout-engine.md) says variable-length fields in aligned mode
get a 4-byte length prefix at a known offset, and "the variable data
lives outside the static layout — either immediately after the fixed
fields (inline length-prefixing) or in a separate data region (offset
indirection)."
The implementation has a bug: `OffsetMap::compute` reserves only 4 bytes
for an inline length-prefixed variable field (the length prefix), but
`TypedefEngine::write_field` for a `String`/`Bytes` field calls
`data_access::write_string` at `range.start`, which writes
`[4-byte length][data]` inline — clobbering every subsequent field. The
`read_field` path has the mirror behavior (reads inline), so the engine
is self-consistent but only works correctly when the variable field is
the last field in the struct (no subsequent field to clobber).
Concretely, `{name: String, id: Uint32}` in aligned mode maps
`name → 0..4`, `id → 4..8`. Writing `"hello"` to `name` writes
`[5,0,0,0,h,e,l,l,o]` at offset 0, overwriting `id`'s range with
`hello`. All existing tests happen to put the variable field last, so
the bug is latent.
The spec's "data region after fixed fields" model (where variable data
lives after all fixed fields) is the correct design for aligned mode,
but implementing it would require a two-region layout (fixed fields +
variable data region) with the `OffsetMap` tracking both the prefix
position and the data position. This is a significant design addition
for a use case that doesn't exist yet — real aligned-format consumers
(metatensor, safetensors) use `maxLength` reservation or
`offset-indirect` encoding for variable data, not inline
length-prefixing.
## Decision
**Reject non-final inline length-prefixed variable fields in aligned
static mode at `OffsetMap::compute` time.**
A variable-length field (`TypeDef:String`, `TypeDef:Bytes`,
`TypeDef:Timestamp`, `TypeDef:Record`) in aligned static mode that uses
the default inline length-prefixing strategy (no `maxLength`, no
`offset-indirect`) must be the last field in its struct. If a non-final
inline length-prefixed variable field is encountered,
`OffsetMap::compute` returns `TypedefError::Offset` with a message
explaining that non-final variable fields in aligned mode require
`maxLength` (fixed-size reservation) or `"encoding": "offset-indirect"`
(offset indirection).
This is a validation-time rejection (schema load time), not a runtime
check. The consumer learns about the problem when compiling the schema,
not when writing data.
### What is NOT rejected
- Inline length-prefixed variable fields that are the last field in
their struct — these are fine (no subsequent field to clobber).
- `maxLength` reservation and `offset-indirect` encoding in any
position — these make the field fixed-size from the layout
perspective (known size at a known offset), so they don't clobber.
- Inline length-prefixed variable fields in packed sequential mode —
packed mode doesn't have fixed offsets; variable fields shift
subsequent fields by design.
## Consequences
### Positive
- **Eliminates a silent data-corruption bug.** A consumer that writes
a non-final string in aligned mode currently clobbers subsequent
fields with no error. After this fix, the schema is rejected at
compile time.
- **Matches real aligned-format usage.** mmap-friendly formats use
`maxLength` or `offset-indirect` for variable data; inline
length-prefixing in aligned mode is only meaningful as the last
field.
- **Simple to implement.** A single check in `compute_struct` (is this
variable field non-final and using inline length-prefixing? → reject).
No two-region layout needed.
- **Defers the two-region design without blocking consumers.** If a
future consumer needs inline length-prefixing in non-final position
in aligned mode, the two-region layout can be implemented then. The
rejection is reversible (remove the check, add the two-region logic).
### Negative
- **A schema that worked before (silently corrupting data) now fails
at compile time.** This is the correct behavior — the schema was
always broken, it just wasn't caught.
- **The "data region after fixed fields" model from the spec is not
implemented.** A consumer that wants inline variable data in a
non-final position must use packed mode or wait for the two-region
layout. This is acceptable for v1 — no current consumer needs it.
## References
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes (aligned static mode's variable-length handling)
- [ADR-097](097-schema-annotations.md) — the three variable-length
encoding strategies (`maxLength`, `offset-indirect`, inline
length-prefixing)
- `../layout-engine.md` §"Variable-length
fields in aligned mode" — the spec's "data region after fixed fields"
description
@@ -0,0 +1,103 @@
# ADR-101: Packed-Mode Read API — Engine as SequentialReader Factory
## Status
Accepted
## Context
`TypedefEngine` stores a `SequentialReader` inside its `Layout::Packed`
variant. The engine exposes it via
`engine.sequential_reader() -> Option<&SequentialReader>`.
The problem: `SequentialReader`'s read methods (`read_next`,
`read_field`, `reset`) all take `&mut self` — they mutate the reader's
internal cursor (`field_index`, `position`). But the engine hands out
`&SequentialReader` (a shared reference), which cannot be used to call
`&mut self` methods. The accessor can only give the consumer
`position()` and `endian()` (the `&self` methods) — the actual read
API is unreachable.
This makes the engine's packed read-side dead API. A consumer that
wants to read a packed buffer must construct their own
`SequentialReader::new(&schema)` from the schema, bypassing the engine
entirely. The stored reader is dead weight.
Three options were considered:
1. **Factory method** — the engine provides a method that returns an
owned fresh `SequentialReader` (reconstructed from the stored
schema). The consumer owns the reader and drives it with `&mut self`.
2. **Interior mutability** — wrap the reader in `Mutex` or `RefCell`
so `&SequentialReader` can be upgraded to `&mut`. Adds overhead and
complexity for mutable cursor state that the consumer legitimately
wants to own.
3. **`sequential_reader_mut()`** — return `&mut SequentialReader`.
Requires `&mut self` on the engine, which is overly restrictive
(the consumer may share the engine across threads or hold it behind
an `Arc`).
## Decision
**The engine is a `SequentialReader` factory.** Replace
`sequential_reader() -> Option<&SequentialReader>` with
`sequential_reader() -> Option<SequentialReader>` — the method returns
an owned fresh reader, reconstructed from the stored schema.
```rust
impl TypedefEngine {
/// Construct a fresh SequentialReader for packed-mode reads.
/// Returns None if compiled in aligned mode.
pub fn sequential_reader(&self) -> Option<SequentialReader>;
}
```
Each call returns a new reader with the cursor at position 0. The
consumer owns the reader and calls `read_next`/`read_field`/`reset` on
it directly. The engine still stores its own reader (used for schema
validation during construction), but no longer exposes it by
reference.
The same applies to `LayoutBuilder`: `layout_builder()` returns
`Option<&LayoutBuilder>` which is fine — `LayoutBuilder::build` takes
`&self`, so the shared reference is usable. No change needed for the
write-side.
### Cost
`SequentialReader::new` clones the top-level struct's field schemas (a
`Vec<(String, Value)>` of the `properties` entries) and clones the
schema itself. This is cheap — a struct has a small number of fields
(SFTP's largest packet has 5). The construction cost is negligible
compared to the cost of reading a buffer.
## Consequences
### Positive
- **The packed read API is now usable.** A consumer calls
`engine.sequential_reader()` to get an owned reader and drives it
directly. No dead API.
- **No interior mutability overhead.** The reader's mutable cursor
state is owned by the consumer, not shared through a lock.
- **Thread-safe engine.** The engine remains `Send + Sync` (it only
exposes `&self` methods). The reader is owned by the calling thread.
- **Simple.** One method signature change. The stored reader in
`Layout::Packed` can be removed (it was only used for schema
validation during construction, which is done by the time the
consumer calls `sequential_reader()`).
### Negative
- **Each call to `sequential_reader()` allocates a new reader.** The
cost is a `Vec` of field schemas + a schema clone. Acceptable for
the use case (one reader per buffer read).
- **The engine no longer holds a live reader.** If a future use case
needs to share a reader's cursor state across calls, the consumer
must manage that themselves. This is the correct separation — cursor
state is consumer-owned, not engine-owned.
## References
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — packed
sequential mode (`SequentialReader` as the read-side)
- `../data-access.md` §"Higher-level
read/write" — the `SequentialReader` API
@@ -0,0 +1,111 @@
# ADR-102: Reject TUnion in Aligned Mode for v1
## Status
Accepted
## Context
The aligned static layout mode (ADR-096) computes fixed byte positions
for each field, enabling random access by field path. `TUnion` in
aligned mode has three implementation problems:
1. **No variant field offsets.** Only the `__discriminator` byte range
is recorded in the `OffsetMap`. Variant field offsets are not
available anywhere in aligned mode — the consumer must recompute
them by hand. This makes `TypedefEngine::read_field` on a union
variant field impossible.
2. **`find_discriminator_field` takes the first variant's offset.** For
a field-name discriminator, the code probes the first variant that
contains the discriminator field and records that offset globally.
If variants order fields differently, the discriminator sits at
different offsets per variant and the recorded range is silently
wrong. The code should validate that the offset is identical across
all variants (or require the discriminator field to be first).
3. **Byte-discriminator union total misaligns the variant.** The union
total is `disc_off + disc_size + variant_max_size`, but the variant
was probed from offset 0 with alignment. A `u8` discriminator
before a `u32`-bearing variant produces a variant region that
starts at an unaligned offset in a mode whose entire purpose is
alignment.
The real question is whether `TUnion` in aligned mode is even needed.
The two consumer profiles are:
- **Protocol consumers** (SFTP, call protocol event types): use packed
sequential mode. `TUnion` with byte-offset discriminators is the
core dispatch mechanism. This is well-supported.
- **mmap consumers** (metatensor, safetensors): use aligned static
mode. These formats are structs and arrays of structs — they don't
use tagged unions. A tensor file has a header struct with tensor
descriptors, not a "which variant is this?" dispatch.
`TUnion` in aligned mode is a combination that no current or planned
consumer needs. Shipping broken semantics for an unused use case is
worse than rejecting it clearly.
## Decision
**Reject `TUnion` in aligned static mode for v1.**
`OffsetMap::compute` returns `TypedefError::Offset` when it encounters
a `TypeDef:Union` field, with a message explaining that unions are not
supported in aligned mode and the consumer should use packed mode (or
restructure as a struct with an explicit discriminator field).
This is a schema-load-time rejection. The consumer learns about the
problem when compiling the schema, not at runtime.
### What is NOT rejected
- `TUnion` in packed sequential mode — this is the core use case
(SFTP `Packet` dispatch, call protocol event types) and is fully
supported by `LayoutBuilder` and `SequentialReader`.
- `TStruct`, `TArray`, and all primitive kinds in aligned mode — these
are the mmap-format primitives and are fully supported.
### Reversal
This is a two-way door. If a future mmap-format consumer needs tagged
unions in aligned mode, the rejection can be lifted and the three
implementation problems fixed. The fix would require:
- Recording per-variant field offsets in the `OffsetMap` (which
variant's offsets to record when variants have different layouts?).
- Validating that field-name discriminators have identical offsets
across all variants.
- Aligning the variant region correctly after the byte discriminator.
These are design questions that should be answered when the use case
arrives, not speculatively now.
## Consequences
### Positive
- **No broken semantics shipped.** The three implementation problems
are removed from the API surface rather than silently producing
wrong offsets.
- **Clear scope boundary.** Aligned mode is for structs and arrays;
packed mode is for protocols (including union dispatch). The
consumer chooses the mode based on the use case.
- **Reversible.** When a real consumer needs aligned-mode unions, the
rejection is lifted and the design questions are worked through with
a concrete use case.
### Negative
- **A schema with a `TUnion` field cannot be compiled in aligned
mode.** A consumer that wants both aligned layout and union dispatch
must use packed mode or restructure. No current consumer needs this.
- **The aligned-mode union code in `offset_map.rs` is dead.** It can
be removed or left as a reference for when the rejection is lifted.
Removing it is cleaner.
## References
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes
- [ADR-097](097-schema-annotations.md) §4 — TUnion discriminators
- `../layout-engine.md` §"TUnion" —
aligned-mode union sizing
+360
View File
@@ -0,0 +1,360 @@
---
status: draft
last_updated: 2026-07-22
---
# alknet-typedef — Layout Engine
The layout engine: offset computation, the two layout modes (packed
sequential vs aligned static), alignment, endianness, and variable-length
field handling. This is the novel code — the recursive walk of the schema
JSON that computes byte positions for each field.
## The Two Layout Modes
The POCs surfaced that protocols and mmap-friendly formats need different
layout strategies. This is the most important architectural finding —
decided in [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md).
### Mode 1: Packed sequential (protocol wire formats)
Fields are packed with no alignment padding. Variable-length fields shift
all subsequent fields. Used by SFTP, channels, TTY, and most binary
protocols.
**Components:**
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `TypeDef:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, TypedefError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, TypedefError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
**How it works:**
For a struct with fields `[u8, u32, string]` where the string is 10 bytes:
```
LayoutBuilder::build(var_sizes: {"payload": 10}):
field[0] u8: offset 0, size 1
field[1] u32: offset 1, size 4
field[2] string: offset 5, size 4 (length prefix) + 10 (data)
total: 19
SequentialReader::read_next (read):
read u8 at offset 0
read u32 at offset 1
read u32 length prefix at offset 5 → data_len
read string data at offset 9, length data_len
next field at offset 9 + data_len
```
There is no alignment padding. The `u32` at offset 1 is unaligned — this
is correct for protocol wire formats, which pack fields tightly.
**Variable-length fields in packed mode:**
The `LayoutBuilder` takes actual data sizes for variable-length fields
to compute correct positions for subsequent fields. The consumer must
know the data sizes before writing — this is inherent to packed layouts.
The `SequentialReader` reads each field's length prefix to determine the
data extent and the position of the next field. The reader walks the
buffer sequentially; it cannot jump to field N without reading fields
0..N-1 first.
### Mode 2: Aligned static (mmap-friendly formats)
Fields have fixed positions with natural alignment padding.
Variable-length fields get a 4-byte length prefix at a known offset; the
variable data is not included in the static layout. Used by metatensor
and safetensors.
**Component:**
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, TypedefError>` (requires `TypeDef:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
**How it works:**
For a struct with fields `[u8, u32, f32]` and natural alignment:
```
OffsetMap:
field[0] u8: offset 0, size 1
field[1] u32: offset 4, size 4 (3 bytes padding after u8)
field[2] f32: offset 8, size 4
total: 12 (struct aligned to 4)
```
The `u32` is aligned to offset 4 (its natural alignment). The consumer
can read `field[1]` at offset 4 without reading `field[0]` first — random
access by field path.
**Variable-length fields in aligned mode:**
Variable-length fields get a 4-byte length prefix at a known offset. The
variable data lives outside the static layout — either immediately after
the fixed fields (inline length-prefixing) or in a separate data region
(offset indirection). The `OffsetMap` records the position of the length
prefix (or the `{offset, length}` pair for offset-indirect fields).
For inline length-prefixing, the variable data follows the fixed fields
but is not included in the `OffsetMap`'s field ranges. The consumer reads
the length prefix from the `OffsetMap`'s known offset, then slices the
data region.
For offset indirection, the field is a struct `{offset: u32, length: u32}`
at a known position in the `OffsetMap`. The consumer reads the offset and
length, then slices the separate data region.
### Inline length-prefixing in aligned mode — non-final field restriction
Inline length-prefixed variable fields in aligned mode are only allowed
as the **last field** in their struct. A non-final inline
length-prefixed variable field is rejected at `OffsetMap::compute` time
with a `TypedefError::Offset` — the `OffsetMap` reserves only 4 bytes
(the length prefix), but `data_access::write_string` writes prefix +
data inline, which would clobber subsequent fields. Non-final variable
fields must use `maxLength` (fixed-size reservation) or
`"encoding": "offset-indirect"`. See
[ADR-100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md).
## Offset Computation Algorithm
The offset computation is a recursive walk of the schema JSON. The
algorithm is the same for both modes; the difference is whether alignment
padding is inserted between fields.
### Fixed-size types
For each fixed-size type, the algorithm:
1. Determines the type's byte size from the `TypeDef:*` kind.
2. In aligned mode: inserts padding to satisfy the type's alignment
(or the field's `align` annotation, or the struct's `align` default).
3. Records the field's `(start, end)` range.
4. Advances the current offset by the type's size.
### Composite types
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
are computed relative to the struct's start offset. The struct's total
size is the sum of its fields' sizes (plus alignment padding in aligned
mode). The struct itself may have an `align` annotation that rounds up
its total size.
**`TUnion`:** TUnion is supported in packed sequential mode only. In
aligned static mode, `OffsetMap::compute` rejects `TUnion` fields with
`TypedefError::Offset` — see
[ADR-102](decisions/102-reject-tunion-in-aligned-mode.md). Unions
are the protocol dispatch pattern (SFTP type bytes, call protocol event
types); mmap-friendly formats use structs and arrays, not tagged unions.
In packed sequential mode, the discriminator occupies
`offset..offset + discriminator_size` bytes. For byte-offset
discriminators, the variant struct starts at `offset + discriminator_size`.
For field-name discriminators, the discriminator is just another field —
its offset is computed like any other field, and the variant struct
follows at the end of the discriminator field.
Variant sizes depend on the actual sizes of variable-length fields within
each variant, which aren't known at schema time. The `LayoutBuilder`
takes the actual variant discriminator value and data sizes at write time,
computes the size of the selected variant, and uses that for the union's
total size. The `SequentialReader` reads the discriminator first, looks
up the variant schema, then reads the variant struct sequentially — it
doesn't need to know the union's total size upfront.
**`TArray` of fixed-size elements:** Element stride = element size (plus
alignment padding in aligned mode). Element `i` starts at
`array_offset + i × stride`. The array's total size is `count × stride`.
**`TArray` of variable-length-element structs:** Deferred for v1
(OQ-069).
### Variable-length types
The typedef engine supports three strategies for variable-length types
(see [schema-layer.md](schema-layer.md) §Variable-length types and
[ADR-097](decisions/097-schema-annotations.md) §3 for the full
annotation shapes).
**Strategy 1: Inline length-prefixing (default).**
1. Records the position of the 4-byte length prefix.
2. In aligned mode: the length prefix is aligned; the variable data is
not included in the static layout.
3. In packed mode: the `LayoutBuilder` takes the actual data size to
compute the length prefix value and the position of subsequent fields.
The `SequentialReader` reads the length prefix to determine the data
extent and the position of the next field.
**Strategy 2: Fixed-size reservation (`maxLength`).**
1. In aligned static mode: reserves `maxLength` bytes at a fixed offset.
Data shorter than `maxLength` is zero-padded. Subsequent fields have
known, unchanging offsets — the field is fixed-size from the layout
perspective. This is the database `VARCHAR(N)` pattern.
2. In packed sequential mode: `maxLength` is a validation constraint
only. The engine uses strategy 1 (inline length-prefixing) because
protocols don't benefit from fixed-size reservation.
**Strategy 3: Offset indirection (`"encoding": "offset-indirect"`).**
1. The field is a struct `{offset: u32, length: u32}`.
2. The `OffsetMap` records the position of this struct.
3. The consumer provides the data region separately. This is the
metatensor blob tensor pattern — the index struct lives in one region,
the blob data lives in another.
### Nested structs and field paths
Nested structs produce dotted field paths: `header.version`,
`header.magic`. The offset computation propagates the field path prefix
during recursion. Both `OffsetMap` and `PackedLayout` store fully-qualified
paths; the `iter()` method of each yields fields in schema `properties`
order, with nested struct fields appearing inline under their parent's
path prefix.
### Endianness
Endianness is per-schema (ADR-097). The offset computation is
endian-agnostic — it computes byte positions, not byte values. The
read/write functions apply endianness when converting between bytes and
typed values. The engine reads the `"endian"` annotation from the schema
and byte-swaps accordingly. All fixed-size types — including `TEnum`
(u32 index) — follow the schema's endianness.
## Mode Selection
The consumer selects the mode at engine construction time via the
`LayoutMode` enum, passed to `TypedefEngine::compile`:
```rust
pub enum LayoutMode {
/// Packed sequential — for protocol wire formats (SFTP, channels, TTY).
Packed,
/// Aligned static — for mmap-friendly formats (metatensor, safetensors).
Aligned,
}
```
The choice is determined by the use case, not by the schema:
- **Protocol consumer** (SFTP, binary call frames, TTY negotiation):
`LayoutMode::Packed` → uses `LayoutBuilder` for writing and
`SequentialReader` for reading.
- **mmap consumer** (metatensor): `LayoutMode::Aligned` → uses `OffsetMap`
for both reading and writing at known offsets.
The same schema can be used in either mode. A schema describing an SFTP
packet can be consumed by a `SequentialReader` (for parsing incoming
frames) and a `LayoutBuilder` (for constructing outgoing frames). A schema
describing a metatensor layout can be consumed by an `OffsetMap` (for
mmap access).
`TypedefEngine` exposes mode-appropriate accessors: `engine.offset_map()`
returns `Some(&OffsetMap)` in aligned mode and `None` in packed mode;
`engine.layout_builder()` returns `Some(&LayoutBuilder)` in packed mode
and `None` in aligned mode. `engine.sequential_reader()` returns
`Option<SequentialReader>` (an owned fresh reader, not a reference — the
reader has mutable cursor state that the consumer owns; see
[ADR-101](decisions/101-packed-mode-read-factory.md)) in packed
mode and `None` in aligned mode. See [validation.md](validation.md)
§"The TypedefEngine struct" for the engine API.
## Public Types
The layout engine produces three public types, one per layout component.
All are re-exported from the crate root.
### `ByteRange` (aligned mode)
```rust
pub struct ByteRange {
pub start: usize, // inclusive
pub end: usize, // exclusive
}
```
A half-open byte range produced by `OffsetMap::compute` for each field.
`end - start` is the field's byte size in the static layout (for
variable-length fields: the length prefix, the `{offset, length}` pair,
or the `maxLength` reservation — not the variable data). `ByteRange`
provides `len()` and `is_empty()`.
### `FieldPosition` (packed mode)
```rust
pub struct FieldPosition {
pub offset: usize,
pub size: usize,
pub kind: TypeDefKind,
}
```
A field's computed position in a packed layout, produced by
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
length prefix); for fixed-size fields, `size` is the type's byte size.
`kind` records the field's `TypeDef:*` kind so the consumer can dispatch
to the correct `data_access` read/write function.
### `PackedLayout` (packed mode)
The result of `LayoutBuilder::build`: a map of `field_path → FieldPosition`
plus the total buffer size needed.
```rust
impl PackedLayout {
pub fn get(&self, field_path: &str) -> Option<&FieldPosition>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = &(String, FieldPosition)>;
}
```
`get` looks up a field by dotted path. For TUnion byte-offset
discriminators, the discriminator is recorded under the synthetic path
`"<union_path>.__discriminator"`. `iter` yields fields in layout order
(schema `properties` order, with nested struct fields appearing inline
under their parent's path prefix).
### `OffsetMap` (aligned mode)
A flat table of `(field_path, byte_range)` pairs computed from a schema.
```rust
impl OffsetMap {
pub fn compute(schema: &Value) -> Result<Self, TypedefError>;
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
}
```
`compute` requires a `TypeDef:Struct` at the top level. `total_size`
includes trailing alignment padding. `iter` yields fields in insertion
order (schema `properties` order, nested struct fields appearing inline).
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Two layout modes | [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential for protocols; aligned static for mmap formats |
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Endianness, alignment, encoding annotations that control layout behavior |
| Non-final inline variable fields | [ADR-100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields); use `maxLength` or `offset-indirect` |
| Packed-mode read factory | [ADR-101](decisions/101-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader, not a reference |
| TUnion in aligned mode | [ADR-102](decisions/102-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
— requires lazy walking logic; blocked on a concrete consumer that
needs it.
## References
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
- [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) —
the two layout modes decision
- [ADR-097](decisions/097-schema-annotations.md) — schema
annotations
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds and their
byte sizes
- [data-access.md](data-access.md) — read/write functions that use the
computed offsets
+104
View File
@@ -0,0 +1,104 @@
---
status: draft
last_updated: 2026-07-22
---
# Open Questions
Each open question lives in its own file under [`questions/`](questions/),
named `NNN-slug.md` (mirroring the ADR convention). This file is the index:
theme-grouped tables for scannability, plus a cross-theme
[Deferred / Blocked](#deferred--blocked) section that surfaces the
safe-exit deferrals with their blocking conditions inline — so "what's
currently parked and why" is answerable at a glance.
**Status values**:
- `open` — Needs to be resolved now. Has a clear path to resolution.
- `resolved` — Decided. The resolution is stated cleanly, without caveats about how it could be changed later.
- `deferred(scope)` — Cannot be resolved yet. The information is genuinely
missing — a crate spec, POC result, or use case that doesn't exist yet.
Has a concrete blocking condition. Not a failure — scope management.
- `deferred(unclear)` — Cannot be resolved yet. The pieces exist (decided
in other ADRs, existing types, existing patterns) but the composition
— how they fit together — isn't clear yet. Resolution requires
investigation (work through examples, maybe POC), not waiting. Has a
concrete investigation target and an impacts field. Not a failure —
honest uncertainty in a poorly-defined problem space.
- `partially resolved` — Some aspects decided, others deferred or open.
- `dissolved` — The question was reframed out of existence (e.g., superseded
by an ADR that retires the premise). Kept for reference.
**Impacts field**: Every unresolved OQ (`open`, `deferred(scope)`,
`deferred(unclear)`, `partially resolved`) should have an `Impacts`
field stating what it blocks downstream. Be specific: "blocks the first
hub deployment because the hub dials workers" not "blocks the hub
crate." This is the triage signal that makes the deferral's urgency
visible.
Door type classifications follow ADR-009 — they describe **reversal cost** (how expensive it is to undo), not urgency:
- **One-way door**: Reversal requires rewriting significant code or permanently closes a capability. Getting it wrong is expensive — requires ADR before implementation.
- **Two-way door**: Reversal is cheap or additive. Getting it wrong is recoverable — decide, implement, revert if needed.
Door type is separate from whether a decision is made. A two-way door is a decision you make now and can revert later, not a decision to defer.
## By Theme
### Layout Engine
| OQ | Title | Status | Door | Pri |
|----|-------|--------|------|-----|
| [OQ-069](questions/069-arrays-of-variable-length-element-structs.md) | Arrays of Variable-Length-Element Structs | deferred(scope) | two | low |
### Platform Support
| OQ | Title | Status | Door | Pri |
|----|-------|--------|------|-----|
| [OQ-070](questions/070-no-std-alloc-support.md) | `no_std` + `alloc` Support | deferred(scope) | two | low |
### Schema Construction
| OQ | Title | Status | Door | Pri |
|----|-------|--------|------|-----|
| [OQ-071](questions/071-builder-api-for-schema-construction.md) | Builder API for Schema Construction | deferred(scope) | two | med |
## Deferred / Blocked
The safe-exit visibility surface. These questions are parked because the
information needed to resolve them does not exist yet — each has a concrete
blocking condition. They are not failures; they are scope management.
This section exists so "what's currently blocking the architect" is
answerable at a glance, not by filtering the tables above.
### OQ-069: Arrays of Variable-Length-Element Structs
- **Blocked on**: A concrete consumer that needs arrays of structs with
variable-length fields, where the elements are interleaved
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
sequentially rather than use a fixed stride. The SFTP `Name` packet
has `Vec<File>` where `File` contains strings, but SFTP serializes
this as a sequence of length-prefixed strings (the serde `SeqAccess`
pattern), not as an array of fixed-stride structs. Arrays of
fixed-size structs are fully supported.
- **Priority**: low
- **Full file**: [OQ-069](questions/069-arrays-of-variable-length-element-structs.md)
### OQ-070: `no_std` + `alloc` Support
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
(e.g., a microcontroller running Rust without `std`). The WASM target
has `std` available via `wasm-bindgen`. The engine's core (offset
computation, read/write) is already allocation-free; the `jsonschema`
dependency is the only `alloc` consumer.
- **Priority**: low
- **Full file**: [OQ-070](questions/070-no-std-alloc-support.md)
### OQ-071: Builder API for Schema Construction
- **Blocked on**: A concrete need for programmatic schema construction
in Rust. The current consumers (SFTP, metatensor, binary call frames,
TTY negotiation) all have schemas that can be hand-written or
generated from TypeBox. A builder API would be a fluent Rust API that
produces the same JSON Schema structure — it would sit on top of the
engine, not inside it.
- **Priority**: medium
- **Full file**: [OQ-071](questions/071-builder-api-for-schema-construction.md)
+203
View File
@@ -0,0 +1,203 @@
---
status: draft
last_updated: 2026-07-22
---
# alknet-typedef — Overview
The binary struct engine: a small Rust crate that takes a JSON Schema
with `TypeDef:*` custom keywords and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
This document covers the crate's purpose, the "schema is the format"
principle, its dependency edges, consumers, and scope boundaries.
Component details are in the sibling documents.
## What
`alknet-typedef` is a library crate that consumes JSON Schemas annotated
with `TypeDef:*` custom keywords (the same kinds defined in TypeBox's
`typedef.ts`, plus `TypeDef:Bytes`, `TypeDef:Int64`, and `TypeDef:Uint64`
as alknet-typedef additions) and produces three capabilities:
1. **An offset map** — walks the schema, computes byte offsets for each
field based on type sizes, field order, and alignment.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that a
buffer's bytes match the schema's type constraints.
The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing). The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are small (a few lines
each, generated from shared macros — see [validation.md](validation.md)).
The crate replaces two prior attempts that built their own jsonschema
engines — typebox-rs (~8,400 lines) and alktype (~5,600 lines) — with
`jsonschema` + an offset map + small custom keyword implementations. See
[ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
## Why
The crate's purpose is to be the binary struct engine for every alknet
component that reads or writes binary data at computed offsets. Instead
of per-protocol serde structs (russh-sftp's 29 packet types), per-handler
wire format code (TTY's 5-byte format parser), or per-format offset
computation (metatensor's tensor access), all of these become instances
of the same engine with different schemas.
The guiding insight:
> **The schema is the format.** A JSON Schema with `TypeDef:Float32`,
> `TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
> the layout spec. No separate format definition, no separate parser, no
> separate validator. One schema, three uses: validate, compute offsets,
> access data.
This is the convergence of three threads identified in the
call-channels-unification research: the `typedef.ts` schema kinds from
TypeBox, the russh-sftp protocol packets, and the metatensor format. The
common pattern: a JSON Schema describes the shape of binary data, and
the binary data is the struct's bytes at computed offsets.
The crate was bumped up in the timeline when the call-channels-unification
research surfaced that channels, TTY, and the binary call protocol are
all variations on the same wire-format family — `[discriminant][length][payload]`.
The typedef engine makes the "channels is call with a binary data plane"
unification concrete: the binary data plane's wire format is the call
protocol's own schema system, just binary-encoded. The `channel_open`
marker says "use binary framing"; the typedef engine says "here's how to
read/write the binary payload."
## The "Schema Is the Format" Principle
A JSON Schema with `TypeDef:*` custom keywords serves three roles
simultaneously:
| Role | Mechanism | When |
|------|-----------|------|
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) |
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
No separate format definition, no separate parser, no separate validator.
The schema is the single source of truth for the binary format. Adding a
new field to a protocol is adding a property to the schema JSON — the
engine computes the new offsets automatically.
This is the same principle as `#[repr(C)]` struct field access, but at
runtime from a portable JSON Schema instead of at compile-time from
language-specific annotations. The schema is the ABI contract.
## Dependencies
```
alknet-typedef
├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support
├── serde_json (with preserve_order) — schema parsing; field order is load-bearing
└── (no tokio, no platform deps) — WASM-clean by construction
```
`alknet-typedef` is dependency-light: `jsonschema` + `serde_json` only.
No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for
browser use. The `jsonschema` crate is already in the workspace at
`/workspace/jsonschema/` but not yet used by any alknet crate — typedef
is the first consumer.
`serde_json` requires the `preserve_order` feature because field order
is load-bearing for binary layouts. The order of properties in the
schema JSON determines the order of fields in the binary struct.
## Consumers
| Consumer | Schema describes | Engine provides |
|----------|-----------------|-----------------|
| russh-sftp | 29 packet structs + Packet union (byte discriminator) | Read/write SFTP frames from bytes |
| metatensor | Model layout (ConvNet struct, tensor refs) | Offset map for mmap'd tensor access |
| binary call frames | `call.requested` / `call.responded` / etc. structs | Read/write binary call frames |
| TTY negotiation | `NegotiateRequest` / `NegotiateResponse` structs | Read/write TTY control frames |
| channels wire | `ChunkHeader { channel_id, length }` | Already trivial (8 bytes, no schema needed) |
The russh-sftp case is the most instructive and the highest-value POC
target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written
dispatch on a type byte followed by serde deserialization. Under typedef,
the dispatch is `TUnion` with a byte-offset discriminator — the schema
says "byte 0 is the discriminator, bytes 1..N are the variant struct."
The engine reads the discriminator, looks up the variant schema, computes
offsets, reads fields. Same result, no per-packet-type code.
## Scope Boundaries (What This Is Not)
These boundaries are decided in [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
is a *format* (8-byte header + JSON header + binary data) that uses the
typedef engine for its offset computation and tensor access.
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
`Value.Convert` — schema evolution — is out of scope for v1. The engine
should not do anything that explicitly blocks adding a Value system
later.
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
concern. The typedef engine consumes schemas; it does not generate them.
- **Not a schema builder.** The typedef engine does not provide a fluent
API for constructing schemas. Schemas are plain JSON — authored in
TypeBox, generated by ujsx components, or hand-written. A builder API
is deferred (OQ-071).
- **Not a serialization framework.** The typedef engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use typedef.
## Architecture (component pointers)
- **[schema-layer.md](schema-layer.md)** — the 19 `TypeDef:*` kinds,
jsonschema custom keyword integration, TypeBox interop, schema
annotations (endianness, alignment, encoding, TUnion discriminators).
- **[layout-engine.md](layout-engine.md)** — offset computation, the two
layout modes (packed sequential vs aligned static), alignment,
endianness, variable-length field handling.
- **[data-access.md](data-access.md)** — read/write functions, TUnion
dispatch, field paths, zero-copy access for fixed-size types,
length-prefix reading for variable-length types.
- **[validation.md](validation.md)** — custom keyword validators for all
19 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time
validation, `TypedefEngine` as the compiled form of a schema.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Purpose, scope, and the jsonschema engine | [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| Two layout modes | [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats |
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) |
| Error handling and validation | [ADR-098](decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Int64/Uint64 kinds | [ADR-099](decisions/099-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (SFTP offsets, metatensor data_offsets) |
| Non-final inline variable fields | [ADR-100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields) |
| Packed-mode read factory | [ADR-101](decisions/101-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader |
| TUnion in aligned mode | [ADR-102](decisions/102-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs.
- **OQ-070** (deferred(scope)): `no_std` + `alloc` support.
- **OQ-071** (deferred(scope)): Builder API for schema construction.
## References
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
passing, two layout modes, TUnion dispatch, endianness)
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine" — the origin of this research
thread
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
@@ -0,0 +1,20 @@
# OQ-069: Arrays of variable-length-element structs
- **Origin**: [../layout-engine.md](../layout-engine.md),
[../data-access.md](../data-access.md);
`docs/research/alknet-typedef/findings.md` §"Problem 3: Nested structs
and arrays of structs"
- **Status**: deferred(scope)
- **Door type**: Two-way (additive — the engine can add lazy walking
logic without changing the existing fixed-stride array support)
- **Priority**: low
- **Impacts**: Blocks any protocol with interleaved variable-length struct arrays (e.g., a protocol where each array element has a string field and elements are packed as `[fixed_0][str_0][fixed_1][str_1]...`). Does NOT block SFTP `Name` packet handling — SFTP serializes this as a sequence of length-prefixed strings (the serde `SeqAccess` pattern), not as an array of fixed-stride structs. Does NOT block any current consumer.
- **Blocked on**: A concrete consumer that needs arrays of structs with
variable-length fields, where the elements are interleaved
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
sequentially rather than use a fixed stride.
- **Resolution**: Not yet decidable. The mechanism (lazy sequential
walking of array elements, reading each element's length prefixes to
find the next element's start) is understood but not needed by any
current consumer. Arrays of fixed-size structs are fully supported.
- **Cross-references**: ADR-096, [layout-engine.md](../layout-engine.md)
@@ -0,0 +1,22 @@
# OQ-070: `no_std` + `alloc` support
- **Origin**: [../overview.md](../overview.md);
`docs/research/alknet-typedef/findings.md` §"Open Questions" (OQ 6)
- **Status**: deferred(scope)
- **Door type**: Two-way (additive — can be added as a feature gate
without changing the existing `std` API)
- **Priority**: low
- **Impacts**: Blocks bare-metal embedded targets (microcontrollers
running Rust without `std`). Does NOT block any current deployment
target. Does NOT block WASM — `wasm32-unknown-unknown` has `std`
available via `wasm-bindgen`; the crate is WASM-clean by construction
(no tokio, no platform deps, `jsonschema` builds for WASM with
`default-features = false`).
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
(e.g., a microcontroller running Rust without `std`).
- **Resolution**: Not yet decidable. Target `std` for v1. If embedded
use cases emerge, `no_std` + `alloc` can be added as a feature gate
later. The engine's core (offset computation, read/write) is already
allocation-free — it operates on `&[u8]` slices. The `jsonschema`
dependency is the only `alloc` consumer.
- **Cross-references**: ADR-095
@@ -0,0 +1,26 @@
# OQ-071: Builder API for schema construction
- **Origin**: [../schema-layer.md](../schema-layer.md),
[../overview.md](../overview.md);
`docs/research/alknet-typedef/findings.md` (the builder API was noted
as the one detail not covered by the POCs)
- **Status**: deferred(scope)
- **Door type**: Two-way (additive — a builder API can be added without
changing the existing JSON-consumption path)
- **Priority**: medium
- **Impacts**: Blocks programmatic schema construction in Rust without a
JS toolchain. Any consumer that wants to build typedef schemas at
runtime from Rust code (rather than loading pre-authored JSON) must
construct the JSON manually or depend on TypeBox. Does NOT block any
current consumer — all v1 consumers (SFTP, metatensor, binary call
frames, TTY negotiation) use pre-authored schemas.
- **Blocked on**: A concrete need for programmatic schema construction
in Rust. The current consumers (SFTP, metatensor, binary call frames,
TTY negotiation) all have schemas that can be hand-written or generated
from TypeBox.
- **Resolution**: Not yet decidable. The builder API is important but
not needed for the initial consumers. The engine's JSON-consumption
path is the primary interface for v1. A builder API would be a fluent
Rust API that produces the same JSON Schema structure — it would sit
on top of the engine, not inside it.
- **Cross-references**: ADR-095, [schema-layer.md](../schema-layer.md)
+506
View File
@@ -0,0 +1,506 @@
---
status: draft
last_updated: 2026-07-22
---
# alknet-typedef — Schema Layer
The schema layer: the 19 `TypeDef:*` custom type kinds, their mapping to
Rust types and byte sizes, the `jsonschema` custom keyword integration,
TypeBox interop, and the concrete JSON shapes for schema-level
annotations.
## The 19 TypeDef Kinds
These are the custom schema kinds defined in TypeBox's `typedef.ts`
(`/workspace/@alkdev/typebox/example/typedef/typedef.ts`, 619 lines) and
ported to Rust via `jsonschema` custom keywords. Each kind carries binary
layout semantics — a known byte size (for fixed-size types) or a known
encoding strategy (for variable-length types).
| Kind | TypeBox key | Rust type | Size | Category |
|------|-------------|-----------|------|----------|
| `TFloat32` | `TypeDef:Float32` | `f32` | 4 | fixed |
| `TFloat64` | `TypeDef:Float64` | `f64` | 8 | fixed |
| `TInt8` | `TypeDef:Int8` | `i8` | 1 | fixed |
| `TInt16` | `TypeDef:Int16` | `i16` | 2 | fixed |
| `TInt32` | `TypeDef:Int32` | `i32` | 4 | fixed |
| `TInt64` | `TypeDef:Int64` | `i64` | 8 | fixed |
| `TUint8` | `TypeDef:Uint8` | `u8` | 1 | fixed |
| `TUint16` | `TypeDef:Uint16` | `u16` | 2 | fixed |
| `TUint32` | `TypeDef:Uint32` | `u32` | 4 | fixed |
| `TUint64` | `TypeDef:Uint64` | `u64` | 8 | fixed |
| `TBoolean` | `TypeDef:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed |
| `TString` | `TypeDef:String` | length-prefixed UTF-8 | variable | variable |
| `TBytes` | `TypeDef:Bytes` | length-prefixed raw bytes | variable | variable |
| `TStruct` | `TypeDef:Struct` | record of fields | sum of field sizes | composite |
| `TUnion` | `TypeDef:Union` | tagged union | discriminator + variant | composite |
| `TArray` | `TypeDef:Array` | repeated element | count × element size | composite |
| `TEnum` | `TypeDef:Enum` | u32 index into enum values | 4 (fixed) | fixed |
| `TRecord` | `TypeDef:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
| `TTimestamp` | `TypeDef:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
`TypeDef:Int64` and `TypeDef:Uint64` are alknet-typedef additions —
TypeBox's `typedef.ts` tops out at 32-bit integers. They are required by
the primary POC targets: SFTP `Read`/`Write` packets have `offset: u64`,
and metatensor `data_offsets` are `u64`. See
[ADR-099](decisions/099-int64-uint64-first-class-kinds.md).
### The `TypeDefKind` enum
The engine represents the 19 kinds as a Rust enum — `TypeDefKind` — with
one variant per kind (`TypeDefKind::Float32`, `TypeDefKind::Struct`, etc.).
The enum provides compile-time exhaustiveness checking and integer
discriminant dispatch (a jump table) instead of string comparison at
every field access. It is `pub` and re-exported from the crate root.
```rust
pub enum TypeDefKind {
Int8, Int16, Int32, Int64,
Uint8, Uint16, Uint32, Uint64,
Float32, Float64,
Boolean, Enum,
String, Bytes, Timestamp,
Struct, Union, Array, Record,
}
```
The enum carries the kind's binary-layout metadata as inherent methods:
| Method | Returns | Notes |
|--------|---------|-------|
| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"TypeDef:Uint8"` |
| `type_size(self)` | `Option<usize>` | `Some(N)` for fixed-size kinds; `None` for variable/composite |
| `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array |
| `is_fixed_size(self)` | `bool` | True for the 12 fixed-size primitive kinds |
| `is_composite(self)` | `bool` | True for Struct, Union, Array, Record |
| `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record |
| `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter |
`TypeDefKind` implements `Display` (renders the keyword string) and
`FromStr` (parses the keyword string back into the variant, returning
`TypedefError::Schema` for unknown kinds). The layout engines and the
validator dispatch on the enum, not on strings.
### Fixed-size types
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset
computation uses these sizes directly. Read/write is zero-copy pointer
cast for these types.
**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other
values are invalid and produce a `TypedefError::Access` on read.
**`TEnum` binary representation:** A `u32` index into the enum's declared
values, in declaration order. The first declared value is index 0, the
second is index 1, etc. The enum's values are declared via the standard
JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`).
The `TypeDef:Enum` custom keyword signals that the type is an enum for
layout purposes; the built-in `enum` keyword provides the value list.
**Design note:** TypeBox's `TEnum` is a string enum (variable-length). The
typedef engine uses a `u32` index instead — a deliberate deviation from
TypeBox fidelity in favor of binary efficiency. Most enums have a small
number of variants (e.g., the call protocol's 5 event types); a `u32`
index is compact, fixed-size, and sufficient for any realistic enum. The
JSON representation (for validation) remains a string; the binary
representation is the `u32` index.
The `u32` index follows the schema's endianness annotation (ADR-097), like
all other fixed-size types. In little-endian mode the index is
`u32::from_le_bytes`; in big-endian mode it is `u32::from_be_bytes`.
### Variable-length types
`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes.
The typedef engine supports three strategies for handling variable-length
types in binary layouts, selected by the `encoding` annotation and the
standard JSON Schema `maxLength` keyword:
| Strategy | Encoding annotation | Layout behavior | Use case |
|----------|-------------------|-----------------|----------|
| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) |
| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) |
| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) |
**Strategy 1: Inline length-prefixing (default).** The field's fixed
portion is a 4-byte length prefix at a computed offset. The variable data
follows immediately after. In packed sequential mode, the length prefix
determines the position of subsequent fields. In aligned static mode, the
length prefix is at a known offset; the variable data is not included in
the static layout. This is the universal pattern used by channels, SFTP,
TTY, and most binary protocols.
**Strategy 2: Fixed-size reservation.** When a variable-length field
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
than `maxLength` is zero-padded; data longer than `maxLength` is a
validation error. This makes the field fixed-size from the layout
perspective — subsequent fields have known, unchanging offsets. This is
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
pattern for fields with known maximum sizes.
In packed sequential mode, `maxLength` is a validation constraint only —
the engine still uses inline length-prefixing (strategy 1) because
protocols don't benefit from fixed-size reservation.
**Strategy 3: Offset indirection.** The field is a struct
`{offset: u32, length: u32}` at a known position. The consumer provides
the data region separately; the engine reads the offset and length, then
slices the data region. This is the metatensor blob tensor pattern — the
index struct lives in one region, the blob data lives in another. Enables
mmap-friendly random access to variable-length data without parsing
length prefixes and without reserving worst-case space.
**Default strategy selection:**
- In packed sequential mode: always strategy 1 (inline length-prefixing).
`maxLength` is a validation constraint only.
- In aligned static mode: strategy 2 (fixed-size reservation) if
`maxLength` is declared; strategy 3 (offset indirection) if
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
length-prefixing) otherwise.
**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3)
respects the schema's `"endian"` annotation (ADR-097). In little-endian
mode, the length is `u32::from_le_bytes`. In big-endian mode, the length
is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have
consistent byte order for both field values and length prefixes.
**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`.
Otherwise identical to `TString` in layout (same three strategies).
**Design note:** `TypeDef:Bytes` is an alknet-typedef addition — it does
not exist in TypeBox's `typedef.ts` (which defines 16 kinds). It is
included because raw byte arrays are a common binary protocol primitive
(SFTP data payloads, channels payloads, tensor data) and are semantically
distinct from UTF-8 strings. In the binary representation, TBytes is raw
bytes with no encoding (not base64, not hex). In the JSON representation
(for validation), TBytes is a string (JSON has no native byte type).
**`TRecord`:** A string-keyed map. The value type is declared via the
schema's `"values"` property (e.g., `"values": { "TypeDef:Float32": true }`).
Binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
The count is the number of entries. Each key is a length-prefixed UTF-8
string. Each value is encoded according to its declared `TypeDef:*` kind
— a `Record<Uint32>` value is 4 raw bytes; a `Record<String>` value is
itself a length-prefixed string; a `Record<Struct>` value is the struct's
fields laid out inline. There is **no separate `value_len` prefix** —
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix). The
count and key-length prefixes respect the schema's endianness. In
aligned static mode with `maxLength`, the entire record is reserved at
`maxLength` bytes (zero-padded).
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
fixed-size reservation (strategy 2 with `maxLength`). The data-access
layer treats timestamps as opaque length-prefixed strings — it does not
parse or validate the timestamp format. The jsonschema custom keyword
validator checks RFC 3339 conformance at the JSON level (see
[validation.md](validation.md)).
`TArray` is variable-length when the element type is variable-length or
when the count is not known at schema time. For fixed-size element arrays
with a known count, the size is `element_size × count`.
**`TArray` count declaration:** The array count is declared via the
standard JSON Schema `"minItems"` and `"maxItems"` keywords. When
`minItems == maxItems`, the array has a fixed count known at schema time.
When they differ or are absent, the count is variable and the array uses
a length-prefixed encoding: `[count: u32][element_0]...[element_N]`.
The count prefix respects the schema's endianness.
### Composite types
`TStruct` and `TUnion` are composite — their size is the sum of their
fields' sizes (plus alignment padding in aligned static mode). The offset
computation recurses into their properties.
## Schema-Layer Public API
The `schema` module exposes the foundational types and functions every
other module depends on. These are re-exported from the crate root.
### `get_typedef_kind` vs `get_typedef_kind_loose`
The engine recognizes a `TypeDef:*` kind on a schema node two ways,
because the keyword value may be either a boolean (`true`) or an
annotation object (`{ "encoding": "..." }`):
| Function | Recognizes | Returns |
|----------|------------|---------|
| `get_typedef_kind(node) -> Option<&str>` | Boolean form only (`{ "TypeDef:String": true }`) | The keyword string, e.g. `"TypeDef:String"` |
| `get_typedef_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string |
| `get_typedef_kind_enum(node) -> Option<TypeDefKind>` | Boolean form only | The parsed enum variant |
| `get_typedef_kind_loose_enum(node) -> Option<TypeDefKind>` | Boolean form **and** object form | The parsed enum variant |
The boolean-form-only functions are used by the validator factories
(which reject the object form as a schema error) and the top-level
kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new`
(which require `TypeDef:Struct` at the root). The "loose" variants are
used by the layout engines during field traversal, so that a variable-
length field with an `encoding` annotation (`{ "TypeDef:String":
{ "encoding": "offset-indirect" } }`) is still recognized as a `String`.
### Annotation parsers
Each schema-level annotation has a dedicated parser that reads it from a
`serde_json::Value` node and returns a sensible default when absent:
| Function | Annotation | Default |
|----------|------------|---------|
| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` |
| `parse_align(node) -> Option<usize>` | `"align"` | `None` |
| `parse_max_length(node) -> Option<usize>` | `"maxLength"` | `None` |
| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` |
| `parse_discriminator(node) -> Result<DiscriminatorKind, TypedefError>` | `"discriminator"` | (required — returns `TypedefError::Schema` if absent) |
### Public enums
```rust
pub enum Endian { Little, Big }
pub enum VariableEncoding { LengthPrefixed, OffsetIndirect }
pub enum DiscriminatorKind {
Byte { offset: usize, disc_type: TypeDefKind },
Field { name: String },
}
```
`DiscriminatorKind::Byte` carries the byte position (`offset`) and the
discriminator's `TypeDef:*` kind (`disc_type`, restricted to `Uint8`/
`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator
field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for
how these drive dispatch.
### `$ref` resolution and normalization
| Function | Purpose |
|----------|---------|
| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/<name>"`. Idempotent. Runs once at `TypedefEngine::compile` time. |
| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. |
| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). |
`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s
JSON Pointer requirement. The layout engines call `resolve_ref_or_inline`
on every `$ref`-bearing node they encounter during traversal.
## jsonschema Custom Keyword Integration
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
via the `with_keyword` API. Each `TypeDef:*` kind is registered as a
custom keyword:
```rust
let validator = jsonschema::options()
.with_keyword("TypeDef:Float32", factory)
.with_keyword("TypeDef:Int32", factory)
.with_keyword("TypeDef:Struct", factory)
// ... all 17 kinds
.build(&schema)?;
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path — enabling cross-keyword awareness. The
`TypeDef:Struct` validator, for example, inspects the parent's
`properties` to validate each field against its declared `TypeDef:*` kind.
Each custom keyword implementation is ~10 lines. The `jsonschema` crate
handles all structural validation (object properties, required fields,
array items, enum values) — the custom keywords only need to validate
the leaf type constraints. See [validation.md](validation.md) for the
validator implementations.
This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side.
Same semantics, different language, same JSON Schema wire format. A
TypeBox schema serialized to JSON feeds into the typedef engine after a
single pre-processing step: normalizing `$ref` values (see below).
## TypeBox Interop
TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox
schema like:
```typescript
const TensorRef = Type.Object({
dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]),
shape: Type.Array(Type.Number()),
data_offsets: Type.Tuple([Type.Number(), Type.Number()])
});
```
serialized to JSON is a standard JSON Schema with `type: "object"`,
`properties`, and `required`. That JSON feeds into the typedef engine
after `$ref` normalization. The `TypeDef:*` custom keywords are added by
TypeBox's `TypeRegistry.Set` — they appear in the serialized JSON as
additional properties on the schema object.
### `$ref` normalization
TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`),
referencing sibling definitions within the same `$defs` block. The
`jsonschema` crate requires full JSON Pointer paths (e.g.,
`"$ref": "#/$defs/Read"`). The typedef engine normalizes TypeBox-style
refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization)
— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to
`"#/$defs/<name>"`. The normalization is idempotent — full JSON Pointer
refs pass through unchanged. It runs once at `TypedefEngine::compile`
time, before the schema is passed to `jsonschema` or the offset
computation.
**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs
with `Resource 'Read' is not present in a registry`. Full JSON Pointer
refs (`#/$defs/Read`) resolve correctly. The normalization step bridges
the gap between TypeBox's output and jsonschema's input.
The typedef engine does not depend on TypeBox or any JS toolchain. It
consumes JSON — whether that JSON was authored in TypeBox, generated by
a ujsx component, or hand-written. The schema is the interface.
## Schema Annotations
Schema-level annotations control binary layout behavior. These are
decided in [ADR-097](decisions/097-schema-annotations.md).
### Endianness
Schema-level annotation with a default of little-endian:
```json
{ "TypeDef:Struct": true, "endian": "big", "properties": { ... } }
```
- `"endian": "little"` (default) — read/write in little-endian byte order.
- `"endian": "big"` — read/write in big-endian byte order.
- Applies to the entire schema and all nested types.
### Alignment
Both struct-level and field-level, with field-level overriding:
```json
{
"TypeDef:Struct": true,
"align": 256,
"properties": {
"weight": { "TypeDef:Float32": true, "align": 16 }
}
}
```
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/
enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix),
1 for struct/union/array.
- Only meaningful in aligned static mode (ADR-096). Ignored in packed
sequential mode.
### Variable-length encoding
The typedef engine supports three strategies for variable-length types
(see §Variable-length types above for full details). The strategy is
selected by the `encoding` annotation and the standard JSON Schema
`maxLength` keyword:
```json
// Strategy 1: Inline length-prefixing (default, shorthand)
{ "TypeDef:String": true }
// Strategy 1: Explicit inline length-prefixing
{ "TypeDef:String": { "encoding": "length-prefixed" } }
// Strategy 2: Fixed-size reservation (uses standard maxLength)
{ "TypeDef:String": true, "maxLength": 256 }
// Strategy 3: Offset indirection (opt-in)
{ "TypeDef:String": { "encoding": "offset-indirect" } }
```
- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at
computed offset, variable data follows immediately. Used by protocol
wire formats.
- `maxLength` (standard JSON Schema keyword) — in aligned static mode,
reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the
field fixed-size from the layout perspective. In packed sequential
mode, `maxLength` is a validation constraint only.
- `"encoding": "offset-indirect"` — field is a struct
`{offset: u32, length: u32}` pointing into a separate data region.
The consumer provides the data region separately. Used by metatensor
blob tensors.
- Applies to all variable-length types: `TypeDef:String`, `TypeDef:Bytes`,
`TypeDef:Array`, `TypeDef:Record`, `TypeDef:Timestamp`.
### TUnion discriminators
Two discriminator kinds: byte-offset (protocol dispatch) and field-name
(typedef.ts pattern).
**Byte-offset discriminator** (SFTP type bytes, call protocol event types):
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "TypeDef:Uint8"
},
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" },
"101": { "$ref": "#/$defs/Status" }
}
}
```
- `"offset"` — byte position of the discriminator.
- `"type"` — the `TypeDef:*` kind of the discriminator (typically
`TypeDef:Uint8`).
- Mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`.
**Field-name discriminator** (typedef.ts pattern):
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "field",
"name": "type"
},
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
}
```
- `"name"` — the field name holding the discriminator value.
- Mapping keys are string values matching the discriminator field's value.
- The discriminator field is just another field in the struct.
Mapping values may be either inline schemas or `$ref` pointers. Both work.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators |
| Int64/Uint64 kinds | [ADR-099](decisions/099-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (required by SFTP offsets and metatensor data_offsets) |
| Purpose and scope | [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-071** (deferred(scope)): Builder API for schema construction.
## References
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- [ADR-097](decisions/097-schema-annotations.md) — schema
annotation shapes
- [validation.md](validation.md) — custom keyword validator implementations
+335
View File
@@ -0,0 +1,335 @@
---
status: draft
last_updated: 2026-07-22
---
# alknet-typedef — Validation
The validation layer: custom keyword validators for all 19 `TypeDef:*`
kinds, the `TypedefError` enum, load-time vs access-time validation
strategy, and the `TypedefEngine` as the compiled form of a schema.
## Validation Strategy
Validation is delegated to the `jsonschema` crate (v0.46.5, Draft
2020-12). The typedef engine does not implement its own validation —
it registers custom keyword validators for each `TypeDef:*` kind and
lets `jsonschema` handle the structural validation (object properties,
required fields, array items, enum values).
The strategy is decided in [ADR-098](decisions/098-error-handling-validation-strategy.md):
1. **Load time:** Parse the schema JSON, build the layout engine, build the
jsonschema validator. This is the `TypedefEngine::compile(schema)` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation.
### What validation validates
The jsonschema validator operates on `serde_json::Value` instances — it
validates JSON representations of data, not raw byte buffers. This is
the correct separation of concerns:
- **JSON validation** (jsonschema): validates that a JSON document
conforms to the schema. Used for validating hand-written schemas,
TypeBox output, JSON payloads, or the JSON representation of a binary
struct after deserialization.
- **Binary access validation** (data access layer): the read/write
functions perform type-level validation at access time — range checks
for integers, UTF-8 validity for strings, buffer bounds checking.
These return `TypedefError::Access` with field paths.
The "schema is the format" principle means the same schema describes
both the JSON shape and the binary layout. The jsonschema validator
checks the JSON shape; the data access layer checks the binary layout.
A consumer that wants to validate a binary buffer end-to-end reads the
buffer into a `Value` tree via the data access layer, then validates
that `Value` against the jsonschema validator. This is a two-step
process, not a single `validate(buffer)` call.
### The `TypedefEngine` struct
The `TypedefEngine` is the compiled form of a schema. It supports both
layout modes (ADR-096) via an internal `Layout` enum:
```rust
pub struct TypedefEngine {
layout: Layout, // packed or aligned (private enum)
validator: jsonschema::Validator, // compiled once at load time
endian: Endian, // parsed from the schema's "endian" annotation
schema: Value, // the normalized schema (refs resolved)
}
// Private — the consumer selects via LayoutMode at compile time.
enum Layout {
Packed { builder: LayoutBuilder },
Aligned { offset_map: OffsetMap },
}
```
The consumer selects the mode at construction time via `LayoutMode`
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout`
enum is private — the engine exposes mode-appropriate accessors instead:
```rust
impl TypedefEngine {
pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result<Self, TypedefError>;
pub fn mode(&self) -> LayoutMode;
pub fn endian(&self) -> Endian;
pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode
pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode
pub fn sequential_reader(&self) -> Option<SequentialReader>; // owned fresh reader (ADR-101)
}
```
`compile` takes `&mut Value` because it normalizes `$ref` values in place
(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization))
before computing the layout and building the validator. The `schema`
field retains the normalized schema for `read_field`'s kind lookup and
for `sequential_reader()`'s factory construction. The validator is
mode-agnostic (it operates on `Value`, not raw bytes).
The `Layout::Packed` variant stores only the `LayoutBuilder` (write-side).
The `SequentialReader` (read-side) is not stored — it has mutable cursor
state that the consumer owns, so `sequential_reader()` constructs a fresh
reader on each call (ADR-101).
The `read_field`/`write_field` methods on `TypedefEngine` are the
aligned-mode data-access API — see [data-access.md](data-access.md)
§"Higher-level read/write".
## Custom Keyword Validators
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
`jsonschema::options().with_keyword(...)`. The validators check leaf
type constraints; `jsonschema` handles all structural validation.
### Numeric type validators
**`TypeDef:Float32` / `TypeDef:Float64`:**
- Value must be a finite number.
- For `Float32`: value must be representable as `f32` (no precision loss
beyond `f32`'s mantissa).
**`TypeDef:Int8` / `TypeDef:Int16` / `TypeDef:Int32`:**
- Value must be an integer within the type's range.
- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
**`TypeDef:Uint8` / `TypeDef:Uint16` / `TypeDef:Uint32`:**
- Value must be a non-negative integer within the type's range.
- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295.
### String and binary validators
**`TypeDef:String`:**
- Value must be a valid UTF-8 string.
- If `maxLength` is specified in the schema, the string's byte length
must not exceed it.
**`TypeDef:Bytes`:**
- Value must be a string (JSON represents binary data as a string — JSON
has no native byte type).
- If `maxLength` is specified, the byte length must not exceed it.
- **Binary representation:** In the binary layout, `TBytes` is raw bytes
with no encoding (not base64, not hex). The JSON representation (for
validation) uses a string; the binary representation (for data access)
uses `&[u8]` directly.
**`TypeDef:Enum`:**
- The `TypeDef:Enum` custom keyword signals that the type is an enum for
*layout* purposes (the engine needs to know it's a fixed-size u32 index,
not a variable-length string). The built-in `enum` keyword provides the
value list and handles value-membership validation. The custom keyword
validator is a no-op beyond the built-in check — it exists solely for
the layout engine to recognize the type.
**`TypeDef:Timestamp`:**
- Value must be a valid RFC 3339 timestamp string (the internet profile
of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
### Composite type validators
**`TypeDef:Struct`:**
- Value must be an object.
- Each property must match its declared `TypeDef:*` kind.
- Required fields must be present.
- The `jsonschema` crate's built-in `properties` and `required` keywords
handle the structural checks — the custom keyword only needs to
validate that each field's value matches its `TypeDef:*` kind.
**`TypeDef:Union`:**
- The discriminator value must be one of the mapping keys.
- The variant struct must match the declared schema for that discriminator
value.
**`TypeDef:Array`:**
- Value must be an array.
- Each element must match the array's declared element type.
- If `minItems`/`maxItems` is specified, the array length must be within
bounds.
### Other validators
**`TypeDef:Boolean`:**
- Value must be `true` or `false`.
**`TypeDef:Record`:**
- Value must be an object.
- All values must match the record's declared value type (specified via
the `"values"` property in the schema, e.g.,
`"values": { "TypeDef:Float32": true }`).
### Validator implementation pattern
Each custom keyword implementation is ~10 lines. Example for
`TypeDef:Float32`:
```rust
struct Float32Validator;
impl Keyword for Float32Validator {
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
match instance {
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
}
}
fn is_valid(&self, instance: &Value) -> bool {
instance.as_f64().map_or(false, |f| f.is_finite())
}
}
```
Registration:
```rust
let validator = jsonschema::options()
.with_keyword("TypeDef:Float32", |parent, value, path| {
Ok(Box::new(Float32Validator))
})
.build(&schema)?;
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path. This enables cross-keyword awareness — for
example, a `TypeDef:Struct` validator can inspect the parent's
`properties` to validate each field against its declared `TypeDef:*` kind.
## TypedefError
A single `TypedefError` enum covers all error conditions across the
engine's three phases (schema parsing, offset computation, read/write)
plus validation. Decided in [ADR-098](decisions/098-error-handling-validation-strategy.md).
```rust
pub enum TypedefError {
/// Schema parsing errors (invalid JSON, missing keywords, unknown TypeDef kinds).
Schema(String),
/// Offset computation errors (field not found, unsupported type).
Offset { field_path: String, reason: String },
/// Read/write errors (buffer too short, invalid UTF-8, value out of range).
Access { field_path: String, reason: String },
/// Validation errors (delegated to jsonschema).
Validation(ValidationError<'static>),
}
```
- **`Schema`** — for errors during `TypedefEngine::compile()`. Invalid
JSON, missing required keywords, unknown `TypeDef:*` kinds.
- **`Offset`** — for errors during offset computation. Field not found
in the schema, type not supported for offset computation, recursive
depth exceeded. Carries the field path.
- **`Access`** — for errors during read/write. Buffer too short, invalid
UTF-8 in a string field, value out of range for the target type.
Carries the field path.
- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The
`'static` lifetime is correct — the validator owns its schema reference
and lives for the lifetime of the `TypedefEngine`.
### Field-path-carrying errors
Read/write and offset errors include the field path for debugging:
```rust
Err(TypedefError::Access {
field_path: "header.version".to_string(),
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
})
```
This makes debugging binary format issues tractable — the error tells
you exactly which field failed and why.
## Validation Timing
### Load time: `TypedefEngine::compile()`
The expensive work happens once at schema load time:
1. Normalize `$ref` values in the schema (`normalize_refs`).
2. Parse the schema's `"endian"` annotation.
3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned).
4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
The result is a `TypedefEngine` that can be used for repeated operations.
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
Validation is opt-in per operation. The consumer calls
`engine.validate_json(instance)` when validation is desired, or
`engine.is_valid_json(instance)` for a boolean check. The jsonschema
validator is already compiled — these are fast checks against the
compiled validator.
```rust
pub fn validate_json(&self, instance: &Value) -> Result<(), TypedefError>;
pub fn is_valid_json(&self, instance: &Value) -> bool;
```
The argument is a `serde_json::Value` (the JSON representation of the
data), not a raw byte buffer — see §"What validation validates" above.
To validate a binary buffer end-to-end, the consumer reads it into a
`Value` tree via the data access layer, then validates that `Value`.
High-throughput paths can skip validation. Security-sensitive paths
(parsing incoming frames from untrusted peers) can validate every frame.
The choice is the consumer's.
## Relationship to Read/Write
Validation and data access are independent operations on the same data.
The consumer can:
1. Validate the JSON representation of a buffer to ensure it conforms to
the schema.
2. Read fields from the binary buffer at computed offsets.
3. Both — validate the JSON representation first, then read the binary
buffer (defense in depth).
The engine does not couple validation and access. A consumer that trusts
its data source can skip validation and go straight to read/write. A
consumer that parses untrusted input can validate the JSON
representation first, then access the binary buffer.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Error handling and validation | [ADR-098](decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Purpose and scope | [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine |
## Open Questions
None specific to validation. The three typedef OQs (OQ-069, OQ-070,
OQ-071) are about layout, platform support, and schema construction —
not validation.
## References
- `docs/research/alknet-typedef/findings.md` §"Validation" — the POC's
custom keyword validators for all 17 kinds
- [ADR-098](decisions/098-error-handling-validation-strategy.md) —
error handling and validation strategy
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds that the
validators check
- [data-access.md](data-access.md) — read/write functions that operate
on the same buffers