docs(architecture): add alknet-typedef crate specs, ADRs 095-098, and OQs 069-071

Add the alknet-typedef architecture specification — the binary struct
engine that takes JSON Schema with TypeDef:* custom keywords and produces
offset maps, read/write functions, and validation.

Specs (docs/architecture/crates/typedef/):
- overview.md: purpose, 'schema is the format' principle, consumers, scope
- schema-layer.md: 17 TypeDef:* kinds, jsonschema integration, annotations
- layout-engine.md: two layout modes, three variable-length strategies
- data-access.md: read/write, TUnion dispatch, field paths, zero-copy
- validation.md: custom keyword validators, TypedefError, TypedefEngine

ADRs:
- 095: Purpose, scope, and the jsonschema engine
- 096: Two layout modes — packed sequential vs aligned static
- 097: Schema annotations — endianness, alignment, encoding, TUnion
- 098: Error handling and validation strategy

OQs (deferred(scope)):
- 069: Arrays of variable-length-element structs
- 070: no_std + alloc support
- 071: Builder API for schema construction

Index updates: README doc table + ADR table, open-questions.md theme
table + Deferred/Blocked section, overview.md crate graph.

Grounded in the alknet-typedef POC (26 tests passing) and the
call-channels-unification research. Reviewed by architecture-reviewer;
all critical issues, warnings, and suggestions addressed.
This commit is contained in:
deepseek-v4-pro committed 2026-07-20 11:57:03 +00:00
1 parent bf0f827bf4
commit 85c5590001
16 files changed
+2230 -1

No files matched your search

+11 -1
View File
@@ -296,6 +296,12 @@ adapter location map is now consistent: all HTTP-backed adapters
| [crates/channels/channels-adapter.md](crates/channels/channels-adapter.md) | draft | `ChannelsAdapter`, `ChannelManager`, demux/mux contracts (REQ-CH-01..04), two-pump pattern (ADR-078) | | [crates/channels/channels-adapter.md](crates/channels/channels-adapter.md) | draft | `ChannelsAdapter`, `ChannelManager`, demux/mux contracts (REQ-CH-01..04), two-pump pattern (ADR-078) |
| [crates/channels/channel-operations.md](crates/channels/channel-operations.md) | draft | `channel/open`/`close`/`control`/`resources/subscribe`, ACL flow, `direction` semantics, hub relay contract (ADR-079) | | [crates/channels/channel-operations.md](crates/channels/channel-operations.md) | draft | `channel/open`/`close`/`control`/`resources/subscribe`, ACL flow, `direction` semantics, hub relay contract (ADR-079) |
| [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved | | [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved |
| [crates/typedef/README.md](crates/typedef/README.md) | draft | alknet-typedef crate — binary struct engine; JSON Schema with `TypeDef:*` custom keywords → offset map + read/write + validation |
| [crates/typedef/overview.md](crates/typedef/overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 16 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md) | draft | Offset computation, two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [crates/typedef/data-access.md](crates/typedef/data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation |
## ADR Table ## ADR Table
@@ -395,10 +401,14 @@ adapter location map is now consistent: all HTTP-backed adapters
| [092](decisions/092-bistream-as-the-handler-leaf.md) | `BiStream` as the Handler Leaf — Unify the Split-Pair `accept_bi` | Accepted (amends ADR-070's `accept_bi` return type; amends ADR-065's `from_stream`/`from_bidi` constructors; amends ADR-074's `ChannelBidiStreamSource::accept_bi` return type; `Connection::from_stream` removed; `from_bidi` is the only public stream constructor) | | [092](decisions/092-bistream-as-the-handler-leaf.md) | `BiStream` as the Handler Leaf — Unify the Split-Pair `accept_bi` | Accepted (amends ADR-070's `accept_bi` return type; amends ADR-065's `from_stream`/`from_bidi` constructors; amends ADR-074's `ChannelBidiStreamSource::accept_bi` return type; `Connection::from_stream` removed; `from_bidi` is the only public stream constructor) |
| [093](decisions/093-channels-pure-channel-multiplexing.md) | alknet-channels — Pure Channel Multiplexing (8-Byte Header, No `stream_type`) | Accepted (amends ADR-071 — 8-byte header; ADR-074 — `into_sub_streams` removed; reverses ADR-077 — TTY always uses its 5-byte format; amends the channels-facing clauses of ADR-072/073/075/076/080/081) | | [093](decisions/093-channels-pure-channel-multiplexing.md) | alknet-channels — Pure Channel Multiplexing (8-Byte Header, No `stream_type`) | Accepted (amends ADR-071 — 8-byte header; ADR-074 — `into_sub_streams` removed; reverses ADR-077 — TTY always uses its 5-byte format; amends the channels-facing clauses of ADR-072/073/075/076/080/081) |
| [094](decisions/094-per-identity-channel-cap.md) | Per-Identity Channel Cap as DoS Defense | Accepted (amends ADR-076 — per-connection `max_channels` reframed as a memory bound; 256 per `PeerId` enforced via `ChannelLifecyclePolicy` in `channels-call`; symmetric; spoke caps hub as direct caller) | | [094](decisions/094-per-identity-channel-cap.md) | Per-Identity Channel Cap as DoS Defense | Accepted (amends ADR-076 — per-connection `max_channels` reframed as a memory bound; 256 per `PeerId` enforced via `ChannelLifecyclePolicy` in `channels-call`; symmetric; spoke caps hub as direct caller) |
| [095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | alknet-typedef — Purpose, Scope, and the jsonschema Engine | Accepted |
| [096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | Accepted |
| [097](decisions/097-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators | Accepted |
| [098](decisions/098-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | Accepted |
## Open Questions ## Open Questions
Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (68 OQs across 20 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](questions/) (`NNN-slug.md`, mirroring the ADR convention). Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (71 OQs across 21 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](questions/) (`NNN-slug.md`, mirroring the ADR convention).
## Document Lifecycle ## Document Lifecycle
+106
View File
@@ -0,0 +1,106 @@
---
status: draft
last_updated: 2026-07-20
---
# alknet-typedef
The binary struct engine: a small Rust crate that takes a JSON Schema
with `TypeDef:*` custom keywords and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
## Documents
| Document | Status | Description |
|----------|--------|-------------|
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [schema-layer.md](schema-layer.md) | draft | The 17 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [validation.md](validation.md) | draft | Custom keyword validators for all 16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation, `TypedefEngine` |
## Applicable ADRs
| ADR | Title | Relevance |
|-----|-------|-----------|
| [095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| [096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` |
| [097](../../decisions/097-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations |
| [098](../../decisions/098-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors |
## Relevant Open Questions
| OQ | Title | Status | Relevance |
|----|-------|--------|-----------|
| OQ-069 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it |
| OQ-070 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case |
| OQ-071 | Builder API for schema construction | deferred(scope) | Schemas are authored in TypeBox or hand-written JSON for v1; blocked on a concrete need |
## Key Design Principles
1. **The schema is the format.** A JSON Schema with `TypeDef:*` custom
keywords is both the validation spec and the layout spec. No separate
format definition, no separate parser, no separate validator. One
schema, three uses: validate, compute offsets, access data. See
[overview.md](overview.md) and [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
2. **jsonschema is the validation engine, not a custom engine.** The
`jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with
custom keyword support. The novel code is the offset computation, not
the validation. This eliminates ~14,000 lines of hand-rolled schema
engines (typebox-rs, alktype). See [schema-layer.md](schema-layer.md)
and [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
3. **Two layout modes for two use cases.** Packed sequential
(`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP,
channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats
(metatensor). The consumer selects the mode; the schema is the same.
See [layout-engine.md](layout-engine.md) and
[ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md).
4. **Variable-length types default to inline length-prefixing.**
`[length: u32][data]` is the universal pattern used by channels, SFTP,
TTY, and most binary protocols. Offset indirection (the metatensor
blob tensor pattern) is opt-in via the `encoding` annotation. See
[layout-engine.md](layout-engine.md) and
[ADR-097](../../decisions/097-schema-annotations.md).
5. **TUnion supports both byte-offset and field-name discriminators.**
Byte-offset for protocol dispatch (SFTP type bytes, call protocol
event types). Field-name for the typedef.ts string pattern. See
[data-access.md](data-access.md) and
[ADR-097](../../decisions/097-schema-annotations.md).
6. **Endianness is per-schema, default little-endian.** The engine reads
the `"endian"` annotation and byte-swaps accordingly. SFTP consumers
specify `"endian": "big"`. See [layout-engine.md](layout-engine.md)
and [ADR-097](../../decisions/097-schema-annotations.md).
7. **Validation is opt-in, built once at load time.** The jsonschema
validator is compiled once at schema load time. Access-time validation
is a fast `is_valid()` check. High-throughput paths can skip
validation; security-sensitive paths can validate every frame. See
[validation.md](validation.md) and
[ADR-098](../../decisions/098-error-handling-validation-strategy.md).
8. **Not a serialization framework.** The typedef engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use typedef. See [overview.md](overview.md) and
[ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
## References
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
passing, two layout modes, TUnion dispatch, endianness)
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine" — the origin of this research
thread
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
@@ -0,0 +1,224 @@
---
status: draft
last_updated: 2026-07-20
---
# alknet-typedef — Data Access
The data access layer: read/write functions, TUnion dispatch, field paths,
zero-copy access for fixed-size types, and length-prefix reading for
variable-length types. This is the consumer-facing API — given a compiled
`TypedefEngine` and a byte buffer, read and write fields at
schema-computed offsets.
## Read/Write Model
The typedef engine operates on raw byte buffers (`&[u8]` for reading,
`&mut [u8]` for writing). There is no intermediate `Value` tree, no
reflection, no dynamic dispatch per field. The engine uses the offset map
(or `LayoutBuilder`/`SequentialReader`) to locate fields, then performs
typed access at the computed positions.
### Fixed-size types
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, etc.) are accessed via
zero-copy pointer casts:
```rust
// Read a u32 at a known offset
fn read_u32(buffer: &[u8], offset: usize, endian: Endian) -> u32 {
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
}
}
// Write a u32 at a known offset
fn write_u32(buffer: &mut [u8], offset: usize, value: u32, endian: Endian) {
let bytes = match endian {
Endian::Little => value.to_le_bytes(),
Endian::Big => value.to_be_bytes(),
};
buffer[offset..offset+4].copy_from_slice(&bytes);
}
```
The engine applies endianness at access time based on the schema's
`"endian"` annotation (ADR-097). The offset computation is
endian-agnostic.
### Variable-length types (inline length-prefixing)
For variable-length types with inline length-prefixing (the default):
```rust
// Read a length-prefixed string
fn read_string<'a>(buffer: &'a [u8], offset: usize) -> &'a str {
let len = u32::from_le_bytes(buffer[offset..offset+4].try_into().unwrap()) as usize;
std::str::from_utf8(&buffer[offset+4..offset+4+len]).unwrap()
}
// Write a length-prefixed string
fn write_string(buffer: &mut [u8], offset: usize, value: &str) {
let data = value.as_bytes();
buffer[offset..offset+4].copy_from_slice(&(data.len() as u32).to_le_bytes());
buffer[offset+4..offset+4+data.len()].copy_from_slice(data);
}
```
The engine reads the 4-byte length prefix at the field's offset, then
slices the data that follows. For writing, the engine writes the length
prefix + data.
In packed sequential mode, the `SequentialReader` uses the length prefix
to determine the position of the next field. In aligned static mode, the
`OffsetMap` records the position of the length prefix; the variable data
is accessed separately.
### Variable-length types (offset indirection)
For variable-length types with offset indirection (opt-in):
```rust
// Read an offset-indirect string
fn read_string_indirect(data_region: &[u8], offset: usize) -> &str {
let ptr_offset = u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()) as usize;
let ptr_length = u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()) as usize;
std::str::from_utf8(&data_region[ptr_offset..ptr_offset+ptr_length]).unwrap()
}
```
The field is a struct `{offset: u32, length: u32}` at a known position
in the `OffsetMap`. The consumer provides the data region separately; the
engine reads the offset and length, then slices the data region.
## TUnion Dispatch
TUnion dispatch reads the discriminator value, looks up the variant
schema, and then reads the variant's fields. The dispatch mechanism
differs by discriminator kind (ADR-097).
### Byte-offset discriminator
```rust
fn read_union(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
let disc = &schema["discriminator"];
let offset = disc["offset"].as_u64().unwrap() as usize;
let disc_type = disc["type"].as_str().unwrap(); // e.g., "TypeDef:Uint8"
// Read the discriminator value
let disc_value: u8 = read_u8(buffer, offset);
let key = disc_value.to_string(); // "5", "6", "101"
// Look up the variant schema
let mapping = &schema["mapping"];
let variant_schema = &mapping[&key];
// Read the variant struct starting at offset + discriminator_size
let variant_offset = offset + 1; // discriminator_size for Uint8
read_struct(buffer, variant_offset, variant_schema)
}
```
The discriminator is a fixed-size integer at a known byte offset. The
mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`.
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
1..N are the variant struct. The call protocol's 5 event types
(`call.requested` → 0x01, etc.) use the same pattern.
### Field-name discriminator
```rust
fn read_union_field(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
let disc = &schema["discriminator"];
let field_name = disc["name"].as_str().unwrap(); // e.g., "type"
// Read the discriminator field like any other field
let disc_value = read_field(buffer, field_name, schema)?;
// Look up the variant schema
let mapping = &schema["mapping"];
let variant_schema = &mapping[disc_value.as_str().unwrap()];
// Read the variant struct
read_struct(buffer, variant_offset, variant_schema)
}
```
The discriminator is a named field within the struct. Its offset is
computed like any other field. The mapping keys are string values.
## Field Paths
Fields are addressed by dotted paths: `"header.version"`, `"payload.data"`.
The `OffsetMap` stores fully-qualified paths. The read/write functions
accept a field path and look up the byte range:
```rust
fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> {
let range = self.offset_map.get(field_path)
.ok_or_else(|| TypedefError::Offset {
field_path: field_path.to_string(),
reason: "field not found in offset map".to_string(),
})?;
if buffer.len() < range.end {
return Err(TypedefError::Access {
field_path: field_path.to_string(),
reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()),
});
}
Ok(read_f32_raw(buffer, range.start, self.endian))
}
```
Nested structs produce nested field paths. The offset computation
propagates the field path prefix during recursion, so the `OffsetMap`
contains entries like `"header.version"` and `"header.magic"`.
## Zero-Copy Access
For fixed-size types, the engine provides zero-copy access — the consumer
gets a reference to the bytes in the buffer, not a copy. This is
important for performance-sensitive paths (metatensor tensor access,
high-throughput protocol parsing).
For variable-length types with inline length-prefixing, the engine
returns a slice of the buffer — the string or byte array data is not
copied. The consumer gets a `&str` or `&[u8]` that borrows from the
input buffer.
For offset-indirect types, the consumer provides the data region; the
engine returns a slice of that region.
## Error Handling
Read/write errors carry the field path for debugging. See
[ADR-098](../../decisions/098-error-handling-validation-strategy.md) and
[validation.md](validation.md) for the full error model.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Two layout modes | [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Determines whether offsets are fixed (OffsetMap) or sequential (SequentialReader) |
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Endianness, encoding, and TUnion discriminator shapes that control data access |
| Error handling | [ADR-098](../../decisions/098-error-handling-validation-strategy.md) | Field-path-carrying errors for read/write operations |
## Open Questions
See [open-questions.md](../../open-questions.md) for full details.
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
— affects the sequential walking logic for array access.
## References
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
(read/write round-trip) and POC 2 (SFTP byte-identical round-trip)
- [layout-engine.md](layout-engine.md) — offset computation that produces
the positions this layer reads/writes at
- [validation.md](validation.md) — validation that runs on the same
buffers
@@ -0,0 +1,254 @@
---
status: draft
last_updated: 2026-07-20
---
# alknet-typedef — Layout Engine
The layout engine: offset computation, the two layout modes (packed
sequential vs aligned static), alignment, endianness, and variable-length
field handling. This is the novel code — the recursive walk of the schema
JSON that computes byte positions for each field.
## The Two Layout Modes
The POCs surfaced that protocols and mmap-friendly formats need different
layout strategies. This is the most important architectural finding —
decided in [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md).
### Mode 1: Packed sequential (protocol wire formats)
Fields are packed with no alignment padding. Variable-length fields shift
all subsequent fields. Used by SFTP, channels, TTY, and most binary
protocols.
**Components:**
- **`LayoutBuilder`** — takes a schema and actual data sizes for
variable-length fields, computes byte positions for each field in a
packed layout. Used at write time when the consumer knows the data
sizes upfront.
- **`SequentialReader`** — walks a buffer field-by-field according to the
schema, reading length prefixes to determine variable-length data
positions. Used at read time when the consumer is parsing an incoming
frame.
**How it works:**
For a struct with fields `[u8, u32, string]`:
```
LayoutBuilder (write):
field[0] u8: offset 0, size 1
field[1] u32: offset 1, size 4
field[2] string: offset 5, size 4 (length prefix) + data_len
total: 9 + data_len
SequentialReader (read):
read u8 at offset 0
read u32 at offset 1
read u32 length prefix at offset 5 → data_len
read string data at offset 9, length data_len
next field at offset 9 + data_len
```
There is no alignment padding. The `u32` at offset 1 is unaligned — this
is correct for protocol wire formats, which pack fields tightly.
**Variable-length fields in packed mode:**
The `LayoutBuilder` takes actual data sizes for variable-length fields
to compute correct positions for subsequent fields. The consumer must
know the data sizes before writing — this is inherent to packed layouts.
The `SequentialReader` reads each field's length prefix to determine the
data extent and the position of the next field. The reader walks the
buffer sequentially; it cannot jump to field N without reading fields
0..N-1 first.
### Mode 2: Aligned static (mmap-friendly formats)
Fields have fixed positions with natural alignment padding.
Variable-length fields get a 4-byte length prefix at a known offset; the
variable data is not included in the static layout. Used by metatensor
and safetensors.
**Component:**
- **`OffsetMap`** — walks the schema once, computes fixed byte positions
for each field based on type sizes and alignment. The output is a flat
table of `(field_path, byte_range)` pairs. Used for both read and write
at known offsets.
**How it works:**
For a struct with fields `[u8, u32, f32]` and natural alignment:
```
OffsetMap:
field[0] u8: offset 0, size 1
field[1] u32: offset 4, size 4 (3 bytes padding after u8)
field[2] f32: offset 8, size 4
total: 12 (struct aligned to 4)
```
The `u32` is aligned to offset 4 (its natural alignment). The consumer
can read `field[1]` at offset 4 without reading `field[0]` first — random
access by field path.
**Variable-length fields in aligned mode:**
Variable-length fields get a 4-byte length prefix at a known offset. The
variable data lives outside the static layout — either immediately after
the fixed fields (inline length-prefixing) or in a separate data region
(offset indirection). The `OffsetMap` records the position of the length
prefix (or the `{offset, length}` pair for offset-indirect fields).
For inline length-prefixing, the variable data follows the fixed fields
but is not included in the `OffsetMap`'s field ranges. The consumer reads
the length prefix from the `OffsetMap`'s known offset, then slices the
data region.
For offset indirection, the field is a struct `{offset: u32, length: u32}`
at a known position in the `OffsetMap`. The consumer reads the offset and
length, then slices the separate data region.
## Offset Computation Algorithm
The offset computation is a recursive walk of the schema JSON. The
algorithm is the same for both modes; the difference is whether alignment
padding is inserted between fields.
### Fixed-size types
For each fixed-size type, the algorithm:
1. Determines the type's byte size from the `TypeDef:*` kind.
2. In aligned mode: inserts padding to satisfy the type's alignment
(or the field's `align` annotation, or the struct's `align` default).
3. Records the field's `(start, end)` range.
4. Advances the current offset by the type's size.
### Composite types
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
are computed relative to the struct's start offset. The struct's total
size is the sum of its fields' sizes (plus alignment padding in aligned
mode). The struct itself may have an `align` annotation that rounds up
its total size.
**`TUnion`:** The discriminator occupies `offset..offset + discriminator_size`
bytes. For byte-offset discriminators, the variant struct starts at
`offset + discriminator_size`. For field-name discriminators, the
discriminator is just another field — its offset is computed like any
other field, and the variant struct follows at the end of the
discriminator field.
In aligned static mode, the union's total size is `discriminator_size +
max(variant_sizes)`, where variant sizes are computed from the schema
(variable-length data lives outside the static layout).
In packed sequential mode, variant sizes depend on the actual sizes of
variable-length fields within each variant, which aren't known at schema
time. The `LayoutBuilder` takes the actual variant discriminator value
and data sizes at write time, computes the size of the selected variant,
and uses that for the union's total size. The `SequentialReader` reads
the discriminator first, looks up the variant schema, then reads the
variant struct sequentially — it doesn't need to know the union's total
size upfront.
**`TArray` of fixed-size elements:** Element stride = element size (plus
alignment padding in aligned mode). Element `i` starts at
`array_offset + i × stride`. The array's total size is `count × stride`.
**`TArray` of variable-length-element structs:** Deferred for v1
(OQ-069).
### Variable-length types
The typedef engine supports three strategies for variable-length types
(see [schema-layer.md](schema-layer.md) §Variable-length types and
[ADR-097](../../decisions/097-schema-annotations.md) §3 for the full
annotation shapes).
**Strategy 1: Inline length-prefixing (default).**
1. Records the position of the 4-byte length prefix.
2. In aligned mode: the length prefix is aligned; the variable data is
not included in the static layout.
3. In packed mode: the `LayoutBuilder` takes the actual data size to
compute the length prefix value and the position of subsequent fields.
The `SequentialReader` reads the length prefix to determine the data
extent and the position of the next field.
**Strategy 2: Fixed-size reservation (`maxLength`).**
1. In aligned static mode: reserves `maxLength` bytes at a fixed offset.
Data shorter than `maxLength` is zero-padded. Subsequent fields have
known, unchanging offsets — the field is fixed-size from the layout
perspective. This is the database `VARCHAR(N)` pattern.
2. In packed sequential mode: `maxLength` is a validation constraint
only. The engine uses strategy 1 (inline length-prefixing) because
protocols don't benefit from fixed-size reservation.
**Strategy 3: Offset indirection (`"encoding": "offset-indirect"`).**
1. The field is a struct `{offset: u32, length: u32}`.
2. The `OffsetMap` records the position of this struct.
3. The consumer provides the data region separately. This is the
metatensor blob tensor pattern — the index struct lives in one region,
the blob data lives in another.
### Nested structs and field paths
Nested structs produce dotted field paths: `header.version`,
`header.magic`. The offset computation propagates the field path prefix
during recursion. The `OffsetMap` stores fully-qualified paths.
### Endianness
Endianness is per-schema (ADR-097). The offset computation is
endian-agnostic — it computes byte positions, not byte values. The
read/write functions apply endianness when converting between bytes and
typed values. The engine reads the `"endian"` annotation from the schema
and byte-swaps accordingly.
## Mode Selection
The consumer selects the mode at engine construction time. The choice is
determined by the use case, not by the schema:
- **Protocol consumer** (SFTP, binary call frames, TTY negotiation):
uses `LayoutBuilder` for writing and `SequentialReader` for reading.
- **mmap consumer** (metatensor): uses `OffsetMap` for both reading and
writing.
The same schema can be used in either mode. A schema describing an SFTP
packet can be consumed by a `SequentialReader` (for parsing incoming
frames) and a `LayoutBuilder` (for constructing outgoing frames). A schema
describing a metatensor layout can be consumed by an `OffsetMap` (for
mmap access).
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Two layout modes | [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential for protocols; aligned static for mmap formats |
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Endianness, alignment, encoding annotations that control layout behavior |
## Open Questions
See [open-questions.md](../../open-questions.md) for full details.
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
— requires lazy walking logic; blocked on a concrete consumer that
needs it.
## References
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
- [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) —
the two layout modes decision
- [ADR-097](../../decisions/097-schema-annotations.md) — schema
annotations
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds and their
byte sizes
- [data-access.md](data-access.md) — read/write functions that use the
computed offsets
@@ -0,0 +1,198 @@
---
status: draft
last_updated: 2026-07-20
---
# alknet-typedef — Overview
The binary struct engine: a small Rust crate that takes a JSON Schema
with `TypeDef:*` custom keywords and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
This document covers the crate's purpose, the "schema is the format"
principle, its dependency edges, consumers, and scope boundaries.
Component details are in the sibling documents.
## What
`alknet-typedef` is a library crate that consumes JSON Schemas annotated
with `TypeDef:*` custom keywords (the same kinds defined in TypeBox's
`typedef.ts`) and produces three capabilities:
1. **An offset map** — walks the schema, computes byte offsets for each
field based on type sizes, field order, and alignment.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that a
buffer's bytes match the schema's type constraints.
The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing). The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are ~10 lines each.
The crate is ~1,900 lines (POC verified, 26 tests passing). It replaces
two prior attempts that built their own jsonschema engines — typebox-rs
(~8,400 lines) and alktype (~5,600 lines) — with `jsonschema` + an
offset map + ~50 lines of custom keyword implementations. See
[ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
## Why
The crate's purpose is to be the binary struct engine for every alknet
component that reads or writes binary data at computed offsets. Instead
of per-protocol serde structs (russh-sftp's 29 packet types), per-handler
wire format code (TTY's 5-byte format parser), or per-format offset
computation (metatensor's tensor access), all of these become instances
of the same engine with different schemas.
The guiding insight:
> **The schema is the format.** A JSON Schema with `TypeDef:Float32`,
> `TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
> the layout spec. No separate format definition, no separate parser, no
> separate validator. One schema, three uses: validate, compute offsets,
> access data.
This is the convergence of three threads identified in the
call-channels-unification research: the `typedef.ts` schema kinds from
TypeBox, the russh-sftp protocol packets, and the metatensor format. The
common pattern: a JSON Schema describes the shape of binary data, and
the binary data is the struct's bytes at computed offsets.
The crate was bumped up in the timeline when the call-channels-unification
research surfaced that channels, TTY, and the binary call protocol are
all variations on the same wire-format family — `[discriminant][length][payload]`.
The typedef engine makes the "channels is call with a binary data plane"
unification concrete: the binary data plane's wire format is the call
protocol's own schema system, just binary-encoded. The `channel_open`
marker says "use binary framing"; the typedef engine says "here's how to
read/write the binary payload."
## The "Schema Is the Format" Principle
A JSON Schema with `TypeDef:*` custom keywords serves three roles
simultaneously:
| Role | Mechanism | When |
|------|-----------|------|
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) |
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
No separate format definition, no separate parser, no separate validator.
The schema is the single source of truth for the binary format. Adding a
new field to a protocol is adding a property to the schema JSON — the
engine computes the new offsets automatically.
This is the same principle as `#[repr(C)]` struct field access, but at
runtime from a portable JSON Schema instead of at compile-time from
language-specific annotations. The schema is the ABI contract.
## Dependencies
```
alknet-typedef
├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support
├── serde_json (with preserve_order) — schema parsing; field order is load-bearing
└── (no tokio, no platform deps) — WASM-clean by construction
```
`alknet-typedef` is dependency-light: `jsonschema` + `serde_json` only.
No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for
browser use. The `jsonschema` crate is already in the workspace at
`/workspace/jsonschema/` but not yet used by any alknet crate — typedef
is the first consumer.
`serde_json` requires the `preserve_order` feature because field order
is load-bearing for binary layouts. The order of properties in the
schema JSON determines the order of fields in the binary struct.
## Consumers
| Consumer | Schema describes | Engine provides |
|----------|-----------------|-----------------|
| russh-sftp | 29 packet structs + Packet union (byte discriminator) | Read/write SFTP frames from bytes |
| metatensor | Model layout (ConvNet struct, tensor refs) | Offset map for mmap'd tensor access |
| binary call frames | `call.requested` / `call.responded` / etc. structs | Read/write binary call frames |
| TTY negotiation | `NegotiateRequest` / `NegotiateResponse` structs | Read/write TTY control frames |
| channels wire | `ChunkHeader { channel_id, length }` | Already trivial (8 bytes, no schema needed) |
The russh-sftp case is the most instructive and the highest-value POC
target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written
dispatch on a type byte followed by serde deserialization. Under typedef,
the dispatch is `TUnion` with a byte-offset discriminator — the schema
says "byte 0 is the discriminator, bytes 1..N are the variant struct."
The engine reads the discriminator, looks up the variant schema, computes
offsets, reads fields. Same result, no per-packet-type code.
## Scope Boundaries (What This Is Not)
These boundaries are decided in [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
is a *format* (8-byte header + JSON header + binary data) that uses the
typedef engine for its offset computation and tensor access.
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
`Value.Convert` — schema evolution — is out of scope for v1. The engine
should not do anything that explicitly blocks adding a Value system
later.
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
concern. The typedef engine consumes schemas; it does not generate them.
- **Not a schema builder.** The typedef engine does not provide a fluent
API for constructing schemas. Schemas are plain JSON — authored in
TypeBox, generated by ujsx components, or hand-written. A builder API
is deferred (OQ-071).
- **Not a serialization framework.** The typedef engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use typedef.
## Architecture (component pointers)
- **[schema-layer.md](schema-layer.md)** — the 17 `TypeDef:*` kinds,
jsonschema custom keyword integration, TypeBox interop, schema
annotations (endianness, alignment, encoding, TUnion discriminators).
- **[layout-engine.md](layout-engine.md)** — offset computation, the two
layout modes (packed sequential vs aligned static), alignment,
endianness, variable-length field handling.
- **[data-access.md](data-access.md)** — read/write functions, TUnion
dispatch, field paths, zero-copy access for fixed-size types,
length-prefix reading for variable-length types.
- **[validation.md](validation.md)** — custom keyword validators for all
16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time
validation, `TypedefEngine` as the compiled form of a schema.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Purpose, scope, and the jsonschema engine | [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| Two layout modes | [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats |
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) |
| Error handling and validation | [ADR-098](../../decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
## Open Questions
See [open-questions.md](../../open-questions.md) for full details.
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs.
- **OQ-070** (deferred(scope)): `no_std` + `alloc` support.
- **OQ-071** (deferred(scope)): Builder API for schema construction.
## References
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
passing, two layout modes, TUnion dispatch, endianness)
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine" — the origin of this research
thread
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
@@ -0,0 +1,348 @@
---
status: draft
last_updated: 2026-07-20
---
# alknet-typedef — Schema Layer
The schema layer: the 16 `TypeDef:*` custom type kinds, their mapping to
Rust types and byte sizes, the `jsonschema` custom keyword integration,
TypeBox interop, and the concrete JSON shapes for schema-level annotations.
## The 17 TypeDef Kinds
These are the custom schema kinds defined in TypeBox's `typedef.ts`
(`/workspace/@alkdev/typebox/example/typedef/typedef.ts`, 619 lines) and
ported to Rust via `jsonschema` custom keywords. Each kind carries binary
layout semantics — a known byte size (for fixed-size types) or a known
encoding strategy (for variable-length types).
| Kind | TypeBox key | Rust type | Size | Category |
|------|-------------|-----------|------|----------|
| `TFloat32` | `TypeDef:Float32` | `f32` | 4 | fixed |
| `TFloat64` | `TypeDef:Float64` | `f64` | 8 | fixed |
| `TInt8` | `TypeDef:Int8` | `i8` | 1 | fixed |
| `TInt16` | `TypeDef:Int16` | `i16` | 2 | fixed |
| `TInt32` | `TypeDef:Int32` | `i32` | 4 | fixed |
| `TUint8` | `TypeDef:Uint8` | `u8` | 1 | fixed |
| `TUint16` | `TypeDef:Uint16` | `u16` | 2 | fixed |
| `TUint32` | `TypeDef:Uint32` | `u32` | 4 | fixed |
| `TBoolean` | `TypeDef:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed |
| `TString` | `TypeDef:String` | length-prefixed UTF-8 | variable | variable |
| `TBytes` | `TypeDef:Bytes` | length-prefixed raw bytes | variable | variable |
| `TStruct` | `TypeDef:Struct` | record of fields | sum of field sizes | composite |
| `TUnion` | `TypeDef:Union` | tagged union | discriminator + variant | composite |
| `TArray` | `TypeDef:Array` | repeated element | count × element size | composite |
| `TEnum` | `TypeDef:Enum` | u32 index into enum values | 4 (fixed) | fixed |
| `TRecord` | `TypeDef:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
| `TTimestamp` | `TypeDef:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
### Fixed-size types
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset
computation uses these sizes directly. Read/write is zero-copy pointer
cast for these types.
**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other
values are invalid and produce a `TypedefError::Access` on read.
**`TEnum` binary representation:** A `u32` index into the enum's declared
values, in declaration order. The first declared value is index 0, the
second is index 1, etc. The enum's values are declared via the standard
JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`).
The `TypeDef:Enum` custom keyword signals that the type is an enum for
layout purposes; the built-in `enum` keyword provides the value list.
The `u32` index is always little-endian (enum indices are not protocol
data — they are internal to the schema). See [ADR-097](../../decisions/097-schema-annotations.md).
### Variable-length types
`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes.
The typedef engine supports three strategies for handling variable-length
types in binary layouts, selected by the `encoding` annotation and the
standard JSON Schema `maxLength` keyword:
| Strategy | Encoding annotation | Layout behavior | Use case |
|----------|-------------------|-----------------|----------|
| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) |
| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) |
| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) |
**Strategy 1: Inline length-prefixing (default).** The field's fixed
portion is a 4-byte length prefix at a computed offset. The variable data
follows immediately after. In packed sequential mode, the length prefix
determines the position of subsequent fields. In aligned static mode, the
length prefix is at a known offset; the variable data is not included in
the static layout. This is the universal pattern used by channels, SFTP,
TTY, and most binary protocols.
**Strategy 2: Fixed-size reservation.** When a variable-length field
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
than `maxLength` is zero-padded; data longer than `maxLength` is a
validation error. This makes the field fixed-size from the layout
perspective — subsequent fields have known, unchanging offsets. This is
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
pattern for fields with known maximum sizes.
In packed sequential mode, `maxLength` is a validation constraint only —
the engine still uses inline length-prefixing (strategy 1) because
protocols don't benefit from fixed-size reservation.
**Strategy 3: Offset indirection.** The field is a struct
`{offset: u32, length: u32}` at a known position. The consumer provides
the data region separately; the engine reads the offset and length, then
slices the data region. This is the metatensor blob tensor pattern — the
index struct lives in one region, the blob data lives in another. Enables
mmap-friendly random access to variable-length data without parsing
length prefixes and without reserving worst-case space.
**Default strategy selection:**
- In packed sequential mode: always strategy 1 (inline length-prefixing).
`maxLength` is a validation constraint only.
- In aligned static mode: strategy 2 (fixed-size reservation) if
`maxLength` is declared; strategy 3 (offset indirection) if
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
length-prefixing) otherwise.
**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3)
respects the schema's `"endian"` annotation (ADR-097). In little-endian
mode, the length is `u32::from_le_bytes`. In big-endian mode, the length
is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have
consistent byte order for both field values and length prefixes.
**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`.
Otherwise identical to `TString` in layout (same three strategies).
**`TRecord`:** A string-keyed map. Binary layout is a count-prefixed
sequence of `(key, value)` pairs: `[count: u32][key_len: u32][key_bytes]
[value_len: u32][value_bytes]...`. The count is the number of entries.
Each key is a length-prefixed UTF-8 string. Each value is the record's
declared value type (specified via the `"values"` property in the schema,
e.g., `"values": { "TypeDef:Float32": true }`). The count prefix respects
the schema's endianness. In aligned static mode with `maxLength`, the
entire record is reserved at `maxLength` bytes (zero-padded).
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
fixed-size reservation (strategy 2 with `maxLength`). The engine does not
parse or validate the timestamp format beyond UTF-8 — the jsonschema
validator checks RFC 3339 conformance at the JSON level.
`TArray` is variable-length when the element type is variable-length or
when the count is not known at schema time. For fixed-size element arrays
with a known count, the size is `element_size × count`.
**`TArray` count declaration:** The array count is declared via the
standard JSON Schema `"minItems"` and `"maxItems"` keywords. When
`minItems == maxItems`, the array has a fixed count known at schema time.
When they differ or are absent, the count is variable and the array uses
a length-prefixed encoding: `[count: u32][element_0]...[element_N]`.
The count prefix respects the schema's endianness.
### Composite types
`TStruct` and `TUnion` are composite — their size is the sum of their
fields' sizes (plus alignment padding in aligned static mode). The offset
computation recurses into their properties.
## jsonschema Custom Keyword Integration
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
via the `with_keyword` API. Each `TypeDef:*` kind is registered as a
custom keyword:
```rust
let validator = jsonschema::options()
.with_keyword("TypeDef:Float32", factory)
.with_keyword("TypeDef:Int32", factory)
.with_keyword("TypeDef:Struct", factory)
// ... all 17 kinds
.build(&schema)?;
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path — enabling cross-keyword awareness. The
`TypeDef:Struct` validator, for example, inspects the parent's
`properties` to validate each field against its declared `TypeDef:*` kind.
Each custom keyword implementation is ~10 lines. The `jsonschema` crate
handles all structural validation (object properties, required fields,
array items, enum values) — the custom keywords only need to validate
the leaf type constraints. See [validation.md](validation.md) for the
validator implementations.
This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side.
Same semantics, different language, same JSON Schema wire format. A
TypeBox schema serialized to JSON feeds directly into
`jsonschema::validator_for(&schema)` on the Rust side — zero translation.
## TypeBox Interop
TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox
schema like:
```typescript
const TensorRef = Type.Object({
dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]),
shape: Type.Array(Type.Number()),
data_offsets: Type.Tuple([Type.Number(), Type.Number()])
});
```
serialized to JSON is a standard JSON Schema with `type: "object"`,
`properties`, and `required`. That JSON feeds directly into the typedef
engine. The `TypeDef:*` custom keywords are added by TypeBox's
`TypeRegistry.Set` — they appear in the serialized JSON as additional
properties on the schema object.
The typedef engine does not depend on TypeBox or any JS toolchain. It
consumes JSON — whether that JSON was authored in TypeBox, generated by
a ujsx component, or hand-written. The schema is the interface.
## Schema Annotations
Schema-level annotations control binary layout behavior. These are
decided in [ADR-097](../../decisions/097-schema-annotations.md).
### Endianness
Schema-level annotation with a default of little-endian:
```json
{ "TypeDef:Struct": true, "endian": "big", "properties": { ... } }
```
- `"endian": "little"` (default) — read/write in little-endian byte order.
- `"endian": "big"` — read/write in big-endian byte order.
- Applies to the entire schema and all nested types.
### Alignment
Both struct-level and field-level, with field-level overriding:
```json
{
"TypeDef:Struct": true,
"align": 256,
"properties": {
"weight": { "TypeDef:Float32": true, "align": 16 }
}
}
```
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/bool, 2 for u16/i16, 4 for u32/i32/f32,
8 for u64/i64/f64, max field alignment for structs.
- Only meaningful in aligned static mode (ADR-096). Ignored in packed
sequential mode.
### Variable-length encoding
The typedef engine supports three strategies for variable-length types
(see §Variable-length types above for full details). The strategy is
selected by the `encoding` annotation and the standard JSON Schema
`maxLength` keyword:
```json
// Strategy 1: Inline length-prefixing (default, shorthand)
{ "TypeDef:String": true }
// Strategy 1: Explicit inline length-prefixing
{ "TypeDef:String": { "encoding": "length-prefixed" } }
// Strategy 2: Fixed-size reservation (uses standard maxLength)
{ "TypeDef:String": true, "maxLength": 256 }
// Strategy 3: Offset indirection (opt-in)
{ "TypeDef:String": { "encoding": "offset-indirect" } }
```
- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at
computed offset, variable data follows immediately. Used by protocol
wire formats.
- `maxLength` (standard JSON Schema keyword) — in aligned static mode,
reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the
field fixed-size from the layout perspective. In packed sequential
mode, `maxLength` is a validation constraint only.
- `"encoding": "offset-indirect"` — field is a struct
`{offset: u32, length: u32}` pointing into a separate data region.
The consumer provides the data region separately. Used by metatensor
blob tensors.
- Applies to all variable-length types: `TypeDef:String`, `TypeDef:Bytes`,
`TypeDef:Array`, `TypeDef:Record`, `TypeDef:Timestamp`.
### TUnion discriminators
Two discriminator kinds: byte-offset (protocol dispatch) and field-name
(typedef.ts pattern).
**Byte-offset discriminator** (SFTP type bytes, call protocol event types):
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "TypeDef:Uint8"
},
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" },
"101": { "$ref": "#/$defs/Status" }
}
}
```
- `"offset"` — byte position of the discriminator.
- `"type"` — the `TypeDef:*` kind of the discriminator (typically
`TypeDef:Uint8`).
- Mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`.
**Field-name discriminator** (typedef.ts pattern):
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "field",
"name": "type"
},
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
}
```
- `"name"` — the field name holding the discriminator value.
- Mapping keys are string values matching the discriminator field's value.
- The discriminator field is just another field in the struct.
Mapping values may be either inline schemas or `$ref` pointers. Both work.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators |
| Purpose and scope | [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle |
## Open Questions
See [open-questions.md](../../open-questions.md) for full details.
- **OQ-071** (deferred(scope)): Builder API for schema construction.
## References
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- [ADR-097](../../decisions/097-schema-annotations.md) — schema
annotation shapes
- [validation.md](validation.md) — custom keyword validator implementations
@@ -0,0 +1,292 @@
---
status: draft
last_updated: 2026-07-20
---
# alknet-typedef — Validation
The validation layer: custom keyword validators for all 17 `TypeDef:*`
kinds, the `TypedefError` enum, load-time vs access-time validation
strategy, and the `TypedefEngine` as the compiled form of a schema.
## Validation Strategy
Validation is delegated to the `jsonschema` crate (v0.46.5, Draft
2020-12). The typedef engine does not implement its own validation —
it registers custom keyword validators for each `TypeDef:*` kind and
lets `jsonschema` handle the structural validation (object properties,
required fields, array items, enum values).
The strategy is decided in [ADR-098](../../decisions/098-error-handling-validation-strategy.md):
1. **Load time:** Parse the schema JSON, build the layout engine, build the
jsonschema validator. This is the `TypedefEngine::compile(schema)` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation.
### What validation validates
The jsonschema validator operates on `serde_json::Value` instances — it
validates JSON representations of data, not raw byte buffers. This is
the correct separation of concerns:
- **JSON validation** (jsonschema): validates that a JSON document
conforms to the schema. Used for validating hand-written schemas,
TypeBox output, JSON payloads, or the JSON representation of a binary
struct after deserialization.
- **Binary access validation** (data access layer): the read/write
functions perform type-level validation at access time — range checks
for integers, UTF-8 validity for strings, buffer bounds checking.
These return `TypedefError::Access` with field paths.
The "schema is the format" principle means the same schema describes
both the JSON shape and the binary layout. The jsonschema validator
checks the JSON shape; the data access layer checks the binary layout.
A consumer that wants to validate a binary buffer end-to-end reads the
buffer into a `Value` tree via the data access layer, then validates
that `Value` against the jsonschema validator. This is a two-step
process, not a single `validate(buffer)` call.
### The `TypedefEngine` struct
The `TypedefEngine` is the compiled form of a schema. It supports both
layout modes (ADR-096) via an internal enum:
```rust
pub struct TypedefEngine {
layout: Layout, // packed or aligned (see below)
validator: jsonschema::Validator, // compiled once at load time
}
enum Layout {
Packed {
builder: LayoutBuilder,
reader: SequentialReader,
},
Aligned {
offset_map: OffsetMap,
},
}
```
The consumer selects the mode at construction time. The `Layout` enum
ensures the engine always has the correct layout strategy for the
consumer's use case — a protocol consumer gets `Packed`, an mmap
consumer gets `Aligned`. The validator is mode-agnostic (it operates on
`Value`, not raw bytes).
## Custom Keyword Validators
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
`jsonschema::options().with_keyword(...)`. The validators check leaf
type constraints; `jsonschema` handles all structural validation.
### Numeric type validators
**`TypeDef:Float32` / `TypeDef:Float64`:**
- Value must be a finite number.
- For `Float32`: value must be representable as `f32` (no precision loss
beyond `f32`'s mantissa).
**`TypeDef:Int8` / `TypeDef:Int16` / `TypeDef:Int32`:**
- Value must be an integer within the type's range.
- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
**`TypeDef:Uint8` / `TypeDef:Uint16` / `TypeDef:Uint32`:**
- Value must be a non-negative integer within the type's range.
- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295.
### String and binary validators
**`TypeDef:String`:**
- Value must be a valid UTF-8 string.
- If `maxLength` is specified in the schema, the string's byte length
must not exceed it.
**`TypeDef:Bytes`:**
- Value must be a string (JSON represents binary data as a string).
- If `maxLength` is specified, the byte length must not exceed it.
**`TypeDef:Enum`:**
- The `TypeDef:Enum` custom keyword signals that the type is an enum for
*layout* purposes (the engine needs to know it's a fixed-size u32 index,
not a variable-length string). The built-in `enum` keyword provides the
value list and handles value-membership validation. The custom keyword
validator is a no-op beyond the built-in check — it exists solely for
the layout engine to recognize the type.
**`TypeDef:Timestamp`:**
- Value must be a valid RFC 3339 timestamp string (the internet profile
of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
### Composite type validators
**`TypeDef:Struct`:**
- Value must be an object.
- Each property must match its declared `TypeDef:*` kind.
- Required fields must be present.
- The `jsonschema` crate's built-in `properties` and `required` keywords
handle the structural checks — the custom keyword only needs to
validate that each field's value matches its `TypeDef:*` kind.
**`TypeDef:Union`:**
- The discriminator value must be one of the mapping keys.
- The variant struct must match the declared schema for that discriminator
value.
**`TypeDef:Array`:**
- Value must be an array.
- Each element must match the array's declared element type.
- If `minItems`/`maxItems` is specified, the array length must be within
bounds.
### Other validators
**`TypeDef:Boolean`:**
- Value must be `true` or `false`.
**`TypeDef:Record`:**
- Value must be an object.
- All values must match the record's declared value type (specified via
the `"values"` property in the schema, e.g.,
`"values": { "TypeDef:Float32": true }`).
### Validator implementation pattern
Each custom keyword implementation is ~10 lines. Example for
`TypeDef:Float32`:
```rust
struct Float32Validator;
impl Keyword for Float32Validator {
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
match instance {
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
}
}
fn is_valid(&self, instance: &Value) -> bool {
instance.as_f64().map_or(false, |f| f.is_finite())
}
}
```
Registration:
```rust
let validator = jsonschema::options()
.with_keyword("TypeDef:Float32", |parent, value, path| {
Ok(Box::new(Float32Validator))
})
.build(&schema)?;
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path. This enables cross-keyword awareness — for
example, a `TypeDef:Struct` validator can inspect the parent's
`properties` to validate each field against its declared `TypeDef:*` kind.
## TypedefError
A single `TypedefError` enum covers all error conditions across the
engine's three phases (schema parsing, offset computation, read/write)
plus validation. Decided in [ADR-098](../../decisions/098-error-handling-validation-strategy.md).
```rust
pub enum TypedefError {
/// Schema parsing errors (invalid JSON, missing keywords, unknown TypeDef kinds).
Schema(String),
/// Offset computation errors (field not found, unsupported type).
Offset { field_path: String, reason: String },
/// Read/write errors (buffer too short, invalid UTF-8, value out of range).
Access { field_path: String, reason: String },
/// Validation errors (delegated to jsonschema).
Validation(ValidationError<'static>),
}
```
- **`Schema`** — for errors during `TypedefEngine::compile()`. Invalid
JSON, missing required keywords, unknown `TypeDef:*` kinds.
- **`Offset`** — for errors during offset computation. Field not found
in the schema, type not supported for offset computation, recursive
depth exceeded. Carries the field path.
- **`Access`** — for errors during read/write. Buffer too short, invalid
UTF-8 in a string field, value out of range for the target type.
Carries the field path.
- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The
`'static` lifetime is correct — the validator owns its schema reference
and lives for the lifetime of the `TypedefEngine`.
### Field-path-carrying errors
Read/write and offset errors include the field path for debugging:
```rust
Err(TypedefError::Access {
field_path: "header.version".to_string(),
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
})
```
This makes debugging binary format issues tractable — the error tells
you exactly which field failed and why.
## Validation Timing
### Load time: `TypedefEngine::compile()`
The expensive work happens once at schema load time:
1. Parse the schema JSON (`serde_json::from_str` with `preserve_order`).
2. Compute the offset map (or `LayoutBuilder`/`SequentialReader`).
3. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
The result is a `TypedefEngine` that can be used for repeated operations.
### Access time: `engine.validate(buffer)`
Validation is opt-in per operation. The consumer calls
`engine.validate(buffer)` when validation is desired. The jsonschema
validator is already compiled — `is_valid()` is a fast check against
the compiled validator.
High-throughput paths can skip validation. Security-sensitive paths
(parsing incoming frames from untrusted peers) can validate every frame.
The choice is the consumer's.
## Relationship to Read/Write
Validation and data access are independent operations on the same buffer.
The consumer can:
1. Validate a buffer to ensure it conforms to the schema.
2. Read fields from the buffer at computed offsets.
3. Both — validate first, then read (defense in depth).
The engine does not couple validation and access. A consumer that trusts
its data source can skip validation and go straight to read/write. A
consumer that parses untrusted input can validate first, then access.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Error handling and validation | [ADR-098](../../decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Purpose and scope | [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine |
## Open Questions
None specific to validation. The three typedef OQs (OQ-069, OQ-070,
OQ-071) are about layout, platform support, and schema construction —
not validation.
## References
- `docs/research/alknet-typedef/findings.md` §"Validation" — the POC's
custom keyword validators for all 17 kinds
- [ADR-098](../../decisions/098-error-handling-validation-strategy.md) —
error handling and validation strategy
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds that the
validators check
- [data-access.md](data-access.md) — read/write functions that operate
on the same buffers
@@ -0,0 +1,168 @@
# ADR-095: alknet-typedef — Purpose, Scope, and the jsonschema Engine
## Status
Accepted
## Context
Three threads in the codebase converge on the same pattern: a JSON Schema
describes the shape of binary data, and the binary data is the struct's
bytes at computed offsets.
1. **typedef.ts** (`/workspace/@alkdev/typebox/example/typedef/typedef.ts`,
619 lines) defines custom TypeBox schema kinds (`TFloat32`, `TStruct`,
`TUnion`, etc.) that carry binary layout semantics. These are registered
via `TypeRegistry.Set` with custom validators.
2. **russh-sftp** has 29 packet types, each a struct with typed fields
(`Read { id: u32, handle: String, offset: u64, len: u32 }`). The wire
format is `[length: u32][type: u8][payload]` where payload is the
struct's serde bytes. The `Packet` enum dispatches on the type byte —
a tagged union of structs. Under the typedef lens, each packet is a
`TStruct`; the `Packet` enum is a `TUnion` with a byte-offset
discriminator.
3. **metatensor** needs an offset map for mmap-friendly tensor access —
given a schema describing a model layout (ConvNet struct, tensor refs),
compute byte offsets for each field so the consumer can read tensor
data at known positions without parsing.
The common pattern: **a JSON Schema with `TypeDef:*` custom keywords
describes the shape of binary data; the binary data is the struct's bytes
at computed offsets.** The schema is the format definition; the engine is
generic.
Two prior attempts built their own jsonschema engines — the fatal flaw:
- **typebox-rs** (`/workspace/@alkimiadev/typebox-rs/`, ~8,400 lines):
a full 26-variant `SchemaKind` enum, a custom `Value` type with typed
arrays, and a 912-line hand-written validator.
- **alktype** (`/workspace/@alkimiadev/alktype/`, ~5,600 lines): a
handler-registry pattern that also implements its own validation for
each type.
The `jsonschema` crate (v0.46.5, Draft 2020-12) is already in the
workspace at `/workspace/jsonschema/`. It handles validation with custom
keyword support — the novel code is the offset computation, not the
validation.
The call-channels-unification research
(`docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine") identified the convergence and
bumped typedef up in the timeline. The POC
(`docs/research/alknet-typedef/findings.md`, 26 tests passing) validated
the approach: a ~1,900-line Rust crate that takes a JSON Schema with
`TypeDef:*` custom keywords and produces an offset map, read/write
functions, and validation — all driven by the schema.
## Decision
**alknet-typedef is a small Rust crate that takes a JSON Schema with
`TypeDef:*` custom keywords and produces three capabilities:**
1. **An offset map** — walks the schema, computes byte offsets for each
field based on type sizes, field order, and alignment.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that a
buffer's bytes match the schema's type constraints.
**The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing).** The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are ~10 lines each.
**The schema is the format.** A JSON Schema with `TypeDef:Float32`,
`TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
the layout spec. No separate format definition, no separate parser, no
separate validator. One schema, three uses: validate, compute offsets,
access data.
**The crate depends on `jsonschema` and `serde_json` (with
`preserve_order`).** No tokio, no platform deps. Compiles to
`wasm32-unknown-unknown` for browser use. The `jsonschema` crate's
`with_keyword("TypeDef:Float32", factory)` API is the integration point
for custom type kinds — each `TypeDef:*` kind maps to a custom keyword
validator in Rust. Same semantics as TypeBox's `TypeRegistry.Set`, same
JSON Schema wire format.
**The crate targets `std` for v1.** The WASM target has `std` available
via `wasm-bindgen`. If embedded use cases emerge, `no_std` + `alloc` can
be added as a feature gate later — the engine's core (offset computation,
read/write) is already allocation-free. See OQ-070.
## Consequences
### Positive
- **Eliminates ~14,000 lines of hand-rolled schema engines.** typebox-rs
and alktype are replaced by `jsonschema` + an offset map + ~50 lines of
custom keyword implementations. The codebase drops from "a port of
TypeBox" to "jsonschema + an offset map."
- **One schema, three uses.** The same JSON Schema validates, computes
offsets, and drives data access. No separate format definition, parser,
or validator per protocol.
- **Schema-driven, not code-driven.** Adding a new SFTP packet type is
adding a variant to the schema JSON, not writing a new Rust struct +
serde impl. The engine is generic; the schema is the configuration.
- **WASM-clean.** `serde_json` + `jsonschema` + byte manipulation. No
tokio, no platform deps. The same typedef schemas work in browser,
Node, Python (via `wasmtime-py`), Go (via `wazero`), and any other
WASM host.
- **TypeBox interop.** TypeBox modules render to standard JSON Schema
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
on the Rust side. Zero translation. The same schema validates in both
ecosystems.
- **Defense in depth.** Schema validation at the byte level — a malformed
binary payload fails validation before any consumer touches it. The
`jsonschema` crate's compiled validators are fast enough to run on
every incoming frame.
### Negative
- **New dependency on `jsonschema`.** The crate is already in the
workspace but not yet used by any alknet crate. This is the first
consumer.
- **`serde_json` with `preserve_order` is required.** Field order is
load-bearing for binary layouts. The `preserve_order` feature adds a
small compile-time cost.
- **Schema authoring is external.** Schemas are authored in TypeBox (JS)
or hand-written JSON. The typedef engine consumes schemas; it does not
generate them. A builder API is deferred (OQ-071).
## Scope Boundaries (What This Is Not)
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
is a *format* (8-byte header + JSON header + binary data) that uses the
typedef engine for its offset computation and tensor access.
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
`Value.Convert` — schema evolution — is out of scope for v1. The engine
should not do anything that explicitly blocks adding a Value system
later.
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
concern. The typedef engine consumes schemas; it does not generate them.
- **Not a schema builder.** The typedef engine does not provide a fluent
API for constructing schemas. Schemas are plain JSON.
- **Not a serialization framework.** The typedef engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use typedef.
## References
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
passing, two layout modes, TUnion dispatch, endianness)
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine" — the origin of this research
thread
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes decision
- [ADR-097](097-schema-annotations.md) — schema annotation shapes
- [ADR-098](098-error-handling-validation-strategy.md) — error handling
and validation strategy
@@ -0,0 +1,137 @@
# ADR-096: Two Layout Modes — Packed Sequential vs Aligned Static
## Status
Accepted
## Context
The POCs surfaced that protocols and mmap-friendly formats need different
layout strategies. POC 1 built an aligned `OffsetMap` with natural
alignment padding — correct for mmap-friendly formats (metatensor) but
wrong for protocol wire formats (SFTP, channels, TTY). POC 2 built a
`LayoutBuilder` and `SequentialReader` for packed sequential layouts —
correct for protocol wire formats but wrong for mmap-friendly formats.
This is the most important architectural finding from the POCs. The
engine must support both modes; a single layout strategy cannot serve
both use cases.
### Packed sequential layout (protocol wire formats)
Protocols pack fields sequentially with no alignment padding.
Variable-length fields shift all subsequent fields. Writing requires
knowing actual data sizes upfront; reading walks the buffer sequentially,
reading length prefixes to determine positions.
This is the layout used by SFTP (all strings and byte arrays are
length-prefixed inline), channels (`[channel_id: u32][size: u32][payload]`),
TTY (`[stream_type: u8][length: u32][payload]`), and most binary protocols.
### Aligned static layout (mmap-friendly formats)
Fields have fixed positions with natural alignment padding.
Variable-length fields get a 4-byte length prefix at a known offset; the
variable data is not included in the static layout. This enables
mmap-friendly random access — the consumer can read field N at a known
offset without parsing the fields before it.
This is the layout used by metatensor (blob tensor pattern: index struct
in one region, blob data in another) and safetensors (header + aligned
tensor data).
## Decision
**The typedef engine supports two layout modes, selected by the consumer
at engine construction time:**
### Mode 1: Packed sequential (`LayoutBuilder` / `SequentialReader`)
For protocol wire formats. Fields are packed with no alignment padding.
Variable-length fields shift all subsequent fields.
- **LayoutBuilder** — takes a schema and actual data sizes for
variable-length fields, computes byte positions for each field in a
packed layout. Used at write time when the consumer knows the data
sizes upfront.
- **SequentialReader** — walks a buffer field-by-field according to the
schema, reading length prefixes to determine variable-length data
positions. Used at read time when the consumer is parsing an incoming
frame.
The `LayoutBuilder` and `SequentialReader` are the primary interface for
protocol consumers (SFTP, binary call frames, TTY negotiation).
### Mode 2: Aligned static (`OffsetMap`)
For mmap-friendly formats. Fields have fixed positions with natural
alignment padding. Variable-length fields get a 4-byte length prefix at
a known offset; the variable data is not included in the static layout.
- **OffsetMap** — walks the schema once, computes fixed byte positions
for each field based on type sizes and alignment. The output is a flat
table of `(field_path, byte_range)` pairs. Used for both read and write
at known offsets.
The `OffsetMap` is the primary interface for mmap consumers (metatensor).
### Variable-length handling in each mode
**Packed sequential mode:** Variable-length fields are inline
length-prefixed by default (`[length: u32][data]`). The `LayoutBuilder`
takes the actual data size to compute the length prefix value and the
position of subsequent fields. The `SequentialReader` reads the length
prefix to determine the data extent and the position of the next field.
**Aligned static mode:** Variable-length fields get a 4-byte length
prefix at a known offset. The variable data lives outside the static
layout — either immediately after the fixed fields (inline
length-prefixing) or in a separate data region (offset indirection, the
metatensor blob tensor pattern). The `OffsetMap` records the position of
the length prefix (or the `{offset, length}` pair for offset-indirect
fields).
### Default for variable-length types
Inline length-prefixing (`[length: u32][data]`) is the default for all
variable-length types in both modes. This is the universal pattern used
by channels, SFTP, TTY, and most binary protocols. Offset indirection is
opt-in via the `encoding` annotation (see ADR-097).
## Consequences
### Positive
- **One engine, two modes.** The same schema can be used in either mode.
A schema describing an SFTP packet can be consumed by a `SequentialReader`
(for parsing incoming frames) and a `LayoutBuilder` (for constructing
outgoing frames). A schema describing a metatensor layout can be
consumed by an `OffsetMap` (for mmap access).
- **Correct for both use cases.** Packed sequential mode produces
byte-identical output to hand-written protocol serialization (validated
by POC 2's russh-sftp round-trip tests). Aligned static mode produces
correct offsets for mmap-friendly access (validated by POC 1's
alignment tests).
- **No mode confusion.** The consumer explicitly selects the mode at
engine construction time. A protocol consumer never accidentally gets
alignment padding; an mmap consumer never accidentally gets
variable-length field shifting.
### Negative
- **Two APIs to learn.** Consumers must choose between
`LayoutBuilder`/`SequentialReader` and `OffsetMap`. The choice is
determined by the use case (protocol vs mmap), not by the schema.
- **Variable-length fields in packed mode require size foreknowledge.**
The `LayoutBuilder` needs actual data sizes for variable-length fields
to compute correct positions for subsequent fields. This is inherent
to packed layouts — the consumer must know the data sizes before
writing.
## References
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-097](097-schema-annotations.md) — schema annotations including
the `encoding` field for variable-length types
@@ -0,0 +1,232 @@
# ADR-097: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
## Status
Accepted
## Context
The typedef engine needs concrete JSON shapes for schema-level
annotations that control binary layout behavior. The POCs validated the
semantics; this ADR pins the shapes.
Four annotation categories need concrete shapes:
1. **Endianness** — safetensors is little-endian, SFTP is big-endian.
The engine needs to know which to use.
2. **Alignment** — different backends have different alignment
requirements (wgpu: 256-byte, protocols: natural, mmap: page).
3. **Variable-length encoding** — inline length-prefixing vs offset
indirection for strings, byte arrays, and other variable-length types.
4. **TUnion discriminators** — byte-offset (protocol dispatch) vs
field-name (typedef.ts pattern).
## Decision
### 1. Endianness
**Schema-level annotation with a default of little-endian.**
```json
{
"TypeDef:Struct": true,
"endian": "big",
"properties": { ... }
}
```
- `"endian": "little"` (default) — read/write in little-endian byte order.
- `"endian": "big"` — read/write in big-endian byte order.
- The annotation applies to the entire schema and all nested types.
- Mixed endianness within one schema is not supported (pathological; no
known protocol requires it).
- The default is little-endian, matching safetensors, wgpu, and most
modern formats. SFTP consumers specify `"endian": "big"`.
### 2. Alignment
**Both struct-level and field-level, with field-level overriding
struct-level.**
```json
{
"TypeDef:Struct": true,
"align": 256,
"properties": {
"header": { "TypeDef:Struct": true, "properties": { ... } },
"weight": { "TypeDef:Float32": true, "align": 16 }
}
}
```
- Struct-level `"align"` sets the default alignment for all fields in
that struct. The struct's total size is rounded up to this alignment.
- Field-level `"align"` overrides the struct default for that specific
field.
- Default alignment (when no annotation is present): 1 for u8/bool, 2
for u16/i16, 4 for u32/i32/f32, 8 for u64/i64/f64, max field alignment
for structs.
- Alignment is only meaningful in aligned static mode (ADR-096). In
packed sequential mode, alignment annotations are ignored — fields are
packed with no padding.
### 3. Variable-length encoding
**Three strategies for variable-length types, selected by the `encoding`
annotation and the standard JSON Schema `maxLength` keyword.**
```json
// Strategy 1: Inline length-prefixing (default, shorthand)
{ "TypeDef:String": true }
// Strategy 1: Explicit inline length-prefixing
{ "TypeDef:String": { "encoding": "length-prefixed" } }
// Strategy 2: Fixed-size reservation (uses standard maxLength)
{ "TypeDef:String": true, "maxLength": 256 }
// Strategy 3: Offset indirection (opt-in)
{ "TypeDef:String": { "encoding": "offset-indirect" } }
```
**Strategy 1: Inline length-prefixing (default).** The field's fixed
portion is a 4-byte length prefix at a computed offset. The variable data
follows immediately after. In packed sequential mode, the length prefix
determines the position of subsequent fields. In aligned static mode, the
length prefix is at a known offset; the variable data is not included in
the static layout. This is the universal pattern used by channels, SFTP,
TTY, and most binary protocols.
**Strategy 2: Fixed-size reservation.** When a variable-length field
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
than `maxLength` is zero-padded; data longer than `maxLength` is a
validation error. This makes the field fixed-size from the layout
perspective — subsequent fields have known, unchanging offsets. This is
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
pattern for fields with known maximum sizes. In packed sequential mode,
`maxLength` is a validation constraint only — the engine still uses
inline length-prefixing (strategy 1).
**Strategy 3: Offset indirection.** The field is a struct
`{offset: u32, length: u32}` that points into a separate data region.
This is the metatensor blob tensor pattern — the index struct lives in
one region, the blob data lives in another. The consumer provides the
data region separately. Enables mmap-friendly random access to
variable-length data without parsing length prefixes and without
reserving worst-case space.
**Default strategy selection:**
- In packed sequential mode: always strategy 1 (inline length-prefixing).
`maxLength` is a validation constraint only.
- In aligned static mode: strategy 2 (fixed-size reservation) if
`maxLength` is declared; strategy 3 (offset indirection) if
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
length-prefixing) otherwise.
- `true` is a shorthand for the default (length-prefixed). This keeps
the common case concise and the override explicit.
- The `encoding` annotation and `maxLength` apply to all variable-length
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
`TypeDef:Record`, `TypeDef:Timestamp`.
### 4. TUnion discriminators
**Two discriminator kinds: byte-offset (protocol dispatch) and
field-name (typedef.ts pattern).**
#### Kind A: Byte-offset discriminator
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "TypeDef:Uint8"
},
"mapping": {
"1": { "$ref": "#/$defs/Init" },
"3": { "$ref": "#/$defs/Open" },
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" },
"101": { "$ref": "#/$defs/Status" }
}
}
```
- The discriminator is a fixed-size integer at a known byte offset.
- `"offset"` is the byte position of the discriminator within the union's
buffer.
- `"type"` is the `TypeDef:*` kind of the discriminator (typically
`TypeDef:Uint8` for protocol type bytes).
- The mapping keys are stringified integers (`"1"`, `"5"`, `"101"`).
The engine parses the key to match the discriminator value.
- The variant struct starts at `offset + discriminator_size`.
- This is the SFTP `Packet` enum pattern and the call protocol's event
type dispatch.
#### Kind B: Field-name discriminator
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "field",
"name": "type"
},
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
}
```
- The discriminator is a named field within the struct.
- `"name"` is the field name that holds the discriminator value.
- The mapping keys are string values matching the discriminator field's
value.
- The discriminator field is just another field in the struct — its
offset is computed like any other field.
- This is the typedef.ts `TUnion` pattern.
#### Mapping values
Mapping values may be either inline schemas or `$ref` pointers. `$ref`
is cleaner for large unions (29 SFTP variants) but requires a `$defs`
section. Inline schemas are simpler for small unions (5 call protocol
event types). Both work.
## Consequences
### Positive
- **Concrete, validated shapes.** All four annotation categories have
concrete JSON shapes that were validated by the POCs.
- **Sensible defaults.** Little-endian, natural alignment, inline
length-prefixing — the common case requires no annotations.
- **Explicit overrides.** Big-endian, custom alignment, offset
indirection — the uncommon case is explicit and self-documenting.
- **TUnion covers both protocol and typedef.ts patterns.** The
byte-offset discriminator handles SFTP type bytes and call protocol
event types. The field-name discriminator handles the typedef.ts string
pattern. No separate union type needed.
### Negative
- **Keyword value shape change.** `"TypeDef:String": true` (boolean) and
`"TypeDef:String": { "encoding": "length-prefixed" }` (object) are both
valid. The engine must handle both shapes. This is a minor parsing
concern — the POC already handles it.
- **Alignment annotations are mode-specific.** Alignment is only
meaningful in aligned static mode. In packed sequential mode, alignment
annotations are ignored. This is documented, not enforced — a consumer
that specifies alignment in packed mode gets no error, just no effect.
## References
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — the
annotation shape questions this ADR resolves
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes (alignment only meaningful in aligned static mode)
@@ -0,0 +1,157 @@
# ADR-098: Error Handling and Validation Strategy
## Status
Accepted
## Context
The typedef engine operates in three phases, each with distinct error
conditions:
1. **Schema parsing** — invalid JSON, missing required keywords, unknown
`TypeDef:*` kinds, malformed annotations.
2. **Offset computation** — field not found, type not supported for
offset computation, recursive schema depth exceeded.
3. **Read/write** — buffer too short, invalid UTF-8, value out of range
for the target type.
4. **Validation** — type constraint violations (range, UTF-8, field
presence, discriminator membership).
The engine also needs a clear strategy for *when* validation happens:
once at schema load time (build the validator) vs repeatedly at access
time (validate each buffer).
## Decision
### Error type: `TypedefError`
A single `TypedefError` enum with variants for each error category:
```rust
pub enum TypedefError {
/// Schema parsing errors.
Schema(String),
/// Offset computation errors.
Offset { field_path: String, reason: String },
/// Read/write errors.
Access { field_path: String, reason: String },
/// Validation errors (delegated to jsonschema).
Validation(ValidationError<'static>),
}
```
- `Schema` — for invalid JSON, missing required keywords, unknown
`TypeDef:*` kinds. The error message describes the problem.
- `Offset` — for field-not-found, unsupported type for offset
computation, etc. Carries the field path for debugging.
- `Access` — for buffer-too-short, invalid UTF-8, value out of range.
Carries the field path for debugging.
- `Validation` — wraps `jsonschema`'s `ValidationError`. The
`jsonschema` crate already provides rich error messages with schema
paths; the typedef engine does not re-wrap or re-interpret them.
The `Validation` variant uses `ValidationError<'static>` because the
validator is built once at schema load time and lives for the lifetime
of the `TypedefEngine`. The `'static` lifetime is correct — the validator
owns its schema reference.
### Validation timing: load-time build, access-time check
The jsonschema validator is built once at schema load time
(`validator_for(&schema)?`) and then called repeatedly
(`validator.is_valid(&instance)`). The typedef engine follows the same
pattern:
1. **Load time:** Parse the schema JSON, build the offset map (or
`LayoutBuilder`/`SequentialReader`), build the jsonschema validator.
This is the `TypedefEngine::compile(schema: &Value) -> Result<Self,
TypedefError>` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation — the consumer calls
`engine.validate(buffer)` when validation is desired.
The `TypedefEngine` struct is the compiled form of a schema:
```rust
pub struct TypedefEngine {
offset_map: OffsetMap, // or LayoutBuilder/SequentialReader
validator: jsonschema::Validator, // compiled once at load time
}
```
### Custom keyword validators
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
`jsonschema::options().with_keyword(...)`. The validators check:
- **Numeric types** (`TypeDef:Float32`, `TypeDef:Int8`, etc.): range
constraints (Int8: -128..127, Uint8: 0..255, etc.), finiteness for
floats.
- **`TypeDef:String`**: UTF-8 validity.
- **`TypeDef:Struct`**: field presence and types (delegated to
jsonschema's structural validation — the custom keyword only needs to
validate that the struct's fields match their declared `TypeDef:*`
kinds).
- **`TypeDef:Union`**: discriminator value membership in the mapping.
- **`TypeDef:Array`**: element type conformance.
- **`TypeDef:Boolean`**: value is `true` or `false`.
- **`TypeDef:Timestamp`**: ISO 8601 string format.
The `jsonschema` crate handles all the structural validation (object
properties, required fields, array items, enum values) — the custom
keywords only need to validate the leaf type constraints. Each custom
keyword implementation is ~10 lines.
### Read/write errors carry field paths
Read/write errors include the field path for debugging:
```rust
// Example: reading a u32 from a buffer that's too short
Err(TypedefError::Access {
field_path: "header.version".to_string(),
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
})
```
This makes debugging binary format issues tractable — the error tells
you exactly which field failed and why.
## Consequences
### Positive
- **Single error type.** Consumers handle one `TypedefError` enum, not
multiple error types from different engine phases.
- **Field-path-carrying errors.** Read/write errors include the field
path, making binary format debugging tractable.
- **Validation is opt-in.** The consumer decides when to validate.
High-throughput paths can skip validation; security-sensitive paths
can validate every frame.
- **jsonschema integration is clean.** The `ValidationError` is wrapped
as-is — no re-interpretation, no information loss.
- **Load-time build, access-time use.** The expensive work (schema
parsing, validator compilation, offset computation) happens once at
load time. Access-time operations are cheap (pointer casts, slice
operations, length-prefix reads).
### Negative
- **`ValidationError<'static>` lifetime.** The `'static` lifetime on the
`Validation` variant means the error cannot borrow from the buffer
being validated. This is correct (the validator owns its schema
reference) but may surprise readers who expect a shorter lifetime.
- **No error recovery.** The engine does not attempt to recover from
partial reads or writes. A buffer-too-short error on field N means
fields N+1.. are also unreadable. This is inherent to binary formats
— there is no "skip to next field" without a schema-driven parser.
## References
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — error
handling strategy question (OQ 8)
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes
- [ADR-097](097-schema-annotations.md) — schema annotations
+42
View File
@@ -211,6 +211,14 @@ Door type is separate from whether a decision is made. A two-way door is a decis
| [OQ-66](questions/066-alknet-register-wire-protocol.md) | `alknet/register` Wire Protocol | deferred(scope) | one | med | | [OQ-66](questions/066-alknet-register-wire-protocol.md) | `alknet/register` Wire Protocol | deferred(scope) | one | med |
| [OQ-67](questions/067-iroh-proxy-support.md) | iroh Proxy Support (Direct-Connection Peer Exposure) | resolved | one | med | | [OQ-67](questions/067-iroh-proxy-support.md) | iroh Proxy Support (Direct-Connection Peer Exposure) | resolved | one | med |
### alknet-typedef
| OQ | Title | Status | Door | Pri |
|----|-------|--------|------|-----|
| [OQ-069](questions/069-arrays-of-variable-length-element-structs.md) | Arrays of Variable-Length-Element Structs | deferred(scope) | two | low |
| [OQ-070](questions/070-no-std-alloc-support.md) | `no_std` + `alloc` Support | deferred(scope) | two | low |
| [OQ-071](questions/071-builder-api-for-schema-construction.md) | Builder API for Schema Construction | deferred(scope) | two | med |
## Deferred / Blocked ## Deferred / Blocked
The safe-exit visibility surface. These questions are parked because the The safe-exit visibility surface. These questions are parked because the
@@ -358,3 +366,37 @@ filtering the tables above.
[OQ-64](questions/064-client-side-tls-helper.md). [OQ-64](questions/064-client-side-tls-helper.md).
- **Full file**: [OQ-64](questions/064-client-side-tls-helper.md) - **Full file**: [OQ-64](questions/064-client-side-tls-helper.md)
### OQ-069: Arrays of Variable-Length-Element Structs
- **Blocked on**: A concrete consumer that needs arrays of structs with
variable-length fields, where the elements are interleaved
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
sequentially rather than use a fixed stride. The SFTP `Name` packet
has `Vec<File>` where `File` contains strings, but SFTP serializes
this as a sequence of length-prefixed strings (the serde `SeqAccess`
pattern), not as an array of fixed-stride structs. Arrays of
fixed-size structs are fully supported.
- **Priority**: low
- **Full file**: [OQ-069](questions/069-arrays-of-variable-length-element-structs.md)
### OQ-070: `no_std` + `alloc` Support
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
(e.g., a microcontroller running Rust without `std`). The WASM target
has `std` available via `wasm-bindgen`. The engine's core (offset
computation, read/write) is already allocation-free; the `jsonschema`
dependency is the only `alloc` consumer.
- **Priority**: low
- **Full file**: [OQ-070](questions/070-no-std-alloc-support.md)
### OQ-071: Builder API for Schema Construction
- **Blocked on**: A concrete need for programmatic schema construction
in Rust. The current consumers (SFTP, metatensor, binary call frames,
TTY negotiation) all have schemas that can be hand-written or
generated from TypeBox. A builder API would be a fluent Rust API that
produces the same JSON Schema structure — it would sit on top of the
engine, not inside it.
- **Priority**: medium
- **Full file**: [OQ-071](questions/071-builder-api-for-schema-construction.md)
+1
View File
@@ -103,6 +103,7 @@ alknet-vault (standalone — foundational to ACL: key derivation, identity)
│ │ ConnectionCredentials + RemoteIdentity per ADR-091) │ │ ConnectionCredentials + RemoteIdentity per ADR-091)
│ ├── alknet-tls TlsServerConfig + TlsClientConfig + FingerprintPinVerifier — shared TLS config across quinn + TCP+TLS + iroh (ADR-082/087; FingerprintPinVerifier per ADR-089 §5) │ ├── alknet-tls TlsServerConfig + TlsClientConfig + FingerprintPinVerifier — shared TLS config across quinn + TCP+TLS + iroh (ADR-082/087; FingerprintPinVerifier per ADR-089 §5)
│ ├── alknet-call CallAdapter on alknet/call, CallClient (spawn_dispatch primary; dial in AlknetClient per ADR-089), OperationRegistry, adapters (no TLS/transport deps) │ ├── alknet-call CallAdapter on alknet/call, CallClient (spawn_dispatch primary; dial in AlknetClient per ADR-089), OperationRegistry, adapters (no TLS/transport deps)
│ ├── alknet-typedef Binary struct engine — JSON Schema with TypeDef:* custom keywords → offset map + read/write + validation (ADR-095–098); depends on jsonschema + serde_json only; WASM-clean
│ ├── alknet-channels │ ├── alknet-channels
│ │ ├── alknet-channels-core pure multiplexer (wire format, demux/mux) — ADR-081 │ │ ├── alknet-channels-core pure multiplexer (wire format, demux/mux) — ADR-081
│ │ └── alknet-channels-call channel 0 pre-negotiation + lifecycle ops — ADR-081 │ │ └── alknet-channels-call channel 0 pre-negotiation + lifecycle ops — ADR-081
@@ -0,0 +1,20 @@
# OQ-069: Arrays of variable-length-element structs
- **Origin**: [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md),
[crates/typedef/data-access.md](crates/typedef/data-access.md);
`docs/research/alknet-typedef/findings.md` §"Problem 3: Nested structs
and arrays of structs"
- **Status**: deferred(scope)
- **Door type**: Two-way (additive — the engine can add lazy walking
logic without changing the existing fixed-stride array support)
- **Priority**: low
- **Impacts**: Blocks any protocol with interleaved variable-length struct arrays (e.g., a protocol where each array element has a string field and elements are packed as `[fixed_0][str_0][fixed_1][str_1]...`). Does NOT block SFTP `Name` packet handling — SFTP serializes this as a sequence of length-prefixed strings (the serde `SeqAccess` pattern), not as an array of fixed-stride structs. Does NOT block any current consumer.
- **Blocked on**: A concrete consumer that needs arrays of structs with
variable-length fields, where the elements are interleaved
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
sequentially rather than use a fixed stride.
- **Resolution**: Not yet decidable. The mechanism (lazy sequential
walking of array elements, reading each element's length prefixes to
find the next element's start) is understood but not needed by any
current consumer. Arrays of fixed-size structs are fully supported.
- **Cross-references**: ADR-096, [layout-engine.md](crates/typedef/layout-engine.md)
@@ -0,0 +1,16 @@
# OQ-070: `no_std` + `alloc` support
- **Origin**: [crates/typedef/overview.md](crates/typedef/overview.md);
`docs/research/alknet-typedef/findings.md` §"Open Questions" (OQ 6)
- **Status**: deferred(scope)
- **Door type**: Two-way (additive — can be added as a feature gate
without changing the existing `std` API)
- **Priority**: low
- **Impacts**: Blocks embedded/WASM-bare-metal deployment targets (microcontrollers, `no_std` environments). Does NOT block WASM-browser (has `std` via `wasm-bindgen`). Does NOT block any current deployment target.
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
(e.g., a microcontroller running Rust without `std`).
- **Resolution**: Not yet decidable. Target `std` for v1. If embedded
use cases emerge, `no_std` + `alloc` can be added as a feature gate
later. The engine's core is already allocation-free; the `jsonschema`
dependency is the only `alloc` consumer.
- **Cross-references**: ADR-095
@@ -0,0 +1,24 @@
# OQ-071: Builder API for schema construction
- **Origin**: [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md),
[crates/typedef/overview.md](crates/typedef/overview.md);
`docs/research/alknet-typedef/findings.md` (the builder API was noted
as the one detail not covered by the POCs)
- **Status**: deferred(scope)
- **Door type**: Two-way (additive — a builder API can be added without
changing the existing JSON-consumption path)
- **Priority**: medium
- **Impacts**: No current consumer. Schemas are authored in TypeBox (JS)
or hand-written JSON for v1. A Rust builder API would enable
programmatic schema construction in Rust without depending on a JS
toolchain, but no current consumer needs this.
- **Blocked on**: A concrete need for programmatic schema construction
in Rust. The current consumers (SFTP, metatensor, binary call frames,
TTY negotiation) all have schemas that can be hand-written or generated
from TypeBox.
- **Resolution**: Not yet decidable. The builder API is important but
not needed for the initial consumers. The engine's JSON-consumption
path is the primary interface for v1. A builder API would be a fluent
Rust API that produces the same JSON Schema structure — it would sit
on top of the engine, not inside it.
- **Cross-references**: ADR-095, [schema-layer.md](crates/typedef/schema-layer.md)