docs(architecture): add alknet-typedef crate specs, ADRs 095-098, and OQs 069-071
Add the alknet-typedef architecture specification — the binary struct engine that takes JSON Schema with TypeDef:* custom keywords and produces offset maps, read/write functions, and validation. Specs (docs/architecture/crates/typedef/): - overview.md: purpose, 'schema is the format' principle, consumers, scope - schema-layer.md: 17 TypeDef:* kinds, jsonschema integration, annotations - layout-engine.md: two layout modes, three variable-length strategies - data-access.md: read/write, TUnion dispatch, field paths, zero-copy - validation.md: custom keyword validators, TypedefError, TypedefEngine ADRs: - 095: Purpose, scope, and the jsonschema engine - 096: Two layout modes — packed sequential vs aligned static - 097: Schema annotations — endianness, alignment, encoding, TUnion - 098: Error handling and validation strategy OQs (deferred(scope)): - 069: Arrays of variable-length-element structs - 070: no_std + alloc support - 071: Builder API for schema construction Index updates: README doc table + ADR table, open-questions.md theme table + Deferred/Blocked section, overview.md crate graph. Grounded in the alknet-typedef POC (26 tests passing) and the call-channels-unification research. Reviewed by architecture-reviewer; all critical issues, warnings, and suggestions addressed.
This commit is contained in:
1 parent
bf0f827bf4
commit
85c5590001
16 files changed
+2230
-1
No files matched your search
@@ -296,6 +296,12 @@ adapter location map is now consistent: all HTTP-backed adapters
|
|||||||
| [crates/channels/channels-adapter.md](crates/channels/channels-adapter.md) | draft | `ChannelsAdapter`, `ChannelManager`, demux/mux contracts (REQ-CH-01..04), two-pump pattern (ADR-078) |
|
| [crates/channels/channels-adapter.md](crates/channels/channels-adapter.md) | draft | `ChannelsAdapter`, `ChannelManager`, demux/mux contracts (REQ-CH-01..04), two-pump pattern (ADR-078) |
|
||||||
| [crates/channels/channel-operations.md](crates/channels/channel-operations.md) | draft | `channel/open`/`close`/`control`/`resources/subscribe`, ACL flow, `direction` semantics, hub relay contract (ADR-079) |
|
| [crates/channels/channel-operations.md](crates/channels/channel-operations.md) | draft | `channel/open`/`close`/`control`/`resources/subscribe`, ACL flow, `direction` semantics, hub relay contract (ADR-079) |
|
||||||
| [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved |
|
| [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved |
|
||||||
|
| [crates/typedef/README.md](crates/typedef/README.md) | draft | alknet-typedef crate — binary struct engine; JSON Schema with `TypeDef:*` custom keywords → offset map + read/write + validation |
|
||||||
|
| [crates/typedef/overview.md](crates/typedef/overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
|
||||||
|
| [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 16 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
|
||||||
|
| [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md) | draft | Offset computation, two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
|
||||||
|
| [crates/typedef/data-access.md](crates/typedef/data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
|
||||||
|
| [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation |
|
||||||
|
|
||||||
## ADR Table
|
## ADR Table
|
||||||
|
|
||||||
@@ -395,10 +401,14 @@ adapter location map is now consistent: all HTTP-backed adapters
|
|||||||
| [092](decisions/092-bistream-as-the-handler-leaf.md) | `BiStream` as the Handler Leaf — Unify the Split-Pair `accept_bi` | Accepted (amends ADR-070's `accept_bi` return type; amends ADR-065's `from_stream`/`from_bidi` constructors; amends ADR-074's `ChannelBidiStreamSource::accept_bi` return type; `Connection::from_stream` removed; `from_bidi` is the only public stream constructor) |
|
| [092](decisions/092-bistream-as-the-handler-leaf.md) | `BiStream` as the Handler Leaf — Unify the Split-Pair `accept_bi` | Accepted (amends ADR-070's `accept_bi` return type; amends ADR-065's `from_stream`/`from_bidi` constructors; amends ADR-074's `ChannelBidiStreamSource::accept_bi` return type; `Connection::from_stream` removed; `from_bidi` is the only public stream constructor) |
|
||||||
| [093](decisions/093-channels-pure-channel-multiplexing.md) | alknet-channels — Pure Channel Multiplexing (8-Byte Header, No `stream_type`) | Accepted (amends ADR-071 — 8-byte header; ADR-074 — `into_sub_streams` removed; reverses ADR-077 — TTY always uses its 5-byte format; amends the channels-facing clauses of ADR-072/073/075/076/080/081) |
|
| [093](decisions/093-channels-pure-channel-multiplexing.md) | alknet-channels — Pure Channel Multiplexing (8-Byte Header, No `stream_type`) | Accepted (amends ADR-071 — 8-byte header; ADR-074 — `into_sub_streams` removed; reverses ADR-077 — TTY always uses its 5-byte format; amends the channels-facing clauses of ADR-072/073/075/076/080/081) |
|
||||||
| [094](decisions/094-per-identity-channel-cap.md) | Per-Identity Channel Cap as DoS Defense | Accepted (amends ADR-076 — per-connection `max_channels` reframed as a memory bound; 256 per `PeerId` enforced via `ChannelLifecyclePolicy` in `channels-call`; symmetric; spoke caps hub as direct caller) |
|
| [094](decisions/094-per-identity-channel-cap.md) | Per-Identity Channel Cap as DoS Defense | Accepted (amends ADR-076 — per-connection `max_channels` reframed as a memory bound; 256 per `PeerId` enforced via `ChannelLifecyclePolicy` in `channels-call`; symmetric; spoke caps hub as direct caller) |
|
||||||
|
| [095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | alknet-typedef — Purpose, Scope, and the jsonschema Engine | Accepted |
|
||||||
|
| [096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | Accepted |
|
||||||
|
| [097](decisions/097-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators | Accepted |
|
||||||
|
| [098](decisions/098-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | Accepted |
|
||||||
|
|
||||||
## Open Questions
|
## Open Questions
|
||||||
|
|
||||||
Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (68 OQs across 20 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](questions/) (`NNN-slug.md`, mirroring the ADR convention).
|
Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (71 OQs across 21 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](questions/) (`NNN-slug.md`, mirroring the ADR convention).
|
||||||
|
|
||||||
## Document Lifecycle
|
## Document Lifecycle
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,106 @@
|
|||||||
|
---
|
||||||
|
status: draft
|
||||||
|
last_updated: 2026-07-20
|
||||||
|
---
|
||||||
|
|
||||||
|
# alknet-typedef
|
||||||
|
|
||||||
|
The binary struct engine: a small Rust crate that takes a JSON Schema
|
||||||
|
with `TypeDef:*` custom keywords and produces an offset map, read/write
|
||||||
|
functions, and validation — all driven by the schema. The schema is the
|
||||||
|
format definition; the engine is generic.
|
||||||
|
|
||||||
|
## Documents
|
||||||
|
|
||||||
|
| Document | Status | Description |
|
||||||
|
|----------|--------|-------------|
|
||||||
|
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
|
||||||
|
| [schema-layer.md](schema-layer.md) | draft | The 17 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
|
||||||
|
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
|
||||||
|
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
|
||||||
|
| [validation.md](validation.md) | draft | Custom keyword validators for all 16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation, `TypedefEngine` |
|
||||||
|
|
||||||
|
## Applicable ADRs
|
||||||
|
|
||||||
|
| ADR | Title | Relevance |
|
||||||
|
|-----|-------|-----------|
|
||||||
|
| [095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
|
||||||
|
| [096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` |
|
||||||
|
| [097](../../decisions/097-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations |
|
||||||
|
| [098](../../decisions/098-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors |
|
||||||
|
|
||||||
|
## Relevant Open Questions
|
||||||
|
|
||||||
|
| OQ | Title | Status | Relevance |
|
||||||
|
|----|-------|--------|-----------|
|
||||||
|
| OQ-069 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it |
|
||||||
|
| OQ-070 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case |
|
||||||
|
| OQ-071 | Builder API for schema construction | deferred(scope) | Schemas are authored in TypeBox or hand-written JSON for v1; blocked on a concrete need |
|
||||||
|
|
||||||
|
## Key Design Principles
|
||||||
|
|
||||||
|
1. **The schema is the format.** A JSON Schema with `TypeDef:*` custom
|
||||||
|
keywords is both the validation spec and the layout spec. No separate
|
||||||
|
format definition, no separate parser, no separate validator. One
|
||||||
|
schema, three uses: validate, compute offsets, access data. See
|
||||||
|
[overview.md](overview.md) and [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||||
|
|
||||||
|
2. **jsonschema is the validation engine, not a custom engine.** The
|
||||||
|
`jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with
|
||||||
|
custom keyword support. The novel code is the offset computation, not
|
||||||
|
the validation. This eliminates ~14,000 lines of hand-rolled schema
|
||||||
|
engines (typebox-rs, alktype). See [schema-layer.md](schema-layer.md)
|
||||||
|
and [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||||
|
|
||||||
|
3. **Two layout modes for two use cases.** Packed sequential
|
||||||
|
(`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP,
|
||||||
|
channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats
|
||||||
|
(metatensor). The consumer selects the mode; the schema is the same.
|
||||||
|
See [layout-engine.md](layout-engine.md) and
|
||||||
|
[ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md).
|
||||||
|
|
||||||
|
4. **Variable-length types default to inline length-prefixing.**
|
||||||
|
`[length: u32][data]` is the universal pattern used by channels, SFTP,
|
||||||
|
TTY, and most binary protocols. Offset indirection (the metatensor
|
||||||
|
blob tensor pattern) is opt-in via the `encoding` annotation. See
|
||||||
|
[layout-engine.md](layout-engine.md) and
|
||||||
|
[ADR-097](../../decisions/097-schema-annotations.md).
|
||||||
|
|
||||||
|
5. **TUnion supports both byte-offset and field-name discriminators.**
|
||||||
|
Byte-offset for protocol dispatch (SFTP type bytes, call protocol
|
||||||
|
event types). Field-name for the typedef.ts string pattern. See
|
||||||
|
[data-access.md](data-access.md) and
|
||||||
|
[ADR-097](../../decisions/097-schema-annotations.md).
|
||||||
|
|
||||||
|
6. **Endianness is per-schema, default little-endian.** The engine reads
|
||||||
|
the `"endian"` annotation and byte-swaps accordingly. SFTP consumers
|
||||||
|
specify `"endian": "big"`. See [layout-engine.md](layout-engine.md)
|
||||||
|
and [ADR-097](../../decisions/097-schema-annotations.md).
|
||||||
|
|
||||||
|
7. **Validation is opt-in, built once at load time.** The jsonschema
|
||||||
|
validator is compiled once at schema load time. Access-time validation
|
||||||
|
is a fast `is_valid()` check. High-throughput paths can skip
|
||||||
|
validation; security-sensitive paths can validate every frame. See
|
||||||
|
[validation.md](validation.md) and
|
||||||
|
[ADR-098](../../decisions/098-error-handling-validation-strategy.md).
|
||||||
|
|
||||||
|
8. **Not a serialization framework.** The typedef engine is not a
|
||||||
|
general-purpose serde replacement. It operates on raw byte buffers at
|
||||||
|
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||||
|
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||||
|
with a known schema, use typedef. See [overview.md](overview.md) and
|
||||||
|
[ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
|
||||||
|
passing, two layout modes, TUnion dispatch, endianness)
|
||||||
|
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||||
|
JSON Schema as the binary struct engine" — the origin of this research
|
||||||
|
thread
|
||||||
|
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||||
|
schema kinds (619 lines)
|
||||||
|
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||||
|
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
|
||||||
|
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
|
||||||
|
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
|
||||||
@@ -0,0 +1,224 @@
|
|||||||
|
---
|
||||||
|
status: draft
|
||||||
|
last_updated: 2026-07-20
|
||||||
|
---
|
||||||
|
|
||||||
|
# alknet-typedef — Data Access
|
||||||
|
|
||||||
|
The data access layer: read/write functions, TUnion dispatch, field paths,
|
||||||
|
zero-copy access for fixed-size types, and length-prefix reading for
|
||||||
|
variable-length types. This is the consumer-facing API — given a compiled
|
||||||
|
`TypedefEngine` and a byte buffer, read and write fields at
|
||||||
|
schema-computed offsets.
|
||||||
|
|
||||||
|
## Read/Write Model
|
||||||
|
|
||||||
|
The typedef engine operates on raw byte buffers (`&[u8]` for reading,
|
||||||
|
`&mut [u8]` for writing). There is no intermediate `Value` tree, no
|
||||||
|
reflection, no dynamic dispatch per field. The engine uses the offset map
|
||||||
|
(or `LayoutBuilder`/`SequentialReader`) to locate fields, then performs
|
||||||
|
typed access at the computed positions.
|
||||||
|
|
||||||
|
### Fixed-size types
|
||||||
|
|
||||||
|
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, etc.) are accessed via
|
||||||
|
zero-copy pointer casts:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
// Read a u32 at a known offset
|
||||||
|
fn read_u32(buffer: &[u8], offset: usize, endian: Endian) -> u32 {
|
||||||
|
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
|
||||||
|
match endian {
|
||||||
|
Endian::Little => u32::from_le_bytes(bytes),
|
||||||
|
Endian::Big => u32::from_be_bytes(bytes),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// Write a u32 at a known offset
|
||||||
|
fn write_u32(buffer: &mut [u8], offset: usize, value: u32, endian: Endian) {
|
||||||
|
let bytes = match endian {
|
||||||
|
Endian::Little => value.to_le_bytes(),
|
||||||
|
Endian::Big => value.to_be_bytes(),
|
||||||
|
};
|
||||||
|
buffer[offset..offset+4].copy_from_slice(&bytes);
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The engine applies endianness at access time based on the schema's
|
||||||
|
`"endian"` annotation (ADR-097). The offset computation is
|
||||||
|
endian-agnostic.
|
||||||
|
|
||||||
|
### Variable-length types (inline length-prefixing)
|
||||||
|
|
||||||
|
For variable-length types with inline length-prefixing (the default):
|
||||||
|
|
||||||
|
```rust
|
||||||
|
// Read a length-prefixed string
|
||||||
|
fn read_string<'a>(buffer: &'a [u8], offset: usize) -> &'a str {
|
||||||
|
let len = u32::from_le_bytes(buffer[offset..offset+4].try_into().unwrap()) as usize;
|
||||||
|
std::str::from_utf8(&buffer[offset+4..offset+4+len]).unwrap()
|
||||||
|
}
|
||||||
|
|
||||||
|
// Write a length-prefixed string
|
||||||
|
fn write_string(buffer: &mut [u8], offset: usize, value: &str) {
|
||||||
|
let data = value.as_bytes();
|
||||||
|
buffer[offset..offset+4].copy_from_slice(&(data.len() as u32).to_le_bytes());
|
||||||
|
buffer[offset+4..offset+4+data.len()].copy_from_slice(data);
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The engine reads the 4-byte length prefix at the field's offset, then
|
||||||
|
slices the data that follows. For writing, the engine writes the length
|
||||||
|
prefix + data.
|
||||||
|
|
||||||
|
In packed sequential mode, the `SequentialReader` uses the length prefix
|
||||||
|
to determine the position of the next field. In aligned static mode, the
|
||||||
|
`OffsetMap` records the position of the length prefix; the variable data
|
||||||
|
is accessed separately.
|
||||||
|
|
||||||
|
### Variable-length types (offset indirection)
|
||||||
|
|
||||||
|
For variable-length types with offset indirection (opt-in):
|
||||||
|
|
||||||
|
```rust
|
||||||
|
// Read an offset-indirect string
|
||||||
|
fn read_string_indirect(data_region: &[u8], offset: usize) -> &str {
|
||||||
|
let ptr_offset = u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()) as usize;
|
||||||
|
let ptr_length = u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()) as usize;
|
||||||
|
std::str::from_utf8(&data_region[ptr_offset..ptr_offset+ptr_length]).unwrap()
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The field is a struct `{offset: u32, length: u32}` at a known position
|
||||||
|
in the `OffsetMap`. The consumer provides the data region separately; the
|
||||||
|
engine reads the offset and length, then slices the data region.
|
||||||
|
|
||||||
|
## TUnion Dispatch
|
||||||
|
|
||||||
|
TUnion dispatch reads the discriminator value, looks up the variant
|
||||||
|
schema, and then reads the variant's fields. The dispatch mechanism
|
||||||
|
differs by discriminator kind (ADR-097).
|
||||||
|
|
||||||
|
### Byte-offset discriminator
|
||||||
|
|
||||||
|
```rust
|
||||||
|
fn read_union(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
|
||||||
|
let disc = &schema["discriminator"];
|
||||||
|
let offset = disc["offset"].as_u64().unwrap() as usize;
|
||||||
|
let disc_type = disc["type"].as_str().unwrap(); // e.g., "TypeDef:Uint8"
|
||||||
|
|
||||||
|
// Read the discriminator value
|
||||||
|
let disc_value: u8 = read_u8(buffer, offset);
|
||||||
|
let key = disc_value.to_string(); // "5", "6", "101"
|
||||||
|
|
||||||
|
// Look up the variant schema
|
||||||
|
let mapping = &schema["mapping"];
|
||||||
|
let variant_schema = &mapping[&key];
|
||||||
|
|
||||||
|
// Read the variant struct starting at offset + discriminator_size
|
||||||
|
let variant_offset = offset + 1; // discriminator_size for Uint8
|
||||||
|
read_struct(buffer, variant_offset, variant_schema)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The discriminator is a fixed-size integer at a known byte offset. The
|
||||||
|
mapping keys are stringified integers. The variant struct starts at
|
||||||
|
`offset + discriminator_size`.
|
||||||
|
|
||||||
|
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
|
||||||
|
1..N are the variant struct. The call protocol's 5 event types
|
||||||
|
(`call.requested` → 0x01, etc.) use the same pattern.
|
||||||
|
|
||||||
|
### Field-name discriminator
|
||||||
|
|
||||||
|
```rust
|
||||||
|
fn read_union_field(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
|
||||||
|
let disc = &schema["discriminator"];
|
||||||
|
let field_name = disc["name"].as_str().unwrap(); // e.g., "type"
|
||||||
|
|
||||||
|
// Read the discriminator field like any other field
|
||||||
|
let disc_value = read_field(buffer, field_name, schema)?;
|
||||||
|
|
||||||
|
// Look up the variant schema
|
||||||
|
let mapping = &schema["mapping"];
|
||||||
|
let variant_schema = &mapping[disc_value.as_str().unwrap()];
|
||||||
|
|
||||||
|
// Read the variant struct
|
||||||
|
read_struct(buffer, variant_offset, variant_schema)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The discriminator is a named field within the struct. Its offset is
|
||||||
|
computed like any other field. The mapping keys are string values.
|
||||||
|
|
||||||
|
## Field Paths
|
||||||
|
|
||||||
|
Fields are addressed by dotted paths: `"header.version"`, `"payload.data"`.
|
||||||
|
The `OffsetMap` stores fully-qualified paths. The read/write functions
|
||||||
|
accept a field path and look up the byte range:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> {
|
||||||
|
let range = self.offset_map.get(field_path)
|
||||||
|
.ok_or_else(|| TypedefError::Offset {
|
||||||
|
field_path: field_path.to_string(),
|
||||||
|
reason: "field not found in offset map".to_string(),
|
||||||
|
})?;
|
||||||
|
if buffer.len() < range.end {
|
||||||
|
return Err(TypedefError::Access {
|
||||||
|
field_path: field_path.to_string(),
|
||||||
|
reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()),
|
||||||
|
});
|
||||||
|
}
|
||||||
|
Ok(read_f32_raw(buffer, range.start, self.endian))
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Nested structs produce nested field paths. The offset computation
|
||||||
|
propagates the field path prefix during recursion, so the `OffsetMap`
|
||||||
|
contains entries like `"header.version"` and `"header.magic"`.
|
||||||
|
|
||||||
|
## Zero-Copy Access
|
||||||
|
|
||||||
|
For fixed-size types, the engine provides zero-copy access — the consumer
|
||||||
|
gets a reference to the bytes in the buffer, not a copy. This is
|
||||||
|
important for performance-sensitive paths (metatensor tensor access,
|
||||||
|
high-throughput protocol parsing).
|
||||||
|
|
||||||
|
For variable-length types with inline length-prefixing, the engine
|
||||||
|
returns a slice of the buffer — the string or byte array data is not
|
||||||
|
copied. The consumer gets a `&str` or `&[u8]` that borrows from the
|
||||||
|
input buffer.
|
||||||
|
|
||||||
|
For offset-indirect types, the consumer provides the data region; the
|
||||||
|
engine returns a slice of that region.
|
||||||
|
|
||||||
|
## Error Handling
|
||||||
|
|
||||||
|
Read/write errors carry the field path for debugging. See
|
||||||
|
[ADR-098](../../decisions/098-error-handling-validation-strategy.md) and
|
||||||
|
[validation.md](validation.md) for the full error model.
|
||||||
|
|
||||||
|
## Design Decisions
|
||||||
|
|
||||||
|
| Decision | ADR | Summary |
|
||||||
|
|----------|-----|---------|
|
||||||
|
| Two layout modes | [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Determines whether offsets are fixed (OffsetMap) or sequential (SequentialReader) |
|
||||||
|
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Endianness, encoding, and TUnion discriminator shapes that control data access |
|
||||||
|
| Error handling | [ADR-098](../../decisions/098-error-handling-validation-strategy.md) | Field-path-carrying errors for read/write operations |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
See [open-questions.md](../../open-questions.md) for full details.
|
||||||
|
|
||||||
|
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
|
||||||
|
— affects the sequential walking logic for array access.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
|
||||||
|
(read/write round-trip) and POC 2 (SFTP byte-identical round-trip)
|
||||||
|
- [layout-engine.md](layout-engine.md) — offset computation that produces
|
||||||
|
the positions this layer reads/writes at
|
||||||
|
- [validation.md](validation.md) — validation that runs on the same
|
||||||
|
buffers
|
||||||
@@ -0,0 +1,254 @@
|
|||||||
|
---
|
||||||
|
status: draft
|
||||||
|
last_updated: 2026-07-20
|
||||||
|
---
|
||||||
|
|
||||||
|
# alknet-typedef — Layout Engine
|
||||||
|
|
||||||
|
The layout engine: offset computation, the two layout modes (packed
|
||||||
|
sequential vs aligned static), alignment, endianness, and variable-length
|
||||||
|
field handling. This is the novel code — the recursive walk of the schema
|
||||||
|
JSON that computes byte positions for each field.
|
||||||
|
|
||||||
|
## The Two Layout Modes
|
||||||
|
|
||||||
|
The POCs surfaced that protocols and mmap-friendly formats need different
|
||||||
|
layout strategies. This is the most important architectural finding —
|
||||||
|
decided in [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md).
|
||||||
|
|
||||||
|
### Mode 1: Packed sequential (protocol wire formats)
|
||||||
|
|
||||||
|
Fields are packed with no alignment padding. Variable-length fields shift
|
||||||
|
all subsequent fields. Used by SFTP, channels, TTY, and most binary
|
||||||
|
protocols.
|
||||||
|
|
||||||
|
**Components:**
|
||||||
|
|
||||||
|
- **`LayoutBuilder`** — takes a schema and actual data sizes for
|
||||||
|
variable-length fields, computes byte positions for each field in a
|
||||||
|
packed layout. Used at write time when the consumer knows the data
|
||||||
|
sizes upfront.
|
||||||
|
- **`SequentialReader`** — walks a buffer field-by-field according to the
|
||||||
|
schema, reading length prefixes to determine variable-length data
|
||||||
|
positions. Used at read time when the consumer is parsing an incoming
|
||||||
|
frame.
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
|
||||||
|
For a struct with fields `[u8, u32, string]`:
|
||||||
|
|
||||||
|
```
|
||||||
|
LayoutBuilder (write):
|
||||||
|
field[0] u8: offset 0, size 1
|
||||||
|
field[1] u32: offset 1, size 4
|
||||||
|
field[2] string: offset 5, size 4 (length prefix) + data_len
|
||||||
|
total: 9 + data_len
|
||||||
|
|
||||||
|
SequentialReader (read):
|
||||||
|
read u8 at offset 0
|
||||||
|
read u32 at offset 1
|
||||||
|
read u32 length prefix at offset 5 → data_len
|
||||||
|
read string data at offset 9, length data_len
|
||||||
|
next field at offset 9 + data_len
|
||||||
|
```
|
||||||
|
|
||||||
|
There is no alignment padding. The `u32` at offset 1 is unaligned — this
|
||||||
|
is correct for protocol wire formats, which pack fields tightly.
|
||||||
|
|
||||||
|
**Variable-length fields in packed mode:**
|
||||||
|
|
||||||
|
The `LayoutBuilder` takes actual data sizes for variable-length fields
|
||||||
|
to compute correct positions for subsequent fields. The consumer must
|
||||||
|
know the data sizes before writing — this is inherent to packed layouts.
|
||||||
|
|
||||||
|
The `SequentialReader` reads each field's length prefix to determine the
|
||||||
|
data extent and the position of the next field. The reader walks the
|
||||||
|
buffer sequentially; it cannot jump to field N without reading fields
|
||||||
|
0..N-1 first.
|
||||||
|
|
||||||
|
### Mode 2: Aligned static (mmap-friendly formats)
|
||||||
|
|
||||||
|
Fields have fixed positions with natural alignment padding.
|
||||||
|
Variable-length fields get a 4-byte length prefix at a known offset; the
|
||||||
|
variable data is not included in the static layout. Used by metatensor
|
||||||
|
and safetensors.
|
||||||
|
|
||||||
|
**Component:**
|
||||||
|
|
||||||
|
- **`OffsetMap`** — walks the schema once, computes fixed byte positions
|
||||||
|
for each field based on type sizes and alignment. The output is a flat
|
||||||
|
table of `(field_path, byte_range)` pairs. Used for both read and write
|
||||||
|
at known offsets.
|
||||||
|
|
||||||
|
**How it works:**
|
||||||
|
|
||||||
|
For a struct with fields `[u8, u32, f32]` and natural alignment:
|
||||||
|
|
||||||
|
```
|
||||||
|
OffsetMap:
|
||||||
|
field[0] u8: offset 0, size 1
|
||||||
|
field[1] u32: offset 4, size 4 (3 bytes padding after u8)
|
||||||
|
field[2] f32: offset 8, size 4
|
||||||
|
total: 12 (struct aligned to 4)
|
||||||
|
```
|
||||||
|
|
||||||
|
The `u32` is aligned to offset 4 (its natural alignment). The consumer
|
||||||
|
can read `field[1]` at offset 4 without reading `field[0]` first — random
|
||||||
|
access by field path.
|
||||||
|
|
||||||
|
**Variable-length fields in aligned mode:**
|
||||||
|
|
||||||
|
Variable-length fields get a 4-byte length prefix at a known offset. The
|
||||||
|
variable data lives outside the static layout — either immediately after
|
||||||
|
the fixed fields (inline length-prefixing) or in a separate data region
|
||||||
|
(offset indirection). The `OffsetMap` records the position of the length
|
||||||
|
prefix (or the `{offset, length}` pair for offset-indirect fields).
|
||||||
|
|
||||||
|
For inline length-prefixing, the variable data follows the fixed fields
|
||||||
|
but is not included in the `OffsetMap`'s field ranges. The consumer reads
|
||||||
|
the length prefix from the `OffsetMap`'s known offset, then slices the
|
||||||
|
data region.
|
||||||
|
|
||||||
|
For offset indirection, the field is a struct `{offset: u32, length: u32}`
|
||||||
|
at a known position in the `OffsetMap`. The consumer reads the offset and
|
||||||
|
length, then slices the separate data region.
|
||||||
|
|
||||||
|
## Offset Computation Algorithm
|
||||||
|
|
||||||
|
The offset computation is a recursive walk of the schema JSON. The
|
||||||
|
algorithm is the same for both modes; the difference is whether alignment
|
||||||
|
padding is inserted between fields.
|
||||||
|
|
||||||
|
### Fixed-size types
|
||||||
|
|
||||||
|
For each fixed-size type, the algorithm:
|
||||||
|
1. Determines the type's byte size from the `TypeDef:*` kind.
|
||||||
|
2. In aligned mode: inserts padding to satisfy the type's alignment
|
||||||
|
(or the field's `align` annotation, or the struct's `align` default).
|
||||||
|
3. Records the field's `(start, end)` range.
|
||||||
|
4. Advances the current offset by the type's size.
|
||||||
|
|
||||||
|
### Composite types
|
||||||
|
|
||||||
|
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
|
||||||
|
are computed relative to the struct's start offset. The struct's total
|
||||||
|
size is the sum of its fields' sizes (plus alignment padding in aligned
|
||||||
|
mode). The struct itself may have an `align` annotation that rounds up
|
||||||
|
its total size.
|
||||||
|
|
||||||
|
**`TUnion`:** The discriminator occupies `offset..offset + discriminator_size`
|
||||||
|
bytes. For byte-offset discriminators, the variant struct starts at
|
||||||
|
`offset + discriminator_size`. For field-name discriminators, the
|
||||||
|
discriminator is just another field — its offset is computed like any
|
||||||
|
other field, and the variant struct follows at the end of the
|
||||||
|
discriminator field.
|
||||||
|
|
||||||
|
In aligned static mode, the union's total size is `discriminator_size +
|
||||||
|
max(variant_sizes)`, where variant sizes are computed from the schema
|
||||||
|
(variable-length data lives outside the static layout).
|
||||||
|
|
||||||
|
In packed sequential mode, variant sizes depend on the actual sizes of
|
||||||
|
variable-length fields within each variant, which aren't known at schema
|
||||||
|
time. The `LayoutBuilder` takes the actual variant discriminator value
|
||||||
|
and data sizes at write time, computes the size of the selected variant,
|
||||||
|
and uses that for the union's total size. The `SequentialReader` reads
|
||||||
|
the discriminator first, looks up the variant schema, then reads the
|
||||||
|
variant struct sequentially — it doesn't need to know the union's total
|
||||||
|
size upfront.
|
||||||
|
|
||||||
|
**`TArray` of fixed-size elements:** Element stride = element size (plus
|
||||||
|
alignment padding in aligned mode). Element `i` starts at
|
||||||
|
`array_offset + i × stride`. The array's total size is `count × stride`.
|
||||||
|
|
||||||
|
**`TArray` of variable-length-element structs:** Deferred for v1
|
||||||
|
(OQ-069).
|
||||||
|
|
||||||
|
### Variable-length types
|
||||||
|
|
||||||
|
The typedef engine supports three strategies for variable-length types
|
||||||
|
(see [schema-layer.md](schema-layer.md) §Variable-length types and
|
||||||
|
[ADR-097](../../decisions/097-schema-annotations.md) §3 for the full
|
||||||
|
annotation shapes).
|
||||||
|
|
||||||
|
**Strategy 1: Inline length-prefixing (default).**
|
||||||
|
1. Records the position of the 4-byte length prefix.
|
||||||
|
2. In aligned mode: the length prefix is aligned; the variable data is
|
||||||
|
not included in the static layout.
|
||||||
|
3. In packed mode: the `LayoutBuilder` takes the actual data size to
|
||||||
|
compute the length prefix value and the position of subsequent fields.
|
||||||
|
The `SequentialReader` reads the length prefix to determine the data
|
||||||
|
extent and the position of the next field.
|
||||||
|
|
||||||
|
**Strategy 2: Fixed-size reservation (`maxLength`).**
|
||||||
|
1. In aligned static mode: reserves `maxLength` bytes at a fixed offset.
|
||||||
|
Data shorter than `maxLength` is zero-padded. Subsequent fields have
|
||||||
|
known, unchanging offsets — the field is fixed-size from the layout
|
||||||
|
perspective. This is the database `VARCHAR(N)` pattern.
|
||||||
|
2. In packed sequential mode: `maxLength` is a validation constraint
|
||||||
|
only. The engine uses strategy 1 (inline length-prefixing) because
|
||||||
|
protocols don't benefit from fixed-size reservation.
|
||||||
|
|
||||||
|
**Strategy 3: Offset indirection (`"encoding": "offset-indirect"`).**
|
||||||
|
1. The field is a struct `{offset: u32, length: u32}`.
|
||||||
|
2. The `OffsetMap` records the position of this struct.
|
||||||
|
3. The consumer provides the data region separately. This is the
|
||||||
|
metatensor blob tensor pattern — the index struct lives in one region,
|
||||||
|
the blob data lives in another.
|
||||||
|
|
||||||
|
### Nested structs and field paths
|
||||||
|
|
||||||
|
Nested structs produce dotted field paths: `header.version`,
|
||||||
|
`header.magic`. The offset computation propagates the field path prefix
|
||||||
|
during recursion. The `OffsetMap` stores fully-qualified paths.
|
||||||
|
|
||||||
|
### Endianness
|
||||||
|
|
||||||
|
Endianness is per-schema (ADR-097). The offset computation is
|
||||||
|
endian-agnostic — it computes byte positions, not byte values. The
|
||||||
|
read/write functions apply endianness when converting between bytes and
|
||||||
|
typed values. The engine reads the `"endian"` annotation from the schema
|
||||||
|
and byte-swaps accordingly.
|
||||||
|
|
||||||
|
## Mode Selection
|
||||||
|
|
||||||
|
The consumer selects the mode at engine construction time. The choice is
|
||||||
|
determined by the use case, not by the schema:
|
||||||
|
|
||||||
|
- **Protocol consumer** (SFTP, binary call frames, TTY negotiation):
|
||||||
|
uses `LayoutBuilder` for writing and `SequentialReader` for reading.
|
||||||
|
- **mmap consumer** (metatensor): uses `OffsetMap` for both reading and
|
||||||
|
writing.
|
||||||
|
|
||||||
|
The same schema can be used in either mode. A schema describing an SFTP
|
||||||
|
packet can be consumed by a `SequentialReader` (for parsing incoming
|
||||||
|
frames) and a `LayoutBuilder` (for constructing outgoing frames). A schema
|
||||||
|
describing a metatensor layout can be consumed by an `OffsetMap` (for
|
||||||
|
mmap access).
|
||||||
|
|
||||||
|
## Design Decisions
|
||||||
|
|
||||||
|
| Decision | ADR | Summary |
|
||||||
|
|----------|-----|---------|
|
||||||
|
| Two layout modes | [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential for protocols; aligned static for mmap formats |
|
||||||
|
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Endianness, alignment, encoding annotations that control layout behavior |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
See [open-questions.md](../../open-questions.md) for full details.
|
||||||
|
|
||||||
|
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
|
||||||
|
— requires lazy walking logic; blocked on a concrete consumer that
|
||||||
|
needs it.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
|
||||||
|
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
|
||||||
|
- [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) —
|
||||||
|
the two layout modes decision
|
||||||
|
- [ADR-097](../../decisions/097-schema-annotations.md) — schema
|
||||||
|
annotations
|
||||||
|
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds and their
|
||||||
|
byte sizes
|
||||||
|
- [data-access.md](data-access.md) — read/write functions that use the
|
||||||
|
computed offsets
|
||||||
@@ -0,0 +1,198 @@
|
|||||||
|
---
|
||||||
|
status: draft
|
||||||
|
last_updated: 2026-07-20
|
||||||
|
---
|
||||||
|
|
||||||
|
# alknet-typedef — Overview
|
||||||
|
|
||||||
|
The binary struct engine: a small Rust crate that takes a JSON Schema
|
||||||
|
with `TypeDef:*` custom keywords and produces an offset map, read/write
|
||||||
|
functions, and validation — all driven by the schema. The schema is the
|
||||||
|
format definition; the engine is generic.
|
||||||
|
|
||||||
|
This document covers the crate's purpose, the "schema is the format"
|
||||||
|
principle, its dependency edges, consumers, and scope boundaries.
|
||||||
|
Component details are in the sibling documents.
|
||||||
|
|
||||||
|
## What
|
||||||
|
|
||||||
|
`alknet-typedef` is a library crate that consumes JSON Schemas annotated
|
||||||
|
with `TypeDef:*` custom keywords (the same kinds defined in TypeBox's
|
||||||
|
`typedef.ts`) and produces three capabilities:
|
||||||
|
|
||||||
|
1. **An offset map** — walks the schema, computes byte offsets for each
|
||||||
|
field based on type sizes, field order, and alignment.
|
||||||
|
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
|
||||||
|
read the field's bytes at its offset (zero-copy for fixed-size types).
|
||||||
|
Given a `&mut [u8]` buffer, write a value at its offset.
|
||||||
|
3. **Validation** — via `jsonschema` custom keywords, validates that a
|
||||||
|
buffer's bytes match the schema's type constraints.
|
||||||
|
|
||||||
|
The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||||
|
`serde_json` (schema parsing). The novel code is the offset computation
|
||||||
|
— a recursive walk of the schema JSON that computes byte positions for
|
||||||
|
each field. The custom keyword implementations are ~10 lines each.
|
||||||
|
|
||||||
|
The crate is ~1,900 lines (POC verified, 26 tests passing). It replaces
|
||||||
|
two prior attempts that built their own jsonschema engines — typebox-rs
|
||||||
|
(~8,400 lines) and alktype (~5,600 lines) — with `jsonschema` + an
|
||||||
|
offset map + ~50 lines of custom keyword implementations. See
|
||||||
|
[ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||||
|
|
||||||
|
## Why
|
||||||
|
|
||||||
|
The crate's purpose is to be the binary struct engine for every alknet
|
||||||
|
component that reads or writes binary data at computed offsets. Instead
|
||||||
|
of per-protocol serde structs (russh-sftp's 29 packet types), per-handler
|
||||||
|
wire format code (TTY's 5-byte format parser), or per-format offset
|
||||||
|
computation (metatensor's tensor access), all of these become instances
|
||||||
|
of the same engine with different schemas.
|
||||||
|
|
||||||
|
The guiding insight:
|
||||||
|
|
||||||
|
> **The schema is the format.** A JSON Schema with `TypeDef:Float32`,
|
||||||
|
> `TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
|
||||||
|
> the layout spec. No separate format definition, no separate parser, no
|
||||||
|
> separate validator. One schema, three uses: validate, compute offsets,
|
||||||
|
> access data.
|
||||||
|
|
||||||
|
This is the convergence of three threads identified in the
|
||||||
|
call-channels-unification research: the `typedef.ts` schema kinds from
|
||||||
|
TypeBox, the russh-sftp protocol packets, and the metatensor format. The
|
||||||
|
common pattern: a JSON Schema describes the shape of binary data, and
|
||||||
|
the binary data is the struct's bytes at computed offsets.
|
||||||
|
|
||||||
|
The crate was bumped up in the timeline when the call-channels-unification
|
||||||
|
research surfaced that channels, TTY, and the binary call protocol are
|
||||||
|
all variations on the same wire-format family — `[discriminant][length][payload]`.
|
||||||
|
The typedef engine makes the "channels is call with a binary data plane"
|
||||||
|
unification concrete: the binary data plane's wire format is the call
|
||||||
|
protocol's own schema system, just binary-encoded. The `channel_open`
|
||||||
|
marker says "use binary framing"; the typedef engine says "here's how to
|
||||||
|
read/write the binary payload."
|
||||||
|
|
||||||
|
## The "Schema Is the Format" Principle
|
||||||
|
|
||||||
|
A JSON Schema with `TypeDef:*` custom keywords serves three roles
|
||||||
|
simultaneously:
|
||||||
|
|
||||||
|
| Role | Mechanism | When |
|
||||||
|
|------|-----------|------|
|
||||||
|
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
|
||||||
|
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) |
|
||||||
|
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
|
||||||
|
|
||||||
|
No separate format definition, no separate parser, no separate validator.
|
||||||
|
The schema is the single source of truth for the binary format. Adding a
|
||||||
|
new field to a protocol is adding a property to the schema JSON — the
|
||||||
|
engine computes the new offsets automatically.
|
||||||
|
|
||||||
|
This is the same principle as `#[repr(C)]` struct field access, but at
|
||||||
|
runtime from a portable JSON Schema instead of at compile-time from
|
||||||
|
language-specific annotations. The schema is the ABI contract.
|
||||||
|
|
||||||
|
## Dependencies
|
||||||
|
|
||||||
|
```
|
||||||
|
alknet-typedef
|
||||||
|
├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support
|
||||||
|
├── serde_json (with preserve_order) — schema parsing; field order is load-bearing
|
||||||
|
└── (no tokio, no platform deps) — WASM-clean by construction
|
||||||
|
```
|
||||||
|
|
||||||
|
`alknet-typedef` is dependency-light: `jsonschema` + `serde_json` only.
|
||||||
|
No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for
|
||||||
|
browser use. The `jsonschema` crate is already in the workspace at
|
||||||
|
`/workspace/jsonschema/` but not yet used by any alknet crate — typedef
|
||||||
|
is the first consumer.
|
||||||
|
|
||||||
|
`serde_json` requires the `preserve_order` feature because field order
|
||||||
|
is load-bearing for binary layouts. The order of properties in the
|
||||||
|
schema JSON determines the order of fields in the binary struct.
|
||||||
|
|
||||||
|
## Consumers
|
||||||
|
|
||||||
|
| Consumer | Schema describes | Engine provides |
|
||||||
|
|----------|-----------------|-----------------|
|
||||||
|
| russh-sftp | 29 packet structs + Packet union (byte discriminator) | Read/write SFTP frames from bytes |
|
||||||
|
| metatensor | Model layout (ConvNet struct, tensor refs) | Offset map for mmap'd tensor access |
|
||||||
|
| binary call frames | `call.requested` / `call.responded` / etc. structs | Read/write binary call frames |
|
||||||
|
| TTY negotiation | `NegotiateRequest` / `NegotiateResponse` structs | Read/write TTY control frames |
|
||||||
|
| channels wire | `ChunkHeader { channel_id, length }` | Already trivial (8 bytes, no schema needed) |
|
||||||
|
|
||||||
|
The russh-sftp case is the most instructive and the highest-value POC
|
||||||
|
target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written
|
||||||
|
dispatch on a type byte followed by serde deserialization. Under typedef,
|
||||||
|
the dispatch is `TUnion` with a byte-offset discriminator — the schema
|
||||||
|
says "byte 0 is the discriminator, bytes 1..N are the variant struct."
|
||||||
|
The engine reads the discriminator, looks up the variant schema, computes
|
||||||
|
offsets, reads fields. Same result, no per-packet-type code.
|
||||||
|
|
||||||
|
## Scope Boundaries (What This Is Not)
|
||||||
|
|
||||||
|
These boundaries are decided in [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||||
|
|
||||||
|
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
|
||||||
|
is a *format* (8-byte header + JSON header + binary data) that uses the
|
||||||
|
typedef engine for its offset computation and tensor access.
|
||||||
|
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
|
||||||
|
`Value.Convert` — schema evolution — is out of scope for v1. The engine
|
||||||
|
should not do anything that explicitly blocks adding a Value system
|
||||||
|
later.
|
||||||
|
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
|
||||||
|
concern. The typedef engine consumes schemas; it does not generate them.
|
||||||
|
- **Not a schema builder.** The typedef engine does not provide a fluent
|
||||||
|
API for constructing schemas. Schemas are plain JSON — authored in
|
||||||
|
TypeBox, generated by ujsx components, or hand-written. A builder API
|
||||||
|
is deferred (OQ-071).
|
||||||
|
- **Not a serialization framework.** The typedef engine is not a
|
||||||
|
general-purpose serde replacement. It operates on raw byte buffers at
|
||||||
|
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||||
|
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||||
|
with a known schema, use typedef.
|
||||||
|
|
||||||
|
## Architecture (component pointers)
|
||||||
|
|
||||||
|
- **[schema-layer.md](schema-layer.md)** — the 17 `TypeDef:*` kinds,
|
||||||
|
jsonschema custom keyword integration, TypeBox interop, schema
|
||||||
|
annotations (endianness, alignment, encoding, TUnion discriminators).
|
||||||
|
- **[layout-engine.md](layout-engine.md)** — offset computation, the two
|
||||||
|
layout modes (packed sequential vs aligned static), alignment,
|
||||||
|
endianness, variable-length field handling.
|
||||||
|
- **[data-access.md](data-access.md)** — read/write functions, TUnion
|
||||||
|
dispatch, field paths, zero-copy access for fixed-size types,
|
||||||
|
length-prefix reading for variable-length types.
|
||||||
|
- **[validation.md](validation.md)** — custom keyword validators for all
|
||||||
|
16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time
|
||||||
|
validation, `TypedefEngine` as the compiled form of a schema.
|
||||||
|
|
||||||
|
## Design Decisions
|
||||||
|
|
||||||
|
| Decision | ADR | Summary |
|
||||||
|
|----------|-----|---------|
|
||||||
|
| Purpose, scope, and the jsonschema engine | [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
|
||||||
|
| Two layout modes | [ADR-096](../../decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats |
|
||||||
|
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) |
|
||||||
|
| Error handling and validation | [ADR-098](../../decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
See [open-questions.md](../../open-questions.md) for full details.
|
||||||
|
|
||||||
|
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs.
|
||||||
|
- **OQ-070** (deferred(scope)): `no_std` + `alloc` support.
|
||||||
|
- **OQ-071** (deferred(scope)): Builder API for schema construction.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
|
||||||
|
passing, two layout modes, TUnion dispatch, endianness)
|
||||||
|
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||||
|
JSON Schema as the binary struct engine" — the origin of this research
|
||||||
|
thread
|
||||||
|
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||||
|
schema kinds (619 lines)
|
||||||
|
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||||
|
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
|
||||||
|
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
|
||||||
|
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
|
||||||
@@ -0,0 +1,348 @@
|
|||||||
|
---
|
||||||
|
status: draft
|
||||||
|
last_updated: 2026-07-20
|
||||||
|
---
|
||||||
|
|
||||||
|
# alknet-typedef — Schema Layer
|
||||||
|
|
||||||
|
The schema layer: the 16 `TypeDef:*` custom type kinds, their mapping to
|
||||||
|
Rust types and byte sizes, the `jsonschema` custom keyword integration,
|
||||||
|
TypeBox interop, and the concrete JSON shapes for schema-level annotations.
|
||||||
|
|
||||||
|
## The 17 TypeDef Kinds
|
||||||
|
|
||||||
|
These are the custom schema kinds defined in TypeBox's `typedef.ts`
|
||||||
|
(`/workspace/@alkdev/typebox/example/typedef/typedef.ts`, 619 lines) and
|
||||||
|
ported to Rust via `jsonschema` custom keywords. Each kind carries binary
|
||||||
|
layout semantics — a known byte size (for fixed-size types) or a known
|
||||||
|
encoding strategy (for variable-length types).
|
||||||
|
|
||||||
|
| Kind | TypeBox key | Rust type | Size | Category |
|
||||||
|
|------|-------------|-----------|------|----------|
|
||||||
|
| `TFloat32` | `TypeDef:Float32` | `f32` | 4 | fixed |
|
||||||
|
| `TFloat64` | `TypeDef:Float64` | `f64` | 8 | fixed |
|
||||||
|
| `TInt8` | `TypeDef:Int8` | `i8` | 1 | fixed |
|
||||||
|
| `TInt16` | `TypeDef:Int16` | `i16` | 2 | fixed |
|
||||||
|
| `TInt32` | `TypeDef:Int32` | `i32` | 4 | fixed |
|
||||||
|
| `TUint8` | `TypeDef:Uint8` | `u8` | 1 | fixed |
|
||||||
|
| `TUint16` | `TypeDef:Uint16` | `u16` | 2 | fixed |
|
||||||
|
| `TUint32` | `TypeDef:Uint32` | `u32` | 4 | fixed |
|
||||||
|
| `TBoolean` | `TypeDef:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed |
|
||||||
|
| `TString` | `TypeDef:String` | length-prefixed UTF-8 | variable | variable |
|
||||||
|
| `TBytes` | `TypeDef:Bytes` | length-prefixed raw bytes | variable | variable |
|
||||||
|
| `TStruct` | `TypeDef:Struct` | record of fields | sum of field sizes | composite |
|
||||||
|
| `TUnion` | `TypeDef:Union` | tagged union | discriminator + variant | composite |
|
||||||
|
| `TArray` | `TypeDef:Array` | repeated element | count × element size | composite |
|
||||||
|
| `TEnum` | `TypeDef:Enum` | u32 index into enum values | 4 (fixed) | fixed |
|
||||||
|
| `TRecord` | `TypeDef:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
|
||||||
|
| `TTimestamp` | `TypeDef:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
|
||||||
|
|
||||||
|
### Fixed-size types
|
||||||
|
|
||||||
|
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
|
||||||
|
`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset
|
||||||
|
computation uses these sizes directly. Read/write is zero-copy pointer
|
||||||
|
cast for these types.
|
||||||
|
|
||||||
|
**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other
|
||||||
|
values are invalid and produce a `TypedefError::Access` on read.
|
||||||
|
|
||||||
|
**`TEnum` binary representation:** A `u32` index into the enum's declared
|
||||||
|
values, in declaration order. The first declared value is index 0, the
|
||||||
|
second is index 1, etc. The enum's values are declared via the standard
|
||||||
|
JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`).
|
||||||
|
The `TypeDef:Enum` custom keyword signals that the type is an enum for
|
||||||
|
layout purposes; the built-in `enum` keyword provides the value list.
|
||||||
|
The `u32` index is always little-endian (enum indices are not protocol
|
||||||
|
data — they are internal to the schema). See [ADR-097](../../decisions/097-schema-annotations.md).
|
||||||
|
|
||||||
|
### Variable-length types
|
||||||
|
|
||||||
|
`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes.
|
||||||
|
The typedef engine supports three strategies for handling variable-length
|
||||||
|
types in binary layouts, selected by the `encoding` annotation and the
|
||||||
|
standard JSON Schema `maxLength` keyword:
|
||||||
|
|
||||||
|
| Strategy | Encoding annotation | Layout behavior | Use case |
|
||||||
|
|----------|-------------------|-----------------|----------|
|
||||||
|
| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) |
|
||||||
|
| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) |
|
||||||
|
| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) |
|
||||||
|
|
||||||
|
**Strategy 1: Inline length-prefixing (default).** The field's fixed
|
||||||
|
portion is a 4-byte length prefix at a computed offset. The variable data
|
||||||
|
follows immediately after. In packed sequential mode, the length prefix
|
||||||
|
determines the position of subsequent fields. In aligned static mode, the
|
||||||
|
length prefix is at a known offset; the variable data is not included in
|
||||||
|
the static layout. This is the universal pattern used by channels, SFTP,
|
||||||
|
TTY, and most binary protocols.
|
||||||
|
|
||||||
|
**Strategy 2: Fixed-size reservation.** When a variable-length field
|
||||||
|
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
|
||||||
|
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
|
||||||
|
than `maxLength` is zero-padded; data longer than `maxLength` is a
|
||||||
|
validation error. This makes the field fixed-size from the layout
|
||||||
|
perspective — subsequent fields have known, unchanging offsets. This is
|
||||||
|
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
|
||||||
|
pattern for fields with known maximum sizes.
|
||||||
|
|
||||||
|
In packed sequential mode, `maxLength` is a validation constraint only —
|
||||||
|
the engine still uses inline length-prefixing (strategy 1) because
|
||||||
|
protocols don't benefit from fixed-size reservation.
|
||||||
|
|
||||||
|
**Strategy 3: Offset indirection.** The field is a struct
|
||||||
|
`{offset: u32, length: u32}` at a known position. The consumer provides
|
||||||
|
the data region separately; the engine reads the offset and length, then
|
||||||
|
slices the data region. This is the metatensor blob tensor pattern — the
|
||||||
|
index struct lives in one region, the blob data lives in another. Enables
|
||||||
|
mmap-friendly random access to variable-length data without parsing
|
||||||
|
length prefixes and without reserving worst-case space.
|
||||||
|
|
||||||
|
**Default strategy selection:**
|
||||||
|
- In packed sequential mode: always strategy 1 (inline length-prefixing).
|
||||||
|
`maxLength` is a validation constraint only.
|
||||||
|
- In aligned static mode: strategy 2 (fixed-size reservation) if
|
||||||
|
`maxLength` is declared; strategy 3 (offset indirection) if
|
||||||
|
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
|
||||||
|
length-prefixing) otherwise.
|
||||||
|
|
||||||
|
**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3)
|
||||||
|
respects the schema's `"endian"` annotation (ADR-097). In little-endian
|
||||||
|
mode, the length is `u32::from_le_bytes`. In big-endian mode, the length
|
||||||
|
is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have
|
||||||
|
consistent byte order for both field values and length prefixes.
|
||||||
|
|
||||||
|
**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`.
|
||||||
|
Otherwise identical to `TString` in layout (same three strategies).
|
||||||
|
|
||||||
|
**`TRecord`:** A string-keyed map. Binary layout is a count-prefixed
|
||||||
|
sequence of `(key, value)` pairs: `[count: u32][key_len: u32][key_bytes]
|
||||||
|
[value_len: u32][value_bytes]...`. The count is the number of entries.
|
||||||
|
Each key is a length-prefixed UTF-8 string. Each value is the record's
|
||||||
|
declared value type (specified via the `"values"` property in the schema,
|
||||||
|
e.g., `"values": { "TypeDef:Float32": true }`). The count prefix respects
|
||||||
|
the schema's endianness. In aligned static mode with `maxLength`, the
|
||||||
|
entire record is reserved at `maxLength` bytes (zero-padded).
|
||||||
|
|
||||||
|
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
|
||||||
|
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
|
||||||
|
fixed-size reservation (strategy 2 with `maxLength`). The engine does not
|
||||||
|
parse or validate the timestamp format beyond UTF-8 — the jsonschema
|
||||||
|
validator checks RFC 3339 conformance at the JSON level.
|
||||||
|
|
||||||
|
`TArray` is variable-length when the element type is variable-length or
|
||||||
|
when the count is not known at schema time. For fixed-size element arrays
|
||||||
|
with a known count, the size is `element_size × count`.
|
||||||
|
|
||||||
|
**`TArray` count declaration:** The array count is declared via the
|
||||||
|
standard JSON Schema `"minItems"` and `"maxItems"` keywords. When
|
||||||
|
`minItems == maxItems`, the array has a fixed count known at schema time.
|
||||||
|
When they differ or are absent, the count is variable and the array uses
|
||||||
|
a length-prefixed encoding: `[count: u32][element_0]...[element_N]`.
|
||||||
|
The count prefix respects the schema's endianness.
|
||||||
|
|
||||||
|
### Composite types
|
||||||
|
|
||||||
|
`TStruct` and `TUnion` are composite — their size is the sum of their
|
||||||
|
fields' sizes (plus alignment padding in aligned static mode). The offset
|
||||||
|
computation recurses into their properties.
|
||||||
|
|
||||||
|
## jsonschema Custom Keyword Integration
|
||||||
|
|
||||||
|
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
|
||||||
|
via the `with_keyword` API. Each `TypeDef:*` kind is registered as a
|
||||||
|
custom keyword:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
let validator = jsonschema::options()
|
||||||
|
.with_keyword("TypeDef:Float32", factory)
|
||||||
|
.with_keyword("TypeDef:Int32", factory)
|
||||||
|
.with_keyword("TypeDef:Struct", factory)
|
||||||
|
// ... all 17 kinds
|
||||||
|
.build(&schema)?;
|
||||||
|
```
|
||||||
|
|
||||||
|
The factory closure receives the parent schema object, the keyword's
|
||||||
|
value, and the schema path — enabling cross-keyword awareness. The
|
||||||
|
`TypeDef:Struct` validator, for example, inspects the parent's
|
||||||
|
`properties` to validate each field against its declared `TypeDef:*` kind.
|
||||||
|
|
||||||
|
Each custom keyword implementation is ~10 lines. The `jsonschema` crate
|
||||||
|
handles all structural validation (object properties, required fields,
|
||||||
|
array items, enum values) — the custom keywords only need to validate
|
||||||
|
the leaf type constraints. See [validation.md](validation.md) for the
|
||||||
|
validator implementations.
|
||||||
|
|
||||||
|
This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side.
|
||||||
|
Same semantics, different language, same JSON Schema wire format. A
|
||||||
|
TypeBox schema serialized to JSON feeds directly into
|
||||||
|
`jsonschema::validator_for(&schema)` on the Rust side — zero translation.
|
||||||
|
|
||||||
|
## TypeBox Interop
|
||||||
|
|
||||||
|
TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox
|
||||||
|
schema like:
|
||||||
|
|
||||||
|
```typescript
|
||||||
|
const TensorRef = Type.Object({
|
||||||
|
dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]),
|
||||||
|
shape: Type.Array(Type.Number()),
|
||||||
|
data_offsets: Type.Tuple([Type.Number(), Type.Number()])
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
serialized to JSON is a standard JSON Schema with `type: "object"`,
|
||||||
|
`properties`, and `required`. That JSON feeds directly into the typedef
|
||||||
|
engine. The `TypeDef:*` custom keywords are added by TypeBox's
|
||||||
|
`TypeRegistry.Set` — they appear in the serialized JSON as additional
|
||||||
|
properties on the schema object.
|
||||||
|
|
||||||
|
The typedef engine does not depend on TypeBox or any JS toolchain. It
|
||||||
|
consumes JSON — whether that JSON was authored in TypeBox, generated by
|
||||||
|
a ujsx component, or hand-written. The schema is the interface.
|
||||||
|
|
||||||
|
## Schema Annotations
|
||||||
|
|
||||||
|
Schema-level annotations control binary layout behavior. These are
|
||||||
|
decided in [ADR-097](../../decisions/097-schema-annotations.md).
|
||||||
|
|
||||||
|
### Endianness
|
||||||
|
|
||||||
|
Schema-level annotation with a default of little-endian:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "TypeDef:Struct": true, "endian": "big", "properties": { ... } }
|
||||||
|
```
|
||||||
|
|
||||||
|
- `"endian": "little"` (default) — read/write in little-endian byte order.
|
||||||
|
- `"endian": "big"` — read/write in big-endian byte order.
|
||||||
|
- Applies to the entire schema and all nested types.
|
||||||
|
|
||||||
|
### Alignment
|
||||||
|
|
||||||
|
Both struct-level and field-level, with field-level overriding:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"TypeDef:Struct": true,
|
||||||
|
"align": 256,
|
||||||
|
"properties": {
|
||||||
|
"weight": { "TypeDef:Float32": true, "align": 16 }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- Struct-level `"align"` sets the default for all fields.
|
||||||
|
- Field-level `"align"` overrides the struct default.
|
||||||
|
- Default alignment: 1 for u8/bool, 2 for u16/i16, 4 for u32/i32/f32,
|
||||||
|
8 for u64/i64/f64, max field alignment for structs.
|
||||||
|
- Only meaningful in aligned static mode (ADR-096). Ignored in packed
|
||||||
|
sequential mode.
|
||||||
|
|
||||||
|
### Variable-length encoding
|
||||||
|
|
||||||
|
The typedef engine supports three strategies for variable-length types
|
||||||
|
(see §Variable-length types above for full details). The strategy is
|
||||||
|
selected by the `encoding` annotation and the standard JSON Schema
|
||||||
|
`maxLength` keyword:
|
||||||
|
|
||||||
|
```json
|
||||||
|
// Strategy 1: Inline length-prefixing (default, shorthand)
|
||||||
|
{ "TypeDef:String": true }
|
||||||
|
|
||||||
|
// Strategy 1: Explicit inline length-prefixing
|
||||||
|
{ "TypeDef:String": { "encoding": "length-prefixed" } }
|
||||||
|
|
||||||
|
// Strategy 2: Fixed-size reservation (uses standard maxLength)
|
||||||
|
{ "TypeDef:String": true, "maxLength": 256 }
|
||||||
|
|
||||||
|
// Strategy 3: Offset indirection (opt-in)
|
||||||
|
{ "TypeDef:String": { "encoding": "offset-indirect" } }
|
||||||
|
```
|
||||||
|
|
||||||
|
- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at
|
||||||
|
computed offset, variable data follows immediately. Used by protocol
|
||||||
|
wire formats.
|
||||||
|
- `maxLength` (standard JSON Schema keyword) — in aligned static mode,
|
||||||
|
reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the
|
||||||
|
field fixed-size from the layout perspective. In packed sequential
|
||||||
|
mode, `maxLength` is a validation constraint only.
|
||||||
|
- `"encoding": "offset-indirect"` — field is a struct
|
||||||
|
`{offset: u32, length: u32}` pointing into a separate data region.
|
||||||
|
The consumer provides the data region separately. Used by metatensor
|
||||||
|
blob tensors.
|
||||||
|
- Applies to all variable-length types: `TypeDef:String`, `TypeDef:Bytes`,
|
||||||
|
`TypeDef:Array`, `TypeDef:Record`, `TypeDef:Timestamp`.
|
||||||
|
|
||||||
|
### TUnion discriminators
|
||||||
|
|
||||||
|
Two discriminator kinds: byte-offset (protocol dispatch) and field-name
|
||||||
|
(typedef.ts pattern).
|
||||||
|
|
||||||
|
**Byte-offset discriminator** (SFTP type bytes, call protocol event types):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"TypeDef:Union": true,
|
||||||
|
"discriminator": {
|
||||||
|
"kind": "byte",
|
||||||
|
"offset": 0,
|
||||||
|
"type": "TypeDef:Uint8"
|
||||||
|
},
|
||||||
|
"mapping": {
|
||||||
|
"5": { "$ref": "#/$defs/Read" },
|
||||||
|
"6": { "$ref": "#/$defs/Write" },
|
||||||
|
"101": { "$ref": "#/$defs/Status" }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `"offset"` — byte position of the discriminator.
|
||||||
|
- `"type"` — the `TypeDef:*` kind of the discriminator (typically
|
||||||
|
`TypeDef:Uint8`).
|
||||||
|
- Mapping keys are stringified integers. The variant struct starts at
|
||||||
|
`offset + discriminator_size`.
|
||||||
|
|
||||||
|
**Field-name discriminator** (typedef.ts pattern):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"TypeDef:Union": true,
|
||||||
|
"discriminator": {
|
||||||
|
"kind": "field",
|
||||||
|
"name": "type"
|
||||||
|
},
|
||||||
|
"mapping": {
|
||||||
|
"read": { "$ref": "#/$defs/Read" },
|
||||||
|
"write": { "$ref": "#/$defs/Write" }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `"name"` — the field name holding the discriminator value.
|
||||||
|
- Mapping keys are string values matching the discriminator field's value.
|
||||||
|
- The discriminator field is just another field in the struct.
|
||||||
|
|
||||||
|
Mapping values may be either inline schemas or `$ref` pointers. Both work.
|
||||||
|
|
||||||
|
## Design Decisions
|
||||||
|
|
||||||
|
| Decision | ADR | Summary |
|
||||||
|
|----------|-----|---------|
|
||||||
|
| Schema annotations | [ADR-097](../../decisions/097-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators |
|
||||||
|
| Purpose and scope | [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
See [open-questions.md](../../open-questions.md) for full details.
|
||||||
|
|
||||||
|
- **OQ-071** (deferred(scope)): Builder API for schema construction.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||||
|
schema kinds (619 lines)
|
||||||
|
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||||
|
- [ADR-097](../../decisions/097-schema-annotations.md) — schema
|
||||||
|
annotation shapes
|
||||||
|
- [validation.md](validation.md) — custom keyword validator implementations
|
||||||
@@ -0,0 +1,292 @@
|
|||||||
|
---
|
||||||
|
status: draft
|
||||||
|
last_updated: 2026-07-20
|
||||||
|
---
|
||||||
|
|
||||||
|
# alknet-typedef — Validation
|
||||||
|
|
||||||
|
The validation layer: custom keyword validators for all 17 `TypeDef:*`
|
||||||
|
kinds, the `TypedefError` enum, load-time vs access-time validation
|
||||||
|
strategy, and the `TypedefEngine` as the compiled form of a schema.
|
||||||
|
|
||||||
|
## Validation Strategy
|
||||||
|
|
||||||
|
Validation is delegated to the `jsonschema` crate (v0.46.5, Draft
|
||||||
|
2020-12). The typedef engine does not implement its own validation —
|
||||||
|
it registers custom keyword validators for each `TypeDef:*` kind and
|
||||||
|
lets `jsonschema` handle the structural validation (object properties,
|
||||||
|
required fields, array items, enum values).
|
||||||
|
|
||||||
|
The strategy is decided in [ADR-098](../../decisions/098-error-handling-validation-strategy.md):
|
||||||
|
|
||||||
|
1. **Load time:** Parse the schema JSON, build the layout engine, build the
|
||||||
|
jsonschema validator. This is the `TypedefEngine::compile(schema)` constructor.
|
||||||
|
2. **Access time:** Use the compiled engine for repeated read/write
|
||||||
|
operations. Validation is opt-in per operation.
|
||||||
|
|
||||||
|
### What validation validates
|
||||||
|
|
||||||
|
The jsonschema validator operates on `serde_json::Value` instances — it
|
||||||
|
validates JSON representations of data, not raw byte buffers. This is
|
||||||
|
the correct separation of concerns:
|
||||||
|
|
||||||
|
- **JSON validation** (jsonschema): validates that a JSON document
|
||||||
|
conforms to the schema. Used for validating hand-written schemas,
|
||||||
|
TypeBox output, JSON payloads, or the JSON representation of a binary
|
||||||
|
struct after deserialization.
|
||||||
|
- **Binary access validation** (data access layer): the read/write
|
||||||
|
functions perform type-level validation at access time — range checks
|
||||||
|
for integers, UTF-8 validity for strings, buffer bounds checking.
|
||||||
|
These return `TypedefError::Access` with field paths.
|
||||||
|
|
||||||
|
The "schema is the format" principle means the same schema describes
|
||||||
|
both the JSON shape and the binary layout. The jsonschema validator
|
||||||
|
checks the JSON shape; the data access layer checks the binary layout.
|
||||||
|
A consumer that wants to validate a binary buffer end-to-end reads the
|
||||||
|
buffer into a `Value` tree via the data access layer, then validates
|
||||||
|
that `Value` against the jsonschema validator. This is a two-step
|
||||||
|
process, not a single `validate(buffer)` call.
|
||||||
|
|
||||||
|
### The `TypedefEngine` struct
|
||||||
|
|
||||||
|
The `TypedefEngine` is the compiled form of a schema. It supports both
|
||||||
|
layout modes (ADR-096) via an internal enum:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
pub struct TypedefEngine {
|
||||||
|
layout: Layout, // packed or aligned (see below)
|
||||||
|
validator: jsonschema::Validator, // compiled once at load time
|
||||||
|
}
|
||||||
|
|
||||||
|
enum Layout {
|
||||||
|
Packed {
|
||||||
|
builder: LayoutBuilder,
|
||||||
|
reader: SequentialReader,
|
||||||
|
},
|
||||||
|
Aligned {
|
||||||
|
offset_map: OffsetMap,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The consumer selects the mode at construction time. The `Layout` enum
|
||||||
|
ensures the engine always has the correct layout strategy for the
|
||||||
|
consumer's use case — a protocol consumer gets `Packed`, an mmap
|
||||||
|
consumer gets `Aligned`. The validator is mode-agnostic (it operates on
|
||||||
|
`Value`, not raw bytes).
|
||||||
|
|
||||||
|
## Custom Keyword Validators
|
||||||
|
|
||||||
|
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
|
||||||
|
`jsonschema::options().with_keyword(...)`. The validators check leaf
|
||||||
|
type constraints; `jsonschema` handles all structural validation.
|
||||||
|
|
||||||
|
### Numeric type validators
|
||||||
|
|
||||||
|
**`TypeDef:Float32` / `TypeDef:Float64`:**
|
||||||
|
- Value must be a finite number.
|
||||||
|
- For `Float32`: value must be representable as `f32` (no precision loss
|
||||||
|
beyond `f32`'s mantissa).
|
||||||
|
|
||||||
|
**`TypeDef:Int8` / `TypeDef:Int16` / `TypeDef:Int32`:**
|
||||||
|
- Value must be an integer within the type's range.
|
||||||
|
- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
|
||||||
|
|
||||||
|
**`TypeDef:Uint8` / `TypeDef:Uint16` / `TypeDef:Uint32`:**
|
||||||
|
- Value must be a non-negative integer within the type's range.
|
||||||
|
- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295.
|
||||||
|
|
||||||
|
### String and binary validators
|
||||||
|
|
||||||
|
**`TypeDef:String`:**
|
||||||
|
- Value must be a valid UTF-8 string.
|
||||||
|
- If `maxLength` is specified in the schema, the string's byte length
|
||||||
|
must not exceed it.
|
||||||
|
|
||||||
|
**`TypeDef:Bytes`:**
|
||||||
|
- Value must be a string (JSON represents binary data as a string).
|
||||||
|
- If `maxLength` is specified, the byte length must not exceed it.
|
||||||
|
|
||||||
|
**`TypeDef:Enum`:**
|
||||||
|
- The `TypeDef:Enum` custom keyword signals that the type is an enum for
|
||||||
|
*layout* purposes (the engine needs to know it's a fixed-size u32 index,
|
||||||
|
not a variable-length string). The built-in `enum` keyword provides the
|
||||||
|
value list and handles value-membership validation. The custom keyword
|
||||||
|
validator is a no-op beyond the built-in check — it exists solely for
|
||||||
|
the layout engine to recognize the type.
|
||||||
|
|
||||||
|
**`TypeDef:Timestamp`:**
|
||||||
|
- Value must be a valid RFC 3339 timestamp string (the internet profile
|
||||||
|
of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
|
||||||
|
|
||||||
|
### Composite type validators
|
||||||
|
|
||||||
|
**`TypeDef:Struct`:**
|
||||||
|
- Value must be an object.
|
||||||
|
- Each property must match its declared `TypeDef:*` kind.
|
||||||
|
- Required fields must be present.
|
||||||
|
- The `jsonschema` crate's built-in `properties` and `required` keywords
|
||||||
|
handle the structural checks — the custom keyword only needs to
|
||||||
|
validate that each field's value matches its `TypeDef:*` kind.
|
||||||
|
|
||||||
|
**`TypeDef:Union`:**
|
||||||
|
- The discriminator value must be one of the mapping keys.
|
||||||
|
- The variant struct must match the declared schema for that discriminator
|
||||||
|
value.
|
||||||
|
|
||||||
|
**`TypeDef:Array`:**
|
||||||
|
- Value must be an array.
|
||||||
|
- Each element must match the array's declared element type.
|
||||||
|
- If `minItems`/`maxItems` is specified, the array length must be within
|
||||||
|
bounds.
|
||||||
|
|
||||||
|
### Other validators
|
||||||
|
|
||||||
|
**`TypeDef:Boolean`:**
|
||||||
|
- Value must be `true` or `false`.
|
||||||
|
|
||||||
|
**`TypeDef:Record`:**
|
||||||
|
- Value must be an object.
|
||||||
|
- All values must match the record's declared value type (specified via
|
||||||
|
the `"values"` property in the schema, e.g.,
|
||||||
|
`"values": { "TypeDef:Float32": true }`).
|
||||||
|
|
||||||
|
### Validator implementation pattern
|
||||||
|
|
||||||
|
Each custom keyword implementation is ~10 lines. Example for
|
||||||
|
`TypeDef:Float32`:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
struct Float32Validator;
|
||||||
|
|
||||||
|
impl Keyword for Float32Validator {
|
||||||
|
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
|
||||||
|
match instance {
|
||||||
|
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
|
||||||
|
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
fn is_valid(&self, instance: &Value) -> bool {
|
||||||
|
instance.as_f64().map_or(false, |f| f.is_finite())
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Registration:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
let validator = jsonschema::options()
|
||||||
|
.with_keyword("TypeDef:Float32", |parent, value, path| {
|
||||||
|
Ok(Box::new(Float32Validator))
|
||||||
|
})
|
||||||
|
.build(&schema)?;
|
||||||
|
```
|
||||||
|
|
||||||
|
The factory closure receives the parent schema object, the keyword's
|
||||||
|
value, and the schema path. This enables cross-keyword awareness — for
|
||||||
|
example, a `TypeDef:Struct` validator can inspect the parent's
|
||||||
|
`properties` to validate each field against its declared `TypeDef:*` kind.
|
||||||
|
|
||||||
|
## TypedefError
|
||||||
|
|
||||||
|
A single `TypedefError` enum covers all error conditions across the
|
||||||
|
engine's three phases (schema parsing, offset computation, read/write)
|
||||||
|
plus validation. Decided in [ADR-098](../../decisions/098-error-handling-validation-strategy.md).
|
||||||
|
|
||||||
|
```rust
|
||||||
|
pub enum TypedefError {
|
||||||
|
/// Schema parsing errors (invalid JSON, missing keywords, unknown TypeDef kinds).
|
||||||
|
Schema(String),
|
||||||
|
/// Offset computation errors (field not found, unsupported type).
|
||||||
|
Offset { field_path: String, reason: String },
|
||||||
|
/// Read/write errors (buffer too short, invalid UTF-8, value out of range).
|
||||||
|
Access { field_path: String, reason: String },
|
||||||
|
/// Validation errors (delegated to jsonschema).
|
||||||
|
Validation(ValidationError<'static>),
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- **`Schema`** — for errors during `TypedefEngine::compile()`. Invalid
|
||||||
|
JSON, missing required keywords, unknown `TypeDef:*` kinds.
|
||||||
|
- **`Offset`** — for errors during offset computation. Field not found
|
||||||
|
in the schema, type not supported for offset computation, recursive
|
||||||
|
depth exceeded. Carries the field path.
|
||||||
|
- **`Access`** — for errors during read/write. Buffer too short, invalid
|
||||||
|
UTF-8 in a string field, value out of range for the target type.
|
||||||
|
Carries the field path.
|
||||||
|
- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The
|
||||||
|
`'static` lifetime is correct — the validator owns its schema reference
|
||||||
|
and lives for the lifetime of the `TypedefEngine`.
|
||||||
|
|
||||||
|
### Field-path-carrying errors
|
||||||
|
|
||||||
|
Read/write and offset errors include the field path for debugging:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
Err(TypedefError::Access {
|
||||||
|
field_path: "header.version".to_string(),
|
||||||
|
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
This makes debugging binary format issues tractable — the error tells
|
||||||
|
you exactly which field failed and why.
|
||||||
|
|
||||||
|
## Validation Timing
|
||||||
|
|
||||||
|
### Load time: `TypedefEngine::compile()`
|
||||||
|
|
||||||
|
The expensive work happens once at schema load time:
|
||||||
|
1. Parse the schema JSON (`serde_json::from_str` with `preserve_order`).
|
||||||
|
2. Compute the offset map (or `LayoutBuilder`/`SequentialReader`).
|
||||||
|
3. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
|
||||||
|
|
||||||
|
The result is a `TypedefEngine` that can be used for repeated operations.
|
||||||
|
|
||||||
|
### Access time: `engine.validate(buffer)`
|
||||||
|
|
||||||
|
Validation is opt-in per operation. The consumer calls
|
||||||
|
`engine.validate(buffer)` when validation is desired. The jsonschema
|
||||||
|
validator is already compiled — `is_valid()` is a fast check against
|
||||||
|
the compiled validator.
|
||||||
|
|
||||||
|
High-throughput paths can skip validation. Security-sensitive paths
|
||||||
|
(parsing incoming frames from untrusted peers) can validate every frame.
|
||||||
|
The choice is the consumer's.
|
||||||
|
|
||||||
|
## Relationship to Read/Write
|
||||||
|
|
||||||
|
Validation and data access are independent operations on the same buffer.
|
||||||
|
The consumer can:
|
||||||
|
|
||||||
|
1. Validate a buffer to ensure it conforms to the schema.
|
||||||
|
2. Read fields from the buffer at computed offsets.
|
||||||
|
3. Both — validate first, then read (defense in depth).
|
||||||
|
|
||||||
|
The engine does not couple validation and access. A consumer that trusts
|
||||||
|
its data source can skip validation and go straight to read/write. A
|
||||||
|
consumer that parses untrusted input can validate first, then access.
|
||||||
|
|
||||||
|
## Design Decisions
|
||||||
|
|
||||||
|
| Decision | ADR | Summary |
|
||||||
|
|----------|-----|---------|
|
||||||
|
| Error handling and validation | [ADR-098](../../decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
|
||||||
|
| Purpose and scope | [ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine |
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
None specific to validation. The three typedef OQs (OQ-069, OQ-070,
|
||||||
|
OQ-071) are about layout, platform support, and schema construction —
|
||||||
|
not validation.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` §"Validation" — the POC's
|
||||||
|
custom keyword validators for all 17 kinds
|
||||||
|
- [ADR-098](../../decisions/098-error-handling-validation-strategy.md) —
|
||||||
|
error handling and validation strategy
|
||||||
|
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds that the
|
||||||
|
validators check
|
||||||
|
- [data-access.md](data-access.md) — read/write functions that operate
|
||||||
|
on the same buffers
|
||||||
@@ -0,0 +1,168 @@
|
|||||||
|
# ADR-095: alknet-typedef — Purpose, Scope, and the jsonschema Engine
|
||||||
|
|
||||||
|
## Status
|
||||||
|
Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Three threads in the codebase converge on the same pattern: a JSON Schema
|
||||||
|
describes the shape of binary data, and the binary data is the struct's
|
||||||
|
bytes at computed offsets.
|
||||||
|
|
||||||
|
1. **typedef.ts** (`/workspace/@alkdev/typebox/example/typedef/typedef.ts`,
|
||||||
|
619 lines) defines custom TypeBox schema kinds (`TFloat32`, `TStruct`,
|
||||||
|
`TUnion`, etc.) that carry binary layout semantics. These are registered
|
||||||
|
via `TypeRegistry.Set` with custom validators.
|
||||||
|
|
||||||
|
2. **russh-sftp** has 29 packet types, each a struct with typed fields
|
||||||
|
(`Read { id: u32, handle: String, offset: u64, len: u32 }`). The wire
|
||||||
|
format is `[length: u32][type: u8][payload]` where payload is the
|
||||||
|
struct's serde bytes. The `Packet` enum dispatches on the type byte —
|
||||||
|
a tagged union of structs. Under the typedef lens, each packet is a
|
||||||
|
`TStruct`; the `Packet` enum is a `TUnion` with a byte-offset
|
||||||
|
discriminator.
|
||||||
|
|
||||||
|
3. **metatensor** needs an offset map for mmap-friendly tensor access —
|
||||||
|
given a schema describing a model layout (ConvNet struct, tensor refs),
|
||||||
|
compute byte offsets for each field so the consumer can read tensor
|
||||||
|
data at known positions without parsing.
|
||||||
|
|
||||||
|
The common pattern: **a JSON Schema with `TypeDef:*` custom keywords
|
||||||
|
describes the shape of binary data; the binary data is the struct's bytes
|
||||||
|
at computed offsets.** The schema is the format definition; the engine is
|
||||||
|
generic.
|
||||||
|
|
||||||
|
Two prior attempts built their own jsonschema engines — the fatal flaw:
|
||||||
|
|
||||||
|
- **typebox-rs** (`/workspace/@alkimiadev/typebox-rs/`, ~8,400 lines):
|
||||||
|
a full 26-variant `SchemaKind` enum, a custom `Value` type with typed
|
||||||
|
arrays, and a 912-line hand-written validator.
|
||||||
|
- **alktype** (`/workspace/@alkimiadev/alktype/`, ~5,600 lines): a
|
||||||
|
handler-registry pattern that also implements its own validation for
|
||||||
|
each type.
|
||||||
|
|
||||||
|
The `jsonschema` crate (v0.46.5, Draft 2020-12) is already in the
|
||||||
|
workspace at `/workspace/jsonschema/`. It handles validation with custom
|
||||||
|
keyword support — the novel code is the offset computation, not the
|
||||||
|
validation.
|
||||||
|
|
||||||
|
The call-channels-unification research
|
||||||
|
(`docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||||
|
JSON Schema as the binary struct engine") identified the convergence and
|
||||||
|
bumped typedef up in the timeline. The POC
|
||||||
|
(`docs/research/alknet-typedef/findings.md`, 26 tests passing) validated
|
||||||
|
the approach: a ~1,900-line Rust crate that takes a JSON Schema with
|
||||||
|
`TypeDef:*` custom keywords and produces an offset map, read/write
|
||||||
|
functions, and validation — all driven by the schema.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**alknet-typedef is a small Rust crate that takes a JSON Schema with
|
||||||
|
`TypeDef:*` custom keywords and produces three capabilities:**
|
||||||
|
|
||||||
|
1. **An offset map** — walks the schema, computes byte offsets for each
|
||||||
|
field based on type sizes, field order, and alignment.
|
||||||
|
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
|
||||||
|
read the field's bytes at its offset (zero-copy for fixed-size types).
|
||||||
|
Given a `&mut [u8]` buffer, write a value at its offset.
|
||||||
|
3. **Validation** — via `jsonschema` custom keywords, validates that a
|
||||||
|
buffer's bytes match the schema's type constraints.
|
||||||
|
|
||||||
|
**The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||||
|
`serde_json` (schema parsing).** The novel code is the offset computation
|
||||||
|
— a recursive walk of the schema JSON that computes byte positions for
|
||||||
|
each field. The custom keyword implementations are ~10 lines each.
|
||||||
|
|
||||||
|
**The schema is the format.** A JSON Schema with `TypeDef:Float32`,
|
||||||
|
`TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
|
||||||
|
the layout spec. No separate format definition, no separate parser, no
|
||||||
|
separate validator. One schema, three uses: validate, compute offsets,
|
||||||
|
access data.
|
||||||
|
|
||||||
|
**The crate depends on `jsonschema` and `serde_json` (with
|
||||||
|
`preserve_order`).** No tokio, no platform deps. Compiles to
|
||||||
|
`wasm32-unknown-unknown` for browser use. The `jsonschema` crate's
|
||||||
|
`with_keyword("TypeDef:Float32", factory)` API is the integration point
|
||||||
|
for custom type kinds — each `TypeDef:*` kind maps to a custom keyword
|
||||||
|
validator in Rust. Same semantics as TypeBox's `TypeRegistry.Set`, same
|
||||||
|
JSON Schema wire format.
|
||||||
|
|
||||||
|
**The crate targets `std` for v1.** The WASM target has `std` available
|
||||||
|
via `wasm-bindgen`. If embedded use cases emerge, `no_std` + `alloc` can
|
||||||
|
be added as a feature gate later — the engine's core (offset computation,
|
||||||
|
read/write) is already allocation-free. See OQ-070.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
### Positive
|
||||||
|
|
||||||
|
- **Eliminates ~14,000 lines of hand-rolled schema engines.** typebox-rs
|
||||||
|
and alktype are replaced by `jsonschema` + an offset map + ~50 lines of
|
||||||
|
custom keyword implementations. The codebase drops from "a port of
|
||||||
|
TypeBox" to "jsonschema + an offset map."
|
||||||
|
- **One schema, three uses.** The same JSON Schema validates, computes
|
||||||
|
offsets, and drives data access. No separate format definition, parser,
|
||||||
|
or validator per protocol.
|
||||||
|
- **Schema-driven, not code-driven.** Adding a new SFTP packet type is
|
||||||
|
adding a variant to the schema JSON, not writing a new Rust struct +
|
||||||
|
serde impl. The engine is generic; the schema is the configuration.
|
||||||
|
- **WASM-clean.** `serde_json` + `jsonschema` + byte manipulation. No
|
||||||
|
tokio, no platform deps. The same typedef schemas work in browser,
|
||||||
|
Node, Python (via `wasmtime-py`), Go (via `wazero`), and any other
|
||||||
|
WASM host.
|
||||||
|
- **TypeBox interop.** TypeBox modules render to standard JSON Schema
|
||||||
|
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
|
||||||
|
on the Rust side. Zero translation. The same schema validates in both
|
||||||
|
ecosystems.
|
||||||
|
- **Defense in depth.** Schema validation at the byte level — a malformed
|
||||||
|
binary payload fails validation before any consumer touches it. The
|
||||||
|
`jsonschema` crate's compiled validators are fast enough to run on
|
||||||
|
every incoming frame.
|
||||||
|
|
||||||
|
### Negative
|
||||||
|
|
||||||
|
- **New dependency on `jsonschema`.** The crate is already in the
|
||||||
|
workspace but not yet used by any alknet crate. This is the first
|
||||||
|
consumer.
|
||||||
|
- **`serde_json` with `preserve_order` is required.** Field order is
|
||||||
|
load-bearing for binary layouts. The `preserve_order` feature adds a
|
||||||
|
small compile-time cost.
|
||||||
|
- **Schema authoring is external.** Schemas are authored in TypeBox (JS)
|
||||||
|
or hand-written JSON. The typedef engine consumes schemas; it does not
|
||||||
|
generate them. A builder API is deferred (OQ-071).
|
||||||
|
|
||||||
|
## Scope Boundaries (What This Is Not)
|
||||||
|
|
||||||
|
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
|
||||||
|
is a *format* (8-byte header + JSON header + binary data) that uses the
|
||||||
|
typedef engine for its offset computation and tensor access.
|
||||||
|
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
|
||||||
|
`Value.Convert` — schema evolution — is out of scope for v1. The engine
|
||||||
|
should not do anything that explicitly blocks adding a Value system
|
||||||
|
later.
|
||||||
|
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
|
||||||
|
concern. The typedef engine consumes schemas; it does not generate them.
|
||||||
|
- **Not a schema builder.** The typedef engine does not provide a fluent
|
||||||
|
API for constructing schemas. Schemas are plain JSON.
|
||||||
|
- **Not a serialization framework.** The typedef engine is not a
|
||||||
|
general-purpose serde replacement. It operates on raw byte buffers at
|
||||||
|
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||||
|
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||||
|
with a known schema, use typedef.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
|
||||||
|
passing, two layout modes, TUnion dispatch, endianness)
|
||||||
|
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||||
|
JSON Schema as the binary struct engine" — the origin of this research
|
||||||
|
thread
|
||||||
|
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||||
|
schema kinds (619 lines)
|
||||||
|
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||||
|
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
|
||||||
|
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||||
|
modes decision
|
||||||
|
- [ADR-097](097-schema-annotations.md) — schema annotation shapes
|
||||||
|
- [ADR-098](098-error-handling-validation-strategy.md) — error handling
|
||||||
|
and validation strategy
|
||||||
@@ -0,0 +1,137 @@
|
|||||||
|
# ADR-096: Two Layout Modes — Packed Sequential vs Aligned Static
|
||||||
|
|
||||||
|
## Status
|
||||||
|
Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The POCs surfaced that protocols and mmap-friendly formats need different
|
||||||
|
layout strategies. POC 1 built an aligned `OffsetMap` with natural
|
||||||
|
alignment padding — correct for mmap-friendly formats (metatensor) but
|
||||||
|
wrong for protocol wire formats (SFTP, channels, TTY). POC 2 built a
|
||||||
|
`LayoutBuilder` and `SequentialReader` for packed sequential layouts —
|
||||||
|
correct for protocol wire formats but wrong for mmap-friendly formats.
|
||||||
|
|
||||||
|
This is the most important architectural finding from the POCs. The
|
||||||
|
engine must support both modes; a single layout strategy cannot serve
|
||||||
|
both use cases.
|
||||||
|
|
||||||
|
### Packed sequential layout (protocol wire formats)
|
||||||
|
|
||||||
|
Protocols pack fields sequentially with no alignment padding.
|
||||||
|
Variable-length fields shift all subsequent fields. Writing requires
|
||||||
|
knowing actual data sizes upfront; reading walks the buffer sequentially,
|
||||||
|
reading length prefixes to determine positions.
|
||||||
|
|
||||||
|
This is the layout used by SFTP (all strings and byte arrays are
|
||||||
|
length-prefixed inline), channels (`[channel_id: u32][size: u32][payload]`),
|
||||||
|
TTY (`[stream_type: u8][length: u32][payload]`), and most binary protocols.
|
||||||
|
|
||||||
|
### Aligned static layout (mmap-friendly formats)
|
||||||
|
|
||||||
|
Fields have fixed positions with natural alignment padding.
|
||||||
|
Variable-length fields get a 4-byte length prefix at a known offset; the
|
||||||
|
variable data is not included in the static layout. This enables
|
||||||
|
mmap-friendly random access — the consumer can read field N at a known
|
||||||
|
offset without parsing the fields before it.
|
||||||
|
|
||||||
|
This is the layout used by metatensor (blob tensor pattern: index struct
|
||||||
|
in one region, blob data in another) and safetensors (header + aligned
|
||||||
|
tensor data).
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**The typedef engine supports two layout modes, selected by the consumer
|
||||||
|
at engine construction time:**
|
||||||
|
|
||||||
|
### Mode 1: Packed sequential (`LayoutBuilder` / `SequentialReader`)
|
||||||
|
|
||||||
|
For protocol wire formats. Fields are packed with no alignment padding.
|
||||||
|
Variable-length fields shift all subsequent fields.
|
||||||
|
|
||||||
|
- **LayoutBuilder** — takes a schema and actual data sizes for
|
||||||
|
variable-length fields, computes byte positions for each field in a
|
||||||
|
packed layout. Used at write time when the consumer knows the data
|
||||||
|
sizes upfront.
|
||||||
|
- **SequentialReader** — walks a buffer field-by-field according to the
|
||||||
|
schema, reading length prefixes to determine variable-length data
|
||||||
|
positions. Used at read time when the consumer is parsing an incoming
|
||||||
|
frame.
|
||||||
|
|
||||||
|
The `LayoutBuilder` and `SequentialReader` are the primary interface for
|
||||||
|
protocol consumers (SFTP, binary call frames, TTY negotiation).
|
||||||
|
|
||||||
|
### Mode 2: Aligned static (`OffsetMap`)
|
||||||
|
|
||||||
|
For mmap-friendly formats. Fields have fixed positions with natural
|
||||||
|
alignment padding. Variable-length fields get a 4-byte length prefix at
|
||||||
|
a known offset; the variable data is not included in the static layout.
|
||||||
|
|
||||||
|
- **OffsetMap** — walks the schema once, computes fixed byte positions
|
||||||
|
for each field based on type sizes and alignment. The output is a flat
|
||||||
|
table of `(field_path, byte_range)` pairs. Used for both read and write
|
||||||
|
at known offsets.
|
||||||
|
|
||||||
|
The `OffsetMap` is the primary interface for mmap consumers (metatensor).
|
||||||
|
|
||||||
|
### Variable-length handling in each mode
|
||||||
|
|
||||||
|
**Packed sequential mode:** Variable-length fields are inline
|
||||||
|
length-prefixed by default (`[length: u32][data]`). The `LayoutBuilder`
|
||||||
|
takes the actual data size to compute the length prefix value and the
|
||||||
|
position of subsequent fields. The `SequentialReader` reads the length
|
||||||
|
prefix to determine the data extent and the position of the next field.
|
||||||
|
|
||||||
|
**Aligned static mode:** Variable-length fields get a 4-byte length
|
||||||
|
prefix at a known offset. The variable data lives outside the static
|
||||||
|
layout — either immediately after the fixed fields (inline
|
||||||
|
length-prefixing) or in a separate data region (offset indirection, the
|
||||||
|
metatensor blob tensor pattern). The `OffsetMap` records the position of
|
||||||
|
the length prefix (or the `{offset, length}` pair for offset-indirect
|
||||||
|
fields).
|
||||||
|
|
||||||
|
### Default for variable-length types
|
||||||
|
|
||||||
|
Inline length-prefixing (`[length: u32][data]`) is the default for all
|
||||||
|
variable-length types in both modes. This is the universal pattern used
|
||||||
|
by channels, SFTP, TTY, and most binary protocols. Offset indirection is
|
||||||
|
opt-in via the `encoding` annotation (see ADR-097).
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
### Positive
|
||||||
|
|
||||||
|
- **One engine, two modes.** The same schema can be used in either mode.
|
||||||
|
A schema describing an SFTP packet can be consumed by a `SequentialReader`
|
||||||
|
(for parsing incoming frames) and a `LayoutBuilder` (for constructing
|
||||||
|
outgoing frames). A schema describing a metatensor layout can be
|
||||||
|
consumed by an `OffsetMap` (for mmap access).
|
||||||
|
- **Correct for both use cases.** Packed sequential mode produces
|
||||||
|
byte-identical output to hand-written protocol serialization (validated
|
||||||
|
by POC 2's russh-sftp round-trip tests). Aligned static mode produces
|
||||||
|
correct offsets for mmap-friendly access (validated by POC 1's
|
||||||
|
alignment tests).
|
||||||
|
- **No mode confusion.** The consumer explicitly selects the mode at
|
||||||
|
engine construction time. A protocol consumer never accidentally gets
|
||||||
|
alignment padding; an mmap consumer never accidentally gets
|
||||||
|
variable-length field shifting.
|
||||||
|
|
||||||
|
### Negative
|
||||||
|
|
||||||
|
- **Two APIs to learn.** Consumers must choose between
|
||||||
|
`LayoutBuilder`/`SequentialReader` and `OffsetMap`. The choice is
|
||||||
|
determined by the use case (protocol vs mmap), not by the schema.
|
||||||
|
- **Variable-length fields in packed mode require size foreknowledge.**
|
||||||
|
The `LayoutBuilder` needs actual data sizes for variable-length fields
|
||||||
|
to compute correct positions for subsequent fields. This is inherent
|
||||||
|
to packed layouts — the consumer must know the data sizes before
|
||||||
|
writing.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
|
||||||
|
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
|
||||||
|
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||||
|
purpose and scope
|
||||||
|
- [ADR-097](097-schema-annotations.md) — schema annotations including
|
||||||
|
the `encoding` field for variable-length types
|
||||||
@@ -0,0 +1,232 @@
|
|||||||
|
# ADR-097: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
|
||||||
|
|
||||||
|
## Status
|
||||||
|
Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The typedef engine needs concrete JSON shapes for schema-level
|
||||||
|
annotations that control binary layout behavior. The POCs validated the
|
||||||
|
semantics; this ADR pins the shapes.
|
||||||
|
|
||||||
|
Four annotation categories need concrete shapes:
|
||||||
|
|
||||||
|
1. **Endianness** — safetensors is little-endian, SFTP is big-endian.
|
||||||
|
The engine needs to know which to use.
|
||||||
|
2. **Alignment** — different backends have different alignment
|
||||||
|
requirements (wgpu: 256-byte, protocols: natural, mmap: page).
|
||||||
|
3. **Variable-length encoding** — inline length-prefixing vs offset
|
||||||
|
indirection for strings, byte arrays, and other variable-length types.
|
||||||
|
4. **TUnion discriminators** — byte-offset (protocol dispatch) vs
|
||||||
|
field-name (typedef.ts pattern).
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
### 1. Endianness
|
||||||
|
|
||||||
|
**Schema-level annotation with a default of little-endian.**
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"TypeDef:Struct": true,
|
||||||
|
"endian": "big",
|
||||||
|
"properties": { ... }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `"endian": "little"` (default) — read/write in little-endian byte order.
|
||||||
|
- `"endian": "big"` — read/write in big-endian byte order.
|
||||||
|
- The annotation applies to the entire schema and all nested types.
|
||||||
|
- Mixed endianness within one schema is not supported (pathological; no
|
||||||
|
known protocol requires it).
|
||||||
|
- The default is little-endian, matching safetensors, wgpu, and most
|
||||||
|
modern formats. SFTP consumers specify `"endian": "big"`.
|
||||||
|
|
||||||
|
### 2. Alignment
|
||||||
|
|
||||||
|
**Both struct-level and field-level, with field-level overriding
|
||||||
|
struct-level.**
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"TypeDef:Struct": true,
|
||||||
|
"align": 256,
|
||||||
|
"properties": {
|
||||||
|
"header": { "TypeDef:Struct": true, "properties": { ... } },
|
||||||
|
"weight": { "TypeDef:Float32": true, "align": 16 }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- Struct-level `"align"` sets the default alignment for all fields in
|
||||||
|
that struct. The struct's total size is rounded up to this alignment.
|
||||||
|
- Field-level `"align"` overrides the struct default for that specific
|
||||||
|
field.
|
||||||
|
- Default alignment (when no annotation is present): 1 for u8/bool, 2
|
||||||
|
for u16/i16, 4 for u32/i32/f32, 8 for u64/i64/f64, max field alignment
|
||||||
|
for structs.
|
||||||
|
- Alignment is only meaningful in aligned static mode (ADR-096). In
|
||||||
|
packed sequential mode, alignment annotations are ignored — fields are
|
||||||
|
packed with no padding.
|
||||||
|
|
||||||
|
### 3. Variable-length encoding
|
||||||
|
|
||||||
|
**Three strategies for variable-length types, selected by the `encoding`
|
||||||
|
annotation and the standard JSON Schema `maxLength` keyword.**
|
||||||
|
|
||||||
|
```json
|
||||||
|
// Strategy 1: Inline length-prefixing (default, shorthand)
|
||||||
|
{ "TypeDef:String": true }
|
||||||
|
|
||||||
|
// Strategy 1: Explicit inline length-prefixing
|
||||||
|
{ "TypeDef:String": { "encoding": "length-prefixed" } }
|
||||||
|
|
||||||
|
// Strategy 2: Fixed-size reservation (uses standard maxLength)
|
||||||
|
{ "TypeDef:String": true, "maxLength": 256 }
|
||||||
|
|
||||||
|
// Strategy 3: Offset indirection (opt-in)
|
||||||
|
{ "TypeDef:String": { "encoding": "offset-indirect" } }
|
||||||
|
```
|
||||||
|
|
||||||
|
**Strategy 1: Inline length-prefixing (default).** The field's fixed
|
||||||
|
portion is a 4-byte length prefix at a computed offset. The variable data
|
||||||
|
follows immediately after. In packed sequential mode, the length prefix
|
||||||
|
determines the position of subsequent fields. In aligned static mode, the
|
||||||
|
length prefix is at a known offset; the variable data is not included in
|
||||||
|
the static layout. This is the universal pattern used by channels, SFTP,
|
||||||
|
TTY, and most binary protocols.
|
||||||
|
|
||||||
|
**Strategy 2: Fixed-size reservation.** When a variable-length field
|
||||||
|
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
|
||||||
|
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
|
||||||
|
than `maxLength` is zero-padded; data longer than `maxLength` is a
|
||||||
|
validation error. This makes the field fixed-size from the layout
|
||||||
|
perspective — subsequent fields have known, unchanging offsets. This is
|
||||||
|
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
|
||||||
|
pattern for fields with known maximum sizes. In packed sequential mode,
|
||||||
|
`maxLength` is a validation constraint only — the engine still uses
|
||||||
|
inline length-prefixing (strategy 1).
|
||||||
|
|
||||||
|
**Strategy 3: Offset indirection.** The field is a struct
|
||||||
|
`{offset: u32, length: u32}` that points into a separate data region.
|
||||||
|
This is the metatensor blob tensor pattern — the index struct lives in
|
||||||
|
one region, the blob data lives in another. The consumer provides the
|
||||||
|
data region separately. Enables mmap-friendly random access to
|
||||||
|
variable-length data without parsing length prefixes and without
|
||||||
|
reserving worst-case space.
|
||||||
|
|
||||||
|
**Default strategy selection:**
|
||||||
|
- In packed sequential mode: always strategy 1 (inline length-prefixing).
|
||||||
|
`maxLength` is a validation constraint only.
|
||||||
|
- In aligned static mode: strategy 2 (fixed-size reservation) if
|
||||||
|
`maxLength` is declared; strategy 3 (offset indirection) if
|
||||||
|
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
|
||||||
|
length-prefixing) otherwise.
|
||||||
|
|
||||||
|
- `true` is a shorthand for the default (length-prefixed). This keeps
|
||||||
|
the common case concise and the override explicit.
|
||||||
|
- The `encoding` annotation and `maxLength` apply to all variable-length
|
||||||
|
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
|
||||||
|
`TypeDef:Record`, `TypeDef:Timestamp`.
|
||||||
|
|
||||||
|
### 4. TUnion discriminators
|
||||||
|
|
||||||
|
**Two discriminator kinds: byte-offset (protocol dispatch) and
|
||||||
|
field-name (typedef.ts pattern).**
|
||||||
|
|
||||||
|
#### Kind A: Byte-offset discriminator
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"TypeDef:Union": true,
|
||||||
|
"discriminator": {
|
||||||
|
"kind": "byte",
|
||||||
|
"offset": 0,
|
||||||
|
"type": "TypeDef:Uint8"
|
||||||
|
},
|
||||||
|
"mapping": {
|
||||||
|
"1": { "$ref": "#/$defs/Init" },
|
||||||
|
"3": { "$ref": "#/$defs/Open" },
|
||||||
|
"5": { "$ref": "#/$defs/Read" },
|
||||||
|
"6": { "$ref": "#/$defs/Write" },
|
||||||
|
"101": { "$ref": "#/$defs/Status" }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- The discriminator is a fixed-size integer at a known byte offset.
|
||||||
|
- `"offset"` is the byte position of the discriminator within the union's
|
||||||
|
buffer.
|
||||||
|
- `"type"` is the `TypeDef:*` kind of the discriminator (typically
|
||||||
|
`TypeDef:Uint8` for protocol type bytes).
|
||||||
|
- The mapping keys are stringified integers (`"1"`, `"5"`, `"101"`).
|
||||||
|
The engine parses the key to match the discriminator value.
|
||||||
|
- The variant struct starts at `offset + discriminator_size`.
|
||||||
|
- This is the SFTP `Packet` enum pattern and the call protocol's event
|
||||||
|
type dispatch.
|
||||||
|
|
||||||
|
#### Kind B: Field-name discriminator
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"TypeDef:Union": true,
|
||||||
|
"discriminator": {
|
||||||
|
"kind": "field",
|
||||||
|
"name": "type"
|
||||||
|
},
|
||||||
|
"mapping": {
|
||||||
|
"read": { "$ref": "#/$defs/Read" },
|
||||||
|
"write": { "$ref": "#/$defs/Write" }
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- The discriminator is a named field within the struct.
|
||||||
|
- `"name"` is the field name that holds the discriminator value.
|
||||||
|
- The mapping keys are string values matching the discriminator field's
|
||||||
|
value.
|
||||||
|
- The discriminator field is just another field in the struct — its
|
||||||
|
offset is computed like any other field.
|
||||||
|
- This is the typedef.ts `TUnion` pattern.
|
||||||
|
|
||||||
|
#### Mapping values
|
||||||
|
|
||||||
|
Mapping values may be either inline schemas or `$ref` pointers. `$ref`
|
||||||
|
is cleaner for large unions (29 SFTP variants) but requires a `$defs`
|
||||||
|
section. Inline schemas are simpler for small unions (5 call protocol
|
||||||
|
event types). Both work.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
### Positive
|
||||||
|
|
||||||
|
- **Concrete, validated shapes.** All four annotation categories have
|
||||||
|
concrete JSON shapes that were validated by the POCs.
|
||||||
|
- **Sensible defaults.** Little-endian, natural alignment, inline
|
||||||
|
length-prefixing — the common case requires no annotations.
|
||||||
|
- **Explicit overrides.** Big-endian, custom alignment, offset
|
||||||
|
indirection — the uncommon case is explicit and self-documenting.
|
||||||
|
- **TUnion covers both protocol and typedef.ts patterns.** The
|
||||||
|
byte-offset discriminator handles SFTP type bytes and call protocol
|
||||||
|
event types. The field-name discriminator handles the typedef.ts string
|
||||||
|
pattern. No separate union type needed.
|
||||||
|
|
||||||
|
### Negative
|
||||||
|
|
||||||
|
- **Keyword value shape change.** `"TypeDef:String": true` (boolean) and
|
||||||
|
`"TypeDef:String": { "encoding": "length-prefixed" }` (object) are both
|
||||||
|
valid. The engine must handle both shapes. This is a minor parsing
|
||||||
|
concern — the POC already handles it.
|
||||||
|
- **Alignment annotations are mode-specific.** Alignment is only
|
||||||
|
meaningful in aligned static mode. In packed sequential mode, alignment
|
||||||
|
annotations are ignored. This is documented, not enforced — a consumer
|
||||||
|
that specifies alignment in packed mode gets no error, just no effect.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — the
|
||||||
|
annotation shape questions this ADR resolves
|
||||||
|
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||||
|
purpose and scope
|
||||||
|
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||||
|
modes (alignment only meaningful in aligned static mode)
|
||||||
@@ -0,0 +1,157 @@
|
|||||||
|
# ADR-098: Error Handling and Validation Strategy
|
||||||
|
|
||||||
|
## Status
|
||||||
|
Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The typedef engine operates in three phases, each with distinct error
|
||||||
|
conditions:
|
||||||
|
|
||||||
|
1. **Schema parsing** — invalid JSON, missing required keywords, unknown
|
||||||
|
`TypeDef:*` kinds, malformed annotations.
|
||||||
|
2. **Offset computation** — field not found, type not supported for
|
||||||
|
offset computation, recursive schema depth exceeded.
|
||||||
|
3. **Read/write** — buffer too short, invalid UTF-8, value out of range
|
||||||
|
for the target type.
|
||||||
|
4. **Validation** — type constraint violations (range, UTF-8, field
|
||||||
|
presence, discriminator membership).
|
||||||
|
|
||||||
|
The engine also needs a clear strategy for *when* validation happens:
|
||||||
|
once at schema load time (build the validator) vs repeatedly at access
|
||||||
|
time (validate each buffer).
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
### Error type: `TypedefError`
|
||||||
|
|
||||||
|
A single `TypedefError` enum with variants for each error category:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
pub enum TypedefError {
|
||||||
|
/// Schema parsing errors.
|
||||||
|
Schema(String),
|
||||||
|
/// Offset computation errors.
|
||||||
|
Offset { field_path: String, reason: String },
|
||||||
|
/// Read/write errors.
|
||||||
|
Access { field_path: String, reason: String },
|
||||||
|
/// Validation errors (delegated to jsonschema).
|
||||||
|
Validation(ValidationError<'static>),
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
- `Schema` — for invalid JSON, missing required keywords, unknown
|
||||||
|
`TypeDef:*` kinds. The error message describes the problem.
|
||||||
|
- `Offset` — for field-not-found, unsupported type for offset
|
||||||
|
computation, etc. Carries the field path for debugging.
|
||||||
|
- `Access` — for buffer-too-short, invalid UTF-8, value out of range.
|
||||||
|
Carries the field path for debugging.
|
||||||
|
- `Validation` — wraps `jsonschema`'s `ValidationError`. The
|
||||||
|
`jsonschema` crate already provides rich error messages with schema
|
||||||
|
paths; the typedef engine does not re-wrap or re-interpret them.
|
||||||
|
|
||||||
|
The `Validation` variant uses `ValidationError<'static>` because the
|
||||||
|
validator is built once at schema load time and lives for the lifetime
|
||||||
|
of the `TypedefEngine`. The `'static` lifetime is correct — the validator
|
||||||
|
owns its schema reference.
|
||||||
|
|
||||||
|
### Validation timing: load-time build, access-time check
|
||||||
|
|
||||||
|
The jsonschema validator is built once at schema load time
|
||||||
|
(`validator_for(&schema)?`) and then called repeatedly
|
||||||
|
(`validator.is_valid(&instance)`). The typedef engine follows the same
|
||||||
|
pattern:
|
||||||
|
|
||||||
|
1. **Load time:** Parse the schema JSON, build the offset map (or
|
||||||
|
`LayoutBuilder`/`SequentialReader`), build the jsonschema validator.
|
||||||
|
This is the `TypedefEngine::compile(schema: &Value) -> Result<Self,
|
||||||
|
TypedefError>` constructor.
|
||||||
|
2. **Access time:** Use the compiled engine for repeated read/write
|
||||||
|
operations. Validation is opt-in per operation — the consumer calls
|
||||||
|
`engine.validate(buffer)` when validation is desired.
|
||||||
|
|
||||||
|
The `TypedefEngine` struct is the compiled form of a schema:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
pub struct TypedefEngine {
|
||||||
|
offset_map: OffsetMap, // or LayoutBuilder/SequentialReader
|
||||||
|
validator: jsonschema::Validator, // compiled once at load time
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Custom keyword validators
|
||||||
|
|
||||||
|
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
|
||||||
|
`jsonschema::options().with_keyword(...)`. The validators check:
|
||||||
|
|
||||||
|
- **Numeric types** (`TypeDef:Float32`, `TypeDef:Int8`, etc.): range
|
||||||
|
constraints (Int8: -128..127, Uint8: 0..255, etc.), finiteness for
|
||||||
|
floats.
|
||||||
|
- **`TypeDef:String`**: UTF-8 validity.
|
||||||
|
- **`TypeDef:Struct`**: field presence and types (delegated to
|
||||||
|
jsonschema's structural validation — the custom keyword only needs to
|
||||||
|
validate that the struct's fields match their declared `TypeDef:*`
|
||||||
|
kinds).
|
||||||
|
- **`TypeDef:Union`**: discriminator value membership in the mapping.
|
||||||
|
- **`TypeDef:Array`**: element type conformance.
|
||||||
|
- **`TypeDef:Boolean`**: value is `true` or `false`.
|
||||||
|
- **`TypeDef:Timestamp`**: ISO 8601 string format.
|
||||||
|
|
||||||
|
The `jsonschema` crate handles all the structural validation (object
|
||||||
|
properties, required fields, array items, enum values) — the custom
|
||||||
|
keywords only need to validate the leaf type constraints. Each custom
|
||||||
|
keyword implementation is ~10 lines.
|
||||||
|
|
||||||
|
### Read/write errors carry field paths
|
||||||
|
|
||||||
|
Read/write errors include the field path for debugging:
|
||||||
|
|
||||||
|
```rust
|
||||||
|
// Example: reading a u32 from a buffer that's too short
|
||||||
|
Err(TypedefError::Access {
|
||||||
|
field_path: "header.version".to_string(),
|
||||||
|
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
This makes debugging binary format issues tractable — the error tells
|
||||||
|
you exactly which field failed and why.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
### Positive
|
||||||
|
|
||||||
|
- **Single error type.** Consumers handle one `TypedefError` enum, not
|
||||||
|
multiple error types from different engine phases.
|
||||||
|
- **Field-path-carrying errors.** Read/write errors include the field
|
||||||
|
path, making binary format debugging tractable.
|
||||||
|
- **Validation is opt-in.** The consumer decides when to validate.
|
||||||
|
High-throughput paths can skip validation; security-sensitive paths
|
||||||
|
can validate every frame.
|
||||||
|
- **jsonschema integration is clean.** The `ValidationError` is wrapped
|
||||||
|
as-is — no re-interpretation, no information loss.
|
||||||
|
- **Load-time build, access-time use.** The expensive work (schema
|
||||||
|
parsing, validator compilation, offset computation) happens once at
|
||||||
|
load time. Access-time operations are cheap (pointer casts, slice
|
||||||
|
operations, length-prefix reads).
|
||||||
|
|
||||||
|
### Negative
|
||||||
|
|
||||||
|
- **`ValidationError<'static>` lifetime.** The `'static` lifetime on the
|
||||||
|
`Validation` variant means the error cannot borrow from the buffer
|
||||||
|
being validated. This is correct (the validator owns its schema
|
||||||
|
reference) but may surprise readers who expect a shorter lifetime.
|
||||||
|
- **No error recovery.** The engine does not attempt to recover from
|
||||||
|
partial reads or writes. A buffer-too-short error on field N means
|
||||||
|
fields N+1.. are also unreadable. This is inherent to binary formats
|
||||||
|
— there is no "skip to next field" without a schema-driven parser.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — error
|
||||||
|
handling strategy question (OQ 8)
|
||||||
|
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||||
|
purpose and scope
|
||||||
|
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||||
|
modes
|
||||||
|
- [ADR-097](097-schema-annotations.md) — schema annotations
|
||||||
@@ -211,6 +211,14 @@ Door type is separate from whether a decision is made. A two-way door is a decis
|
|||||||
| [OQ-66](questions/066-alknet-register-wire-protocol.md) | `alknet/register` Wire Protocol | deferred(scope) | one | med |
|
| [OQ-66](questions/066-alknet-register-wire-protocol.md) | `alknet/register` Wire Protocol | deferred(scope) | one | med |
|
||||||
| [OQ-67](questions/067-iroh-proxy-support.md) | iroh Proxy Support (Direct-Connection Peer Exposure) | resolved | one | med |
|
| [OQ-67](questions/067-iroh-proxy-support.md) | iroh Proxy Support (Direct-Connection Peer Exposure) | resolved | one | med |
|
||||||
|
|
||||||
|
### alknet-typedef
|
||||||
|
|
||||||
|
| OQ | Title | Status | Door | Pri |
|
||||||
|
|----|-------|--------|------|-----|
|
||||||
|
| [OQ-069](questions/069-arrays-of-variable-length-element-structs.md) | Arrays of Variable-Length-Element Structs | deferred(scope) | two | low |
|
||||||
|
| [OQ-070](questions/070-no-std-alloc-support.md) | `no_std` + `alloc` Support | deferred(scope) | two | low |
|
||||||
|
| [OQ-071](questions/071-builder-api-for-schema-construction.md) | Builder API for Schema Construction | deferred(scope) | two | med |
|
||||||
|
|
||||||
## Deferred / Blocked
|
## Deferred / Blocked
|
||||||
|
|
||||||
The safe-exit visibility surface. These questions are parked because the
|
The safe-exit visibility surface. These questions are parked because the
|
||||||
@@ -358,3 +366,37 @@ filtering the tables above.
|
|||||||
[OQ-64](questions/064-client-side-tls-helper.md).
|
[OQ-64](questions/064-client-side-tls-helper.md).
|
||||||
- **Full file**: [OQ-64](questions/064-client-side-tls-helper.md)
|
- **Full file**: [OQ-64](questions/064-client-side-tls-helper.md)
|
||||||
|
|
||||||
|
### OQ-069: Arrays of Variable-Length-Element Structs
|
||||||
|
|
||||||
|
- **Blocked on**: A concrete consumer that needs arrays of structs with
|
||||||
|
variable-length fields, where the elements are interleaved
|
||||||
|
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
|
||||||
|
sequentially rather than use a fixed stride. The SFTP `Name` packet
|
||||||
|
has `Vec<File>` where `File` contains strings, but SFTP serializes
|
||||||
|
this as a sequence of length-prefixed strings (the serde `SeqAccess`
|
||||||
|
pattern), not as an array of fixed-stride structs. Arrays of
|
||||||
|
fixed-size structs are fully supported.
|
||||||
|
- **Priority**: low
|
||||||
|
- **Full file**: [OQ-069](questions/069-arrays-of-variable-length-element-structs.md)
|
||||||
|
|
||||||
|
### OQ-070: `no_std` + `alloc` Support
|
||||||
|
|
||||||
|
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
|
||||||
|
(e.g., a microcontroller running Rust without `std`). The WASM target
|
||||||
|
has `std` available via `wasm-bindgen`. The engine's core (offset
|
||||||
|
computation, read/write) is already allocation-free; the `jsonschema`
|
||||||
|
dependency is the only `alloc` consumer.
|
||||||
|
- **Priority**: low
|
||||||
|
- **Full file**: [OQ-070](questions/070-no-std-alloc-support.md)
|
||||||
|
|
||||||
|
### OQ-071: Builder API for Schema Construction
|
||||||
|
|
||||||
|
- **Blocked on**: A concrete need for programmatic schema construction
|
||||||
|
in Rust. The current consumers (SFTP, metatensor, binary call frames,
|
||||||
|
TTY negotiation) all have schemas that can be hand-written or
|
||||||
|
generated from TypeBox. A builder API would be a fluent Rust API that
|
||||||
|
produces the same JSON Schema structure — it would sit on top of the
|
||||||
|
engine, not inside it.
|
||||||
|
- **Priority**: medium
|
||||||
|
- **Full file**: [OQ-071](questions/071-builder-api-for-schema-construction.md)
|
||||||
|
|
||||||
@@ -103,6 +103,7 @@ alknet-vault (standalone — foundational to ACL: key derivation, identity)
|
|||||||
│ │ ConnectionCredentials + RemoteIdentity per ADR-091)
|
│ │ ConnectionCredentials + RemoteIdentity per ADR-091)
|
||||||
│ ├── alknet-tls TlsServerConfig + TlsClientConfig + FingerprintPinVerifier — shared TLS config across quinn + TCP+TLS + iroh (ADR-082/087; FingerprintPinVerifier per ADR-089 §5)
|
│ ├── alknet-tls TlsServerConfig + TlsClientConfig + FingerprintPinVerifier — shared TLS config across quinn + TCP+TLS + iroh (ADR-082/087; FingerprintPinVerifier per ADR-089 §5)
|
||||||
│ ├── alknet-call CallAdapter on alknet/call, CallClient (spawn_dispatch primary; dial in AlknetClient per ADR-089), OperationRegistry, adapters (no TLS/transport deps)
|
│ ├── alknet-call CallAdapter on alknet/call, CallClient (spawn_dispatch primary; dial in AlknetClient per ADR-089), OperationRegistry, adapters (no TLS/transport deps)
|
||||||
|
│ ├── alknet-typedef Binary struct engine — JSON Schema with TypeDef:* custom keywords → offset map + read/write + validation (ADR-095–098); depends on jsonschema + serde_json only; WASM-clean
|
||||||
│ ├── alknet-channels
|
│ ├── alknet-channels
|
||||||
│ │ ├── alknet-channels-core pure multiplexer (wire format, demux/mux) — ADR-081
|
│ │ ├── alknet-channels-core pure multiplexer (wire format, demux/mux) — ADR-081
|
||||||
│ │ └── alknet-channels-call channel 0 pre-negotiation + lifecycle ops — ADR-081
|
│ │ └── alknet-channels-call channel 0 pre-negotiation + lifecycle ops — ADR-081
|
||||||
|
|||||||
@@ -0,0 +1,20 @@
|
|||||||
|
# OQ-069: Arrays of variable-length-element structs
|
||||||
|
|
||||||
|
- **Origin**: [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md),
|
||||||
|
[crates/typedef/data-access.md](crates/typedef/data-access.md);
|
||||||
|
`docs/research/alknet-typedef/findings.md` §"Problem 3: Nested structs
|
||||||
|
and arrays of structs"
|
||||||
|
- **Status**: deferred(scope)
|
||||||
|
- **Door type**: Two-way (additive — the engine can add lazy walking
|
||||||
|
logic without changing the existing fixed-stride array support)
|
||||||
|
- **Priority**: low
|
||||||
|
- **Impacts**: Blocks any protocol with interleaved variable-length struct arrays (e.g., a protocol where each array element has a string field and elements are packed as `[fixed_0][str_0][fixed_1][str_1]...`). Does NOT block SFTP `Name` packet handling — SFTP serializes this as a sequence of length-prefixed strings (the serde `SeqAccess` pattern), not as an array of fixed-stride structs. Does NOT block any current consumer.
|
||||||
|
- **Blocked on**: A concrete consumer that needs arrays of structs with
|
||||||
|
variable-length fields, where the elements are interleaved
|
||||||
|
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
|
||||||
|
sequentially rather than use a fixed stride.
|
||||||
|
- **Resolution**: Not yet decidable. The mechanism (lazy sequential
|
||||||
|
walking of array elements, reading each element's length prefixes to
|
||||||
|
find the next element's start) is understood but not needed by any
|
||||||
|
current consumer. Arrays of fixed-size structs are fully supported.
|
||||||
|
- **Cross-references**: ADR-096, [layout-engine.md](crates/typedef/layout-engine.md)
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
# OQ-070: `no_std` + `alloc` support
|
||||||
|
|
||||||
|
- **Origin**: [crates/typedef/overview.md](crates/typedef/overview.md);
|
||||||
|
`docs/research/alknet-typedef/findings.md` §"Open Questions" (OQ 6)
|
||||||
|
- **Status**: deferred(scope)
|
||||||
|
- **Door type**: Two-way (additive — can be added as a feature gate
|
||||||
|
without changing the existing `std` API)
|
||||||
|
- **Priority**: low
|
||||||
|
- **Impacts**: Blocks embedded/WASM-bare-metal deployment targets (microcontrollers, `no_std` environments). Does NOT block WASM-browser (has `std` via `wasm-bindgen`). Does NOT block any current deployment target.
|
||||||
|
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
|
||||||
|
(e.g., a microcontroller running Rust without `std`).
|
||||||
|
- **Resolution**: Not yet decidable. Target `std` for v1. If embedded
|
||||||
|
use cases emerge, `no_std` + `alloc` can be added as a feature gate
|
||||||
|
later. The engine's core is already allocation-free; the `jsonschema`
|
||||||
|
dependency is the only `alloc` consumer.
|
||||||
|
- **Cross-references**: ADR-095
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
# OQ-071: Builder API for schema construction
|
||||||
|
|
||||||
|
- **Origin**: [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md),
|
||||||
|
[crates/typedef/overview.md](crates/typedef/overview.md);
|
||||||
|
`docs/research/alknet-typedef/findings.md` (the builder API was noted
|
||||||
|
as the one detail not covered by the POCs)
|
||||||
|
- **Status**: deferred(scope)
|
||||||
|
- **Door type**: Two-way (additive — a builder API can be added without
|
||||||
|
changing the existing JSON-consumption path)
|
||||||
|
- **Priority**: medium
|
||||||
|
- **Impacts**: No current consumer. Schemas are authored in TypeBox (JS)
|
||||||
|
or hand-written JSON for v1. A Rust builder API would enable
|
||||||
|
programmatic schema construction in Rust without depending on a JS
|
||||||
|
toolchain, but no current consumer needs this.
|
||||||
|
- **Blocked on**: A concrete need for programmatic schema construction
|
||||||
|
in Rust. The current consumers (SFTP, metatensor, binary call frames,
|
||||||
|
TTY negotiation) all have schemas that can be hand-written or generated
|
||||||
|
from TypeBox.
|
||||||
|
- **Resolution**: Not yet decidable. The builder API is important but
|
||||||
|
not needed for the initial consumers. The engine's JSON-consumption
|
||||||
|
path is the primary interface for v1. A builder API would be a fluent
|
||||||
|
Rust API that produces the same JSON Schema structure — it would sit
|
||||||
|
on top of the engine, not inside it.
|
||||||
|
- **Cross-references**: ADR-095, [schema-layer.md](crates/typedef/schema-layer.md)
|
||||||
Reference in new issue
Block a user