docs(typedef): sync specs with the implemented API surface

The code introduced concrete public types and a unified API during
implementation that the specs described only conceptually. Sync the
specs to match the code:

- schema-layer.md: document the TypeDefKind enum and its inherent methods;
  add a Schema-Layer Public API section (get_typedef_kind vs
  get_typedef_kind_loose, annotation parsers, Endian/VariableEncoding/
  DiscriminatorKind, normalize_refs/resolve_ref/resolve_ref_or_inline);
  fix the TRecord layout (values are encoded by their declared kind, not
  universally value_len-prefixed); fix the alignment default list.
- layout-engine.md: document the LayoutMode enum and the ByteRange/
  FieldPosition/PackedLayout/OffsetMap public types with their actual
  signatures; update the LayoutBuilder/SequentialReader/OffsetMap
  component descriptions with the real new/build/compute signatures.
- data-access.md: document the FieldValue enum; add a Higher-level
  read/write section (TypedefEngine::read_field/write_field,
  SequentialReader::read_next/read_field); fix primitive signatures to
  include field_path and Endian; replace the wrong
  read_union_discriminator pseudo-code with the actual tunion module API
  (read_byte_discriminator/read_field_discriminator/resolve_variant/
  discriminator_size) and the UnionDispatch struct.
- validation.md: fix the TypedefEngine struct (add endian/schema fields,
  mark Layout as private); add the real compile signature (&mut Value,
  LayoutMode) and mode-appropriate accessors; fix engine.validate(buffer)
  -> engine.validate_json(&Value)/is_valid_json (the validator operates
  on serde_json::Value, not byte buffers — matches ADR-098).
- overview.md: remove the stale ~1,900 lines / 26 tests line count.
- ADR-097 §3a: correct the TRecord layout (no separate value_len prefix;
  the value is encoded by its declared TypeDef:* kind).
This commit is contained in:
glm-5.2 committed 2026-07-21 13:40:12 +00:00
1 parent 14d9cf281f
commit 7603f98799
7 files changed
+544 -279

No files matched your search

+1 -1
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-20
last_updated: 2026-07-21
---
# alknet-typedef
+233 -194
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-20
last_updated: 2026-07-21
---
# alknet-typedef — Data Access
@@ -11,6 +11,48 @@ variable-length types. This is the consumer-facing API — given a compiled
`TypedefEngine` and a byte buffer, read and write fields at
schema-computed offsets.
This document covers two layers:
- **Primitive read/write functions** in the `data_access` module —
typed reads/writes at a caller-provided offset. These are the building
blocks used by the layout types (`OffsetMap`, `LayoutBuilder`,
`SequentialReader`) and the `TypedefEngine`. Each operates on a raw
byte buffer at a known offset and returns a `TypedefError::Access`
carrying the field path on bounds or encoding failures.
- **The `FieldValue` enum and the higher-level APIs** —
`TypedefEngine::read_field`/`write_field` (aligned mode) and
`SequentialReader::read_next`/`read_field` (packed mode) — which look
up a field's offset via the layout and dispatch to the primitive
functions, returning a unified `FieldValue<'a>`.
## The `FieldValue` enum
The higher-level read APIs return a single unified type — `FieldValue<'a>`
— so one method can read any field kind without the caller dispatching on
schema kind first. The variant carries the typed value; the lifetime
borrows from the input buffer for variable-length kinds (zero-copy).
```rust
pub enum FieldValue<'a> {
I8(i8), I16(i16), I32(i32),
U8(u8), U16(u16), U32(u32),
F32(f32), F64(f64),
Bool(bool),
Enum(u32), // u32 index into the schema's "enum" array
String(&'a str), // borrows from the buffer
Bytes(&'a [u8]), // borrows from the buffer
Struct { start: usize, end: usize }, // consumer recurses with a fresh reader
Union { discriminator: String, variant_start: usize },
Array { count: u32, element_start: usize, element_stride: usize },
}
```
For composite kinds (`Struct`, `Union`, `Array`), `FieldValue` returns a
layout descriptor, not the decoded contents — the consumer recurses with
a fresh `SequentialReader` (or a sub-range read) scoped to the reported
byte range. `Array`'s `element_stride` is `0` for variable-length element
types, signalling the consumer must walk each element sequentially.
## Read/Write Model
The typedef engine operates on raw byte buffers (`&[u8]` for reading,
@@ -19,48 +61,100 @@ reflection, no dynamic dispatch per field. The engine uses the offset map
(or `LayoutBuilder`/`SequentialReader`) to locate fields, then performs
typed access at the computed positions.
### Fixed-size types
### Higher-level read/write
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, `TEnum`, etc.) are accessed via
zero-copy pointer casts:
The `TypedefEngine` and `SequentialReader` provide the primary
consumer-facing read/write APIs. They look up a field's offset via the
layout and dispatch to the primitive `data_access` functions, returning
`FieldValue` (read) or accepting `&FieldValue` (write).
```rust
// Read a u32 at a known offset
fn read_u32(buffer: &[u8], offset: usize, endian: Endian) -> u32 {
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
}
impl TypedefEngine {
// Aligned mode: looks up the field's ByteRange in the OffsetMap,
// dispatches to the right data_access function by TypeDefKind.
// Returns TypedefError::Access if compiled in packed mode
// (use sequential_reader() for packed mode).
pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str)
-> Result<FieldValue<'a>, TypedefError>;
pub fn write_field(&self, buffer: &mut [u8], field_path: &str,
value: &FieldValue<'_>) -> Result<(), TypedefError>;
}
// Write a u32 at a known offset
fn write_u32(buffer: &mut [u8], offset: usize, value: u32, endian: Endian) {
impl SequentialReader {
// Packed mode: walks the buffer field-by-field, reading length
// prefixes to find each field's position. read_field walks all
// preceding fields to reach the target.
pub fn read_next<'a>(&mut self, buffer: &'a [u8])
-> Result<Option<(String, FieldValue<'a>)>, TypedefError>;
pub fn read_field<'a>(&mut self, buffer: &'a [u8], field_path: &str)
-> Result<FieldValue<'a>, TypedefError>;
pub fn reset(&mut self);
pub fn position(&self) -> usize;
pub fn endian(&self) -> Endian;
}
```
`read_field`/`write_field` on `TypedefEngine` work for the fixed-size
primitive kinds and the length-prefixed `String`/`Bytes`/`Timestamp`
fields. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a
`FieldValue` carrying a layout descriptor (byte range, variant start,
or array stride) for the consumer to recurse on — see §"FieldValue" above.
For writing in packed mode, the consumer uses `LayoutBuilder::build` to
compute positions, then calls the primitive `data_access::write_*`
functions at the computed offsets. There is no packed-mode
`engine.write_field` — the layout depends on the actual data sizes,
which the builder consumes at `build` time.
### Primitive read/write functions
The `data_access` module exposes typed read/write functions for each
primitive kind. Each takes `field_path: &str` for error attribution
(produces a `TypedefError::Access` carrying the path on bounds or
encoding failures) and, for multi-byte types, an `Endian` parameter.
### Fixed-size types
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, `TEnum`, etc.) are
accessed via zero-copy reads of N bytes at the offset:
```rust
// Read a u32 at a known offset, applying endianness. Bounds-checked.
fn read_u32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
-> Result<u32, TypedefError> {
let bytes: [u8; 4] = read_array(buffer, offset, field_path)?;
Ok(match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
})
}
// Write a u32 at a known offset, applying endianness. Bounds-checked.
fn write_u32(buffer: &mut [u8], offset: usize, value: u32,
field_path: &str, endian: Endian) -> Result<(), TypedefError> {
let bytes = match endian {
Endian::Little => value.to_le_bytes(),
Endian::Big => value.to_be_bytes(),
};
buffer[offset..offset+4].copy_from_slice(&bytes);
write_array(buffer, offset, bytes, field_path)
}
```
The engine applies endianness at access time based on the schema's
`"endian"` annotation (ADR-097). The offset computation is
endian-agnostic.
endian-agnostic. The `read_array`/`write_array` helpers perform the
bounds check and produce `TypedefError::Access` with the field path on
failure.
### TEnum access
`TEnum` is a fixed-size type (4 bytes, `u32` index). Read/write follows
the same pattern as other fixed-size types — the engine reads/writes a
`u32` at the field's computed offset, applying the schema's endianness:
`TEnum` is a fixed-size type (4 bytes, `u32` index). Read/write delegates
to the `u32` primitives, applying the schema's endianness:
```rust
fn read_enum(buffer: &[u8], offset: usize, endian: Endian) -> u32 {
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
}
pub fn read_enum(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
-> Result<u32, TypedefError> {
read_u32(buffer, offset, field_path, endian)
}
```
@@ -72,43 +166,28 @@ corresponds to a valid enum value at the JSON level.
### Variable-length types (inline length-prefixing)
For variable-length types with inline length-prefixing (the default):
For variable-length types with inline length-prefixing (the default),
the `data_access` module provides `read_string`/`write_string`/
`read_bytes`/`write_bytes`. Each takes `field_path: &str` for error
attribution and `endian` for the length prefix:
```rust
// Read a length-prefixed string
fn read_string<'a>(buffer: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
let len_bytes: [u8; 4] = buffer[offset..offset+4].try_into()
.map_err(|_| TypedefError::Access { /* ... */ })?;
let len = match endian {
Endian::Little => u32::from_le_bytes(len_bytes),
Endian::Big => u32::from_be_bytes(len_bytes),
} as usize;
let data = buffer.get(offset+4..offset+4+len)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
std::str::from_utf8(data)
.map_err(|e| TypedefError::Access { /* ... */ })
}
// Read a length-prefixed string, borrowing from the buffer.
fn read_string<'a>(buffer: &'a [u8], offset: usize,
field_path: &str, endian: Endian) -> Result<&'a str, TypedefError>;
// Write a length-prefixed string
fn write_string(buffer: &mut [u8], offset: usize, value: &str, endian: Endian) -> Result<(), TypedefError> {
let data = value.as_bytes();
let len_bytes = match endian {
Endian::Little => (data.len() as u32).to_le_bytes(),
Endian::Big => (data.len() as u32).to_be_bytes(),
};
buffer.get_mut(offset..offset+4)
.ok_or_else(|| TypedefError::Access { /* ... */ })?
.copy_from_slice(&len_bytes);
buffer.get_mut(offset+4..offset+4+data.len())
.ok_or_else(|| TypedefError::Access { /* ... */ })?
.copy_from_slice(data);
Ok(())
}
// Write a length-prefixed string. Returns total bytes written (4 + data.len()).
fn write_string(buffer: &mut [u8], offset: usize, value: &str,
field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
// read_bytes / write_bytes have the same shape — raw bytes, no UTF-8 check.
```
The engine reads the 4-byte length prefix at the field's offset, then
slices the data that follows. For writing, the engine writes the length
prefix + data.
prefix + data. `read_string` validates UTF-8 and returns a `&str`
borrowing from the input buffer (zero-copy); `read_bytes` returns a
`&[u8]` slice with no encoding check.
In packed sequential mode, the `SequentialReader` uses the length prefix
to determine the position of the next field. In aligned static mode, the
@@ -117,24 +196,19 @@ is accessed separately.
### Variable-length types (offset indirection)
For variable-length types with offset indirection (opt-in):
For variable-length types with offset indirection (opt-in), the
`data_access` module provides `read_string_indirect`/`read_bytes_indirect`.
The 8-byte struct at `buffer[offset..offset+8]` is
`{ data_offset: u32, data_length: u32 }` (endian-aware); the actual
bytes live in a separate `data_region`:
```rust
// Read an offset-indirect string
fn read_string_indirect<'a>(data_region: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
let ptr_offset = match endian {
Endian::Little => u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()),
Endian::Big => u32::from_be_bytes(data_region[offset..offset+4].try_into().unwrap()),
} as usize;
let ptr_length = match endian {
Endian::Little => u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
Endian::Big => u32::from_be_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
} as usize;
let data = data_region.get(ptr_offset..ptr_offset+ptr_length)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
std::str::from_utf8(data)
.map_err(|e| TypedefError::Access { /* ... */ })
}
fn read_string_indirect<'a>(buffer: &'a [u8], offset: usize,
data_region: &'a [u8], field_path: &str,
endian: Endian) -> Result<&'a str, TypedefError>;
fn read_bytes_indirect<'a>(buffer: &'a [u8], offset: usize,
data_region: &'a [u8], field_path: &str,
endian: Endian) -> Result<&'a [u8], TypedefError>;
```
The field is a struct `{offset: u32, length: u32}` at a known position
@@ -143,157 +217,122 @@ engine reads the offset and length, then slices the data region.
## TUnion Dispatch
TUnion dispatch reads the discriminator value, looks up the variant
schema, and then reads the variant's fields. The dispatch mechanism
differs by discriminator kind (ADR-097).
The `tunion` module provides TUnion discriminator dispatch — reading the
discriminator value from a byte buffer, looking up the variant schema in
the union's `mapping`, and reporting the offset where the variant struct
begins. All reads go through the `data_access` primitives so bounds checks
and endianness handling are uniform with the rest of the engine.
The result of dispatch is a `UnionDispatch` struct:
```rust
pub struct UnionDispatch {
pub key: String, // mapping key (stringified disc value)
pub variant_offset: usize, // byte offset where the variant struct starts
pub discriminator_size: usize, // discriminator's byte size
}
```
After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)`
to get the variant schema, then reads the variant's fields at
`dispatch.variant_offset` using the normal `data_access` functions (or a
fresh `SequentialReader` scoped to the variant).
### Byte-offset discriminator
```rust
/// Read the discriminator value from a byte-offset TUnion.
/// Returns the mapping key (as a string) so the consumer can look up
/// the variant schema and read the variant's fields.
fn read_union_discriminator(
/// Read the discriminator value from a byte-offset TUnion. The discriminator
/// is a fixed-size integer (TypeDef:Uint8/Uint16/Uint32) at a known byte
/// offset. Returns the mapping key (stringified integer) and the variant
/// struct offset.
pub fn read_byte_discriminator(
buffer: &[u8],
schema: &Value,
union_schema: &Value,
endian: Endian,
) -> Result<String, TypedefError> {
let disc = schema["discriminator"].as_object()
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
let offset = disc["offset"].as_u64().unwrap_or(0) as usize;
let disc_type = disc["type"].as_str().unwrap_or("TypeDef:Uint8");
let (disc_value, disc_size) = match disc_type {
"TypeDef:Uint8" => {
let b = *buffer.get(offset)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
(b as u32, 1)
}
"TypeDef:Uint16" => {
let bytes: [u8; 2] = buffer[offset..offset+2].try_into().unwrap();
let v = match endian {
Endian::Little => u16::from_le_bytes(bytes),
Endian::Big => u16::from_be_bytes(bytes),
};
(v as u32, 2)
}
"TypeDef:Uint32" => {
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
let v = match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
};
(v, 4)
}
_ => return Err(TypedefError::Schema(format!("unsupported discriminator type: {disc_type}"))),
};
let key = disc_value.to_string();
if schema["mapping"].as_object().map_or(false, |m| m.contains_key(&key)) {
Ok(key)
} else {
Err(TypedefError::Access {
field_path: "__discriminator".into(),
reason: format!("unknown discriminator value: {disc_value}"),
})
}
}
) -> Result<UnionDispatch, TypedefError>;
```
The discriminator is a fixed-size integer at a known byte offset. The
mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`. After reading the discriminator, the
consumer looks up the variant schema and reads the variant's fields
using the normal typed read functions (e.g., `read_u32`, `read_string`)
at `offset + discriminator_size`.
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
1..N are the variant struct. The call protocol's 5 event types
(`call.requested` → 0x01, etc.) use the same pattern.
(`call.requested` → 0x01, etc.) use the same pattern. The variant struct
starts at `offset + discriminator_size`.
### Field-name discriminator
```rust
/// Read the discriminator value from a field-name TUnion.
/// The discriminator is a named field within the struct — its offset
/// is computed like any other field. The consumer reads the field's
/// value, looks up the variant schema, then reads the variant's fields.
fn read_union_field_discriminator(
/// Read the discriminator value from a field-name TUnion. The
/// discriminator is a named field within the struct — the consumer
/// provides the field's computed offset (from the OffsetMap or
/// LayoutBuilder). Supports TypeDef:String, Uint8, and Enum discriminator
/// fields.
pub fn read_field_discriminator(
buffer: &[u8],
schema: &Value,
offset_map: &OffsetMap,
union_schema: &Value,
disc_field_offset: usize,
endian: Endian,
) -> Result<String, TypedefError> {
let disc = schema["discriminator"].as_object()
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
let field_name = disc["name"].as_str()
.ok_or_else(|| TypedefError::Schema("discriminator has no 'name'".into()))?;
// Read the discriminator field at its computed offset.
// The field's TypeDef kind determines how to read it (typically a string).
let field_schema = schema["properties"].get(field_name)
.ok_or_else(|| TypedefError::Schema(format!("discriminator field '{field_name}' not found")))?;
let kind = get_typedef_kind(field_schema)
.ok_or_else(|| TypedefError::Schema("discriminator field has no TypeDef kind".into()))?;
match kind {
"TypeDef:String" => {
let range = offset_map.get(field_name)
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
read_string(buffer, range.start, endian)
.map(|s| s.to_string())
}
"TypeDef:Uint8" => {
let range = offset_map.get(field_name)
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
Ok(buffer[range.start].to_string())
}
_ => Err(TypedefError::Schema(format!(
"unsupported discriminator field type: {kind}"
))),
}
}
) -> Result<UnionDispatch, TypedefError>;
```
The discriminator is a named field within the struct. Its offset is
computed like any other field. The mapping keys are string values.
After reading the discriminator, the consumer looks up the variant
schema and reads the variant's fields starting at the end of the
discriminator field (or at the start of the union buffer if the
discriminator is the first field).
computed like any other field (the consumer passes it in as
`disc_field_offset`). The mapping keys are string values. After reading
the discriminator, the consumer looks up the variant schema and reads
the variant's fields starting at the end of the discriminator field.
### Variant resolution
```rust
/// Look up a variant schema from the union's mapping. Inline schemas
/// are returned directly. $ref pointers of the form "#/$defs/<name>"
/// are resolved against the union schema's own $defs block.
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str)
-> Result<&'a Value, TypedefError>;
/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a
/// byte-offset TUnion. Field-name discriminators have no fixed size
/// and produce a TypedefError::Schema.
pub fn discriminator_size(union_schema: &Value) -> Result<usize, TypedefError>;
```
### TUnion in the layout engines
The `LayoutBuilder` and `SequentialReader` also handle TUnion fields
inline during traversal (the consumer does not need to call the `tunion`
functions for a union field reached during a sequential walk). For
`LayoutBuilder`, the consumer supplies the discriminator value (byte-offset)
or variant index (field-name) in `var_sizes` under the synthetic key
`"<union_path>.__discriminator"` or `"<union_path>.__variant"`. For
`SequentialReader`, a union field yields
`FieldValue::Union { discriminator, variant_start }`. The standalone
`tunion` functions are for dispatch outside the layout walk — e.g., a
consumer that receives a bare union buffer and needs to identify the
variant before recursing.
## Field Paths
Fields are addressed by dotted paths: `"header.version"`, `"payload.data"`.
The `OffsetMap` stores fully-qualified paths. The read/write functions
accept a field path and look up the byte range:
Both `OffsetMap` and `PackedLayout` store fully-qualified paths (nested
struct fields appear under their parent's path prefix). The higher-level
APIs (`TypedefEngine::read_field`/`write_field`, `SequentialReader::read_field`)
accept a field path, look up the byte range/position in the layout, and
dispatch to the primitive `data_access` function for the field's kind.
```rust
/// Read an f32 field by path. This is the aligned-mode path (uses OffsetMap).
/// In packed mode, the consumer uses SequentialReader instead.
fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> {
let range = self.offset_map.get(field_path)
.ok_or_else(|| TypedefError::Offset {
field_path: field_path.to_string(),
reason: "field not found in offset map".to_string(),
})?;
if buffer.len() < range.end {
return Err(TypedefError::Access {
field_path: field_path.to_string(),
reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()),
});
}
let bytes: [u8; 4] = buffer[range.start..range.end].try_into().unwrap();
Ok(match self.endian {
Endian::Little => f32::from_le_bytes(bytes),
Endian::Big => f32::from_be_bytes(bytes),
})
}
```
For aligned-mode access, `TypedefEngine::read_field(&buffer, "header.version")`
returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds
the field's `TypeDef:*` kind in the schema, and calls the matching
`data_access::read_*` function. `write_field` is the mirror. Composite
kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue`
carrying a layout descriptor; the consumer recurses with a fresh reader
or sub-range read.
For packed-mode access, `SequentialReader::read_field(&buffer, "c")` walks
all preceding fields to reach the target (sequential access is inherent
to packed layouts). `read_next` walks fields in declaration order.
Nested structs produce nested field paths. The offset computation
propagates the field path prefix during recursion, so the `OffsetMap`
contains entries like `"header.version"` and `"header.magic"`.
and `PackedLayout` contain entries like `"header.version"` and
`"header.magic"`.
## Zero-Copy Access
+108 -24
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-20
last_updated: 2026-07-21
---
# alknet-typedef — Layout Engine
@@ -24,27 +24,21 @@ protocols.
**Components:**
- **`LayoutBuilder`** — takes a schema and actual data sizes for
variable-length fields, computes byte positions for each field in a
packed layout. Used at write time when the consumer knows the data
sizes upfront.
- **`SequentialReader`** — walks a buffer field-by-field according to the
schema, reading length prefixes to determine variable-length data
positions. Used at read time when the consumer is parsing an incoming
frame.
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `TypeDef:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, TypedefError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, TypedefError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
**How it works:**
For a struct with fields `[u8, u32, string]`:
For a struct with fields `[u8, u32, string]` where the string is 10 bytes:
```
LayoutBuilder (write):
LayoutBuilder::build(var_sizes: {"payload": 10}):
field[0] u8: offset 0, size 1
field[1] u32: offset 1, size 4
field[2] string: offset 5, size 4 (length prefix) + data_len
total: 9 + data_len
field[2] string: offset 5, size 4 (length prefix) + 10 (data)
total: 19
SequentialReader (read):
SequentialReader::read_next (read):
read u8 at offset 0
read u32 at offset 1
read u32 length prefix at offset 5 → data_len
@@ -75,10 +69,7 @@ and safetensors.
**Component:**
- **`OffsetMap`** — walks the schema once, computes fixed byte positions
for each field based on type sizes and alignment. The output is a flat
table of `(field_path, byte_range)` pairs. Used for both read and write
at known offsets.
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, TypedefError>` (requires `TypeDef:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
**How it works:**
@@ -199,7 +190,10 @@ annotation shapes).
Nested structs produce dotted field paths: `header.version`,
`header.magic`. The offset computation propagates the field path prefix
during recursion. The `OffsetMap` stores fully-qualified paths.
during recursion. Both `OffsetMap` and `PackedLayout` store fully-qualified
paths; the `iter()` method of each yields fields in schema `properties`
order, with nested struct fields appearing inline under their parent's
path prefix.
### Endianness
@@ -212,13 +206,25 @@ and byte-swaps accordingly. All fixed-size types — including `TEnum`
## Mode Selection
The consumer selects the mode at engine construction time. The choice is
determined by the use case, not by the schema:
The consumer selects the mode at engine construction time via the
`LayoutMode` enum, passed to `TypedefEngine::compile`:
```rust
pub enum LayoutMode {
/// Packed sequential — for protocol wire formats (SFTP, channels, TTY).
Packed,
/// Aligned static — for mmap-friendly formats (metatensor, safetensors).
Aligned,
}
```
The choice is determined by the use case, not by the schema:
- **Protocol consumer** (SFTP, binary call frames, TTY negotiation):
uses `LayoutBuilder` for writing and `SequentialReader` for reading.
- **mmap consumer** (metatensor): uses `OffsetMap` for both reading and
writing.
`LayoutMode::Packed` → uses `LayoutBuilder` for writing and
`SequentialReader` for reading.
- **mmap consumer** (metatensor): `LayoutMode::Aligned` → uses `OffsetMap`
for both reading and writing at known offsets.
The same schema can be used in either mode. A schema describing an SFTP
packet can be consumed by a `SequentialReader` (for parsing incoming
@@ -226,6 +232,84 @@ frames) and a `LayoutBuilder` (for constructing outgoing frames). A schema
describing a metatensor layout can be consumed by an `OffsetMap` (for
mmap access).
`TypedefEngine` exposes mode-appropriate accessors: `engine.offset_map()`
returns `Some` in aligned mode and `None` in packed mode;
`engine.layout_builder()` and `engine.sequential_reader()` return `Some`
in packed mode and `None` in aligned mode. See [validation.md](validation.md)
§"The TypedefEngine struct" for the engine API.
## Public Types
The layout engine produces three public types, one per layout component.
All are re-exported from the crate root.
### `ByteRange` (aligned mode)
```rust
pub struct ByteRange {
pub start: usize, // inclusive
pub end: usize, // exclusive
}
```
A half-open byte range produced by `OffsetMap::compute` for each field.
`end - start` is the field's byte size in the static layout (for
variable-length fields: the length prefix, the `{offset, length}` pair,
or the `maxLength` reservation — not the variable data). `ByteRange`
provides `len()` and `is_empty()`.
### `FieldPosition` (packed mode)
```rust
pub struct FieldPosition {
pub offset: usize,
pub size: usize,
pub kind: TypeDefKind,
}
```
A field's computed position in a packed layout, produced by
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
length prefix); for fixed-size fields, `size` is the type's byte size.
`kind` records the field's `TypeDef:*` kind so the consumer can dispatch
to the correct `data_access` read/write function.
### `PackedLayout` (packed mode)
The result of `LayoutBuilder::build`: a map of `field_path → FieldPosition`
plus the total buffer size needed.
```rust
impl PackedLayout {
pub fn get(&self, field_path: &str) -> Option<&FieldPosition>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = &(String, FieldPosition)>;
}
```
`get` looks up a field by dotted path. For TUnion byte-offset
discriminators, the discriminator is recorded under the synthetic path
`"<union_path>.__discriminator"`. `iter` yields fields in layout order
(schema `properties` order, with nested struct fields appearing inline
under their parent's path prefix).
### `OffsetMap` (aligned mode)
A flat table of `(field_path, byte_range)` pairs computed from a schema.
```rust
impl OffsetMap {
pub fn compute(schema: &Value) -> Result<Self, TypedefError>;
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
}
```
`compute` requires a `TypeDef:Struct` at the top level. `total_size`
includes trailing alignment padding. `iter` yields fields in insertion
order (schema `properties` order, nested struct fields appearing inline).
## Design Decisions
| Decision | ADR | Summary |
+6 -6
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-20
last_updated: 2026-07-21
---
# alknet-typedef — Overview
@@ -31,12 +31,12 @@ with `TypeDef:*` custom keywords (the same kinds defined in TypeBox's
The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing). The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are ~10 lines each.
each field. The custom keyword implementations are small (a few lines
each, generated from shared macros — see [validation.md](validation.md)).
The crate is ~1,900 lines (POC verified, 26 tests passing). It replaces
two prior attempts that built their own jsonschema engines — typebox-rs
(~8,400 lines) and alktype (~5,600 lines) — with `jsonschema` + an
offset map + ~50 lines of custom keyword implementations. See
The crate replaces two prior attempts that built their own jsonschema
engines — typebox-rs (~8,400 lines) and alktype (~5,600 lines) — with
`jsonschema` + an offset map + small custom keyword implementations. See
[ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
## Why
+128 -25
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-20
last_updated: 2026-07-21
---
# alknet-typedef — Schema Layer
@@ -37,6 +37,42 @@ encoding strategy (for variable-length types).
| `TRecord` | `TypeDef:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
| `TTimestamp` | `TypeDef:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
### The `TypeDefKind` enum
The engine represents the 17 kinds as a Rust enum — `TypeDefKind` — with
one variant per kind (`TypeDefKind::Float32`, `TypeDefKind::Struct`, etc.).
The enum provides compile-time exhaustiveness checking and integer
discriminant dispatch (a jump table) instead of string comparison at
every field access. It is `pub` and re-exported from the crate root.
```rust
pub enum TypeDefKind {
Int8, Int16, Int32,
Uint8, Uint16, Uint32,
Float32, Float64,
Boolean, Enum,
String, Bytes, Timestamp,
Struct, Union, Array, Record,
}
```
The enum carries the kind's binary-layout metadata as inherent methods:
| Method | Returns | Notes |
|--------|---------|-------|
| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"TypeDef:Uint8"` |
| `type_size(self)` | `Option<usize>` | `Some(N)` for fixed-size kinds; `None` for variable/composite |
| `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array |
| `is_fixed_size(self)` | `bool` | True for the 10 fixed-size primitive kinds |
| `is_composite(self)` | `bool` | True for Struct, Union, Array, Record |
| `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record |
| `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter |
`TypeDefKind` implements `Display` (renders the keyword string) and
`FromStr` (parses the keyword string back into the variant, returning
`TypedefError::Schema` for unknown kinds). The layout engines and the
validator dispatch on the enum, not on strings.
### Fixed-size types
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
@@ -132,14 +168,20 @@ distinct from UTF-8 strings. In the binary representation, TBytes is raw
bytes with no encoding (not base64, not hex). In the JSON representation
(for validation), TBytes is a string (JSON has no native byte type).
**`TRecord`:** A string-keyed map. Binary layout is a count-prefixed
sequence of `(key, value)` pairs: `[count: u32][key_len: u32][key_bytes]
[value_len: u32][value_bytes]...`. The count is the number of entries.
Each key is a length-prefixed UTF-8 string. Each value is the record's
declared value type (specified via the `"values"` property in the schema,
e.g., `"values": { "TypeDef:Float32": true }`). The count prefix respects
the schema's endianness. In aligned static mode with `maxLength`, the
entire record is reserved at `maxLength` bytes (zero-padded).
**`TRecord`:** A string-keyed map. The value type is declared via the
schema's `"values"` property (e.g., `"values": { "TypeDef:Float32": true }`).
Binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
The count is the number of entries. Each key is a length-prefixed UTF-8
string. Each value is encoded according to its declared `TypeDef:*` kind
— a `Record<Uint32>` value is 4 raw bytes; a `Record<String>` value is
itself a length-prefixed string; a `Record<Struct>` value is the struct's
fields laid out inline. There is **no separate `value_len` prefix** —
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix). The
count and key-length prefixes respect the schema's endianness. In
aligned static mode with `maxLength`, the entire record is reserved at
`maxLength` bytes (zero-padded).
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
@@ -166,6 +208,74 @@ The count prefix respects the schema's endianness.
fields' sizes (plus alignment padding in aligned static mode). The offset
computation recurses into their properties.
## Schema-Layer Public API
The `schema` module exposes the foundational types and functions every
other module depends on. These are re-exported from the crate root.
### `get_typedef_kind` vs `get_typedef_kind_loose`
The engine recognizes a `TypeDef:*` kind on a schema node two ways,
because the keyword value may be either a boolean (`true`) or an
annotation object (`{ "encoding": "..." }`):
| Function | Recognizes | Returns |
|----------|------------|---------|
| `get_typedef_kind(node) -> Option<&str>` | Boolean form only (`{ "TypeDef:String": true }`) | The keyword string, e.g. `"TypeDef:String"` |
| `get_typedef_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string |
| `get_typedef_kind_enum(node) -> Option<TypeDefKind>` | Boolean form only | The parsed enum variant |
| `get_typedef_kind_loose_enum(node) -> Option<TypeDefKind>` | Boolean form **and** object form | The parsed enum variant |
The boolean-form-only functions are used by the validator factories
(which reject the object form as a schema error) and the top-level
kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new`
(which require `TypeDef:Struct` at the root). The "loose" variants are
used by the layout engines during field traversal, so that a variable-
length field with an `encoding` annotation (`{ "TypeDef:String":
{ "encoding": "offset-indirect" } }`) is still recognized as a `String`.
### Annotation parsers
Each schema-level annotation has a dedicated parser that reads it from a
`serde_json::Value` node and returns a sensible default when absent:
| Function | Annotation | Default |
|----------|------------|---------|
| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` |
| `parse_align(node) -> Option<usize>` | `"align"` | `None` |
| `parse_max_length(node) -> Option<usize>` | `"maxLength"` | `None` |
| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` |
| `parse_discriminator(node) -> Result<DiscriminatorKind, TypedefError>` | `"discriminator"` | (required — returns `TypedefError::Schema` if absent) |
### Public enums
```rust
pub enum Endian { Little, Big }
pub enum VariableEncoding { LengthPrefixed, OffsetIndirect }
pub enum DiscriminatorKind {
Byte { offset: usize, disc_type: TypeDefKind },
Field { name: String },
}
```
`DiscriminatorKind::Byte` carries the byte position (`offset`) and the
discriminator's `TypeDef:*` kind (`disc_type`, restricted to `Uint8`/
`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator
field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for
how these drive dispatch.
### `$ref` resolution and normalization
| Function | Purpose |
|----------|---------|
| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/<name>"`. Idempotent. Runs once at `TypedefEngine::compile` time. |
| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. |
| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). |
`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s
JSON Pointer requirement. The layout engines call `resolve_ref_or_inline`
on every `$ref`-bearing node they encounter during traversal.
## jsonschema Custom Keyword Integration
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
@@ -222,20 +332,12 @@ TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`),
referencing sibling definitions within the same `$defs` block. The
`jsonschema` crate requires full JSON Pointer paths (e.g.,
`"$ref": "#/$defs/Read"`). The typedef engine normalizes TypeBox-style
refs at schema load time:
```rust
fn normalize_refs(schema: &mut Value) {
// Walk the schema tree. For every "$ref" whose value is a bare name
// (no "#" prefix), rewrite it to "#/$defs/<name>".
// "$ref": "Read" → "$ref": "#/$defs/Read"
}
```
This is a ~20-line recursive walk of the schema JSON. It runs once at
load time, before the schema is passed to `jsonschema::validator_for`
or the offset computation. The normalization is idempotent — full JSON
Pointer refs pass through unchanged.
refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization)
— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to
`"#/$defs/<name>"`. The normalization is idempotent — full JSON Pointer
refs pass through unchanged. It runs once at `TypedefEngine::compile`
time, before the schema is passed to `jsonschema` or the offset
computation.
**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs
with `Resource 'Read' is not present in a registry`. Full JSON Pointer
@@ -279,8 +381,9 @@ Both struct-level and field-level, with field-level overriding:
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/bool, 2 for u16/i16, 4 for u32/i32/f32,
8 for u64/i64/f64, max field alignment for structs.
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/
enum, 8 for f64, 4 for variable-length (the u32 length prefix), 1 for
struct/union/array.
- Only meaningful in aligned static mode (ADR-096). Ignored in packed
sequential mode.
+59 -27
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-20
last_updated: 2026-07-21
---
# alknet-typedef — Validation
@@ -50,30 +50,47 @@ process, not a single `validate(buffer)` call.
### The `TypedefEngine` struct
The `TypedefEngine` is the compiled form of a schema. It supports both
layout modes (ADR-096) via an internal enum:
layout modes (ADR-096) via an internal `Layout` enum:
```rust
pub struct TypedefEngine {
layout: Layout, // packed or aligned (see below)
layout: Layout, // packed or aligned (private enum)
validator: jsonschema::Validator, // compiled once at load time
endian: Endian, // parsed from the schema's "endian" annotation
schema: Value, // the normalized schema (refs resolved)
}
// Private — the consumer selects via LayoutMode at compile time.
enum Layout {
Packed {
builder: LayoutBuilder,
reader: SequentialReader,
},
Aligned {
offset_map: OffsetMap,
},
Packed { builder: LayoutBuilder, reader: SequentialReader },
Aligned { offset_map: OffsetMap },
}
```
The consumer selects the mode at construction time. The `Layout` enum
ensures the engine always has the correct layout strategy for the
consumer's use case — a protocol consumer gets `Packed`, an mmap
consumer gets `Aligned`. The validator is mode-agnostic (it operates on
`Value`, not raw bytes).
The consumer selects the mode at construction time via `LayoutMode`
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout`
enum is private — the engine exposes mode-appropriate accessors instead:
```rust
impl TypedefEngine {
pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result<Self, TypedefError>;
pub fn mode(&self) -> LayoutMode;
pub fn endian(&self) -> Endian;
pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode
pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode
pub fn sequential_reader(&self) -> Option<&SequentialReader>; // Some in packed mode
}
```
`compile` takes `&mut Value` because it normalizes `$ref` values in place
(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization))
before computing the layout and building the validator. The `schema`
field retains the normalized schema for `read_field`'s kind lookup. The
validator is mode-agnostic (it operates on `Value`, not raw bytes).
The `read_field`/`write_field` methods on `TypedefEngine` are the
aligned-mode data-access API — see [data-access.md](data-access.md)
§"Higher-level read/write".
## Custom Keyword Validators
@@ -242,18 +259,30 @@ you exactly which field failed and why.
### Load time: `TypedefEngine::compile()`
The expensive work happens once at schema load time:
1. Parse the schema JSON (`serde_json::from_str` with `preserve_order`).
2. Compute the offset map (or `LayoutBuilder`/`SequentialReader`).
3. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
1. Normalize `$ref` values in the schema (`normalize_refs`).
2. Parse the schema's `"endian"` annotation.
3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned).
4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
The result is a `TypedefEngine` that can be used for repeated operations.
### Access time: `engine.validate(buffer)`
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
Validation is opt-in per operation. The consumer calls
`engine.validate(buffer)` when validation is desired. The jsonschema
validator is already compiled — `is_valid()` is a fast check against
the compiled validator.
`engine.validate_json(instance)` when validation is desired, or
`engine.is_valid_json(instance)` for a boolean check. The jsonschema
validator is already compiled — these are fast checks against the
compiled validator.
```rust
pub fn validate_json(&self, instance: &Value) -> Result<(), TypedefError>;
pub fn is_valid_json(&self, instance: &Value) -> bool;
```
The argument is a `serde_json::Value` (the JSON representation of the
data), not a raw byte buffer — see §"What validation validates" above.
To validate a binary buffer end-to-end, the consumer reads it into a
`Value` tree via the data access layer, then validates that `Value`.
High-throughput paths can skip validation. Security-sensitive paths
(parsing incoming frames from untrusted peers) can validate every frame.
@@ -261,16 +290,19 @@ The choice is the consumer's.
## Relationship to Read/Write
Validation and data access are independent operations on the same buffer.
Validation and data access are independent operations on the same data.
The consumer can:
1. Validate a buffer to ensure it conforms to the schema.
2. Read fields from the buffer at computed offsets.
3. Both — validate first, then read (defense in depth).
1. Validate the JSON representation of a buffer to ensure it conforms to
the schema.
2. Read fields from the binary buffer at computed offsets.
3. Both — validate the JSON representation first, then read the binary
buffer (defense in depth).
The engine does not couple validation and access. A consumer that trusts
its data source can skip validation and go straight to read/write. A
consumer that parses untrusted input can validate first, then access.
consumer that parses untrusted input can validate the JSON
representation first, then access the binary buffer.
## Design Decisions
@@ -144,8 +144,15 @@ the `"values"` property in the schema:
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
values in the record. All values share the same type.
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value_len: u32][value_bytes]...`.
- The count prefix respects the schema's endianness.
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count`
times. Each key is a length-prefixed UTF-8 string. Each value is
encoded according to its declared `TypeDef:*` kind — a `Record<Uint32>`
value is 4 raw bytes; a `Record<String>` value is itself a
length-prefixed string; a `Record<Struct>` value is the struct's
fields laid out inline. There is **no separate `value_len` prefix** —
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix).
- The count and key-length prefixes respect the schema's endianness.
- In aligned static mode with `maxLength`, the entire record is reserved
at `maxLength` bytes (zero-padded).