docs(typedef): sync specs with the implemented API surface
The code introduced concrete public types and a unified API during implementation that the specs described only conceptually. Sync the specs to match the code: - schema-layer.md: document the TypeDefKind enum and its inherent methods; add a Schema-Layer Public API section (get_typedef_kind vs get_typedef_kind_loose, annotation parsers, Endian/VariableEncoding/ DiscriminatorKind, normalize_refs/resolve_ref/resolve_ref_or_inline); fix the TRecord layout (values are encoded by their declared kind, not universally value_len-prefixed); fix the alignment default list. - layout-engine.md: document the LayoutMode enum and the ByteRange/ FieldPosition/PackedLayout/OffsetMap public types with their actual signatures; update the LayoutBuilder/SequentialReader/OffsetMap component descriptions with the real new/build/compute signatures. - data-access.md: document the FieldValue enum; add a Higher-level read/write section (TypedefEngine::read_field/write_field, SequentialReader::read_next/read_field); fix primitive signatures to include field_path and Endian; replace the wrong read_union_discriminator pseudo-code with the actual tunion module API (read_byte_discriminator/read_field_discriminator/resolve_variant/ discriminator_size) and the UnionDispatch struct. - validation.md: fix the TypedefEngine struct (add endian/schema fields, mark Layout as private); add the real compile signature (&mut Value, LayoutMode) and mode-appropriate accessors; fix engine.validate(buffer) -> engine.validate_json(&Value)/is_valid_json (the validator operates on serde_json::Value, not byte buffers — matches ADR-098). - overview.md: remove the stale ~1,900 lines / 26 tests line count. - ADR-097 §3a: correct the TRecord layout (no separate value_len prefix; the value is encoded by its declared TypeDef:* kind).
This commit is contained in:
1 parent
14d9cf281f
commit
7603f98799
7 files changed
+544
-279
No files matched your search
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-20
|
||||
last_updated: 2026-07-21
|
||||
---
|
||||
|
||||
# alknet-typedef
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-20
|
||||
last_updated: 2026-07-21
|
||||
---
|
||||
|
||||
# alknet-typedef — Data Access
|
||||
@@ -11,6 +11,48 @@ variable-length types. This is the consumer-facing API — given a compiled
|
||||
`TypedefEngine` and a byte buffer, read and write fields at
|
||||
schema-computed offsets.
|
||||
|
||||
This document covers two layers:
|
||||
|
||||
- **Primitive read/write functions** in the `data_access` module —
|
||||
typed reads/writes at a caller-provided offset. These are the building
|
||||
blocks used by the layout types (`OffsetMap`, `LayoutBuilder`,
|
||||
`SequentialReader`) and the `TypedefEngine`. Each operates on a raw
|
||||
byte buffer at a known offset and returns a `TypedefError::Access`
|
||||
carrying the field path on bounds or encoding failures.
|
||||
- **The `FieldValue` enum and the higher-level APIs** —
|
||||
`TypedefEngine::read_field`/`write_field` (aligned mode) and
|
||||
`SequentialReader::read_next`/`read_field` (packed mode) — which look
|
||||
up a field's offset via the layout and dispatch to the primitive
|
||||
functions, returning a unified `FieldValue<'a>`.
|
||||
|
||||
## The `FieldValue` enum
|
||||
|
||||
The higher-level read APIs return a single unified type — `FieldValue<'a>`
|
||||
— so one method can read any field kind without the caller dispatching on
|
||||
schema kind first. The variant carries the typed value; the lifetime
|
||||
borrows from the input buffer for variable-length kinds (zero-copy).
|
||||
|
||||
```rust
|
||||
pub enum FieldValue<'a> {
|
||||
I8(i8), I16(i16), I32(i32),
|
||||
U8(u8), U16(u16), U32(u32),
|
||||
F32(f32), F64(f64),
|
||||
Bool(bool),
|
||||
Enum(u32), // u32 index into the schema's "enum" array
|
||||
String(&'a str), // borrows from the buffer
|
||||
Bytes(&'a [u8]), // borrows from the buffer
|
||||
Struct { start: usize, end: usize }, // consumer recurses with a fresh reader
|
||||
Union { discriminator: String, variant_start: usize },
|
||||
Array { count: u32, element_start: usize, element_stride: usize },
|
||||
}
|
||||
```
|
||||
|
||||
For composite kinds (`Struct`, `Union`, `Array`), `FieldValue` returns a
|
||||
layout descriptor, not the decoded contents — the consumer recurses with
|
||||
a fresh `SequentialReader` (or a sub-range read) scoped to the reported
|
||||
byte range. `Array`'s `element_stride` is `0` for variable-length element
|
||||
types, signalling the consumer must walk each element sequentially.
|
||||
|
||||
## Read/Write Model
|
||||
|
||||
The typedef engine operates on raw byte buffers (`&[u8]` for reading,
|
||||
@@ -19,48 +61,100 @@ reflection, no dynamic dispatch per field. The engine uses the offset map
|
||||
(or `LayoutBuilder`/`SequentialReader`) to locate fields, then performs
|
||||
typed access at the computed positions.
|
||||
|
||||
### Fixed-size types
|
||||
### Higher-level read/write
|
||||
|
||||
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, `TEnum`, etc.) are accessed via
|
||||
zero-copy pointer casts:
|
||||
The `TypedefEngine` and `SequentialReader` provide the primary
|
||||
consumer-facing read/write APIs. They look up a field's offset via the
|
||||
layout and dispatch to the primitive `data_access` functions, returning
|
||||
`FieldValue` (read) or accepting `&FieldValue` (write).
|
||||
|
||||
```rust
|
||||
// Read a u32 at a known offset
|
||||
fn read_u32(buffer: &[u8], offset: usize, endian: Endian) -> u32 {
|
||||
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
|
||||
match endian {
|
||||
Endian::Little => u32::from_le_bytes(bytes),
|
||||
Endian::Big => u32::from_be_bytes(bytes),
|
||||
}
|
||||
impl TypedefEngine {
|
||||
// Aligned mode: looks up the field's ByteRange in the OffsetMap,
|
||||
// dispatches to the right data_access function by TypeDefKind.
|
||||
// Returns TypedefError::Access if compiled in packed mode
|
||||
// (use sequential_reader() for packed mode).
|
||||
pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str)
|
||||
-> Result<FieldValue<'a>, TypedefError>;
|
||||
pub fn write_field(&self, buffer: &mut [u8], field_path: &str,
|
||||
value: &FieldValue<'_>) -> Result<(), TypedefError>;
|
||||
}
|
||||
|
||||
// Write a u32 at a known offset
|
||||
fn write_u32(buffer: &mut [u8], offset: usize, value: u32, endian: Endian) {
|
||||
impl SequentialReader {
|
||||
// Packed mode: walks the buffer field-by-field, reading length
|
||||
// prefixes to find each field's position. read_field walks all
|
||||
// preceding fields to reach the target.
|
||||
pub fn read_next<'a>(&mut self, buffer: &'a [u8])
|
||||
-> Result<Option<(String, FieldValue<'a>)>, TypedefError>;
|
||||
pub fn read_field<'a>(&mut self, buffer: &'a [u8], field_path: &str)
|
||||
-> Result<FieldValue<'a>, TypedefError>;
|
||||
pub fn reset(&mut self);
|
||||
pub fn position(&self) -> usize;
|
||||
pub fn endian(&self) -> Endian;
|
||||
}
|
||||
```
|
||||
|
||||
`read_field`/`write_field` on `TypedefEngine` work for the fixed-size
|
||||
primitive kinds and the length-prefixed `String`/`Bytes`/`Timestamp`
|
||||
fields. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a
|
||||
`FieldValue` carrying a layout descriptor (byte range, variant start,
|
||||
or array stride) for the consumer to recurse on — see §"FieldValue" above.
|
||||
|
||||
For writing in packed mode, the consumer uses `LayoutBuilder::build` to
|
||||
compute positions, then calls the primitive `data_access::write_*`
|
||||
functions at the computed offsets. There is no packed-mode
|
||||
`engine.write_field` — the layout depends on the actual data sizes,
|
||||
which the builder consumes at `build` time.
|
||||
|
||||
### Primitive read/write functions
|
||||
|
||||
The `data_access` module exposes typed read/write functions for each
|
||||
primitive kind. Each takes `field_path: &str` for error attribution
|
||||
(produces a `TypedefError::Access` carrying the path on bounds or
|
||||
encoding failures) and, for multi-byte types, an `Endian` parameter.
|
||||
|
||||
### Fixed-size types
|
||||
|
||||
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, `TEnum`, etc.) are
|
||||
accessed via zero-copy reads of N bytes at the offset:
|
||||
|
||||
```rust
|
||||
// Read a u32 at a known offset, applying endianness. Bounds-checked.
|
||||
fn read_u32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
|
||||
-> Result<u32, TypedefError> {
|
||||
let bytes: [u8; 4] = read_array(buffer, offset, field_path)?;
|
||||
Ok(match endian {
|
||||
Endian::Little => u32::from_le_bytes(bytes),
|
||||
Endian::Big => u32::from_be_bytes(bytes),
|
||||
})
|
||||
}
|
||||
|
||||
// Write a u32 at a known offset, applying endianness. Bounds-checked.
|
||||
fn write_u32(buffer: &mut [u8], offset: usize, value: u32,
|
||||
field_path: &str, endian: Endian) -> Result<(), TypedefError> {
|
||||
let bytes = match endian {
|
||||
Endian::Little => value.to_le_bytes(),
|
||||
Endian::Big => value.to_be_bytes(),
|
||||
};
|
||||
buffer[offset..offset+4].copy_from_slice(&bytes);
|
||||
write_array(buffer, offset, bytes, field_path)
|
||||
}
|
||||
```
|
||||
|
||||
The engine applies endianness at access time based on the schema's
|
||||
`"endian"` annotation (ADR-097). The offset computation is
|
||||
endian-agnostic.
|
||||
endian-agnostic. The `read_array`/`write_array` helpers perform the
|
||||
bounds check and produce `TypedefError::Access` with the field path on
|
||||
failure.
|
||||
|
||||
### TEnum access
|
||||
|
||||
`TEnum` is a fixed-size type (4 bytes, `u32` index). Read/write follows
|
||||
the same pattern as other fixed-size types — the engine reads/writes a
|
||||
`u32` at the field's computed offset, applying the schema's endianness:
|
||||
`TEnum` is a fixed-size type (4 bytes, `u32` index). Read/write delegates
|
||||
to the `u32` primitives, applying the schema's endianness:
|
||||
|
||||
```rust
|
||||
fn read_enum(buffer: &[u8], offset: usize, endian: Endian) -> u32 {
|
||||
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
|
||||
match endian {
|
||||
Endian::Little => u32::from_le_bytes(bytes),
|
||||
Endian::Big => u32::from_be_bytes(bytes),
|
||||
}
|
||||
pub fn read_enum(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
|
||||
-> Result<u32, TypedefError> {
|
||||
read_u32(buffer, offset, field_path, endian)
|
||||
}
|
||||
```
|
||||
|
||||
@@ -72,43 +166,28 @@ corresponds to a valid enum value at the JSON level.
|
||||
|
||||
### Variable-length types (inline length-prefixing)
|
||||
|
||||
For variable-length types with inline length-prefixing (the default):
|
||||
For variable-length types with inline length-prefixing (the default),
|
||||
the `data_access` module provides `read_string`/`write_string`/
|
||||
`read_bytes`/`write_bytes`. Each takes `field_path: &str` for error
|
||||
attribution and `endian` for the length prefix:
|
||||
|
||||
```rust
|
||||
// Read a length-prefixed string
|
||||
fn read_string<'a>(buffer: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
|
||||
let len_bytes: [u8; 4] = buffer[offset..offset+4].try_into()
|
||||
.map_err(|_| TypedefError::Access { /* ... */ })?;
|
||||
let len = match endian {
|
||||
Endian::Little => u32::from_le_bytes(len_bytes),
|
||||
Endian::Big => u32::from_be_bytes(len_bytes),
|
||||
} as usize;
|
||||
let data = buffer.get(offset+4..offset+4+len)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
|
||||
std::str::from_utf8(data)
|
||||
.map_err(|e| TypedefError::Access { /* ... */ })
|
||||
}
|
||||
// Read a length-prefixed string, borrowing from the buffer.
|
||||
fn read_string<'a>(buffer: &'a [u8], offset: usize,
|
||||
field_path: &str, endian: Endian) -> Result<&'a str, TypedefError>;
|
||||
|
||||
// Write a length-prefixed string
|
||||
fn write_string(buffer: &mut [u8], offset: usize, value: &str, endian: Endian) -> Result<(), TypedefError> {
|
||||
let data = value.as_bytes();
|
||||
let len_bytes = match endian {
|
||||
Endian::Little => (data.len() as u32).to_le_bytes(),
|
||||
Endian::Big => (data.len() as u32).to_be_bytes(),
|
||||
};
|
||||
buffer.get_mut(offset..offset+4)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?
|
||||
.copy_from_slice(&len_bytes);
|
||||
buffer.get_mut(offset+4..offset+4+data.len())
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?
|
||||
.copy_from_slice(data);
|
||||
Ok(())
|
||||
}
|
||||
// Write a length-prefixed string. Returns total bytes written (4 + data.len()).
|
||||
fn write_string(buffer: &mut [u8], offset: usize, value: &str,
|
||||
field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
|
||||
|
||||
// read_bytes / write_bytes have the same shape — raw bytes, no UTF-8 check.
|
||||
```
|
||||
|
||||
The engine reads the 4-byte length prefix at the field's offset, then
|
||||
slices the data that follows. For writing, the engine writes the length
|
||||
prefix + data.
|
||||
prefix + data. `read_string` validates UTF-8 and returns a `&str`
|
||||
borrowing from the input buffer (zero-copy); `read_bytes` returns a
|
||||
`&[u8]` slice with no encoding check.
|
||||
|
||||
In packed sequential mode, the `SequentialReader` uses the length prefix
|
||||
to determine the position of the next field. In aligned static mode, the
|
||||
@@ -117,24 +196,19 @@ is accessed separately.
|
||||
|
||||
### Variable-length types (offset indirection)
|
||||
|
||||
For variable-length types with offset indirection (opt-in):
|
||||
For variable-length types with offset indirection (opt-in), the
|
||||
`data_access` module provides `read_string_indirect`/`read_bytes_indirect`.
|
||||
The 8-byte struct at `buffer[offset..offset+8]` is
|
||||
`{ data_offset: u32, data_length: u32 }` (endian-aware); the actual
|
||||
bytes live in a separate `data_region`:
|
||||
|
||||
```rust
|
||||
// Read an offset-indirect string
|
||||
fn read_string_indirect<'a>(data_region: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
|
||||
let ptr_offset = match endian {
|
||||
Endian::Little => u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()),
|
||||
Endian::Big => u32::from_be_bytes(data_region[offset..offset+4].try_into().unwrap()),
|
||||
} as usize;
|
||||
let ptr_length = match endian {
|
||||
Endian::Little => u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
|
||||
Endian::Big => u32::from_be_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
|
||||
} as usize;
|
||||
let data = data_region.get(ptr_offset..ptr_offset+ptr_length)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
|
||||
std::str::from_utf8(data)
|
||||
.map_err(|e| TypedefError::Access { /* ... */ })
|
||||
}
|
||||
fn read_string_indirect<'a>(buffer: &'a [u8], offset: usize,
|
||||
data_region: &'a [u8], field_path: &str,
|
||||
endian: Endian) -> Result<&'a str, TypedefError>;
|
||||
fn read_bytes_indirect<'a>(buffer: &'a [u8], offset: usize,
|
||||
data_region: &'a [u8], field_path: &str,
|
||||
endian: Endian) -> Result<&'a [u8], TypedefError>;
|
||||
```
|
||||
|
||||
The field is a struct `{offset: u32, length: u32}` at a known position
|
||||
@@ -143,157 +217,122 @@ engine reads the offset and length, then slices the data region.
|
||||
|
||||
## TUnion Dispatch
|
||||
|
||||
TUnion dispatch reads the discriminator value, looks up the variant
|
||||
schema, and then reads the variant's fields. The dispatch mechanism
|
||||
differs by discriminator kind (ADR-097).
|
||||
The `tunion` module provides TUnion discriminator dispatch — reading the
|
||||
discriminator value from a byte buffer, looking up the variant schema in
|
||||
the union's `mapping`, and reporting the offset where the variant struct
|
||||
begins. All reads go through the `data_access` primitives so bounds checks
|
||||
and endianness handling are uniform with the rest of the engine.
|
||||
|
||||
The result of dispatch is a `UnionDispatch` struct:
|
||||
|
||||
```rust
|
||||
pub struct UnionDispatch {
|
||||
pub key: String, // mapping key (stringified disc value)
|
||||
pub variant_offset: usize, // byte offset where the variant struct starts
|
||||
pub discriminator_size: usize, // discriminator's byte size
|
||||
}
|
||||
```
|
||||
|
||||
After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)`
|
||||
to get the variant schema, then reads the variant's fields at
|
||||
`dispatch.variant_offset` using the normal `data_access` functions (or a
|
||||
fresh `SequentialReader` scoped to the variant).
|
||||
|
||||
### Byte-offset discriminator
|
||||
|
||||
```rust
|
||||
/// Read the discriminator value from a byte-offset TUnion.
|
||||
/// Returns the mapping key (as a string) so the consumer can look up
|
||||
/// the variant schema and read the variant's fields.
|
||||
fn read_union_discriminator(
|
||||
/// Read the discriminator value from a byte-offset TUnion. The discriminator
|
||||
/// is a fixed-size integer (TypeDef:Uint8/Uint16/Uint32) at a known byte
|
||||
/// offset. Returns the mapping key (stringified integer) and the variant
|
||||
/// struct offset.
|
||||
pub fn read_byte_discriminator(
|
||||
buffer: &[u8],
|
||||
schema: &Value,
|
||||
union_schema: &Value,
|
||||
endian: Endian,
|
||||
) -> Result<String, TypedefError> {
|
||||
let disc = schema["discriminator"].as_object()
|
||||
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
|
||||
let offset = disc["offset"].as_u64().unwrap_or(0) as usize;
|
||||
let disc_type = disc["type"].as_str().unwrap_or("TypeDef:Uint8");
|
||||
|
||||
let (disc_value, disc_size) = match disc_type {
|
||||
"TypeDef:Uint8" => {
|
||||
let b = *buffer.get(offset)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
|
||||
(b as u32, 1)
|
||||
}
|
||||
"TypeDef:Uint16" => {
|
||||
let bytes: [u8; 2] = buffer[offset..offset+2].try_into().unwrap();
|
||||
let v = match endian {
|
||||
Endian::Little => u16::from_le_bytes(bytes),
|
||||
Endian::Big => u16::from_be_bytes(bytes),
|
||||
};
|
||||
(v as u32, 2)
|
||||
}
|
||||
"TypeDef:Uint32" => {
|
||||
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
|
||||
let v = match endian {
|
||||
Endian::Little => u32::from_le_bytes(bytes),
|
||||
Endian::Big => u32::from_be_bytes(bytes),
|
||||
};
|
||||
(v, 4)
|
||||
}
|
||||
_ => return Err(TypedefError::Schema(format!("unsupported discriminator type: {disc_type}"))),
|
||||
};
|
||||
|
||||
let key = disc_value.to_string();
|
||||
if schema["mapping"].as_object().map_or(false, |m| m.contains_key(&key)) {
|
||||
Ok(key)
|
||||
} else {
|
||||
Err(TypedefError::Access {
|
||||
field_path: "__discriminator".into(),
|
||||
reason: format!("unknown discriminator value: {disc_value}"),
|
||||
})
|
||||
}
|
||||
}
|
||||
) -> Result<UnionDispatch, TypedefError>;
|
||||
```
|
||||
|
||||
The discriminator is a fixed-size integer at a known byte offset. The
|
||||
mapping keys are stringified integers. The variant struct starts at
|
||||
`offset + discriminator_size`. After reading the discriminator, the
|
||||
consumer looks up the variant schema and reads the variant's fields
|
||||
using the normal typed read functions (e.g., `read_u32`, `read_string`)
|
||||
at `offset + discriminator_size`.
|
||||
|
||||
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
|
||||
1..N are the variant struct. The call protocol's 5 event types
|
||||
(`call.requested` → 0x01, etc.) use the same pattern.
|
||||
(`call.requested` → 0x01, etc.) use the same pattern. The variant struct
|
||||
starts at `offset + discriminator_size`.
|
||||
|
||||
### Field-name discriminator
|
||||
|
||||
```rust
|
||||
/// Read the discriminator value from a field-name TUnion.
|
||||
/// The discriminator is a named field within the struct — its offset
|
||||
/// is computed like any other field. The consumer reads the field's
|
||||
/// value, looks up the variant schema, then reads the variant's fields.
|
||||
fn read_union_field_discriminator(
|
||||
/// Read the discriminator value from a field-name TUnion. The
|
||||
/// discriminator is a named field within the struct — the consumer
|
||||
/// provides the field's computed offset (from the OffsetMap or
|
||||
/// LayoutBuilder). Supports TypeDef:String, Uint8, and Enum discriminator
|
||||
/// fields.
|
||||
pub fn read_field_discriminator(
|
||||
buffer: &[u8],
|
||||
schema: &Value,
|
||||
offset_map: &OffsetMap,
|
||||
union_schema: &Value,
|
||||
disc_field_offset: usize,
|
||||
endian: Endian,
|
||||
) -> Result<String, TypedefError> {
|
||||
let disc = schema["discriminator"].as_object()
|
||||
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
|
||||
let field_name = disc["name"].as_str()
|
||||
.ok_or_else(|| TypedefError::Schema("discriminator has no 'name'".into()))?;
|
||||
|
||||
// Read the discriminator field at its computed offset.
|
||||
// The field's TypeDef kind determines how to read it (typically a string).
|
||||
let field_schema = schema["properties"].get(field_name)
|
||||
.ok_or_else(|| TypedefError::Schema(format!("discriminator field '{field_name}' not found")))?;
|
||||
let kind = get_typedef_kind(field_schema)
|
||||
.ok_or_else(|| TypedefError::Schema("discriminator field has no TypeDef kind".into()))?;
|
||||
|
||||
match kind {
|
||||
"TypeDef:String" => {
|
||||
let range = offset_map.get(field_name)
|
||||
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
|
||||
read_string(buffer, range.start, endian)
|
||||
.map(|s| s.to_string())
|
||||
}
|
||||
"TypeDef:Uint8" => {
|
||||
let range = offset_map.get(field_name)
|
||||
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
|
||||
Ok(buffer[range.start].to_string())
|
||||
}
|
||||
_ => Err(TypedefError::Schema(format!(
|
||||
"unsupported discriminator field type: {kind}"
|
||||
))),
|
||||
}
|
||||
}
|
||||
) -> Result<UnionDispatch, TypedefError>;
|
||||
```
|
||||
|
||||
The discriminator is a named field within the struct. Its offset is
|
||||
computed like any other field. The mapping keys are string values.
|
||||
After reading the discriminator, the consumer looks up the variant
|
||||
schema and reads the variant's fields starting at the end of the
|
||||
discriminator field (or at the start of the union buffer if the
|
||||
discriminator is the first field).
|
||||
computed like any other field (the consumer passes it in as
|
||||
`disc_field_offset`). The mapping keys are string values. After reading
|
||||
the discriminator, the consumer looks up the variant schema and reads
|
||||
the variant's fields starting at the end of the discriminator field.
|
||||
|
||||
### Variant resolution
|
||||
|
||||
```rust
|
||||
/// Look up a variant schema from the union's mapping. Inline schemas
|
||||
/// are returned directly. $ref pointers of the form "#/$defs/<name>"
|
||||
/// are resolved against the union schema's own $defs block.
|
||||
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str)
|
||||
-> Result<&'a Value, TypedefError>;
|
||||
|
||||
/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a
|
||||
/// byte-offset TUnion. Field-name discriminators have no fixed size
|
||||
/// and produce a TypedefError::Schema.
|
||||
pub fn discriminator_size(union_schema: &Value) -> Result<usize, TypedefError>;
|
||||
```
|
||||
|
||||
### TUnion in the layout engines
|
||||
|
||||
The `LayoutBuilder` and `SequentialReader` also handle TUnion fields
|
||||
inline during traversal (the consumer does not need to call the `tunion`
|
||||
functions for a union field reached during a sequential walk). For
|
||||
`LayoutBuilder`, the consumer supplies the discriminator value (byte-offset)
|
||||
or variant index (field-name) in `var_sizes` under the synthetic key
|
||||
`"<union_path>.__discriminator"` or `"<union_path>.__variant"`. For
|
||||
`SequentialReader`, a union field yields
|
||||
`FieldValue::Union { discriminator, variant_start }`. The standalone
|
||||
`tunion` functions are for dispatch outside the layout walk — e.g., a
|
||||
consumer that receives a bare union buffer and needs to identify the
|
||||
variant before recursing.
|
||||
|
||||
## Field Paths
|
||||
|
||||
Fields are addressed by dotted paths: `"header.version"`, `"payload.data"`.
|
||||
The `OffsetMap` stores fully-qualified paths. The read/write functions
|
||||
accept a field path and look up the byte range:
|
||||
Both `OffsetMap` and `PackedLayout` store fully-qualified paths (nested
|
||||
struct fields appear under their parent's path prefix). The higher-level
|
||||
APIs (`TypedefEngine::read_field`/`write_field`, `SequentialReader::read_field`)
|
||||
accept a field path, look up the byte range/position in the layout, and
|
||||
dispatch to the primitive `data_access` function for the field's kind.
|
||||
|
||||
```rust
|
||||
/// Read an f32 field by path. This is the aligned-mode path (uses OffsetMap).
|
||||
/// In packed mode, the consumer uses SequentialReader instead.
|
||||
fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> {
|
||||
let range = self.offset_map.get(field_path)
|
||||
.ok_or_else(|| TypedefError::Offset {
|
||||
field_path: field_path.to_string(),
|
||||
reason: "field not found in offset map".to_string(),
|
||||
})?;
|
||||
if buffer.len() < range.end {
|
||||
return Err(TypedefError::Access {
|
||||
field_path: field_path.to_string(),
|
||||
reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()),
|
||||
});
|
||||
}
|
||||
let bytes: [u8; 4] = buffer[range.start..range.end].try_into().unwrap();
|
||||
Ok(match self.endian {
|
||||
Endian::Little => f32::from_le_bytes(bytes),
|
||||
Endian::Big => f32::from_be_bytes(bytes),
|
||||
})
|
||||
}
|
||||
```
|
||||
For aligned-mode access, `TypedefEngine::read_field(&buffer, "header.version")`
|
||||
returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds
|
||||
the field's `TypeDef:*` kind in the schema, and calls the matching
|
||||
`data_access::read_*` function. `write_field` is the mirror. Composite
|
||||
kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue`
|
||||
carrying a layout descriptor; the consumer recurses with a fresh reader
|
||||
or sub-range read.
|
||||
|
||||
For packed-mode access, `SequentialReader::read_field(&buffer, "c")` walks
|
||||
all preceding fields to reach the target (sequential access is inherent
|
||||
to packed layouts). `read_next` walks fields in declaration order.
|
||||
|
||||
Nested structs produce nested field paths. The offset computation
|
||||
propagates the field path prefix during recursion, so the `OffsetMap`
|
||||
contains entries like `"header.version"` and `"header.magic"`.
|
||||
and `PackedLayout` contain entries like `"header.version"` and
|
||||
`"header.magic"`.
|
||||
|
||||
## Zero-Copy Access
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-20
|
||||
last_updated: 2026-07-21
|
||||
---
|
||||
|
||||
# alknet-typedef — Layout Engine
|
||||
@@ -24,27 +24,21 @@ protocols.
|
||||
|
||||
**Components:**
|
||||
|
||||
- **`LayoutBuilder`** — takes a schema and actual data sizes for
|
||||
variable-length fields, computes byte positions for each field in a
|
||||
packed layout. Used at write time when the consumer knows the data
|
||||
sizes upfront.
|
||||
- **`SequentialReader`** — walks a buffer field-by-field according to the
|
||||
schema, reading length prefixes to determine variable-length data
|
||||
positions. Used at read time when the consumer is parsing an incoming
|
||||
frame.
|
||||
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `TypeDef:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, TypedefError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
|
||||
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, TypedefError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
|
||||
|
||||
**How it works:**
|
||||
|
||||
For a struct with fields `[u8, u32, string]`:
|
||||
For a struct with fields `[u8, u32, string]` where the string is 10 bytes:
|
||||
|
||||
```
|
||||
LayoutBuilder (write):
|
||||
LayoutBuilder::build(var_sizes: {"payload": 10}):
|
||||
field[0] u8: offset 0, size 1
|
||||
field[1] u32: offset 1, size 4
|
||||
field[2] string: offset 5, size 4 (length prefix) + data_len
|
||||
total: 9 + data_len
|
||||
field[2] string: offset 5, size 4 (length prefix) + 10 (data)
|
||||
total: 19
|
||||
|
||||
SequentialReader (read):
|
||||
SequentialReader::read_next (read):
|
||||
read u8 at offset 0
|
||||
read u32 at offset 1
|
||||
read u32 length prefix at offset 5 → data_len
|
||||
@@ -75,10 +69,7 @@ and safetensors.
|
||||
|
||||
**Component:**
|
||||
|
||||
- **`OffsetMap`** — walks the schema once, computes fixed byte positions
|
||||
for each field based on type sizes and alignment. The output is a flat
|
||||
table of `(field_path, byte_range)` pairs. Used for both read and write
|
||||
at known offsets.
|
||||
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, TypedefError>` (requires `TypeDef:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
|
||||
|
||||
**How it works:**
|
||||
|
||||
@@ -199,7 +190,10 @@ annotation shapes).
|
||||
|
||||
Nested structs produce dotted field paths: `header.version`,
|
||||
`header.magic`. The offset computation propagates the field path prefix
|
||||
during recursion. The `OffsetMap` stores fully-qualified paths.
|
||||
during recursion. Both `OffsetMap` and `PackedLayout` store fully-qualified
|
||||
paths; the `iter()` method of each yields fields in schema `properties`
|
||||
order, with nested struct fields appearing inline under their parent's
|
||||
path prefix.
|
||||
|
||||
### Endianness
|
||||
|
||||
@@ -212,13 +206,25 @@ and byte-swaps accordingly. All fixed-size types — including `TEnum`
|
||||
|
||||
## Mode Selection
|
||||
|
||||
The consumer selects the mode at engine construction time. The choice is
|
||||
determined by the use case, not by the schema:
|
||||
The consumer selects the mode at engine construction time via the
|
||||
`LayoutMode` enum, passed to `TypedefEngine::compile`:
|
||||
|
||||
```rust
|
||||
pub enum LayoutMode {
|
||||
/// Packed sequential — for protocol wire formats (SFTP, channels, TTY).
|
||||
Packed,
|
||||
/// Aligned static — for mmap-friendly formats (metatensor, safetensors).
|
||||
Aligned,
|
||||
}
|
||||
```
|
||||
|
||||
The choice is determined by the use case, not by the schema:
|
||||
|
||||
- **Protocol consumer** (SFTP, binary call frames, TTY negotiation):
|
||||
uses `LayoutBuilder` for writing and `SequentialReader` for reading.
|
||||
- **mmap consumer** (metatensor): uses `OffsetMap` for both reading and
|
||||
writing.
|
||||
`LayoutMode::Packed` → uses `LayoutBuilder` for writing and
|
||||
`SequentialReader` for reading.
|
||||
- **mmap consumer** (metatensor): `LayoutMode::Aligned` → uses `OffsetMap`
|
||||
for both reading and writing at known offsets.
|
||||
|
||||
The same schema can be used in either mode. A schema describing an SFTP
|
||||
packet can be consumed by a `SequentialReader` (for parsing incoming
|
||||
@@ -226,6 +232,84 @@ frames) and a `LayoutBuilder` (for constructing outgoing frames). A schema
|
||||
describing a metatensor layout can be consumed by an `OffsetMap` (for
|
||||
mmap access).
|
||||
|
||||
`TypedefEngine` exposes mode-appropriate accessors: `engine.offset_map()`
|
||||
returns `Some` in aligned mode and `None` in packed mode;
|
||||
`engine.layout_builder()` and `engine.sequential_reader()` return `Some`
|
||||
in packed mode and `None` in aligned mode. See [validation.md](validation.md)
|
||||
§"The TypedefEngine struct" for the engine API.
|
||||
|
||||
## Public Types
|
||||
|
||||
The layout engine produces three public types, one per layout component.
|
||||
All are re-exported from the crate root.
|
||||
|
||||
### `ByteRange` (aligned mode)
|
||||
|
||||
```rust
|
||||
pub struct ByteRange {
|
||||
pub start: usize, // inclusive
|
||||
pub end: usize, // exclusive
|
||||
}
|
||||
```
|
||||
|
||||
A half-open byte range produced by `OffsetMap::compute` for each field.
|
||||
`end - start` is the field's byte size in the static layout (for
|
||||
variable-length fields: the length prefix, the `{offset, length}` pair,
|
||||
or the `maxLength` reservation — not the variable data). `ByteRange`
|
||||
provides `len()` and `is_empty()`.
|
||||
|
||||
### `FieldPosition` (packed mode)
|
||||
|
||||
```rust
|
||||
pub struct FieldPosition {
|
||||
pub offset: usize,
|
||||
pub size: usize,
|
||||
pub kind: TypeDefKind,
|
||||
}
|
||||
```
|
||||
|
||||
A field's computed position in a packed layout, produced by
|
||||
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
|
||||
length prefix); for fixed-size fields, `size` is the type's byte size.
|
||||
`kind` records the field's `TypeDef:*` kind so the consumer can dispatch
|
||||
to the correct `data_access` read/write function.
|
||||
|
||||
### `PackedLayout` (packed mode)
|
||||
|
||||
The result of `LayoutBuilder::build`: a map of `field_path → FieldPosition`
|
||||
plus the total buffer size needed.
|
||||
|
||||
```rust
|
||||
impl PackedLayout {
|
||||
pub fn get(&self, field_path: &str) -> Option<&FieldPosition>;
|
||||
pub fn total_size(&self) -> usize;
|
||||
pub fn iter(&self) -> impl Iterator<Item = &(String, FieldPosition)>;
|
||||
}
|
||||
```
|
||||
|
||||
`get` looks up a field by dotted path. For TUnion byte-offset
|
||||
discriminators, the discriminator is recorded under the synthetic path
|
||||
`"<union_path>.__discriminator"`. `iter` yields fields in layout order
|
||||
(schema `properties` order, with nested struct fields appearing inline
|
||||
under their parent's path prefix).
|
||||
|
||||
### `OffsetMap` (aligned mode)
|
||||
|
||||
A flat table of `(field_path, byte_range)` pairs computed from a schema.
|
||||
|
||||
```rust
|
||||
impl OffsetMap {
|
||||
pub fn compute(schema: &Value) -> Result<Self, TypedefError>;
|
||||
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
|
||||
pub fn total_size(&self) -> usize;
|
||||
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
|
||||
}
|
||||
```
|
||||
|
||||
`compute` requires a `TypeDef:Struct` at the top level. `total_size`
|
||||
includes trailing alignment padding. `iter` yields fields in insertion
|
||||
order (schema `properties` order, nested struct fields appearing inline).
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-20
|
||||
last_updated: 2026-07-21
|
||||
---
|
||||
|
||||
# alknet-typedef — Overview
|
||||
@@ -31,12 +31,12 @@ with `TypeDef:*` custom keywords (the same kinds defined in TypeBox's
|
||||
The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||
`serde_json` (schema parsing). The novel code is the offset computation
|
||||
— a recursive walk of the schema JSON that computes byte positions for
|
||||
each field. The custom keyword implementations are ~10 lines each.
|
||||
each field. The custom keyword implementations are small (a few lines
|
||||
each, generated from shared macros — see [validation.md](validation.md)).
|
||||
|
||||
The crate is ~1,900 lines (POC verified, 26 tests passing). It replaces
|
||||
two prior attempts that built their own jsonschema engines — typebox-rs
|
||||
(~8,400 lines) and alktype (~5,600 lines) — with `jsonschema` + an
|
||||
offset map + ~50 lines of custom keyword implementations. See
|
||||
The crate replaces two prior attempts that built their own jsonschema
|
||||
engines — typebox-rs (~8,400 lines) and alktype (~5,600 lines) — with
|
||||
`jsonschema` + an offset map + small custom keyword implementations. See
|
||||
[ADR-095](../../decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
## Why
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-20
|
||||
last_updated: 2026-07-21
|
||||
---
|
||||
|
||||
# alknet-typedef — Schema Layer
|
||||
@@ -37,6 +37,42 @@ encoding strategy (for variable-length types).
|
||||
| `TRecord` | `TypeDef:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
|
||||
| `TTimestamp` | `TypeDef:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
|
||||
|
||||
### The `TypeDefKind` enum
|
||||
|
||||
The engine represents the 17 kinds as a Rust enum — `TypeDefKind` — with
|
||||
one variant per kind (`TypeDefKind::Float32`, `TypeDefKind::Struct`, etc.).
|
||||
The enum provides compile-time exhaustiveness checking and integer
|
||||
discriminant dispatch (a jump table) instead of string comparison at
|
||||
every field access. It is `pub` and re-exported from the crate root.
|
||||
|
||||
```rust
|
||||
pub enum TypeDefKind {
|
||||
Int8, Int16, Int32,
|
||||
Uint8, Uint16, Uint32,
|
||||
Float32, Float64,
|
||||
Boolean, Enum,
|
||||
String, Bytes, Timestamp,
|
||||
Struct, Union, Array, Record,
|
||||
}
|
||||
```
|
||||
|
||||
The enum carries the kind's binary-layout metadata as inherent methods:
|
||||
|
||||
| Method | Returns | Notes |
|
||||
|--------|---------|-------|
|
||||
| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"TypeDef:Uint8"` |
|
||||
| `type_size(self)` | `Option<usize>` | `Some(N)` for fixed-size kinds; `None` for variable/composite |
|
||||
| `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array |
|
||||
| `is_fixed_size(self)` | `bool` | True for the 10 fixed-size primitive kinds |
|
||||
| `is_composite(self)` | `bool` | True for Struct, Union, Array, Record |
|
||||
| `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record |
|
||||
| `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter |
|
||||
|
||||
`TypeDefKind` implements `Display` (renders the keyword string) and
|
||||
`FromStr` (parses the keyword string back into the variant, returning
|
||||
`TypedefError::Schema` for unknown kinds). The layout engines and the
|
||||
validator dispatch on the enum, not on strings.
|
||||
|
||||
### Fixed-size types
|
||||
|
||||
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
|
||||
@@ -132,14 +168,20 @@ distinct from UTF-8 strings. In the binary representation, TBytes is raw
|
||||
bytes with no encoding (not base64, not hex). In the JSON representation
|
||||
(for validation), TBytes is a string (JSON has no native byte type).
|
||||
|
||||
**`TRecord`:** A string-keyed map. Binary layout is a count-prefixed
|
||||
sequence of `(key, value)` pairs: `[count: u32][key_len: u32][key_bytes]
|
||||
[value_len: u32][value_bytes]...`. The count is the number of entries.
|
||||
Each key is a length-prefixed UTF-8 string. Each value is the record's
|
||||
declared value type (specified via the `"values"` property in the schema,
|
||||
e.g., `"values": { "TypeDef:Float32": true }`). The count prefix respects
|
||||
the schema's endianness. In aligned static mode with `maxLength`, the
|
||||
entire record is reserved at `maxLength` bytes (zero-padded).
|
||||
**`TRecord`:** A string-keyed map. The value type is declared via the
|
||||
schema's `"values"` property (e.g., `"values": { "TypeDef:Float32": true }`).
|
||||
Binary layout is a count-prefixed sequence of `(key, value)` pairs:
|
||||
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
|
||||
The count is the number of entries. Each key is a length-prefixed UTF-8
|
||||
string. Each value is encoded according to its declared `TypeDef:*` kind
|
||||
— a `Record<Uint32>` value is 4 raw bytes; a `Record<String>` value is
|
||||
itself a length-prefixed string; a `Record<Struct>` value is the struct's
|
||||
fields laid out inline. There is **no separate `value_len` prefix** —
|
||||
the value's size is determined by its kind (fixed-size kinds have a
|
||||
known size; variable-length kinds carry their own length prefix). The
|
||||
count and key-length prefixes respect the schema's endianness. In
|
||||
aligned static mode with `maxLength`, the entire record is reserved at
|
||||
`maxLength` bytes (zero-padded).
|
||||
|
||||
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
|
||||
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
|
||||
@@ -166,6 +208,74 @@ The count prefix respects the schema's endianness.
|
||||
fields' sizes (plus alignment padding in aligned static mode). The offset
|
||||
computation recurses into their properties.
|
||||
|
||||
## Schema-Layer Public API
|
||||
|
||||
The `schema` module exposes the foundational types and functions every
|
||||
other module depends on. These are re-exported from the crate root.
|
||||
|
||||
### `get_typedef_kind` vs `get_typedef_kind_loose`
|
||||
|
||||
The engine recognizes a `TypeDef:*` kind on a schema node two ways,
|
||||
because the keyword value may be either a boolean (`true`) or an
|
||||
annotation object (`{ "encoding": "..." }`):
|
||||
|
||||
| Function | Recognizes | Returns |
|
||||
|----------|------------|---------|
|
||||
| `get_typedef_kind(node) -> Option<&str>` | Boolean form only (`{ "TypeDef:String": true }`) | The keyword string, e.g. `"TypeDef:String"` |
|
||||
| `get_typedef_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string |
|
||||
| `get_typedef_kind_enum(node) -> Option<TypeDefKind>` | Boolean form only | The parsed enum variant |
|
||||
| `get_typedef_kind_loose_enum(node) -> Option<TypeDefKind>` | Boolean form **and** object form | The parsed enum variant |
|
||||
|
||||
The boolean-form-only functions are used by the validator factories
|
||||
(which reject the object form as a schema error) and the top-level
|
||||
kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new`
|
||||
(which require `TypeDef:Struct` at the root). The "loose" variants are
|
||||
used by the layout engines during field traversal, so that a variable-
|
||||
length field with an `encoding` annotation (`{ "TypeDef:String":
|
||||
{ "encoding": "offset-indirect" } }`) is still recognized as a `String`.
|
||||
|
||||
### Annotation parsers
|
||||
|
||||
Each schema-level annotation has a dedicated parser that reads it from a
|
||||
`serde_json::Value` node and returns a sensible default when absent:
|
||||
|
||||
| Function | Annotation | Default |
|
||||
|----------|------------|---------|
|
||||
| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` |
|
||||
| `parse_align(node) -> Option<usize>` | `"align"` | `None` |
|
||||
| `parse_max_length(node) -> Option<usize>` | `"maxLength"` | `None` |
|
||||
| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` |
|
||||
| `parse_discriminator(node) -> Result<DiscriminatorKind, TypedefError>` | `"discriminator"` | (required — returns `TypedefError::Schema` if absent) |
|
||||
|
||||
### Public enums
|
||||
|
||||
```rust
|
||||
pub enum Endian { Little, Big }
|
||||
pub enum VariableEncoding { LengthPrefixed, OffsetIndirect }
|
||||
pub enum DiscriminatorKind {
|
||||
Byte { offset: usize, disc_type: TypeDefKind },
|
||||
Field { name: String },
|
||||
}
|
||||
```
|
||||
|
||||
`DiscriminatorKind::Byte` carries the byte position (`offset`) and the
|
||||
discriminator's `TypeDef:*` kind (`disc_type`, restricted to `Uint8`/
|
||||
`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator
|
||||
field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for
|
||||
how these drive dispatch.
|
||||
|
||||
### `$ref` resolution and normalization
|
||||
|
||||
| Function | Purpose |
|
||||
|----------|---------|
|
||||
| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/<name>"`. Idempotent. Runs once at `TypedefEngine::compile` time. |
|
||||
| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. |
|
||||
| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). |
|
||||
|
||||
`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s
|
||||
JSON Pointer requirement. The layout engines call `resolve_ref_or_inline`
|
||||
on every `$ref`-bearing node they encounter during traversal.
|
||||
|
||||
## jsonschema Custom Keyword Integration
|
||||
|
||||
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
|
||||
@@ -222,20 +332,12 @@ TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`),
|
||||
referencing sibling definitions within the same `$defs` block. The
|
||||
`jsonschema` crate requires full JSON Pointer paths (e.g.,
|
||||
`"$ref": "#/$defs/Read"`). The typedef engine normalizes TypeBox-style
|
||||
refs at schema load time:
|
||||
|
||||
```rust
|
||||
fn normalize_refs(schema: &mut Value) {
|
||||
// Walk the schema tree. For every "$ref" whose value is a bare name
|
||||
// (no "#" prefix), rewrite it to "#/$defs/<name>".
|
||||
// "$ref": "Read" → "$ref": "#/$defs/Read"
|
||||
}
|
||||
```
|
||||
|
||||
This is a ~20-line recursive walk of the schema JSON. It runs once at
|
||||
load time, before the schema is passed to `jsonschema::validator_for`
|
||||
or the offset computation. The normalization is idempotent — full JSON
|
||||
Pointer refs pass through unchanged.
|
||||
refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization)
|
||||
— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to
|
||||
`"#/$defs/<name>"`. The normalization is idempotent — full JSON Pointer
|
||||
refs pass through unchanged. It runs once at `TypedefEngine::compile`
|
||||
time, before the schema is passed to `jsonschema` or the offset
|
||||
computation.
|
||||
|
||||
**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs
|
||||
with `Resource 'Read' is not present in a registry`. Full JSON Pointer
|
||||
@@ -279,8 +381,9 @@ Both struct-level and field-level, with field-level overriding:
|
||||
|
||||
- Struct-level `"align"` sets the default for all fields.
|
||||
- Field-level `"align"` overrides the struct default.
|
||||
- Default alignment: 1 for u8/bool, 2 for u16/i16, 4 for u32/i32/f32,
|
||||
8 for u64/i64/f64, max field alignment for structs.
|
||||
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/
|
||||
enum, 8 for f64, 4 for variable-length (the u32 length prefix), 1 for
|
||||
struct/union/array.
|
||||
- Only meaningful in aligned static mode (ADR-096). Ignored in packed
|
||||
sequential mode.
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-20
|
||||
last_updated: 2026-07-21
|
||||
---
|
||||
|
||||
# alknet-typedef — Validation
|
||||
@@ -50,30 +50,47 @@ process, not a single `validate(buffer)` call.
|
||||
### The `TypedefEngine` struct
|
||||
|
||||
The `TypedefEngine` is the compiled form of a schema. It supports both
|
||||
layout modes (ADR-096) via an internal enum:
|
||||
layout modes (ADR-096) via an internal `Layout` enum:
|
||||
|
||||
```rust
|
||||
pub struct TypedefEngine {
|
||||
layout: Layout, // packed or aligned (see below)
|
||||
layout: Layout, // packed or aligned (private enum)
|
||||
validator: jsonschema::Validator, // compiled once at load time
|
||||
endian: Endian, // parsed from the schema's "endian" annotation
|
||||
schema: Value, // the normalized schema (refs resolved)
|
||||
}
|
||||
|
||||
// Private — the consumer selects via LayoutMode at compile time.
|
||||
enum Layout {
|
||||
Packed {
|
||||
builder: LayoutBuilder,
|
||||
reader: SequentialReader,
|
||||
},
|
||||
Aligned {
|
||||
offset_map: OffsetMap,
|
||||
},
|
||||
Packed { builder: LayoutBuilder, reader: SequentialReader },
|
||||
Aligned { offset_map: OffsetMap },
|
||||
}
|
||||
```
|
||||
|
||||
The consumer selects the mode at construction time. The `Layout` enum
|
||||
ensures the engine always has the correct layout strategy for the
|
||||
consumer's use case — a protocol consumer gets `Packed`, an mmap
|
||||
consumer gets `Aligned`. The validator is mode-agnostic (it operates on
|
||||
`Value`, not raw bytes).
|
||||
The consumer selects the mode at construction time via `LayoutMode`
|
||||
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout`
|
||||
enum is private — the engine exposes mode-appropriate accessors instead:
|
||||
|
||||
```rust
|
||||
impl TypedefEngine {
|
||||
pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result<Self, TypedefError>;
|
||||
pub fn mode(&self) -> LayoutMode;
|
||||
pub fn endian(&self) -> Endian;
|
||||
pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode
|
||||
pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode
|
||||
pub fn sequential_reader(&self) -> Option<&SequentialReader>; // Some in packed mode
|
||||
}
|
||||
```
|
||||
|
||||
`compile` takes `&mut Value` because it normalizes `$ref` values in place
|
||||
(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization))
|
||||
before computing the layout and building the validator. The `schema`
|
||||
field retains the normalized schema for `read_field`'s kind lookup. The
|
||||
validator is mode-agnostic (it operates on `Value`, not raw bytes).
|
||||
|
||||
The `read_field`/`write_field` methods on `TypedefEngine` are the
|
||||
aligned-mode data-access API — see [data-access.md](data-access.md)
|
||||
§"Higher-level read/write".
|
||||
|
||||
## Custom Keyword Validators
|
||||
|
||||
@@ -242,18 +259,30 @@ you exactly which field failed and why.
|
||||
### Load time: `TypedefEngine::compile()`
|
||||
|
||||
The expensive work happens once at schema load time:
|
||||
1. Parse the schema JSON (`serde_json::from_str` with `preserve_order`).
|
||||
2. Compute the offset map (or `LayoutBuilder`/`SequentialReader`).
|
||||
3. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
|
||||
1. Normalize `$ref` values in the schema (`normalize_refs`).
|
||||
2. Parse the schema's `"endian"` annotation.
|
||||
3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned).
|
||||
4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
|
||||
|
||||
The result is a `TypedefEngine` that can be used for repeated operations.
|
||||
|
||||
### Access time: `engine.validate(buffer)`
|
||||
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
|
||||
|
||||
Validation is opt-in per operation. The consumer calls
|
||||
`engine.validate(buffer)` when validation is desired. The jsonschema
|
||||
validator is already compiled — `is_valid()` is a fast check against
|
||||
the compiled validator.
|
||||
`engine.validate_json(instance)` when validation is desired, or
|
||||
`engine.is_valid_json(instance)` for a boolean check. The jsonschema
|
||||
validator is already compiled — these are fast checks against the
|
||||
compiled validator.
|
||||
|
||||
```rust
|
||||
pub fn validate_json(&self, instance: &Value) -> Result<(), TypedefError>;
|
||||
pub fn is_valid_json(&self, instance: &Value) -> bool;
|
||||
```
|
||||
|
||||
The argument is a `serde_json::Value` (the JSON representation of the
|
||||
data), not a raw byte buffer — see §"What validation validates" above.
|
||||
To validate a binary buffer end-to-end, the consumer reads it into a
|
||||
`Value` tree via the data access layer, then validates that `Value`.
|
||||
|
||||
High-throughput paths can skip validation. Security-sensitive paths
|
||||
(parsing incoming frames from untrusted peers) can validate every frame.
|
||||
@@ -261,16 +290,19 @@ The choice is the consumer's.
|
||||
|
||||
## Relationship to Read/Write
|
||||
|
||||
Validation and data access are independent operations on the same buffer.
|
||||
Validation and data access are independent operations on the same data.
|
||||
The consumer can:
|
||||
|
||||
1. Validate a buffer to ensure it conforms to the schema.
|
||||
2. Read fields from the buffer at computed offsets.
|
||||
3. Both — validate first, then read (defense in depth).
|
||||
1. Validate the JSON representation of a buffer to ensure it conforms to
|
||||
the schema.
|
||||
2. Read fields from the binary buffer at computed offsets.
|
||||
3. Both — validate the JSON representation first, then read the binary
|
||||
buffer (defense in depth).
|
||||
|
||||
The engine does not couple validation and access. A consumer that trusts
|
||||
its data source can skip validation and go straight to read/write. A
|
||||
consumer that parses untrusted input can validate first, then access.
|
||||
consumer that parses untrusted input can validate the JSON
|
||||
representation first, then access the binary buffer.
|
||||
|
||||
## Design Decisions
|
||||
|
||||
|
||||
@@ -144,8 +144,15 @@ the `"values"` property in the schema:
|
||||
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
|
||||
values in the record. All values share the same type.
|
||||
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
|
||||
`[count: u32][key_len: u32][key_bytes][value_len: u32][value_bytes]...`.
|
||||
- The count prefix respects the schema's endianness.
|
||||
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count`
|
||||
times. Each key is a length-prefixed UTF-8 string. Each value is
|
||||
encoded according to its declared `TypeDef:*` kind — a `Record<Uint32>`
|
||||
value is 4 raw bytes; a `Record<String>` value is itself a
|
||||
length-prefixed string; a `Record<Struct>` value is the struct's
|
||||
fields laid out inline. There is **no separate `value_len` prefix** —
|
||||
the value's size is determined by its kind (fixed-size kinds have a
|
||||
known size; variable-length kinds carry their own length prefix).
|
||||
- The count and key-length prefixes respect the schema's endianness.
|
||||
- In aligned static mode with `maxLength`, the entire record is reserved
|
||||
at `maxLength` bytes (zero-padded).
|
||||
|
||||
|
||||
Reference in new issue
Block a user