Port alknet-typedef crate from alknet
Copy the binary struct engine (src/, tests/) verbatim from alknet/crates/alknet-typedef and create a standalone Cargo.toml (workspace-inherited fields inlined). Port the architecture docs (specs, ADRs 095-102, OQs 069-071) from alknet's nested multi-crate layout to a flat single-crate layout, fixing relative link paths. Build, 295 tests, and clippy all pass clean.
This commit is contained in:
1 parent
eac7ad88b3
commit
2c4a4994dc
36 files changed
+13805
No files matched your search
@@ -0,0 +1,110 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
---
|
||||
|
||||
# alknet-typedef
|
||||
|
||||
The binary struct engine: a small Rust crate that takes a JSON Schema
|
||||
with `TypeDef:*` custom keywords and produces an offset map, read/write
|
||||
functions, and validation — all driven by the schema. The schema is the
|
||||
format definition; the engine is generic.
|
||||
|
||||
## Documents
|
||||
|
||||
| Document | Status | Description |
|
||||
|----------|--------|-------------|
|
||||
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
|
||||
| [schema-layer.md](schema-layer.md) | draft | The 19 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
|
||||
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
|
||||
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
|
||||
| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation, `TypedefEngine` |
|
||||
|
||||
## Applicable ADRs
|
||||
|
||||
| ADR | Title | Relevance |
|
||||
|-----|-------|-----------|
|
||||
| [095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
|
||||
| [096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` |
|
||||
| [097](decisions/097-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations |
|
||||
| [098](decisions/098-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors |
|
||||
| [099](decisions/099-int64-uint64-first-class-kinds.md) | Int64/Uint64 as First-Class Kinds | 64-bit integers (SFTP offsets, metatensor data_offsets); JSON precision caveat |
|
||||
| [100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode | Prevents silent data corruption (inline variable data clobbering subsequent fields) |
|
||||
| [101](decisions/101-packed-mode-read-factory.md) | Packed-Mode Read API — Engine as SequentialReader Factory | `engine.sequential_reader()` returns an owned reader, not a reference |
|
||||
| [102](decisions/102-reject-tunion-in-aligned-mode.md) | Reject TUnion in Aligned Mode for v1 | Unions are the protocol pattern; aligned-mode union semantics were broken |
|
||||
|
||||
## Relevant Open Questions
|
||||
|
||||
| OQ | Title | Status | Relevance |
|
||||
|----|-------|--------|-----------|
|
||||
| OQ-069 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it |
|
||||
| OQ-070 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case |
|
||||
| OQ-071 | Builder API for schema construction | deferred(scope) | Schemas are authored in TypeBox or hand-written JSON for v1; blocked on a concrete need |
|
||||
|
||||
## Key Design Principles
|
||||
|
||||
1. **The schema is the format.** A JSON Schema with `TypeDef:*` custom
|
||||
keywords is both the validation spec and the layout spec. No separate
|
||||
format definition, no separate parser, no separate validator. One
|
||||
schema, three uses: validate, compute offsets, access data. See
|
||||
[overview.md](overview.md) and [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
2. **jsonschema is the validation engine, not a custom engine.** The
|
||||
`jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with
|
||||
custom keyword support. The novel code is the offset computation, not
|
||||
the validation. This eliminates ~14,000 lines of hand-rolled schema
|
||||
engines (typebox-rs, alktype). See [schema-layer.md](schema-layer.md)
|
||||
and [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
3. **Two layout modes for two use cases.** Packed sequential
|
||||
(`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP,
|
||||
channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats
|
||||
(metatensor). The consumer selects the mode; the schema is the same.
|
||||
See [layout-engine.md](layout-engine.md) and
|
||||
[ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md).
|
||||
|
||||
4. **Variable-length types default to inline length-prefixing.**
|
||||
`[length: u32][data]` is the universal pattern used by channels, SFTP,
|
||||
TTY, and most binary protocols. Offset indirection (the metatensor
|
||||
blob tensor pattern) is opt-in via the `encoding` annotation. See
|
||||
[layout-engine.md](layout-engine.md) and
|
||||
[ADR-097](decisions/097-schema-annotations.md).
|
||||
|
||||
5. **TUnion supports both byte-offset and field-name discriminators.**
|
||||
Byte-offset for protocol dispatch (SFTP type bytes, call protocol
|
||||
event types). Field-name for the typedef.ts string pattern. See
|
||||
[data-access.md](data-access.md) and
|
||||
[ADR-097](decisions/097-schema-annotations.md).
|
||||
|
||||
6. **Endianness is per-schema, default little-endian.** The engine reads
|
||||
the `"endian"` annotation and byte-swaps accordingly. SFTP consumers
|
||||
specify `"endian": "big"`. See [layout-engine.md](layout-engine.md)
|
||||
and [ADR-097](decisions/097-schema-annotations.md).
|
||||
|
||||
7. **Validation is opt-in, built once at load time.** The jsonschema
|
||||
validator is compiled once at schema load time. Access-time validation
|
||||
is a fast `is_valid()` check. High-throughput paths can skip
|
||||
validation; security-sensitive paths can validate every frame. See
|
||||
[validation.md](validation.md) and
|
||||
[ADR-098](decisions/098-error-handling-validation-strategy.md).
|
||||
|
||||
8. **Not a serialization framework.** The typedef engine is not a
|
||||
general-purpose serde replacement. It operates on raw byte buffers at
|
||||
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||
with a known schema, use typedef. See [overview.md](overview.md) and
|
||||
[ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
|
||||
passing, two layout modes, TUnion dispatch, endianness)
|
||||
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||
JSON Schema as the binary struct engine" — the origin of this research
|
||||
thread
|
||||
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||
schema kinds (619 lines)
|
||||
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
|
||||
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
|
||||
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
|
||||
@@ -0,0 +1,385 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
---
|
||||
|
||||
# alknet-typedef — Data Access
|
||||
|
||||
The data access layer: read/write functions, TUnion dispatch, field paths,
|
||||
zero-copy access for fixed-size types, and length-prefix reading for
|
||||
variable-length types. This is the consumer-facing API — given a compiled
|
||||
`TypedefEngine` and a byte buffer, read and write fields at
|
||||
schema-computed offsets.
|
||||
|
||||
This document covers two layers:
|
||||
|
||||
- **Primitive read/write functions** in the `data_access` module —
|
||||
typed reads/writes at a caller-provided offset. These are the building
|
||||
blocks used by the layout types (`OffsetMap`, `LayoutBuilder`,
|
||||
`SequentialReader`) and the `TypedefEngine`. Each operates on a raw
|
||||
byte buffer at a known offset and returns a `TypedefError::Access`
|
||||
carrying the field path on bounds or encoding failures.
|
||||
- **The `FieldValue` enum and the higher-level APIs** —
|
||||
`TypedefEngine::read_field`/`write_field` (aligned mode) and
|
||||
`SequentialReader::read_next`/`read_field` (packed mode) — which look
|
||||
up a field's offset via the layout and dispatch to the primitive
|
||||
functions, returning a unified `FieldValue<'a>`.
|
||||
|
||||
## The `FieldValue` enum
|
||||
|
||||
The higher-level read APIs return a single unified type — `FieldValue<'a>`
|
||||
— so one method can read any field kind without the caller dispatching on
|
||||
schema kind first. The variant carries the typed value; the lifetime
|
||||
borrows from the input buffer for variable-length kinds (zero-copy).
|
||||
|
||||
```rust
|
||||
pub enum FieldValue<'a> {
|
||||
I8(i8), I16(i16), I32(i32), I64(i64),
|
||||
U8(u8), U16(u16), U32(u32), U64(u64),
|
||||
F32(f32), F64(f64),
|
||||
Bool(bool),
|
||||
Enum(u32), // u32 index into the schema's "enum" array
|
||||
String(&'a str), // borrows from the buffer
|
||||
Bytes(&'a [u8]), // borrows from the buffer
|
||||
Struct { start: usize, end: usize }, // consumer recurses with a fresh reader
|
||||
Union { discriminator: String, variant_start: usize },
|
||||
Array { count: u32, element_start: usize, element_stride: usize },
|
||||
}
|
||||
```
|
||||
|
||||
For composite kinds (`Struct`, `Union`, `Array`), `FieldValue` returns a
|
||||
layout descriptor, not the decoded contents — the consumer recurses with
|
||||
a fresh `SequentialReader` (or a sub-range read) scoped to the reported
|
||||
byte range. `Array`'s `element_stride` is `0` for variable-length element
|
||||
types, signalling the consumer must walk each element sequentially.
|
||||
|
||||
## Read/Write Model
|
||||
|
||||
The typedef engine operates on raw byte buffers (`&[u8]` for reading,
|
||||
`&mut [u8]` for writing). There is no intermediate `Value` tree, no
|
||||
reflection, no dynamic dispatch per field. The engine uses the offset map
|
||||
(or `LayoutBuilder`/`SequentialReader`) to locate fields, then performs
|
||||
typed access at the computed positions.
|
||||
|
||||
### Higher-level read/write
|
||||
|
||||
The `TypedefEngine` and `SequentialReader` provide the primary
|
||||
consumer-facing read/write APIs. They look up a field's offset via the
|
||||
layout and dispatch to the primitive `data_access` functions, returning
|
||||
`FieldValue` (read) or accepting `&FieldValue` (write).
|
||||
|
||||
```rust
|
||||
impl TypedefEngine {
|
||||
// Aligned mode: looks up the field's ByteRange in the OffsetMap,
|
||||
// dispatches to the right data_access function by TypeDefKind.
|
||||
// Returns TypedefError::Access if compiled in packed mode
|
||||
// (use sequential_reader() for packed mode).
|
||||
pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str)
|
||||
-> Result<FieldValue<'a>, TypedefError>;
|
||||
pub fn write_field(&self, buffer: &mut [u8], field_path: &str,
|
||||
value: &FieldValue<'_>) -> Result<(), TypedefError>;
|
||||
|
||||
// Packed mode: returns an owned fresh SequentialReader (ADR-101).
|
||||
// Each call returns a new reader with the cursor at position 0.
|
||||
// The consumer owns the reader and drives read_next/read_field/reset.
|
||||
pub fn sequential_reader(&self) -> Option<SequentialReader>;
|
||||
}
|
||||
|
||||
impl SequentialReader {
|
||||
// Packed mode: walks the buffer field-by-field, reading length
|
||||
// prefixes to find each field's position. read_field walks all
|
||||
// preceding fields to reach the target.
|
||||
pub fn read_next<'a>(&mut self, buffer: &'a [u8])
|
||||
-> Result<Option<(String, FieldValue<'a>)>, TypedefError>;
|
||||
pub fn read_field<'a>(&mut self, buffer: &'a [u8], field_path: &str)
|
||||
-> Result<FieldValue<'a>, TypedefError>;
|
||||
pub fn reset(&mut self);
|
||||
pub fn position(&self) -> usize;
|
||||
pub fn endian(&self) -> Endian;
|
||||
}
|
||||
```
|
||||
|
||||
`read_field`/`write_field` on `TypedefEngine` work for the fixed-size
|
||||
primitive kinds and the length-prefixed `String`/`Bytes`/`Timestamp`
|
||||
fields. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a
|
||||
`FieldValue` carrying a layout descriptor (byte range, variant start,
|
||||
or array stride) for the consumer to recurse on — see §"FieldValue" above.
|
||||
|
||||
For writing in packed mode, the consumer uses `LayoutBuilder::build` to
|
||||
compute positions, then calls the primitive `data_access::write_*`
|
||||
functions at the computed offsets. There is no packed-mode
|
||||
`engine.write_field` — the layout depends on the actual data sizes,
|
||||
which the builder consumes at `build` time.
|
||||
|
||||
### Primitive read/write functions
|
||||
|
||||
The `data_access` module exposes typed read/write functions for each
|
||||
primitive kind. Each takes `field_path: &str` for error attribution
|
||||
(produces a `TypedefError::Access` carrying the path on bounds or
|
||||
encoding failures) and, for multi-byte types, an `Endian` parameter.
|
||||
|
||||
### Fixed-size types
|
||||
|
||||
Fixed-size types (`TFloat32`, `TInt32`, `TUint8`, `TEnum`, etc.) are
|
||||
accessed via zero-copy reads of N bytes at the offset:
|
||||
|
||||
```rust
|
||||
// Read a u32 at a known offset, applying endianness. Bounds-checked.
|
||||
fn read_u32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
|
||||
-> Result<u32, TypedefError> {
|
||||
let bytes: [u8; 4] = read_array(buffer, offset, field_path)?;
|
||||
Ok(match endian {
|
||||
Endian::Little => u32::from_le_bytes(bytes),
|
||||
Endian::Big => u32::from_be_bytes(bytes),
|
||||
})
|
||||
}
|
||||
|
||||
// Write a u32 at a known offset, applying endianness. Bounds-checked.
|
||||
fn write_u32(buffer: &mut [u8], offset: usize, value: u32,
|
||||
field_path: &str, endian: Endian) -> Result<(), TypedefError> {
|
||||
let bytes = match endian {
|
||||
Endian::Little => value.to_le_bytes(),
|
||||
Endian::Big => value.to_be_bytes(),
|
||||
};
|
||||
write_array(buffer, offset, bytes, field_path)
|
||||
}
|
||||
```
|
||||
|
||||
The engine applies endianness at access time based on the schema's
|
||||
`"endian"` annotation (ADR-097). The offset computation is
|
||||
endian-agnostic. The `read_array`/`write_array` helpers perform the
|
||||
bounds check and produce `TypedefError::Access` with the field path on
|
||||
failure.
|
||||
|
||||
### TEnum access
|
||||
|
||||
`TEnum` is a fixed-size type (4 bytes, `u32` index). Read/write delegates
|
||||
to the `u32` primitives, applying the schema's endianness:
|
||||
|
||||
```rust
|
||||
pub fn read_enum(buffer: &[u8], offset: usize, field_path: &str, endian: Endian)
|
||||
-> Result<u32, TypedefError> {
|
||||
read_u32(buffer, offset, field_path, endian)
|
||||
}
|
||||
```
|
||||
|
||||
The consumer maps the `u32` index back to the enum's string values using
|
||||
the schema's `"enum"` array (index 0 → first value, index 1 → second
|
||||
value, etc.). The engine does not perform this mapping — it operates on
|
||||
the raw `u32` index. The jsonschema validator checks that the index
|
||||
corresponds to a valid enum value at the JSON level.
|
||||
|
||||
### Variable-length types (inline length-prefixing)
|
||||
|
||||
For variable-length types with inline length-prefixing (the default),
|
||||
the `data_access` module provides `read_string`/`write_string`/
|
||||
`read_bytes`/`write_bytes`. Each takes `field_path: &str` for error
|
||||
attribution and `endian` for the length prefix:
|
||||
|
||||
```rust
|
||||
// Read a length-prefixed string, borrowing from the buffer.
|
||||
fn read_string<'a>(buffer: &'a [u8], offset: usize,
|
||||
field_path: &str, endian: Endian) -> Result<&'a str, TypedefError>;
|
||||
|
||||
// Write a length-prefixed string. Returns total bytes written (4 + data.len()).
|
||||
fn write_string(buffer: &mut [u8], offset: usize, value: &str,
|
||||
field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
|
||||
|
||||
// read_bytes / write_bytes have the same shape — raw bytes, no UTF-8 check.
|
||||
```
|
||||
|
||||
The engine reads the 4-byte length prefix at the field's offset, then
|
||||
slices the data that follows. For writing, the engine writes the length
|
||||
prefix + data. `read_string` validates UTF-8 and returns a `&str`
|
||||
borrowing from the input buffer (zero-copy); `read_bytes` returns a
|
||||
`&[u8]` slice with no encoding check.
|
||||
|
||||
In packed sequential mode, the `SequentialReader` uses the length prefix
|
||||
to determine the position of the next field. In aligned static mode, the
|
||||
`OffsetMap` records the position of the length prefix; the variable data
|
||||
is accessed separately.
|
||||
|
||||
### Variable-length types (offset indirection)
|
||||
|
||||
For variable-length types with offset indirection (opt-in), the
|
||||
`data_access` module provides `read_string_indirect`/`read_bytes_indirect`.
|
||||
The 8-byte struct at `buffer[offset..offset+8]` is
|
||||
`{ data_offset: u32, data_length: u32 }` (endian-aware); the actual
|
||||
bytes live in a separate `data_region`:
|
||||
|
||||
```rust
|
||||
fn read_string_indirect<'a>(buffer: &'a [u8], offset: usize,
|
||||
data_region: &'a [u8], field_path: &str,
|
||||
endian: Endian) -> Result<&'a str, TypedefError>;
|
||||
fn read_bytes_indirect<'a>(buffer: &'a [u8], offset: usize,
|
||||
data_region: &'a [u8], field_path: &str,
|
||||
endian: Endian) -> Result<&'a [u8], TypedefError>;
|
||||
```
|
||||
|
||||
The field is a struct `{offset: u32, length: u32}` at a known position
|
||||
in the `OffsetMap`. The consumer provides the data region separately; the
|
||||
engine reads the offset and length, then slices the data region.
|
||||
|
||||
## TUnion Dispatch
|
||||
|
||||
The `tunion` module provides TUnion discriminator dispatch — reading the
|
||||
discriminator value from a byte buffer, looking up the variant schema in
|
||||
the union's `mapping`, and reporting the offset where the variant struct
|
||||
begins. All reads go through the `data_access` primitives so bounds checks
|
||||
and endianness handling are uniform with the rest of the engine.
|
||||
|
||||
The result of dispatch is a `UnionDispatch` struct:
|
||||
|
||||
```rust
|
||||
pub struct UnionDispatch {
|
||||
pub key: String, // mapping key (stringified disc value)
|
||||
pub variant_offset: usize, // byte offset where the variant struct starts
|
||||
pub discriminator_size: usize, // discriminator's byte size
|
||||
}
|
||||
```
|
||||
|
||||
After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)`
|
||||
to get the variant schema, then reads the variant's fields at
|
||||
`dispatch.variant_offset` using the normal `data_access` functions (or a
|
||||
fresh `SequentialReader` scoped to the variant).
|
||||
|
||||
### Byte-offset discriminator
|
||||
|
||||
```rust
|
||||
/// Read the discriminator value from a byte-offset TUnion. The discriminator
|
||||
/// is a fixed-size integer (TypeDef:Uint8/Uint16/Uint32) at a known byte
|
||||
/// offset. Returns the mapping key (stringified integer) and the variant
|
||||
/// struct offset.
|
||||
pub fn read_byte_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &Value,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, TypedefError>;
|
||||
```
|
||||
|
||||
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
|
||||
1..N are the variant struct. The call protocol's 5 event types
|
||||
(`call.requested` → 0x01, etc.) use the same pattern. The variant struct
|
||||
starts at `offset + discriminator_size`.
|
||||
|
||||
### Field-name discriminator
|
||||
|
||||
```rust
|
||||
/// Read the discriminator value from a field-name TUnion. The
|
||||
/// discriminator is a named field within the struct — the consumer
|
||||
/// provides the field's computed offset (from the OffsetMap or
|
||||
/// LayoutBuilder). Supports TypeDef:String, Uint8, and Enum discriminator
|
||||
/// fields.
|
||||
pub fn read_field_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &Value,
|
||||
disc_field_offset: usize,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, TypedefError>;
|
||||
```
|
||||
|
||||
The discriminator is a named field within the struct. Its offset is
|
||||
computed like any other field (the consumer passes it in as
|
||||
`disc_field_offset`). The mapping keys are string values. After reading
|
||||
the discriminator, the consumer looks up the variant schema and reads
|
||||
the variant's fields starting at the end of the discriminator field.
|
||||
|
||||
### Variant resolution
|
||||
|
||||
```rust
|
||||
/// Look up a variant schema from the union's mapping. Inline schemas
|
||||
/// are returned directly. $ref pointers of the form "#/$defs/<name>"
|
||||
/// are resolved against the union schema's own $defs block.
|
||||
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str)
|
||||
-> Result<&'a Value, TypedefError>;
|
||||
|
||||
/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a
|
||||
/// byte-offset TUnion. Field-name discriminators have no fixed size
|
||||
/// and produce a TypedefError::Schema.
|
||||
pub fn discriminator_size(union_schema: &Value) -> Result<usize, TypedefError>;
|
||||
```
|
||||
|
||||
### TUnion in the layout engines
|
||||
|
||||
The `LayoutBuilder` and `SequentialReader` also handle TUnion fields
|
||||
inline during traversal (the consumer does not need to call the `tunion`
|
||||
functions for a union field reached during a sequential walk). For
|
||||
`LayoutBuilder`, the consumer supplies the discriminator value (byte-offset)
|
||||
or variant index (field-name) in `var_sizes` under the synthetic key
|
||||
`"<union_path>.__discriminator"` or `"<union_path>.__variant"`. For
|
||||
`SequentialReader`, a union field yields
|
||||
`FieldValue::Union { discriminator, variant_start }`. The standalone
|
||||
`tunion` functions are for dispatch outside the layout walk — e.g., a
|
||||
consumer that receives a bare union buffer and needs to identify the
|
||||
variant before recursing.
|
||||
|
||||
## Field Paths
|
||||
|
||||
Fields are addressed by dotted paths: `"header.version"`, `"payload.data"`.
|
||||
Both `OffsetMap` and `PackedLayout` store fully-qualified paths (nested
|
||||
struct fields appear under their parent's path prefix). The higher-level
|
||||
APIs (`TypedefEngine::read_field`/`write_field`, `SequentialReader::read_field`)
|
||||
accept a field path, look up the byte range/position in the layout, and
|
||||
dispatch to the primitive `data_access` function for the field's kind.
|
||||
|
||||
For aligned-mode access, `TypedefEngine::read_field(&buffer, "header.version")`
|
||||
returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds
|
||||
the field's `TypeDef:*` kind in the schema, and calls the matching
|
||||
`data_access::read_*` function. `write_field` is the mirror. Composite
|
||||
kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue`
|
||||
carrying a layout descriptor; the consumer recurses with a fresh reader
|
||||
or sub-range read.
|
||||
|
||||
For packed-mode access, `SequentialReader::read_field(&buffer, "c")` walks
|
||||
all preceding fields to reach the target (sequential access is inherent
|
||||
to packed layouts). `read_next` walks fields in declaration order.
|
||||
|
||||
Nested structs produce nested field paths. The offset computation
|
||||
propagates the field path prefix during recursion, so the `OffsetMap`
|
||||
and `PackedLayout` contain entries like `"header.version"` and
|
||||
`"header.magic"`.
|
||||
|
||||
## Zero-Copy Access
|
||||
|
||||
For fixed-size types, the engine provides zero-copy access — the consumer
|
||||
gets a reference to the bytes in the buffer, not a copy. This is
|
||||
important for performance-sensitive paths (metatensor tensor access,
|
||||
high-throughput protocol parsing).
|
||||
|
||||
For variable-length types with inline length-prefixing, the engine
|
||||
returns a slice of the buffer — the string or byte array data is not
|
||||
copied. The consumer gets a `&str` or `&[u8]` that borrows from the
|
||||
input buffer.
|
||||
|
||||
For offset-indirect types, the consumer provides the data region; the
|
||||
engine returns a slice of that region.
|
||||
|
||||
## Error Handling
|
||||
|
||||
Read/write errors carry the field path for debugging. See
|
||||
[ADR-098](decisions/098-error-handling-validation-strategy.md) and
|
||||
[validation.md](validation.md) for the full error model.
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Two layout modes | [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Determines whether offsets are fixed (OffsetMap) or sequential (SequentialReader) |
|
||||
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Endianness, encoding, and TUnion discriminator shapes that control data access |
|
||||
| Error handling | [ADR-098](decisions/098-error-handling-validation-strategy.md) | Field-path-carrying errors for read/write operations |
|
||||
|
||||
## Open Questions
|
||||
|
||||
See [open-questions.md](open-questions.md) for full details.
|
||||
|
||||
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
|
||||
— affects the sequential walking logic for array access.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
|
||||
(read/write round-trip) and POC 2 (SFTP byte-identical round-trip)
|
||||
- [layout-engine.md](layout-engine.md) — offset computation that produces
|
||||
the positions this layer reads/writes at
|
||||
- [validation.md](validation.md) — validation that runs on the same
|
||||
buffers
|
||||
@@ -0,0 +1,173 @@
|
||||
# ADR-095: alknet-typedef — Purpose, Scope, and the jsonschema Engine
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Three threads in the codebase converge on the same pattern: a JSON Schema
|
||||
describes the shape of binary data, and the binary data is the struct's
|
||||
bytes at computed offsets.
|
||||
|
||||
1. **typedef.ts** (`/workspace/@alkdev/typebox/example/typedef/typedef.ts`,
|
||||
619 lines) defines custom TypeBox schema kinds (`TFloat32`, `TStruct`,
|
||||
`TUnion`, etc.) that carry binary layout semantics. These are registered
|
||||
via `TypeRegistry.Set` with custom validators.
|
||||
|
||||
2. **russh-sftp** has 29 packet types, each a struct with typed fields
|
||||
(`Read { id: u32, handle: String, offset: u64, len: u32 }`). The wire
|
||||
format is `[length: u32][type: u8][payload]` where payload is the
|
||||
struct's serde bytes. The `Packet` enum dispatches on the type byte —
|
||||
a tagged union of structs. Under the typedef lens, each packet is a
|
||||
`TStruct`; the `Packet` enum is a `TUnion` with a byte-offset
|
||||
discriminator.
|
||||
|
||||
3. **metatensor** needs an offset map for mmap-friendly tensor access —
|
||||
given a schema describing a model layout (ConvNet struct, tensor refs),
|
||||
compute byte offsets for each field so the consumer can read tensor
|
||||
data at known positions without parsing.
|
||||
|
||||
The common pattern: **a JSON Schema with `TypeDef:*` custom keywords
|
||||
describes the shape of binary data; the binary data is the struct's bytes
|
||||
at computed offsets.** The schema is the format definition; the engine is
|
||||
generic.
|
||||
|
||||
Two prior attempts built their own jsonschema engines — the fatal flaw:
|
||||
|
||||
- **typebox-rs** (`/workspace/@alkimiadev/typebox-rs/`, ~8,400 lines):
|
||||
a full 26-variant `SchemaKind` enum, a custom `Value` type with typed
|
||||
arrays, and a 912-line hand-written validator.
|
||||
- **alktype** (`/workspace/@alkimiadev/alktype/`, ~5,600 lines): a
|
||||
handler-registry pattern that also implements its own validation for
|
||||
each type.
|
||||
|
||||
The `jsonschema` crate (v0.46.5, Draft 2020-12) is already in the
|
||||
workspace at `/workspace/jsonschema/`. It handles validation with custom
|
||||
keyword support — the novel code is the offset computation, not the
|
||||
validation.
|
||||
|
||||
The call-channels-unification research
|
||||
(`docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||
JSON Schema as the binary struct engine") identified the convergence and
|
||||
bumped typedef up in the timeline. The POC
|
||||
(`docs/research/alknet-typedef/findings.md`, 26 tests passing) validated
|
||||
the approach: a ~1,900-line Rust crate that takes a JSON Schema with
|
||||
`TypeDef:*` custom keywords and produces an offset map, read/write
|
||||
functions, and validation — all driven by the schema.
|
||||
|
||||
## Decision
|
||||
|
||||
**alknet-typedef is a small Rust crate that takes a JSON Schema with
|
||||
`TypeDef:*` custom keywords and produces three capabilities:**
|
||||
|
||||
1. **An offset map** — walks the schema, computes byte offsets for each
|
||||
field based on type sizes, field order, and alignment.
|
||||
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
|
||||
read the field's bytes at its offset (zero-copy for fixed-size types).
|
||||
Given a `&mut [u8]` buffer, write a value at its offset.
|
||||
3. **Validation** — via `jsonschema` custom keywords, validates that data
|
||||
conforms to the schema's type constraints. The jsonschema validator
|
||||
operates on `serde_json::Value` instances (JSON representations), not
|
||||
raw byte buffers directly. A consumer that wants to validate a binary
|
||||
buffer reads it into a `Value` tree via the data access layer, then
|
||||
validates that `Value` against the jsonschema validator.
|
||||
|
||||
**The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||
`serde_json` (schema parsing).** The novel code is the offset computation
|
||||
— a recursive walk of the schema JSON that computes byte positions for
|
||||
each field. The custom keyword implementations are ~10 lines each.
|
||||
|
||||
**The schema is the format.** A JSON Schema with `TypeDef:Float32`,
|
||||
`TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
|
||||
the layout spec. No separate format definition, no separate parser, no
|
||||
separate validator. One schema, three uses: validate, compute offsets,
|
||||
access data.
|
||||
|
||||
**The crate depends on `jsonschema` and `serde_json` (with
|
||||
`preserve_order`).** No tokio, no platform deps. Compiles to
|
||||
`wasm32-unknown-unknown` for browser use. The `jsonschema` crate's
|
||||
`with_keyword("TypeDef:Float32", factory)` API is the integration point
|
||||
for custom type kinds — each `TypeDef:*` kind maps to a custom keyword
|
||||
validator in Rust. Same semantics as TypeBox's `TypeRegistry.Set`, same
|
||||
JSON Schema wire format.
|
||||
|
||||
**The crate targets `std` for v1.** The WASM target has `std` available
|
||||
via `wasm-bindgen`. If embedded use cases emerge, `no_std` + `alloc` can
|
||||
be added as a feature gate later — the engine's core (offset computation,
|
||||
read/write) is already allocation-free. See OQ-070.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Eliminates ~14,000 lines of hand-rolled schema engines.** typebox-rs
|
||||
and alktype are replaced by `jsonschema` + an offset map + ~50 lines of
|
||||
custom keyword implementations. The codebase drops from "a port of
|
||||
TypeBox" to "jsonschema + an offset map."
|
||||
- **One schema, three uses.** The same JSON Schema validates, computes
|
||||
offsets, and drives data access. No separate format definition, parser,
|
||||
or validator per protocol.
|
||||
- **Schema-driven, not code-driven.** Adding a new SFTP packet type is
|
||||
adding a variant to the schema JSON, not writing a new Rust struct +
|
||||
serde impl. The engine is generic; the schema is the configuration.
|
||||
- **WASM-clean.** `serde_json` + `jsonschema` + byte manipulation. No
|
||||
tokio, no platform deps. The same typedef schemas work in browser,
|
||||
Node, Python (via `wasmtime-py`), Go (via `wazero`), and any other
|
||||
WASM host.
|
||||
- **TypeBox interop.** TypeBox modules render to standard JSON Schema
|
||||
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
|
||||
on the Rust side. Zero translation. The same schema validates in both
|
||||
ecosystems.
|
||||
- **Defense in depth.** Schema validation via jsonschema custom keywords —
|
||||
a malformed binary payload can be read into a `Value` tree via the data
|
||||
access layer and validated against the schema before any consumer
|
||||
touches it. The `jsonschema` crate's compiled validators are fast enough
|
||||
to run on every incoming frame.
|
||||
|
||||
### Negative
|
||||
|
||||
- **New dependency on `jsonschema`.** The crate is already in the
|
||||
workspace but not yet used by any alknet crate. This is the first
|
||||
consumer.
|
||||
- **`serde_json` with `preserve_order` is required.** Field order is
|
||||
load-bearing for binary layouts. The `preserve_order` feature adds a
|
||||
small compile-time cost.
|
||||
- **Schema authoring is external.** Schemas are authored in TypeBox (JS)
|
||||
or hand-written JSON. The typedef engine consumes schemas; it does not
|
||||
generate them. A builder API is deferred (OQ-071).
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
|
||||
is a *format* (8-byte header + JSON header + binary data) that uses the
|
||||
typedef engine for its offset computation and tensor access.
|
||||
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
|
||||
`Value.Convert` — schema evolution — is out of scope for v1. The engine
|
||||
should not do anything that explicitly blocks adding a Value system
|
||||
later.
|
||||
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
|
||||
concern. The typedef engine consumes schemas; it does not generate them.
|
||||
- **Not a schema builder.** The typedef engine does not provide a fluent
|
||||
API for constructing schemas. Schemas are plain JSON.
|
||||
- **Not a serialization framework.** The typedef engine is not a
|
||||
general-purpose serde replacement. It operates on raw byte buffers at
|
||||
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||
with a known schema, use typedef.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
|
||||
passing, two layout modes, TUnion dispatch, endianness)
|
||||
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||
JSON Schema as the binary struct engine" — the origin of this research
|
||||
thread
|
||||
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||
schema kinds (619 lines)
|
||||
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes decision
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotation shapes
|
||||
- [ADR-098](098-error-handling-validation-strategy.md) — error handling
|
||||
and validation strategy
|
||||
@@ -0,0 +1,137 @@
|
||||
# ADR-096: Two Layout Modes — Packed Sequential vs Aligned Static
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The POCs surfaced that protocols and mmap-friendly formats need different
|
||||
layout strategies. POC 1 built an aligned `OffsetMap` with natural
|
||||
alignment padding — correct for mmap-friendly formats (metatensor) but
|
||||
wrong for protocol wire formats (SFTP, channels, TTY). POC 2 built a
|
||||
`LayoutBuilder` and `SequentialReader` for packed sequential layouts —
|
||||
correct for protocol wire formats but wrong for mmap-friendly formats.
|
||||
|
||||
This is the most important architectural finding from the POCs. The
|
||||
engine must support both modes; a single layout strategy cannot serve
|
||||
both use cases.
|
||||
|
||||
### Packed sequential layout (protocol wire formats)
|
||||
|
||||
Protocols pack fields sequentially with no alignment padding.
|
||||
Variable-length fields shift all subsequent fields. Writing requires
|
||||
knowing actual data sizes upfront; reading walks the buffer sequentially,
|
||||
reading length prefixes to determine positions.
|
||||
|
||||
This is the layout used by SFTP (all strings and byte arrays are
|
||||
length-prefixed inline), channels (`[channel_id: u32][size: u32][payload]`),
|
||||
TTY (`[stream_type: u8][length: u32][payload]`), and most binary protocols.
|
||||
|
||||
### Aligned static layout (mmap-friendly formats)
|
||||
|
||||
Fields have fixed positions with natural alignment padding.
|
||||
Variable-length fields get a 4-byte length prefix at a known offset; the
|
||||
variable data is not included in the static layout. This enables
|
||||
mmap-friendly random access — the consumer can read field N at a known
|
||||
offset without parsing the fields before it.
|
||||
|
||||
This is the layout used by metatensor (blob tensor pattern: index struct
|
||||
in one region, blob data in another) and safetensors (header + aligned
|
||||
tensor data).
|
||||
|
||||
## Decision
|
||||
|
||||
**The typedef engine supports two layout modes, selected by the consumer
|
||||
at engine construction time:**
|
||||
|
||||
### Mode 1: Packed sequential (`LayoutBuilder` / `SequentialReader`)
|
||||
|
||||
For protocol wire formats. Fields are packed with no alignment padding.
|
||||
Variable-length fields shift all subsequent fields.
|
||||
|
||||
- **LayoutBuilder** — takes a schema and actual data sizes for
|
||||
variable-length fields, computes byte positions for each field in a
|
||||
packed layout. Used at write time when the consumer knows the data
|
||||
sizes upfront.
|
||||
- **SequentialReader** — walks a buffer field-by-field according to the
|
||||
schema, reading length prefixes to determine variable-length data
|
||||
positions. Used at read time when the consumer is parsing an incoming
|
||||
frame.
|
||||
|
||||
The `LayoutBuilder` and `SequentialReader` are the primary interface for
|
||||
protocol consumers (SFTP, binary call frames, TTY negotiation).
|
||||
|
||||
### Mode 2: Aligned static (`OffsetMap`)
|
||||
|
||||
For mmap-friendly formats. Fields have fixed positions with natural
|
||||
alignment padding. Variable-length fields get a 4-byte length prefix at
|
||||
a known offset; the variable data is not included in the static layout.
|
||||
|
||||
- **OffsetMap** — walks the schema once, computes fixed byte positions
|
||||
for each field based on type sizes and alignment. The output is a flat
|
||||
table of `(field_path, byte_range)` pairs. Used for both read and write
|
||||
at known offsets.
|
||||
|
||||
The `OffsetMap` is the primary interface for mmap consumers (metatensor).
|
||||
|
||||
### Variable-length handling in each mode
|
||||
|
||||
**Packed sequential mode:** Variable-length fields are inline
|
||||
length-prefixed by default (`[length: u32][data]`). The `LayoutBuilder`
|
||||
takes the actual data size to compute the length prefix value and the
|
||||
position of subsequent fields. The `SequentialReader` reads the length
|
||||
prefix to determine the data extent and the position of the next field.
|
||||
|
||||
**Aligned static mode:** Variable-length fields get a 4-byte length
|
||||
prefix at a known offset. The variable data lives outside the static
|
||||
layout — either immediately after the fixed fields (inline
|
||||
length-prefixing) or in a separate data region (offset indirection, the
|
||||
metatensor blob tensor pattern). The `OffsetMap` records the position of
|
||||
the length prefix (or the `{offset, length}` pair for offset-indirect
|
||||
fields).
|
||||
|
||||
### Default for variable-length types
|
||||
|
||||
Inline length-prefixing (`[length: u32][data]`) is the default for all
|
||||
variable-length types in both modes. This is the universal pattern used
|
||||
by channels, SFTP, TTY, and most binary protocols. Offset indirection is
|
||||
opt-in via the `encoding` annotation (see ADR-097).
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **One engine, two modes.** The same schema can be used in either mode.
|
||||
A schema describing an SFTP packet can be consumed by a `SequentialReader`
|
||||
(for parsing incoming frames) and a `LayoutBuilder` (for constructing
|
||||
outgoing frames). A schema describing a metatensor layout can be
|
||||
consumed by an `OffsetMap` (for mmap access).
|
||||
- **Correct for both use cases.** Packed sequential mode produces
|
||||
byte-identical output to hand-written protocol serialization (validated
|
||||
by POC 2's russh-sftp round-trip tests). Aligned static mode produces
|
||||
correct offsets for mmap-friendly access (validated by POC 1's
|
||||
alignment tests).
|
||||
- **No mode confusion.** The consumer explicitly selects the mode at
|
||||
engine construction time. A protocol consumer never accidentally gets
|
||||
alignment padding; an mmap consumer never accidentally gets
|
||||
variable-length field shifting.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Two APIs to learn.** Consumers must choose between
|
||||
`LayoutBuilder`/`SequentialReader` and `OffsetMap`. The choice is
|
||||
determined by the use case (protocol vs mmap), not by the schema.
|
||||
- **Variable-length fields in packed mode require size foreknowledge.**
|
||||
The `LayoutBuilder` needs actual data sizes for variable-length fields
|
||||
to compute correct positions for subsequent fields. This is inherent
|
||||
to packed layouts — the consumer must know the data sizes before
|
||||
writing.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
|
||||
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotations including
|
||||
the `encoding` field for variable-length types
|
||||
@@ -0,0 +1,259 @@
|
||||
# ADR-097: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The typedef engine needs concrete JSON shapes for schema-level
|
||||
annotations that control binary layout behavior. The POCs validated the
|
||||
semantics; this ADR pins the shapes.
|
||||
|
||||
Four annotation categories need concrete shapes:
|
||||
|
||||
1. **Endianness** — safetensors is little-endian, SFTP is big-endian.
|
||||
The engine needs to know which to use.
|
||||
2. **Alignment** — different backends have different alignment
|
||||
requirements (wgpu: 256-byte, protocols: natural, mmap: page).
|
||||
3. **Variable-length encoding** — inline length-prefixing vs offset
|
||||
indirection for strings, byte arrays, and other variable-length types.
|
||||
4. **TUnion discriminators** — byte-offset (protocol dispatch) vs
|
||||
field-name (typedef.ts pattern).
|
||||
|
||||
## Decision
|
||||
|
||||
### 1. Endianness
|
||||
|
||||
**Schema-level annotation with a default of little-endian.**
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Struct": true,
|
||||
"endian": "big",
|
||||
"properties": { ... }
|
||||
}
|
||||
```
|
||||
|
||||
- `"endian": "little"` (default) — read/write in little-endian byte order.
|
||||
- `"endian": "big"` — read/write in big-endian byte order.
|
||||
- The annotation applies to the entire schema and all nested types.
|
||||
- Mixed endianness within one schema is not supported (pathological; no
|
||||
known protocol requires it).
|
||||
- The default is little-endian, matching safetensors, wgpu, and most
|
||||
modern formats. SFTP consumers specify `"endian": "big"`.
|
||||
|
||||
### 2. Alignment
|
||||
|
||||
**Both struct-level and field-level, with field-level overriding
|
||||
struct-level.**
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Struct": true,
|
||||
"align": 256,
|
||||
"properties": {
|
||||
"header": { "TypeDef:Struct": true, "properties": { ... } },
|
||||
"weight": { "TypeDef:Float32": true, "align": 16 }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- Struct-level `"align"` sets the default alignment for all fields in
|
||||
that struct. The struct's total size is rounded up to this alignment.
|
||||
- Field-level `"align"` overrides the struct default for that specific
|
||||
field.
|
||||
- Default alignment (when no annotation is present): 1 for u8/bool, 2
|
||||
for u16/i16, 4 for u32/i32/f32, 8 for u64/i64/f64, max field alignment
|
||||
for structs.
|
||||
- Alignment is only meaningful in aligned static mode (ADR-096). In
|
||||
packed sequential mode, alignment annotations are ignored — fields are
|
||||
packed with no padding.
|
||||
|
||||
### 3. Variable-length encoding
|
||||
|
||||
**Three strategies for variable-length types, selected by the `encoding`
|
||||
annotation and the standard JSON Schema `maxLength` keyword.**
|
||||
|
||||
```json
|
||||
// Strategy 1: Inline length-prefixing (default, shorthand)
|
||||
{ "TypeDef:String": true }
|
||||
|
||||
// Strategy 1: Explicit inline length-prefixing
|
||||
{ "TypeDef:String": { "encoding": "length-prefixed" } }
|
||||
|
||||
// Strategy 2: Fixed-size reservation (uses standard maxLength)
|
||||
{ "TypeDef:String": true, "maxLength": 256 }
|
||||
|
||||
// Strategy 3: Offset indirection (opt-in)
|
||||
{ "TypeDef:String": { "encoding": "offset-indirect" } }
|
||||
```
|
||||
|
||||
**Strategy 1: Inline length-prefixing (default).** The field's fixed
|
||||
portion is a 4-byte length prefix at a computed offset. The variable data
|
||||
follows immediately after. In packed sequential mode, the length prefix
|
||||
determines the position of subsequent fields. In aligned static mode, the
|
||||
length prefix is at a known offset; the variable data is not included in
|
||||
the static layout. This is the universal pattern used by channels, SFTP,
|
||||
TTY, and most binary protocols.
|
||||
|
||||
**Strategy 2: Fixed-size reservation.** When a variable-length field
|
||||
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
|
||||
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
|
||||
than `maxLength` is zero-padded; data longer than `maxLength` is a
|
||||
validation error. This makes the field fixed-size from the layout
|
||||
perspective — subsequent fields have known, unchanging offsets. This is
|
||||
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
|
||||
pattern for fields with known maximum sizes. In packed sequential mode,
|
||||
`maxLength` is a validation constraint only — the engine still uses
|
||||
inline length-prefixing (strategy 1).
|
||||
|
||||
**Strategy 3: Offset indirection.** The field is a struct
|
||||
`{offset: u32, length: u32}` that points into a separate data region.
|
||||
This is the metatensor blob tensor pattern — the index struct lives in
|
||||
one region, the blob data lives in another. The consumer provides the
|
||||
data region separately. Enables mmap-friendly random access to
|
||||
variable-length data without parsing length prefixes and without
|
||||
reserving worst-case space.
|
||||
|
||||
**Default strategy selection:**
|
||||
- In packed sequential mode: always strategy 1 (inline length-prefixing).
|
||||
`maxLength` is a validation constraint only.
|
||||
- In aligned static mode: strategy 2 (fixed-size reservation) if
|
||||
`maxLength` is declared; strategy 3 (offset indirection) if
|
||||
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
|
||||
length-prefixing) otherwise.
|
||||
|
||||
- `true` is a shorthand for the default (length-prefixed). This keeps
|
||||
the common case concise and the override explicit.
|
||||
- The `encoding` annotation and `maxLength` apply to all variable-length
|
||||
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
|
||||
`TypeDef:Record`, `TypeDef:Timestamp`.
|
||||
|
||||
### 3a. TRecord value type
|
||||
|
||||
`TypeDef:Record` is a string-keyed map. The value type is declared via
|
||||
the `"values"` property in the schema:
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Record": true,
|
||||
"values": { "TypeDef:Float32": true }
|
||||
}
|
||||
```
|
||||
|
||||
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
|
||||
values in the record. All values share the same type.
|
||||
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
|
||||
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count`
|
||||
times. Each key is a length-prefixed UTF-8 string. Each value is
|
||||
encoded according to its declared `TypeDef:*` kind — a `Record<Uint32>`
|
||||
value is 4 raw bytes; a `Record<String>` value is itself a
|
||||
length-prefixed string; a `Record<Struct>` value is the struct's
|
||||
fields laid out inline. There is **no separate `value_len` prefix** —
|
||||
the value's size is determined by its kind (fixed-size kinds have a
|
||||
known size; variable-length kinds carry their own length prefix).
|
||||
- The count and key-length prefixes respect the schema's endianness.
|
||||
- In aligned static mode with `maxLength`, the entire record is reserved
|
||||
at `maxLength` bytes (zero-padded).
|
||||
|
||||
### 4. TUnion discriminators
|
||||
|
||||
**Two discriminator kinds: byte-offset (protocol dispatch) and
|
||||
field-name (typedef.ts pattern).**
|
||||
|
||||
#### Kind A: Byte-offset discriminator
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "TypeDef:Uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"1": { "$ref": "#/$defs/Init" },
|
||||
"3": { "$ref": "#/$defs/Open" },
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" },
|
||||
"101": { "$ref": "#/$defs/Status" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- The discriminator is a fixed-size integer at a known byte offset.
|
||||
- `"offset"` is the byte position of the discriminator within the union's
|
||||
buffer.
|
||||
- `"type"` is the `TypeDef:*` kind of the discriminator (typically
|
||||
`TypeDef:Uint8` for protocol type bytes).
|
||||
- The mapping keys are stringified integers (`"1"`, `"5"`, `"101"`).
|
||||
The engine parses the key to match the discriminator value.
|
||||
- The variant struct starts at `offset + discriminator_size`.
|
||||
- This is the SFTP `Packet` enum pattern and the call protocol's event
|
||||
type dispatch.
|
||||
|
||||
#### Kind B: Field-name discriminator
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "field",
|
||||
"name": "type"
|
||||
},
|
||||
"mapping": {
|
||||
"read": { "$ref": "#/$defs/Read" },
|
||||
"write": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- The discriminator is a named field within the struct.
|
||||
- `"name"` is the field name that holds the discriminator value.
|
||||
- The mapping keys are string values matching the discriminator field's
|
||||
value.
|
||||
- The discriminator field is just another field in the struct — its
|
||||
offset is computed like any other field.
|
||||
- This is the typedef.ts `TUnion` pattern.
|
||||
|
||||
#### Mapping values
|
||||
|
||||
Mapping values may be either inline schemas or `$ref` pointers. `$ref`
|
||||
is cleaner for large unions (29 SFTP variants) but requires a `$defs`
|
||||
section. Inline schemas are simpler for small unions (5 call protocol
|
||||
event types). Both work.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Concrete, validated shapes.** All four annotation categories have
|
||||
concrete JSON shapes that were validated by the POCs.
|
||||
- **Sensible defaults.** Little-endian, natural alignment, inline
|
||||
length-prefixing — the common case requires no annotations.
|
||||
- **Explicit overrides.** Big-endian, custom alignment, offset
|
||||
indirection — the uncommon case is explicit and self-documenting.
|
||||
- **TUnion covers both protocol and typedef.ts patterns.** The
|
||||
byte-offset discriminator handles SFTP type bytes and call protocol
|
||||
event types. The field-name discriminator handles the typedef.ts string
|
||||
pattern. No separate union type needed.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Keyword value shape change.** `"TypeDef:String": true` (boolean) and
|
||||
`"TypeDef:String": { "encoding": "length-prefixed" }` (object) are both
|
||||
valid. The engine must handle both shapes. This is a minor parsing
|
||||
concern — the POC already handles it.
|
||||
- **Alignment annotations are mode-specific.** Alignment is only
|
||||
meaningful in aligned static mode. In packed sequential mode, alignment
|
||||
annotations are ignored. This is documented, not enforced — a consumer
|
||||
that specifies alignment in packed mode gets no error, just no effect.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — the
|
||||
annotation shape questions this ADR resolves
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes (alignment only meaningful in aligned static mode)
|
||||
@@ -0,0 +1,157 @@
|
||||
# ADR-098: Error Handling and Validation Strategy
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The typedef engine operates in three phases, each with distinct error
|
||||
conditions:
|
||||
|
||||
1. **Schema parsing** — invalid JSON, missing required keywords, unknown
|
||||
`TypeDef:*` kinds, malformed annotations.
|
||||
2. **Offset computation** — field not found, type not supported for
|
||||
offset computation, recursive schema depth exceeded.
|
||||
3. **Read/write** — buffer too short, invalid UTF-8, value out of range
|
||||
for the target type.
|
||||
4. **Validation** — type constraint violations (range, UTF-8, field
|
||||
presence, discriminator membership).
|
||||
|
||||
The engine also needs a clear strategy for *when* validation happens:
|
||||
once at schema load time (build the validator) vs repeatedly at access
|
||||
time (validate each buffer).
|
||||
|
||||
## Decision
|
||||
|
||||
### Error type: `TypedefError`
|
||||
|
||||
A single `TypedefError` enum with variants for each error category:
|
||||
|
||||
```rust
|
||||
pub enum TypedefError {
|
||||
/// Schema parsing errors.
|
||||
Schema(String),
|
||||
/// Offset computation errors.
|
||||
Offset { field_path: String, reason: String },
|
||||
/// Read/write errors.
|
||||
Access { field_path: String, reason: String },
|
||||
/// Validation errors (delegated to jsonschema).
|
||||
Validation(ValidationError<'static>),
|
||||
}
|
||||
```
|
||||
|
||||
- `Schema` — for invalid JSON, missing required keywords, unknown
|
||||
`TypeDef:*` kinds. The error message describes the problem.
|
||||
- `Offset` — for field-not-found, unsupported type for offset
|
||||
computation, etc. Carries the field path for debugging.
|
||||
- `Access` — for buffer-too-short, invalid UTF-8, value out of range.
|
||||
Carries the field path for debugging.
|
||||
- `Validation` — wraps `jsonschema`'s `ValidationError`. The
|
||||
`jsonschema` crate already provides rich error messages with schema
|
||||
paths; the typedef engine does not re-wrap or re-interpret them.
|
||||
|
||||
The `Validation` variant uses `ValidationError<'static>` because the
|
||||
validator is built once at schema load time and lives for the lifetime
|
||||
of the `TypedefEngine`. The `'static` lifetime is correct — the validator
|
||||
owns its schema reference.
|
||||
|
||||
### Validation timing: load-time build, access-time check
|
||||
|
||||
The jsonschema validator is built once at schema load time
|
||||
(`validator_for(&schema)?`) and then called repeatedly
|
||||
(`validator.is_valid(&instance)`). The typedef engine follows the same
|
||||
pattern:
|
||||
|
||||
1. **Load time:** Parse the schema JSON, build the offset map (or
|
||||
`LayoutBuilder`/`SequentialReader`), build the jsonschema validator.
|
||||
This is the `TypedefEngine::compile(schema: &Value) -> Result<Self,
|
||||
TypedefError>` constructor.
|
||||
2. **Access time:** Use the compiled engine for repeated read/write
|
||||
operations. Validation is opt-in per operation — the consumer calls
|
||||
`engine.validate(buffer)` when validation is desired.
|
||||
|
||||
The `TypedefEngine` struct is the compiled form of a schema:
|
||||
|
||||
```rust
|
||||
pub struct TypedefEngine {
|
||||
offset_map: OffsetMap, // or LayoutBuilder/SequentialReader
|
||||
validator: jsonschema::Validator, // compiled once at load time
|
||||
}
|
||||
```
|
||||
|
||||
### Custom keyword validators
|
||||
|
||||
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
|
||||
`jsonschema::options().with_keyword(...)`. The validators check:
|
||||
|
||||
- **Numeric types** (`TypeDef:Float32`, `TypeDef:Int8`, etc.): range
|
||||
constraints (Int8: -128..127, Uint8: 0..255, etc.), finiteness for
|
||||
floats.
|
||||
- **`TypeDef:String`**: UTF-8 validity.
|
||||
- **`TypeDef:Struct`**: field presence and types (delegated to
|
||||
jsonschema's structural validation — the custom keyword only needs to
|
||||
validate that the struct's fields match their declared `TypeDef:*`
|
||||
kinds).
|
||||
- **`TypeDef:Union`**: discriminator value membership in the mapping.
|
||||
- **`TypeDef:Array`**: element type conformance.
|
||||
- **`TypeDef:Boolean`**: value is `true` or `false`.
|
||||
- **`TypeDef:Timestamp`**: RFC 3339 string format (the internet profile of ISO 8601).
|
||||
|
||||
The `jsonschema` crate handles all the structural validation (object
|
||||
properties, required fields, array items, enum values) — the custom
|
||||
keywords only need to validate the leaf type constraints. Each custom
|
||||
keyword implementation is ~10 lines.
|
||||
|
||||
### Read/write errors carry field paths
|
||||
|
||||
Read/write errors include the field path for debugging:
|
||||
|
||||
```rust
|
||||
// Example: reading a u32 from a buffer that's too short
|
||||
Err(TypedefError::Access {
|
||||
field_path: "header.version".to_string(),
|
||||
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
|
||||
})
|
||||
```
|
||||
|
||||
This makes debugging binary format issues tractable — the error tells
|
||||
you exactly which field failed and why.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Single error type.** Consumers handle one `TypedefError` enum, not
|
||||
multiple error types from different engine phases.
|
||||
- **Field-path-carrying errors.** Read/write errors include the field
|
||||
path, making binary format debugging tractable.
|
||||
- **Validation is opt-in.** The consumer decides when to validate.
|
||||
High-throughput paths can skip validation; security-sensitive paths
|
||||
can validate every frame.
|
||||
- **jsonschema integration is clean.** The `ValidationError` is wrapped
|
||||
as-is — no re-interpretation, no information loss.
|
||||
- **Load-time build, access-time use.** The expensive work (schema
|
||||
parsing, validator compilation, offset computation) happens once at
|
||||
load time. Access-time operations are cheap (pointer casts, slice
|
||||
operations, length-prefix reads).
|
||||
|
||||
### Negative
|
||||
|
||||
- **`ValidationError<'static>` lifetime.** The `'static` lifetime on the
|
||||
`Validation` variant means the error cannot borrow from the buffer
|
||||
being validated. This is correct (the validator owns its schema
|
||||
reference) but may surprise readers who expect a shorter lifetime.
|
||||
- **No error recovery.** The engine does not attempt to recover from
|
||||
partial reads or writes. A buffer-too-short error on field N means
|
||||
fields N+1.. are also unreadable. This is inherent to binary formats
|
||||
— there is no "skip to next field" without a schema-driven parser.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — error
|
||||
handling strategy question (OQ 8)
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotations
|
||||
@@ -0,0 +1,114 @@
|
||||
# ADR-099: Int64/Uint64 as First-Class Kinds
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The typedef engine's kind set (ADR-095, ADR-097) tops out at 32-bit
|
||||
integers. The POC included `u64` read/write primitives, and the
|
||||
call-channels-unification research's own SFTP schema example uses
|
||||
`"TypeDef:Uint64"` for the `offset` field (`Read`/`Write` packets have
|
||||
`offset: u64`). Metatensor/safetensors `data_offsets` are also `u64`.
|
||||
|
||||
A `TypeDef:Uint64` variant was added to the `TypeDefKind` enum during
|
||||
implementation (the task decomposition correctly identified the gap),
|
||||
but without an ADR the addition was half-finished: `type_size()`
|
||||
returned `None`, the layout engines couldn't compute offsets for it, and
|
||||
the validator didn't register a `TypeDef:Uint64` keyword. The variant
|
||||
was then removed (commit `14d9cf2`) on the grounds that it was
|
||||
unintended and a latent panic — but the underlying gap is real: SFTP and
|
||||
metatensor, the two primary POC targets, both require 64-bit integers.
|
||||
|
||||
The presumed reason 64-bit integers were left out of the original
|
||||
specification is a JSON-level concern: `serde_json::Number` loses
|
||||
precision past 2^53 when parsing from JSON text. This is a
|
||||
*validation-layer* caveat, not a *layout-layer* one — the binary layout
|
||||
is 8 raw bytes, and `from_le_bytes`/`from_be_bytes` work correctly for
|
||||
the full `u64`/`i64` range. The validation concern is handled by
|
||||
accepting integer-form JSON values (the `jsonschema` crate's
|
||||
`as_i64`/`as_u64` methods handle the common range; values past 2^53 are
|
||||
a JSON representation limitation, not a typedef limitation).
|
||||
|
||||
## Decision
|
||||
|
||||
**Add `TypeDef:Int64` and `TypeDef:Uint64` as first-class kinds.**
|
||||
|
||||
Both are fixed-size (8 bytes), with natural alignment 8. They follow
|
||||
the schema's endianness annotation like all other fixed-size types.
|
||||
Read/write is via `data_access::read_i64`/`write_i64`/`read_u64`/
|
||||
`write_u64` (endian-aware, 8 bytes).
|
||||
|
||||
### Kind table additions
|
||||
|
||||
| Kind | TypeBox key | Rust type | Size | Alignment |
|
||||
|------|-------------|-----------|------|-----------|
|
||||
| `TInt64` | `TypeDef:Int64` | `i64` | 8 | 8 |
|
||||
| `TUint64` | `TypeDef:Uint64` | `u64` | 8 | 8 |
|
||||
|
||||
### Validation
|
||||
|
||||
The custom keyword validators check:
|
||||
- `TypeDef:Int64`: value must be an integer in `i64::MIN..=i64::MAX`
|
||||
(`-9223372036854775808` to `9223372036854775807`).
|
||||
- `TypeDef:Uint64`: value must be a non-negative integer in
|
||||
`0..=u64::MAX` (`0` to `18446744073709551615`).
|
||||
|
||||
The `jsonschema` crate's `as_i64`/`as_u64` handle the common range.
|
||||
JSON numbers past 2^53 lose precision in the JSON representation —
|
||||
this is a JSON limitation, not a typedef limitation. The binary
|
||||
representation (8 raw bytes) is always exact. A consumer that needs
|
||||
to validate the full 64-bit range from JSON should provide the value
|
||||
as a JSON integer (which `serde_json` preserves for values up to
|
||||
`u64::MAX`/`i64::MIN` when the `arbitrary_precision` feature is
|
||||
enabled, or when the value fits in `i64`/`u64` without the feature).
|
||||
|
||||
### `FieldValue` additions
|
||||
|
||||
`FieldValue::I64(i64)` and `FieldValue::U64(u64)` are added to the
|
||||
unified return type. The `SequentialReader`, `TypedefEngine::read_field`,
|
||||
and `TypedefEngine::write_field` dispatch on the new kinds.
|
||||
|
||||
### Kind count
|
||||
|
||||
The engine now has **19** first-class kinds (17 + Int64 + Uint64).
|
||||
`TypeDefKind::is_fixed_size()` returns `true` for both new kinds.
|
||||
`type_size()` returns `Some(8)`. `natural_alignment()` returns `8`.
|
||||
`needs_endian()` returns `true`.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Unblocks the two primary POC targets.** SFTP `Read`/`Write` packets
|
||||
(`offset: u64`) and metatensor `data_offsets` (`u64`) are now
|
||||
expressible in typedef schemas.
|
||||
- **Completes the half-finished addition.** The `TypeDefKind` enum,
|
||||
`data_access` primitives, and `FieldValue` variants for 64-bit
|
||||
integers now have matching layout, validator, and engine support.
|
||||
- **No new design surface.** Int64/Uint64 are fixed-size types that
|
||||
follow all existing patterns (endianness, alignment, zero-copy
|
||||
read/write). They are mechanical additions.
|
||||
|
||||
### Negative
|
||||
|
||||
- **JSON precision caveat.** Values past 2^53 lose precision in the
|
||||
JSON representation (not in the binary representation). This is a
|
||||
JSON limitation, not a typedef limitation, but it means the
|
||||
validation layer cannot perfectly round-trip the full 64-bit range
|
||||
through JSON `Number` without `arbitrary_precision`. In practice,
|
||||
SFTP offsets and tensor data offsets are well within 2^53.
|
||||
- **Two more kinds to maintain.** The kind table, validator
|
||||
registration, `FieldValue` enum, and dispatch arms all grow by two
|
||||
variants. This is the cost of completeness.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/call-channels-unification/findings.md` §"russh-sftp" —
|
||||
the SFTP schema with `"offset": { "TypeDef:Uint64": true }`
|
||||
- `docs/research/alknet-typedef/findings.md` §"POC 1" — the POC included
|
||||
u64 read/write
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope (the kind set)
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotations
|
||||
(endianness applies to the new kinds)
|
||||
+112
@@ -0,0 +1,112 @@
|
||||
# ADR-100: Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The aligned static layout mode (ADR-096) is designed for mmap-friendly
|
||||
formats: fields have fixed positions with natural alignment padding,
|
||||
enabling random access by field path without parsing preceding fields.
|
||||
|
||||
The spec (layout-engine.md) says variable-length fields in aligned mode
|
||||
get a 4-byte length prefix at a known offset, and "the variable data
|
||||
lives outside the static layout — either immediately after the fixed
|
||||
fields (inline length-prefixing) or in a separate data region (offset
|
||||
indirection)."
|
||||
|
||||
The implementation has a bug: `OffsetMap::compute` reserves only 4 bytes
|
||||
for an inline length-prefixed variable field (the length prefix), but
|
||||
`TypedefEngine::write_field` for a `String`/`Bytes` field calls
|
||||
`data_access::write_string` at `range.start`, which writes
|
||||
`[4-byte length][data]` inline — clobbering every subsequent field. The
|
||||
`read_field` path has the mirror behavior (reads inline), so the engine
|
||||
is self-consistent but only works correctly when the variable field is
|
||||
the last field in the struct (no subsequent field to clobber).
|
||||
|
||||
Concretely, `{name: String, id: Uint32}` in aligned mode maps
|
||||
`name → 0..4`, `id → 4..8`. Writing `"hello"` to `name` writes
|
||||
`[5,0,0,0,h,e,l,l,o]` at offset 0, overwriting `id`'s range with
|
||||
`hello`. All existing tests happen to put the variable field last, so
|
||||
the bug is latent.
|
||||
|
||||
The spec's "data region after fixed fields" model (where variable data
|
||||
lives after all fixed fields) is the correct design for aligned mode,
|
||||
but implementing it would require a two-region layout (fixed fields +
|
||||
variable data region) with the `OffsetMap` tracking both the prefix
|
||||
position and the data position. This is a significant design addition
|
||||
for a use case that doesn't exist yet — real aligned-format consumers
|
||||
(metatensor, safetensors) use `maxLength` reservation or
|
||||
`offset-indirect` encoding for variable data, not inline
|
||||
length-prefixing.
|
||||
|
||||
## Decision
|
||||
|
||||
**Reject non-final inline length-prefixed variable fields in aligned
|
||||
static mode at `OffsetMap::compute` time.**
|
||||
|
||||
A variable-length field (`TypeDef:String`, `TypeDef:Bytes`,
|
||||
`TypeDef:Timestamp`, `TypeDef:Record`) in aligned static mode that uses
|
||||
the default inline length-prefixing strategy (no `maxLength`, no
|
||||
`offset-indirect`) must be the last field in its struct. If a non-final
|
||||
inline length-prefixed variable field is encountered,
|
||||
`OffsetMap::compute` returns `TypedefError::Offset` with a message
|
||||
explaining that non-final variable fields in aligned mode require
|
||||
`maxLength` (fixed-size reservation) or `"encoding": "offset-indirect"`
|
||||
(offset indirection).
|
||||
|
||||
This is a validation-time rejection (schema load time), not a runtime
|
||||
check. The consumer learns about the problem when compiling the schema,
|
||||
not when writing data.
|
||||
|
||||
### What is NOT rejected
|
||||
|
||||
- Inline length-prefixed variable fields that are the last field in
|
||||
their struct — these are fine (no subsequent field to clobber).
|
||||
- `maxLength` reservation and `offset-indirect` encoding in any
|
||||
position — these make the field fixed-size from the layout
|
||||
perspective (known size at a known offset), so they don't clobber.
|
||||
- Inline length-prefixed variable fields in packed sequential mode —
|
||||
packed mode doesn't have fixed offsets; variable fields shift
|
||||
subsequent fields by design.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Eliminates a silent data-corruption bug.** A consumer that writes
|
||||
a non-final string in aligned mode currently clobbers subsequent
|
||||
fields with no error. After this fix, the schema is rejected at
|
||||
compile time.
|
||||
- **Matches real aligned-format usage.** mmap-friendly formats use
|
||||
`maxLength` or `offset-indirect` for variable data; inline
|
||||
length-prefixing in aligned mode is only meaningful as the last
|
||||
field.
|
||||
- **Simple to implement.** A single check in `compute_struct` (is this
|
||||
variable field non-final and using inline length-prefixing? → reject).
|
||||
No two-region layout needed.
|
||||
- **Defers the two-region design without blocking consumers.** If a
|
||||
future consumer needs inline length-prefixing in non-final position
|
||||
in aligned mode, the two-region layout can be implemented then. The
|
||||
rejection is reversible (remove the check, add the two-region logic).
|
||||
|
||||
### Negative
|
||||
|
||||
- **A schema that worked before (silently corrupting data) now fails
|
||||
at compile time.** This is the correct behavior — the schema was
|
||||
always broken, it just wasn't caught.
|
||||
- **The "data region after fixed fields" model from the spec is not
|
||||
implemented.** A consumer that wants inline variable data in a
|
||||
non-final position must use packed mode or wait for the two-region
|
||||
layout. This is acceptable for v1 — no current consumer needs it.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes (aligned static mode's variable-length handling)
|
||||
- [ADR-097](097-schema-annotations.md) — the three variable-length
|
||||
encoding strategies (`maxLength`, `offset-indirect`, inline
|
||||
length-prefixing)
|
||||
- `../layout-engine.md` §"Variable-length
|
||||
fields in aligned mode" — the spec's "data region after fixed fields"
|
||||
description
|
||||
@@ -0,0 +1,103 @@
|
||||
# ADR-101: Packed-Mode Read API — Engine as SequentialReader Factory
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
`TypedefEngine` stores a `SequentialReader` inside its `Layout::Packed`
|
||||
variant. The engine exposes it via
|
||||
`engine.sequential_reader() -> Option<&SequentialReader>`.
|
||||
|
||||
The problem: `SequentialReader`'s read methods (`read_next`,
|
||||
`read_field`, `reset`) all take `&mut self` — they mutate the reader's
|
||||
internal cursor (`field_index`, `position`). But the engine hands out
|
||||
`&SequentialReader` (a shared reference), which cannot be used to call
|
||||
`&mut self` methods. The accessor can only give the consumer
|
||||
`position()` and `endian()` (the `&self` methods) — the actual read
|
||||
API is unreachable.
|
||||
|
||||
This makes the engine's packed read-side dead API. A consumer that
|
||||
wants to read a packed buffer must construct their own
|
||||
`SequentialReader::new(&schema)` from the schema, bypassing the engine
|
||||
entirely. The stored reader is dead weight.
|
||||
|
||||
Three options were considered:
|
||||
1. **Factory method** — the engine provides a method that returns an
|
||||
owned fresh `SequentialReader` (reconstructed from the stored
|
||||
schema). The consumer owns the reader and drives it with `&mut self`.
|
||||
2. **Interior mutability** — wrap the reader in `Mutex` or `RefCell`
|
||||
so `&SequentialReader` can be upgraded to `&mut`. Adds overhead and
|
||||
complexity for mutable cursor state that the consumer legitimately
|
||||
wants to own.
|
||||
3. **`sequential_reader_mut()`** — return `&mut SequentialReader`.
|
||||
Requires `&mut self` on the engine, which is overly restrictive
|
||||
(the consumer may share the engine across threads or hold it behind
|
||||
an `Arc`).
|
||||
|
||||
## Decision
|
||||
|
||||
**The engine is a `SequentialReader` factory.** Replace
|
||||
`sequential_reader() -> Option<&SequentialReader>` with
|
||||
`sequential_reader() -> Option<SequentialReader>` — the method returns
|
||||
an owned fresh reader, reconstructed from the stored schema.
|
||||
|
||||
```rust
|
||||
impl TypedefEngine {
|
||||
/// Construct a fresh SequentialReader for packed-mode reads.
|
||||
/// Returns None if compiled in aligned mode.
|
||||
pub fn sequential_reader(&self) -> Option<SequentialReader>;
|
||||
}
|
||||
```
|
||||
|
||||
Each call returns a new reader with the cursor at position 0. The
|
||||
consumer owns the reader and calls `read_next`/`read_field`/`reset` on
|
||||
it directly. The engine still stores its own reader (used for schema
|
||||
validation during construction), but no longer exposes it by
|
||||
reference.
|
||||
|
||||
The same applies to `LayoutBuilder`: `layout_builder()` returns
|
||||
`Option<&LayoutBuilder>` which is fine — `LayoutBuilder::build` takes
|
||||
`&self`, so the shared reference is usable. No change needed for the
|
||||
write-side.
|
||||
|
||||
### Cost
|
||||
|
||||
`SequentialReader::new` clones the top-level struct's field schemas (a
|
||||
`Vec<(String, Value)>` of the `properties` entries) and clones the
|
||||
schema itself. This is cheap — a struct has a small number of fields
|
||||
(SFTP's largest packet has 5). The construction cost is negligible
|
||||
compared to the cost of reading a buffer.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **The packed read API is now usable.** A consumer calls
|
||||
`engine.sequential_reader()` to get an owned reader and drives it
|
||||
directly. No dead API.
|
||||
- **No interior mutability overhead.** The reader's mutable cursor
|
||||
state is owned by the consumer, not shared through a lock.
|
||||
- **Thread-safe engine.** The engine remains `Send + Sync` (it only
|
||||
exposes `&self` methods). The reader is owned by the calling thread.
|
||||
- **Simple.** One method signature change. The stored reader in
|
||||
`Layout::Packed` can be removed (it was only used for schema
|
||||
validation during construction, which is done by the time the
|
||||
consumer calls `sequential_reader()`).
|
||||
|
||||
### Negative
|
||||
|
||||
- **Each call to `sequential_reader()` allocates a new reader.** The
|
||||
cost is a `Vec` of field schemas + a schema clone. Acceptable for
|
||||
the use case (one reader per buffer read).
|
||||
- **The engine no longer holds a live reader.** If a future use case
|
||||
needs to share a reader's cursor state across calls, the consumer
|
||||
must manage that themselves. This is the correct separation — cursor
|
||||
state is consumer-owned, not engine-owned.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — packed
|
||||
sequential mode (`SequentialReader` as the read-side)
|
||||
- `../data-access.md` §"Higher-level
|
||||
read/write" — the `SequentialReader` API
|
||||
@@ -0,0 +1,111 @@
|
||||
# ADR-102: Reject TUnion in Aligned Mode for v1
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The aligned static layout mode (ADR-096) computes fixed byte positions
|
||||
for each field, enabling random access by field path. `TUnion` in
|
||||
aligned mode has three implementation problems:
|
||||
|
||||
1. **No variant field offsets.** Only the `__discriminator` byte range
|
||||
is recorded in the `OffsetMap`. Variant field offsets are not
|
||||
available anywhere in aligned mode — the consumer must recompute
|
||||
them by hand. This makes `TypedefEngine::read_field` on a union
|
||||
variant field impossible.
|
||||
|
||||
2. **`find_discriminator_field` takes the first variant's offset.** For
|
||||
a field-name discriminator, the code probes the first variant that
|
||||
contains the discriminator field and records that offset globally.
|
||||
If variants order fields differently, the discriminator sits at
|
||||
different offsets per variant and the recorded range is silently
|
||||
wrong. The code should validate that the offset is identical across
|
||||
all variants (or require the discriminator field to be first).
|
||||
|
||||
3. **Byte-discriminator union total misaligns the variant.** The union
|
||||
total is `disc_off + disc_size + variant_max_size`, but the variant
|
||||
was probed from offset 0 with alignment. A `u8` discriminator
|
||||
before a `u32`-bearing variant produces a variant region that
|
||||
starts at an unaligned offset in a mode whose entire purpose is
|
||||
alignment.
|
||||
|
||||
The real question is whether `TUnion` in aligned mode is even needed.
|
||||
The two consumer profiles are:
|
||||
|
||||
- **Protocol consumers** (SFTP, call protocol event types): use packed
|
||||
sequential mode. `TUnion` with byte-offset discriminators is the
|
||||
core dispatch mechanism. This is well-supported.
|
||||
- **mmap consumers** (metatensor, safetensors): use aligned static
|
||||
mode. These formats are structs and arrays of structs — they don't
|
||||
use tagged unions. A tensor file has a header struct with tensor
|
||||
descriptors, not a "which variant is this?" dispatch.
|
||||
|
||||
`TUnion` in aligned mode is a combination that no current or planned
|
||||
consumer needs. Shipping broken semantics for an unused use case is
|
||||
worse than rejecting it clearly.
|
||||
|
||||
## Decision
|
||||
|
||||
**Reject `TUnion` in aligned static mode for v1.**
|
||||
|
||||
`OffsetMap::compute` returns `TypedefError::Offset` when it encounters
|
||||
a `TypeDef:Union` field, with a message explaining that unions are not
|
||||
supported in aligned mode and the consumer should use packed mode (or
|
||||
restructure as a struct with an explicit discriminator field).
|
||||
|
||||
This is a schema-load-time rejection. The consumer learns about the
|
||||
problem when compiling the schema, not at runtime.
|
||||
|
||||
### What is NOT rejected
|
||||
|
||||
- `TUnion` in packed sequential mode — this is the core use case
|
||||
(SFTP `Packet` dispatch, call protocol event types) and is fully
|
||||
supported by `LayoutBuilder` and `SequentialReader`.
|
||||
- `TStruct`, `TArray`, and all primitive kinds in aligned mode — these
|
||||
are the mmap-format primitives and are fully supported.
|
||||
|
||||
### Reversal
|
||||
|
||||
This is a two-way door. If a future mmap-format consumer needs tagged
|
||||
unions in aligned mode, the rejection can be lifted and the three
|
||||
implementation problems fixed. The fix would require:
|
||||
- Recording per-variant field offsets in the `OffsetMap` (which
|
||||
variant's offsets to record when variants have different layouts?).
|
||||
- Validating that field-name discriminators have identical offsets
|
||||
across all variants.
|
||||
- Aligning the variant region correctly after the byte discriminator.
|
||||
|
||||
These are design questions that should be answered when the use case
|
||||
arrives, not speculatively now.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **No broken semantics shipped.** The three implementation problems
|
||||
are removed from the API surface rather than silently producing
|
||||
wrong offsets.
|
||||
- **Clear scope boundary.** Aligned mode is for structs and arrays;
|
||||
packed mode is for protocols (including union dispatch). The
|
||||
consumer chooses the mode based on the use case.
|
||||
- **Reversible.** When a real consumer needs aligned-mode unions, the
|
||||
rejection is lifted and the design questions are worked through with
|
||||
a concrete use case.
|
||||
|
||||
### Negative
|
||||
|
||||
- **A schema with a `TUnion` field cannot be compiled in aligned
|
||||
mode.** A consumer that wants both aligned layout and union dispatch
|
||||
must use packed mode or restructure. No current consumer needs this.
|
||||
- **The aligned-mode union code in `offset_map.rs` is dead.** It can
|
||||
be removed or left as a reference for when the rejection is lifted.
|
||||
Removing it is cleaner.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes
|
||||
- [ADR-097](097-schema-annotations.md) §4 — TUnion discriminators
|
||||
- `../layout-engine.md` §"TUnion" —
|
||||
aligned-mode union sizing
|
||||
@@ -0,0 +1,360 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
---
|
||||
|
||||
# alknet-typedef — Layout Engine
|
||||
|
||||
The layout engine: offset computation, the two layout modes (packed
|
||||
sequential vs aligned static), alignment, endianness, and variable-length
|
||||
field handling. This is the novel code — the recursive walk of the schema
|
||||
JSON that computes byte positions for each field.
|
||||
|
||||
## The Two Layout Modes
|
||||
|
||||
The POCs surfaced that protocols and mmap-friendly formats need different
|
||||
layout strategies. This is the most important architectural finding —
|
||||
decided in [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md).
|
||||
|
||||
### Mode 1: Packed sequential (protocol wire formats)
|
||||
|
||||
Fields are packed with no alignment padding. Variable-length fields shift
|
||||
all subsequent fields. Used by SFTP, channels, TTY, and most binary
|
||||
protocols.
|
||||
|
||||
**Components:**
|
||||
|
||||
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `TypeDef:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, TypedefError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
|
||||
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, TypedefError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
|
||||
|
||||
**How it works:**
|
||||
|
||||
For a struct with fields `[u8, u32, string]` where the string is 10 bytes:
|
||||
|
||||
```
|
||||
LayoutBuilder::build(var_sizes: {"payload": 10}):
|
||||
field[0] u8: offset 0, size 1
|
||||
field[1] u32: offset 1, size 4
|
||||
field[2] string: offset 5, size 4 (length prefix) + 10 (data)
|
||||
total: 19
|
||||
|
||||
SequentialReader::read_next (read):
|
||||
read u8 at offset 0
|
||||
read u32 at offset 1
|
||||
read u32 length prefix at offset 5 → data_len
|
||||
read string data at offset 9, length data_len
|
||||
next field at offset 9 + data_len
|
||||
```
|
||||
|
||||
There is no alignment padding. The `u32` at offset 1 is unaligned — this
|
||||
is correct for protocol wire formats, which pack fields tightly.
|
||||
|
||||
**Variable-length fields in packed mode:**
|
||||
|
||||
The `LayoutBuilder` takes actual data sizes for variable-length fields
|
||||
to compute correct positions for subsequent fields. The consumer must
|
||||
know the data sizes before writing — this is inherent to packed layouts.
|
||||
|
||||
The `SequentialReader` reads each field's length prefix to determine the
|
||||
data extent and the position of the next field. The reader walks the
|
||||
buffer sequentially; it cannot jump to field N without reading fields
|
||||
0..N-1 first.
|
||||
|
||||
### Mode 2: Aligned static (mmap-friendly formats)
|
||||
|
||||
Fields have fixed positions with natural alignment padding.
|
||||
Variable-length fields get a 4-byte length prefix at a known offset; the
|
||||
variable data is not included in the static layout. Used by metatensor
|
||||
and safetensors.
|
||||
|
||||
**Component:**
|
||||
|
||||
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, TypedefError>` (requires `TypeDef:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
|
||||
|
||||
**How it works:**
|
||||
|
||||
For a struct with fields `[u8, u32, f32]` and natural alignment:
|
||||
|
||||
```
|
||||
OffsetMap:
|
||||
field[0] u8: offset 0, size 1
|
||||
field[1] u32: offset 4, size 4 (3 bytes padding after u8)
|
||||
field[2] f32: offset 8, size 4
|
||||
total: 12 (struct aligned to 4)
|
||||
```
|
||||
|
||||
The `u32` is aligned to offset 4 (its natural alignment). The consumer
|
||||
can read `field[1]` at offset 4 without reading `field[0]` first — random
|
||||
access by field path.
|
||||
|
||||
**Variable-length fields in aligned mode:**
|
||||
|
||||
Variable-length fields get a 4-byte length prefix at a known offset. The
|
||||
variable data lives outside the static layout — either immediately after
|
||||
the fixed fields (inline length-prefixing) or in a separate data region
|
||||
(offset indirection). The `OffsetMap` records the position of the length
|
||||
prefix (or the `{offset, length}` pair for offset-indirect fields).
|
||||
|
||||
For inline length-prefixing, the variable data follows the fixed fields
|
||||
but is not included in the `OffsetMap`'s field ranges. The consumer reads
|
||||
the length prefix from the `OffsetMap`'s known offset, then slices the
|
||||
data region.
|
||||
|
||||
For offset indirection, the field is a struct `{offset: u32, length: u32}`
|
||||
at a known position in the `OffsetMap`. The consumer reads the offset and
|
||||
length, then slices the separate data region.
|
||||
|
||||
### Inline length-prefixing in aligned mode — non-final field restriction
|
||||
|
||||
Inline length-prefixed variable fields in aligned mode are only allowed
|
||||
as the **last field** in their struct. A non-final inline
|
||||
length-prefixed variable field is rejected at `OffsetMap::compute` time
|
||||
with a `TypedefError::Offset` — the `OffsetMap` reserves only 4 bytes
|
||||
(the length prefix), but `data_access::write_string` writes prefix +
|
||||
data inline, which would clobber subsequent fields. Non-final variable
|
||||
fields must use `maxLength` (fixed-size reservation) or
|
||||
`"encoding": "offset-indirect"`. See
|
||||
[ADR-100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md).
|
||||
|
||||
## Offset Computation Algorithm
|
||||
|
||||
The offset computation is a recursive walk of the schema JSON. The
|
||||
algorithm is the same for both modes; the difference is whether alignment
|
||||
padding is inserted between fields.
|
||||
|
||||
### Fixed-size types
|
||||
|
||||
For each fixed-size type, the algorithm:
|
||||
1. Determines the type's byte size from the `TypeDef:*` kind.
|
||||
2. In aligned mode: inserts padding to satisfy the type's alignment
|
||||
(or the field's `align` annotation, or the struct's `align` default).
|
||||
3. Records the field's `(start, end)` range.
|
||||
4. Advances the current offset by the type's size.
|
||||
|
||||
### Composite types
|
||||
|
||||
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
|
||||
are computed relative to the struct's start offset. The struct's total
|
||||
size is the sum of its fields' sizes (plus alignment padding in aligned
|
||||
mode). The struct itself may have an `align` annotation that rounds up
|
||||
its total size.
|
||||
|
||||
**`TUnion`:** TUnion is supported in packed sequential mode only. In
|
||||
aligned static mode, `OffsetMap::compute` rejects `TUnion` fields with
|
||||
`TypedefError::Offset` — see
|
||||
[ADR-102](decisions/102-reject-tunion-in-aligned-mode.md). Unions
|
||||
are the protocol dispatch pattern (SFTP type bytes, call protocol event
|
||||
types); mmap-friendly formats use structs and arrays, not tagged unions.
|
||||
|
||||
In packed sequential mode, the discriminator occupies
|
||||
`offset..offset + discriminator_size` bytes. For byte-offset
|
||||
discriminators, the variant struct starts at `offset + discriminator_size`.
|
||||
For field-name discriminators, the discriminator is just another field —
|
||||
its offset is computed like any other field, and the variant struct
|
||||
follows at the end of the discriminator field.
|
||||
|
||||
Variant sizes depend on the actual sizes of variable-length fields within
|
||||
each variant, which aren't known at schema time. The `LayoutBuilder`
|
||||
takes the actual variant discriminator value and data sizes at write time,
|
||||
computes the size of the selected variant, and uses that for the union's
|
||||
total size. The `SequentialReader` reads the discriminator first, looks
|
||||
up the variant schema, then reads the variant struct sequentially — it
|
||||
doesn't need to know the union's total size upfront.
|
||||
|
||||
**`TArray` of fixed-size elements:** Element stride = element size (plus
|
||||
alignment padding in aligned mode). Element `i` starts at
|
||||
`array_offset + i × stride`. The array's total size is `count × stride`.
|
||||
|
||||
**`TArray` of variable-length-element structs:** Deferred for v1
|
||||
(OQ-069).
|
||||
|
||||
### Variable-length types
|
||||
|
||||
The typedef engine supports three strategies for variable-length types
|
||||
(see [schema-layer.md](schema-layer.md) §Variable-length types and
|
||||
[ADR-097](decisions/097-schema-annotations.md) §3 for the full
|
||||
annotation shapes).
|
||||
|
||||
**Strategy 1: Inline length-prefixing (default).**
|
||||
1. Records the position of the 4-byte length prefix.
|
||||
2. In aligned mode: the length prefix is aligned; the variable data is
|
||||
not included in the static layout.
|
||||
3. In packed mode: the `LayoutBuilder` takes the actual data size to
|
||||
compute the length prefix value and the position of subsequent fields.
|
||||
The `SequentialReader` reads the length prefix to determine the data
|
||||
extent and the position of the next field.
|
||||
|
||||
**Strategy 2: Fixed-size reservation (`maxLength`).**
|
||||
1. In aligned static mode: reserves `maxLength` bytes at a fixed offset.
|
||||
Data shorter than `maxLength` is zero-padded. Subsequent fields have
|
||||
known, unchanging offsets — the field is fixed-size from the layout
|
||||
perspective. This is the database `VARCHAR(N)` pattern.
|
||||
2. In packed sequential mode: `maxLength` is a validation constraint
|
||||
only. The engine uses strategy 1 (inline length-prefixing) because
|
||||
protocols don't benefit from fixed-size reservation.
|
||||
|
||||
**Strategy 3: Offset indirection (`"encoding": "offset-indirect"`).**
|
||||
1. The field is a struct `{offset: u32, length: u32}`.
|
||||
2. The `OffsetMap` records the position of this struct.
|
||||
3. The consumer provides the data region separately. This is the
|
||||
metatensor blob tensor pattern — the index struct lives in one region,
|
||||
the blob data lives in another.
|
||||
|
||||
### Nested structs and field paths
|
||||
|
||||
Nested structs produce dotted field paths: `header.version`,
|
||||
`header.magic`. The offset computation propagates the field path prefix
|
||||
during recursion. Both `OffsetMap` and `PackedLayout` store fully-qualified
|
||||
paths; the `iter()` method of each yields fields in schema `properties`
|
||||
order, with nested struct fields appearing inline under their parent's
|
||||
path prefix.
|
||||
|
||||
### Endianness
|
||||
|
||||
Endianness is per-schema (ADR-097). The offset computation is
|
||||
endian-agnostic — it computes byte positions, not byte values. The
|
||||
read/write functions apply endianness when converting between bytes and
|
||||
typed values. The engine reads the `"endian"` annotation from the schema
|
||||
and byte-swaps accordingly. All fixed-size types — including `TEnum`
|
||||
(u32 index) — follow the schema's endianness.
|
||||
|
||||
## Mode Selection
|
||||
|
||||
The consumer selects the mode at engine construction time via the
|
||||
`LayoutMode` enum, passed to `TypedefEngine::compile`:
|
||||
|
||||
```rust
|
||||
pub enum LayoutMode {
|
||||
/// Packed sequential — for protocol wire formats (SFTP, channels, TTY).
|
||||
Packed,
|
||||
/// Aligned static — for mmap-friendly formats (metatensor, safetensors).
|
||||
Aligned,
|
||||
}
|
||||
```
|
||||
|
||||
The choice is determined by the use case, not by the schema:
|
||||
|
||||
- **Protocol consumer** (SFTP, binary call frames, TTY negotiation):
|
||||
`LayoutMode::Packed` → uses `LayoutBuilder` for writing and
|
||||
`SequentialReader` for reading.
|
||||
- **mmap consumer** (metatensor): `LayoutMode::Aligned` → uses `OffsetMap`
|
||||
for both reading and writing at known offsets.
|
||||
|
||||
The same schema can be used in either mode. A schema describing an SFTP
|
||||
packet can be consumed by a `SequentialReader` (for parsing incoming
|
||||
frames) and a `LayoutBuilder` (for constructing outgoing frames). A schema
|
||||
describing a metatensor layout can be consumed by an `OffsetMap` (for
|
||||
mmap access).
|
||||
|
||||
`TypedefEngine` exposes mode-appropriate accessors: `engine.offset_map()`
|
||||
returns `Some(&OffsetMap)` in aligned mode and `None` in packed mode;
|
||||
`engine.layout_builder()` returns `Some(&LayoutBuilder)` in packed mode
|
||||
and `None` in aligned mode. `engine.sequential_reader()` returns
|
||||
`Option<SequentialReader>` (an owned fresh reader, not a reference — the
|
||||
reader has mutable cursor state that the consumer owns; see
|
||||
[ADR-101](decisions/101-packed-mode-read-factory.md)) in packed
|
||||
mode and `None` in aligned mode. See [validation.md](validation.md)
|
||||
§"The TypedefEngine struct" for the engine API.
|
||||
|
||||
## Public Types
|
||||
|
||||
The layout engine produces three public types, one per layout component.
|
||||
All are re-exported from the crate root.
|
||||
|
||||
### `ByteRange` (aligned mode)
|
||||
|
||||
```rust
|
||||
pub struct ByteRange {
|
||||
pub start: usize, // inclusive
|
||||
pub end: usize, // exclusive
|
||||
}
|
||||
```
|
||||
|
||||
A half-open byte range produced by `OffsetMap::compute` for each field.
|
||||
`end - start` is the field's byte size in the static layout (for
|
||||
variable-length fields: the length prefix, the `{offset, length}` pair,
|
||||
or the `maxLength` reservation — not the variable data). `ByteRange`
|
||||
provides `len()` and `is_empty()`.
|
||||
|
||||
### `FieldPosition` (packed mode)
|
||||
|
||||
```rust
|
||||
pub struct FieldPosition {
|
||||
pub offset: usize,
|
||||
pub size: usize,
|
||||
pub kind: TypeDefKind,
|
||||
}
|
||||
```
|
||||
|
||||
A field's computed position in a packed layout, produced by
|
||||
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
|
||||
length prefix); for fixed-size fields, `size` is the type's byte size.
|
||||
`kind` records the field's `TypeDef:*` kind so the consumer can dispatch
|
||||
to the correct `data_access` read/write function.
|
||||
|
||||
### `PackedLayout` (packed mode)
|
||||
|
||||
The result of `LayoutBuilder::build`: a map of `field_path → FieldPosition`
|
||||
plus the total buffer size needed.
|
||||
|
||||
```rust
|
||||
impl PackedLayout {
|
||||
pub fn get(&self, field_path: &str) -> Option<&FieldPosition>;
|
||||
pub fn total_size(&self) -> usize;
|
||||
pub fn iter(&self) -> impl Iterator<Item = &(String, FieldPosition)>;
|
||||
}
|
||||
```
|
||||
|
||||
`get` looks up a field by dotted path. For TUnion byte-offset
|
||||
discriminators, the discriminator is recorded under the synthetic path
|
||||
`"<union_path>.__discriminator"`. `iter` yields fields in layout order
|
||||
(schema `properties` order, with nested struct fields appearing inline
|
||||
under their parent's path prefix).
|
||||
|
||||
### `OffsetMap` (aligned mode)
|
||||
|
||||
A flat table of `(field_path, byte_range)` pairs computed from a schema.
|
||||
|
||||
```rust
|
||||
impl OffsetMap {
|
||||
pub fn compute(schema: &Value) -> Result<Self, TypedefError>;
|
||||
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
|
||||
pub fn total_size(&self) -> usize;
|
||||
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
|
||||
}
|
||||
```
|
||||
|
||||
`compute` requires a `TypeDef:Struct` at the top level. `total_size`
|
||||
includes trailing alignment padding. `iter` yields fields in insertion
|
||||
order (schema `properties` order, nested struct fields appearing inline).
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Two layout modes | [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential for protocols; aligned static for mmap formats |
|
||||
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Endianness, alignment, encoding annotations that control layout behavior |
|
||||
| Non-final inline variable fields | [ADR-100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields); use `maxLength` or `offset-indirect` |
|
||||
| Packed-mode read factory | [ADR-101](decisions/101-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader, not a reference |
|
||||
| TUnion in aligned mode | [ADR-102](decisions/102-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) |
|
||||
|
||||
## Open Questions
|
||||
|
||||
See [open-questions.md](open-questions.md) for full details.
|
||||
|
||||
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs
|
||||
— requires lazy walking logic; blocked on a concrete consumer that
|
||||
needs it.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
|
||||
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
|
||||
- [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) —
|
||||
the two layout modes decision
|
||||
- [ADR-097](decisions/097-schema-annotations.md) — schema
|
||||
annotations
|
||||
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds and their
|
||||
byte sizes
|
||||
- [data-access.md](data-access.md) — read/write functions that use the
|
||||
computed offsets
|
||||
@@ -0,0 +1,104 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
---
|
||||
|
||||
# Open Questions
|
||||
|
||||
Each open question lives in its own file under [`questions/`](questions/),
|
||||
named `NNN-slug.md` (mirroring the ADR convention). This file is the index:
|
||||
theme-grouped tables for scannability, plus a cross-theme
|
||||
[Deferred / Blocked](#deferred--blocked) section that surfaces the
|
||||
safe-exit deferrals with their blocking conditions inline — so "what's
|
||||
currently parked and why" is answerable at a glance.
|
||||
|
||||
**Status values**:
|
||||
- `open` — Needs to be resolved now. Has a clear path to resolution.
|
||||
- `resolved` — Decided. The resolution is stated cleanly, without caveats about how it could be changed later.
|
||||
- `deferred(scope)` — Cannot be resolved yet. The information is genuinely
|
||||
missing — a crate spec, POC result, or use case that doesn't exist yet.
|
||||
Has a concrete blocking condition. Not a failure — scope management.
|
||||
- `deferred(unclear)` — Cannot be resolved yet. The pieces exist (decided
|
||||
in other ADRs, existing types, existing patterns) but the composition
|
||||
— how they fit together — isn't clear yet. Resolution requires
|
||||
investigation (work through examples, maybe POC), not waiting. Has a
|
||||
concrete investigation target and an impacts field. Not a failure —
|
||||
honest uncertainty in a poorly-defined problem space.
|
||||
- `partially resolved` — Some aspects decided, others deferred or open.
|
||||
- `dissolved` — The question was reframed out of existence (e.g., superseded
|
||||
by an ADR that retires the premise). Kept for reference.
|
||||
|
||||
**Impacts field**: Every unresolved OQ (`open`, `deferred(scope)`,
|
||||
`deferred(unclear)`, `partially resolved`) should have an `Impacts`
|
||||
field stating what it blocks downstream. Be specific: "blocks the first
|
||||
hub deployment because the hub dials workers" not "blocks the hub
|
||||
crate." This is the triage signal that makes the deferral's urgency
|
||||
visible.
|
||||
|
||||
Door type classifications follow ADR-009 — they describe **reversal cost** (how expensive it is to undo), not urgency:
|
||||
- **One-way door**: Reversal requires rewriting significant code or permanently closes a capability. Getting it wrong is expensive — requires ADR before implementation.
|
||||
- **Two-way door**: Reversal is cheap or additive. Getting it wrong is recoverable — decide, implement, revert if needed.
|
||||
|
||||
Door type is separate from whether a decision is made. A two-way door is a decision you make now and can revert later, not a decision to defer.
|
||||
|
||||
## By Theme
|
||||
|
||||
### Layout Engine
|
||||
|
||||
| OQ | Title | Status | Door | Pri |
|
||||
|----|-------|--------|------|-----|
|
||||
| [OQ-069](questions/069-arrays-of-variable-length-element-structs.md) | Arrays of Variable-Length-Element Structs | deferred(scope) | two | low |
|
||||
|
||||
### Platform Support
|
||||
|
||||
| OQ | Title | Status | Door | Pri |
|
||||
|----|-------|--------|------|-----|
|
||||
| [OQ-070](questions/070-no-std-alloc-support.md) | `no_std` + `alloc` Support | deferred(scope) | two | low |
|
||||
|
||||
### Schema Construction
|
||||
|
||||
| OQ | Title | Status | Door | Pri |
|
||||
|----|-------|--------|------|-----|
|
||||
| [OQ-071](questions/071-builder-api-for-schema-construction.md) | Builder API for Schema Construction | deferred(scope) | two | med |
|
||||
|
||||
## Deferred / Blocked
|
||||
|
||||
The safe-exit visibility surface. These questions are parked because the
|
||||
information needed to resolve them does not exist yet — each has a concrete
|
||||
blocking condition. They are not failures; they are scope management.
|
||||
This section exists so "what's currently blocking the architect" is
|
||||
answerable at a glance, not by filtering the tables above.
|
||||
|
||||
### OQ-069: Arrays of Variable-Length-Element Structs
|
||||
|
||||
- **Blocked on**: A concrete consumer that needs arrays of structs with
|
||||
variable-length fields, where the elements are interleaved
|
||||
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
|
||||
sequentially rather than use a fixed stride. The SFTP `Name` packet
|
||||
has `Vec<File>` where `File` contains strings, but SFTP serializes
|
||||
this as a sequence of length-prefixed strings (the serde `SeqAccess`
|
||||
pattern), not as an array of fixed-stride structs. Arrays of
|
||||
fixed-size structs are fully supported.
|
||||
- **Priority**: low
|
||||
- **Full file**: [OQ-069](questions/069-arrays-of-variable-length-element-structs.md)
|
||||
|
||||
### OQ-070: `no_std` + `alloc` Support
|
||||
|
||||
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
|
||||
(e.g., a microcontroller running Rust without `std`). The WASM target
|
||||
has `std` available via `wasm-bindgen`. The engine's core (offset
|
||||
computation, read/write) is already allocation-free; the `jsonschema`
|
||||
dependency is the only `alloc` consumer.
|
||||
- **Priority**: low
|
||||
- **Full file**: [OQ-070](questions/070-no-std-alloc-support.md)
|
||||
|
||||
### OQ-071: Builder API for Schema Construction
|
||||
|
||||
- **Blocked on**: A concrete need for programmatic schema construction
|
||||
in Rust. The current consumers (SFTP, metatensor, binary call frames,
|
||||
TTY negotiation) all have schemas that can be hand-written or
|
||||
generated from TypeBox. A builder API would be a fluent Rust API that
|
||||
produces the same JSON Schema structure — it would sit on top of the
|
||||
engine, not inside it.
|
||||
- **Priority**: medium
|
||||
- **Full file**: [OQ-071](questions/071-builder-api-for-schema-construction.md)
|
||||
@@ -0,0 +1,203 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
---
|
||||
|
||||
# alknet-typedef — Overview
|
||||
|
||||
The binary struct engine: a small Rust crate that takes a JSON Schema
|
||||
with `TypeDef:*` custom keywords and produces an offset map, read/write
|
||||
functions, and validation — all driven by the schema. The schema is the
|
||||
format definition; the engine is generic.
|
||||
|
||||
This document covers the crate's purpose, the "schema is the format"
|
||||
principle, its dependency edges, consumers, and scope boundaries.
|
||||
Component details are in the sibling documents.
|
||||
|
||||
## What
|
||||
|
||||
`alknet-typedef` is a library crate that consumes JSON Schemas annotated
|
||||
with `TypeDef:*` custom keywords (the same kinds defined in TypeBox's
|
||||
`typedef.ts`, plus `TypeDef:Bytes`, `TypeDef:Int64`, and `TypeDef:Uint64`
|
||||
as alknet-typedef additions) and produces three capabilities:
|
||||
|
||||
1. **An offset map** — walks the schema, computes byte offsets for each
|
||||
field based on type sizes, field order, and alignment.
|
||||
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
|
||||
read the field's bytes at its offset (zero-copy for fixed-size types).
|
||||
Given a `&mut [u8]` buffer, write a value at its offset.
|
||||
3. **Validation** — via `jsonschema` custom keywords, validates that a
|
||||
buffer's bytes match the schema's type constraints.
|
||||
|
||||
The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||
`serde_json` (schema parsing). The novel code is the offset computation
|
||||
— a recursive walk of the schema JSON that computes byte positions for
|
||||
each field. The custom keyword implementations are small (a few lines
|
||||
each, generated from shared macros — see [validation.md](validation.md)).
|
||||
|
||||
The crate replaces two prior attempts that built their own jsonschema
|
||||
engines — typebox-rs (~8,400 lines) and alktype (~5,600 lines) — with
|
||||
`jsonschema` + an offset map + small custom keyword implementations. See
|
||||
[ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
## Why
|
||||
|
||||
The crate's purpose is to be the binary struct engine for every alknet
|
||||
component that reads or writes binary data at computed offsets. Instead
|
||||
of per-protocol serde structs (russh-sftp's 29 packet types), per-handler
|
||||
wire format code (TTY's 5-byte format parser), or per-format offset
|
||||
computation (metatensor's tensor access), all of these become instances
|
||||
of the same engine with different schemas.
|
||||
|
||||
The guiding insight:
|
||||
|
||||
> **The schema is the format.** A JSON Schema with `TypeDef:Float32`,
|
||||
> `TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
|
||||
> the layout spec. No separate format definition, no separate parser, no
|
||||
> separate validator. One schema, three uses: validate, compute offsets,
|
||||
> access data.
|
||||
|
||||
This is the convergence of three threads identified in the
|
||||
call-channels-unification research: the `typedef.ts` schema kinds from
|
||||
TypeBox, the russh-sftp protocol packets, and the metatensor format. The
|
||||
common pattern: a JSON Schema describes the shape of binary data, and
|
||||
the binary data is the struct's bytes at computed offsets.
|
||||
|
||||
The crate was bumped up in the timeline when the call-channels-unification
|
||||
research surfaced that channels, TTY, and the binary call protocol are
|
||||
all variations on the same wire-format family — `[discriminant][length][payload]`.
|
||||
The typedef engine makes the "channels is call with a binary data plane"
|
||||
unification concrete: the binary data plane's wire format is the call
|
||||
protocol's own schema system, just binary-encoded. The `channel_open`
|
||||
marker says "use binary framing"; the typedef engine says "here's how to
|
||||
read/write the binary payload."
|
||||
|
||||
## The "Schema Is the Format" Principle
|
||||
|
||||
A JSON Schema with `TypeDef:*` custom keywords serves three roles
|
||||
simultaneously:
|
||||
|
||||
| Role | Mechanism | When |
|
||||
|------|-----------|------|
|
||||
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
|
||||
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) |
|
||||
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
|
||||
|
||||
No separate format definition, no separate parser, no separate validator.
|
||||
The schema is the single source of truth for the binary format. Adding a
|
||||
new field to a protocol is adding a property to the schema JSON — the
|
||||
engine computes the new offsets automatically.
|
||||
|
||||
This is the same principle as `#[repr(C)]` struct field access, but at
|
||||
runtime from a portable JSON Schema instead of at compile-time from
|
||||
language-specific annotations. The schema is the ABI contract.
|
||||
|
||||
## Dependencies
|
||||
|
||||
```
|
||||
alknet-typedef
|
||||
├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support
|
||||
├── serde_json (with preserve_order) — schema parsing; field order is load-bearing
|
||||
└── (no tokio, no platform deps) — WASM-clean by construction
|
||||
```
|
||||
|
||||
`alknet-typedef` is dependency-light: `jsonschema` + `serde_json` only.
|
||||
No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for
|
||||
browser use. The `jsonschema` crate is already in the workspace at
|
||||
`/workspace/jsonschema/` but not yet used by any alknet crate — typedef
|
||||
is the first consumer.
|
||||
|
||||
`serde_json` requires the `preserve_order` feature because field order
|
||||
is load-bearing for binary layouts. The order of properties in the
|
||||
schema JSON determines the order of fields in the binary struct.
|
||||
|
||||
## Consumers
|
||||
|
||||
| Consumer | Schema describes | Engine provides |
|
||||
|----------|-----------------|-----------------|
|
||||
| russh-sftp | 29 packet structs + Packet union (byte discriminator) | Read/write SFTP frames from bytes |
|
||||
| metatensor | Model layout (ConvNet struct, tensor refs) | Offset map for mmap'd tensor access |
|
||||
| binary call frames | `call.requested` / `call.responded` / etc. structs | Read/write binary call frames |
|
||||
| TTY negotiation | `NegotiateRequest` / `NegotiateResponse` structs | Read/write TTY control frames |
|
||||
| channels wire | `ChunkHeader { channel_id, length }` | Already trivial (8 bytes, no schema needed) |
|
||||
|
||||
The russh-sftp case is the most instructive and the highest-value POC
|
||||
target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written
|
||||
dispatch on a type byte followed by serde deserialization. Under typedef,
|
||||
the dispatch is `TUnion` with a byte-offset discriminator — the schema
|
||||
says "byte 0 is the discriminator, bytes 1..N are the variant struct."
|
||||
The engine reads the discriminator, looks up the variant schema, computes
|
||||
offsets, reads fields. Same result, no per-packet-type code.
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
These boundaries are decided in [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md).
|
||||
|
||||
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
|
||||
is a *format* (8-byte header + JSON header + binary data) that uses the
|
||||
typedef engine for its offset computation and tensor access.
|
||||
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
|
||||
`Value.Convert` — schema evolution — is out of scope for v1. The engine
|
||||
should not do anything that explicitly blocks adding a Value system
|
||||
later.
|
||||
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
|
||||
concern. The typedef engine consumes schemas; it does not generate them.
|
||||
- **Not a schema builder.** The typedef engine does not provide a fluent
|
||||
API for constructing schemas. Schemas are plain JSON — authored in
|
||||
TypeBox, generated by ujsx components, or hand-written. A builder API
|
||||
is deferred (OQ-071).
|
||||
- **Not a serialization framework.** The typedef engine is not a
|
||||
general-purpose serde replacement. It operates on raw byte buffers at
|
||||
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||
with a known schema, use typedef.
|
||||
|
||||
## Architecture (component pointers)
|
||||
|
||||
- **[schema-layer.md](schema-layer.md)** — the 19 `TypeDef:*` kinds,
|
||||
jsonschema custom keyword integration, TypeBox interop, schema
|
||||
annotations (endianness, alignment, encoding, TUnion discriminators).
|
||||
- **[layout-engine.md](layout-engine.md)** — offset computation, the two
|
||||
layout modes (packed sequential vs aligned static), alignment,
|
||||
endianness, variable-length field handling.
|
||||
- **[data-access.md](data-access.md)** — read/write functions, TUnion
|
||||
dispatch, field paths, zero-copy access for fixed-size types,
|
||||
length-prefix reading for variable-length types.
|
||||
- **[validation.md](validation.md)** — custom keyword validators for all
|
||||
19 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time
|
||||
validation, `TypedefEngine` as the compiled form of a schema.
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Purpose, scope, and the jsonschema engine | [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
|
||||
| Two layout modes | [ADR-096](decisions/096-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats |
|
||||
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) |
|
||||
| Error handling and validation | [ADR-098](decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
|
||||
| Int64/Uint64 kinds | [ADR-099](decisions/099-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (SFTP offsets, metatensor data_offsets) |
|
||||
| Non-final inline variable fields | [ADR-100](decisions/100-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields) |
|
||||
| Packed-mode read factory | [ADR-101](decisions/101-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader |
|
||||
| TUnion in aligned mode | [ADR-102](decisions/102-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) |
|
||||
|
||||
## Open Questions
|
||||
|
||||
See [open-questions.md](open-questions.md) for full details.
|
||||
|
||||
- **OQ-069** (deferred(scope)): Arrays of variable-length-element structs.
|
||||
- **OQ-070** (deferred(scope)): `no_std` + `alloc` support.
|
||||
- **OQ-071** (deferred(scope)): Builder API for schema construction.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
|
||||
passing, two layout modes, TUnion dispatch, endianness)
|
||||
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||
JSON Schema as the binary struct engine" — the origin of this research
|
||||
thread
|
||||
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||
schema kinds (619 lines)
|
||||
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
|
||||
- `/workspace/@alkimiadev/typebox-rs/` — prior attempt, replaced by typedef
|
||||
- `/workspace/@alkimiadev/alktype/` — prior attempt, replaced by typedef
|
||||
@@ -0,0 +1,20 @@
|
||||
# OQ-069: Arrays of variable-length-element structs
|
||||
|
||||
- **Origin**: [../layout-engine.md](../layout-engine.md),
|
||||
[../data-access.md](../data-access.md);
|
||||
`docs/research/alknet-typedef/findings.md` §"Problem 3: Nested structs
|
||||
and arrays of structs"
|
||||
- **Status**: deferred(scope)
|
||||
- **Door type**: Two-way (additive — the engine can add lazy walking
|
||||
logic without changing the existing fixed-stride array support)
|
||||
- **Priority**: low
|
||||
- **Impacts**: Blocks any protocol with interleaved variable-length struct arrays (e.g., a protocol where each array element has a string field and elements are packed as `[fixed_0][str_0][fixed_1][str_1]...`). Does NOT block SFTP `Name` packet handling — SFTP serializes this as a sequence of length-prefixed strings (the serde `SeqAccess` pattern), not as an array of fixed-stride structs. Does NOT block any current consumer.
|
||||
- **Blocked on**: A concrete consumer that needs arrays of structs with
|
||||
variable-length fields, where the elements are interleaved
|
||||
(`[fixed_0][str_0][fixed_1][str_1]...`) and the engine must walk
|
||||
sequentially rather than use a fixed stride.
|
||||
- **Resolution**: Not yet decidable. The mechanism (lazy sequential
|
||||
walking of array elements, reading each element's length prefixes to
|
||||
find the next element's start) is understood but not needed by any
|
||||
current consumer. Arrays of fixed-size structs are fully supported.
|
||||
- **Cross-references**: ADR-096, [layout-engine.md](../layout-engine.md)
|
||||
@@ -0,0 +1,22 @@
|
||||
# OQ-070: `no_std` + `alloc` support
|
||||
|
||||
- **Origin**: [../overview.md](../overview.md);
|
||||
`docs/research/alknet-typedef/findings.md` §"Open Questions" (OQ 6)
|
||||
- **Status**: deferred(scope)
|
||||
- **Door type**: Two-way (additive — can be added as a feature gate
|
||||
without changing the existing `std` API)
|
||||
- **Priority**: low
|
||||
- **Impacts**: Blocks bare-metal embedded targets (microcontrollers
|
||||
running Rust without `std`). Does NOT block any current deployment
|
||||
target. Does NOT block WASM — `wasm32-unknown-unknown` has `std`
|
||||
available via `wasm-bindgen`; the crate is WASM-clean by construction
|
||||
(no tokio, no platform deps, `jsonschema` builds for WASM with
|
||||
`default-features = false`).
|
||||
- **Blocked on**: An embedded use case that requires `no_std` + `alloc`
|
||||
(e.g., a microcontroller running Rust without `std`).
|
||||
- **Resolution**: Not yet decidable. Target `std` for v1. If embedded
|
||||
use cases emerge, `no_std` + `alloc` can be added as a feature gate
|
||||
later. The engine's core (offset computation, read/write) is already
|
||||
allocation-free — it operates on `&[u8]` slices. The `jsonschema`
|
||||
dependency is the only `alloc` consumer.
|
||||
- **Cross-references**: ADR-095
|
||||
@@ -0,0 +1,26 @@
|
||||
# OQ-071: Builder API for schema construction
|
||||
|
||||
- **Origin**: [../schema-layer.md](../schema-layer.md),
|
||||
[../overview.md](../overview.md);
|
||||
`docs/research/alknet-typedef/findings.md` (the builder API was noted
|
||||
as the one detail not covered by the POCs)
|
||||
- **Status**: deferred(scope)
|
||||
- **Door type**: Two-way (additive — a builder API can be added without
|
||||
changing the existing JSON-consumption path)
|
||||
- **Priority**: medium
|
||||
- **Impacts**: Blocks programmatic schema construction in Rust without a
|
||||
JS toolchain. Any consumer that wants to build typedef schemas at
|
||||
runtime from Rust code (rather than loading pre-authored JSON) must
|
||||
construct the JSON manually or depend on TypeBox. Does NOT block any
|
||||
current consumer — all v1 consumers (SFTP, metatensor, binary call
|
||||
frames, TTY negotiation) use pre-authored schemas.
|
||||
- **Blocked on**: A concrete need for programmatic schema construction
|
||||
in Rust. The current consumers (SFTP, metatensor, binary call frames,
|
||||
TTY negotiation) all have schemas that can be hand-written or generated
|
||||
from TypeBox.
|
||||
- **Resolution**: Not yet decidable. The builder API is important but
|
||||
not needed for the initial consumers. The engine's JSON-consumption
|
||||
path is the primary interface for v1. A builder API would be a fluent
|
||||
Rust API that produces the same JSON Schema structure — it would sit
|
||||
on top of the engine, not inside it.
|
||||
- **Cross-references**: ADR-095, [schema-layer.md](../schema-layer.md)
|
||||
@@ -0,0 +1,506 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
---
|
||||
|
||||
# alknet-typedef — Schema Layer
|
||||
|
||||
The schema layer: the 19 `TypeDef:*` custom type kinds, their mapping to
|
||||
Rust types and byte sizes, the `jsonschema` custom keyword integration,
|
||||
TypeBox interop, and the concrete JSON shapes for schema-level
|
||||
annotations.
|
||||
|
||||
## The 19 TypeDef Kinds
|
||||
|
||||
These are the custom schema kinds defined in TypeBox's `typedef.ts`
|
||||
(`/workspace/@alkdev/typebox/example/typedef/typedef.ts`, 619 lines) and
|
||||
ported to Rust via `jsonschema` custom keywords. Each kind carries binary
|
||||
layout semantics — a known byte size (for fixed-size types) or a known
|
||||
encoding strategy (for variable-length types).
|
||||
|
||||
| Kind | TypeBox key | Rust type | Size | Category |
|
||||
|------|-------------|-----------|------|----------|
|
||||
| `TFloat32` | `TypeDef:Float32` | `f32` | 4 | fixed |
|
||||
| `TFloat64` | `TypeDef:Float64` | `f64` | 8 | fixed |
|
||||
| `TInt8` | `TypeDef:Int8` | `i8` | 1 | fixed |
|
||||
| `TInt16` | `TypeDef:Int16` | `i16` | 2 | fixed |
|
||||
| `TInt32` | `TypeDef:Int32` | `i32` | 4 | fixed |
|
||||
| `TInt64` | `TypeDef:Int64` | `i64` | 8 | fixed |
|
||||
| `TUint8` | `TypeDef:Uint8` | `u8` | 1 | fixed |
|
||||
| `TUint16` | `TypeDef:Uint16` | `u16` | 2 | fixed |
|
||||
| `TUint32` | `TypeDef:Uint32` | `u32` | 4 | fixed |
|
||||
| `TUint64` | `TypeDef:Uint64` | `u64` | 8 | fixed |
|
||||
| `TBoolean` | `TypeDef:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed |
|
||||
| `TString` | `TypeDef:String` | length-prefixed UTF-8 | variable | variable |
|
||||
| `TBytes` | `TypeDef:Bytes` | length-prefixed raw bytes | variable | variable |
|
||||
| `TStruct` | `TypeDef:Struct` | record of fields | sum of field sizes | composite |
|
||||
| `TUnion` | `TypeDef:Union` | tagged union | discriminator + variant | composite |
|
||||
| `TArray` | `TypeDef:Array` | repeated element | count × element size | composite |
|
||||
| `TEnum` | `TypeDef:Enum` | u32 index into enum values | 4 (fixed) | fixed |
|
||||
| `TRecord` | `TypeDef:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
|
||||
| `TTimestamp` | `TypeDef:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
|
||||
|
||||
`TypeDef:Int64` and `TypeDef:Uint64` are alknet-typedef additions —
|
||||
TypeBox's `typedef.ts` tops out at 32-bit integers. They are required by
|
||||
the primary POC targets: SFTP `Read`/`Write` packets have `offset: u64`,
|
||||
and metatensor `data_offsets` are `u64`. See
|
||||
[ADR-099](decisions/099-int64-uint64-first-class-kinds.md).
|
||||
|
||||
### The `TypeDefKind` enum
|
||||
|
||||
The engine represents the 19 kinds as a Rust enum — `TypeDefKind` — with
|
||||
one variant per kind (`TypeDefKind::Float32`, `TypeDefKind::Struct`, etc.).
|
||||
The enum provides compile-time exhaustiveness checking and integer
|
||||
discriminant dispatch (a jump table) instead of string comparison at
|
||||
every field access. It is `pub` and re-exported from the crate root.
|
||||
|
||||
```rust
|
||||
pub enum TypeDefKind {
|
||||
Int8, Int16, Int32, Int64,
|
||||
Uint8, Uint16, Uint32, Uint64,
|
||||
Float32, Float64,
|
||||
Boolean, Enum,
|
||||
String, Bytes, Timestamp,
|
||||
Struct, Union, Array, Record,
|
||||
}
|
||||
```
|
||||
|
||||
The enum carries the kind's binary-layout metadata as inherent methods:
|
||||
|
||||
| Method | Returns | Notes |
|
||||
|--------|---------|-------|
|
||||
| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"TypeDef:Uint8"` |
|
||||
| `type_size(self)` | `Option<usize>` | `Some(N)` for fixed-size kinds; `None` for variable/composite |
|
||||
| `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array |
|
||||
| `is_fixed_size(self)` | `bool` | True for the 12 fixed-size primitive kinds |
|
||||
| `is_composite(self)` | `bool` | True for Struct, Union, Array, Record |
|
||||
| `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record |
|
||||
| `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter |
|
||||
|
||||
`TypeDefKind` implements `Display` (renders the keyword string) and
|
||||
`FromStr` (parses the keyword string back into the variant, returning
|
||||
`TypedefError::Schema` for unknown kinds). The layout engines and the
|
||||
validator dispatch on the enum, not on strings.
|
||||
|
||||
### Fixed-size types
|
||||
|
||||
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
|
||||
`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset
|
||||
computation uses these sizes directly. Read/write is zero-copy pointer
|
||||
cast for these types.
|
||||
|
||||
**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other
|
||||
values are invalid and produce a `TypedefError::Access` on read.
|
||||
|
||||
**`TEnum` binary representation:** A `u32` index into the enum's declared
|
||||
values, in declaration order. The first declared value is index 0, the
|
||||
second is index 1, etc. The enum's values are declared via the standard
|
||||
JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`).
|
||||
The `TypeDef:Enum` custom keyword signals that the type is an enum for
|
||||
layout purposes; the built-in `enum` keyword provides the value list.
|
||||
|
||||
**Design note:** TypeBox's `TEnum` is a string enum (variable-length). The
|
||||
typedef engine uses a `u32` index instead — a deliberate deviation from
|
||||
TypeBox fidelity in favor of binary efficiency. Most enums have a small
|
||||
number of variants (e.g., the call protocol's 5 event types); a `u32`
|
||||
index is compact, fixed-size, and sufficient for any realistic enum. The
|
||||
JSON representation (for validation) remains a string; the binary
|
||||
representation is the `u32` index.
|
||||
The `u32` index follows the schema's endianness annotation (ADR-097), like
|
||||
all other fixed-size types. In little-endian mode the index is
|
||||
`u32::from_le_bytes`; in big-endian mode it is `u32::from_be_bytes`.
|
||||
|
||||
### Variable-length types
|
||||
|
||||
`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes.
|
||||
The typedef engine supports three strategies for handling variable-length
|
||||
types in binary layouts, selected by the `encoding` annotation and the
|
||||
standard JSON Schema `maxLength` keyword:
|
||||
|
||||
| Strategy | Encoding annotation | Layout behavior | Use case |
|
||||
|----------|-------------------|-----------------|----------|
|
||||
| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) |
|
||||
| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) |
|
||||
| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) |
|
||||
|
||||
**Strategy 1: Inline length-prefixing (default).** The field's fixed
|
||||
portion is a 4-byte length prefix at a computed offset. The variable data
|
||||
follows immediately after. In packed sequential mode, the length prefix
|
||||
determines the position of subsequent fields. In aligned static mode, the
|
||||
length prefix is at a known offset; the variable data is not included in
|
||||
the static layout. This is the universal pattern used by channels, SFTP,
|
||||
TTY, and most binary protocols.
|
||||
|
||||
**Strategy 2: Fixed-size reservation.** When a variable-length field
|
||||
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
|
||||
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
|
||||
than `maxLength` is zero-padded; data longer than `maxLength` is a
|
||||
validation error. This makes the field fixed-size from the layout
|
||||
perspective — subsequent fields have known, unchanging offsets. This is
|
||||
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
|
||||
pattern for fields with known maximum sizes.
|
||||
|
||||
In packed sequential mode, `maxLength` is a validation constraint only —
|
||||
the engine still uses inline length-prefixing (strategy 1) because
|
||||
protocols don't benefit from fixed-size reservation.
|
||||
|
||||
**Strategy 3: Offset indirection.** The field is a struct
|
||||
`{offset: u32, length: u32}` at a known position. The consumer provides
|
||||
the data region separately; the engine reads the offset and length, then
|
||||
slices the data region. This is the metatensor blob tensor pattern — the
|
||||
index struct lives in one region, the blob data lives in another. Enables
|
||||
mmap-friendly random access to variable-length data without parsing
|
||||
length prefixes and without reserving worst-case space.
|
||||
|
||||
**Default strategy selection:**
|
||||
- In packed sequential mode: always strategy 1 (inline length-prefixing).
|
||||
`maxLength` is a validation constraint only.
|
||||
- In aligned static mode: strategy 2 (fixed-size reservation) if
|
||||
`maxLength` is declared; strategy 3 (offset indirection) if
|
||||
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
|
||||
length-prefixing) otherwise.
|
||||
|
||||
**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3)
|
||||
respects the schema's `"endian"` annotation (ADR-097). In little-endian
|
||||
mode, the length is `u32::from_le_bytes`. In big-endian mode, the length
|
||||
is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have
|
||||
consistent byte order for both field values and length prefixes.
|
||||
|
||||
**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`.
|
||||
Otherwise identical to `TString` in layout (same three strategies).
|
||||
|
||||
**Design note:** `TypeDef:Bytes` is an alknet-typedef addition — it does
|
||||
not exist in TypeBox's `typedef.ts` (which defines 16 kinds). It is
|
||||
included because raw byte arrays are a common binary protocol primitive
|
||||
(SFTP data payloads, channels payloads, tensor data) and are semantically
|
||||
distinct from UTF-8 strings. In the binary representation, TBytes is raw
|
||||
bytes with no encoding (not base64, not hex). In the JSON representation
|
||||
(for validation), TBytes is a string (JSON has no native byte type).
|
||||
|
||||
**`TRecord`:** A string-keyed map. The value type is declared via the
|
||||
schema's `"values"` property (e.g., `"values": { "TypeDef:Float32": true }`).
|
||||
Binary layout is a count-prefixed sequence of `(key, value)` pairs:
|
||||
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
|
||||
The count is the number of entries. Each key is a length-prefixed UTF-8
|
||||
string. Each value is encoded according to its declared `TypeDef:*` kind
|
||||
— a `Record<Uint32>` value is 4 raw bytes; a `Record<String>` value is
|
||||
itself a length-prefixed string; a `Record<Struct>` value is the struct's
|
||||
fields laid out inline. There is **no separate `value_len` prefix** —
|
||||
the value's size is determined by its kind (fixed-size kinds have a
|
||||
known size; variable-length kinds carry their own length prefix). The
|
||||
count and key-length prefixes respect the schema's endianness. In
|
||||
aligned static mode with `maxLength`, the entire record is reserved at
|
||||
`maxLength` bytes (zero-padded).
|
||||
|
||||
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
|
||||
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
|
||||
fixed-size reservation (strategy 2 with `maxLength`). The data-access
|
||||
layer treats timestamps as opaque length-prefixed strings — it does not
|
||||
parse or validate the timestamp format. The jsonschema custom keyword
|
||||
validator checks RFC 3339 conformance at the JSON level (see
|
||||
[validation.md](validation.md)).
|
||||
|
||||
`TArray` is variable-length when the element type is variable-length or
|
||||
when the count is not known at schema time. For fixed-size element arrays
|
||||
with a known count, the size is `element_size × count`.
|
||||
|
||||
**`TArray` count declaration:** The array count is declared via the
|
||||
standard JSON Schema `"minItems"` and `"maxItems"` keywords. When
|
||||
`minItems == maxItems`, the array has a fixed count known at schema time.
|
||||
When they differ or are absent, the count is variable and the array uses
|
||||
a length-prefixed encoding: `[count: u32][element_0]...[element_N]`.
|
||||
The count prefix respects the schema's endianness.
|
||||
|
||||
### Composite types
|
||||
|
||||
`TStruct` and `TUnion` are composite — their size is the sum of their
|
||||
fields' sizes (plus alignment padding in aligned static mode). The offset
|
||||
computation recurses into their properties.
|
||||
|
||||
## Schema-Layer Public API
|
||||
|
||||
The `schema` module exposes the foundational types and functions every
|
||||
other module depends on. These are re-exported from the crate root.
|
||||
|
||||
### `get_typedef_kind` vs `get_typedef_kind_loose`
|
||||
|
||||
The engine recognizes a `TypeDef:*` kind on a schema node two ways,
|
||||
because the keyword value may be either a boolean (`true`) or an
|
||||
annotation object (`{ "encoding": "..." }`):
|
||||
|
||||
| Function | Recognizes | Returns |
|
||||
|----------|------------|---------|
|
||||
| `get_typedef_kind(node) -> Option<&str>` | Boolean form only (`{ "TypeDef:String": true }`) | The keyword string, e.g. `"TypeDef:String"` |
|
||||
| `get_typedef_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string |
|
||||
| `get_typedef_kind_enum(node) -> Option<TypeDefKind>` | Boolean form only | The parsed enum variant |
|
||||
| `get_typedef_kind_loose_enum(node) -> Option<TypeDefKind>` | Boolean form **and** object form | The parsed enum variant |
|
||||
|
||||
The boolean-form-only functions are used by the validator factories
|
||||
(which reject the object form as a schema error) and the top-level
|
||||
kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new`
|
||||
(which require `TypeDef:Struct` at the root). The "loose" variants are
|
||||
used by the layout engines during field traversal, so that a variable-
|
||||
length field with an `encoding` annotation (`{ "TypeDef:String":
|
||||
{ "encoding": "offset-indirect" } }`) is still recognized as a `String`.
|
||||
|
||||
### Annotation parsers
|
||||
|
||||
Each schema-level annotation has a dedicated parser that reads it from a
|
||||
`serde_json::Value` node and returns a sensible default when absent:
|
||||
|
||||
| Function | Annotation | Default |
|
||||
|----------|------------|---------|
|
||||
| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` |
|
||||
| `parse_align(node) -> Option<usize>` | `"align"` | `None` |
|
||||
| `parse_max_length(node) -> Option<usize>` | `"maxLength"` | `None` |
|
||||
| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` |
|
||||
| `parse_discriminator(node) -> Result<DiscriminatorKind, TypedefError>` | `"discriminator"` | (required — returns `TypedefError::Schema` if absent) |
|
||||
|
||||
### Public enums
|
||||
|
||||
```rust
|
||||
pub enum Endian { Little, Big }
|
||||
pub enum VariableEncoding { LengthPrefixed, OffsetIndirect }
|
||||
pub enum DiscriminatorKind {
|
||||
Byte { offset: usize, disc_type: TypeDefKind },
|
||||
Field { name: String },
|
||||
}
|
||||
```
|
||||
|
||||
`DiscriminatorKind::Byte` carries the byte position (`offset`) and the
|
||||
discriminator's `TypeDef:*` kind (`disc_type`, restricted to `Uint8`/
|
||||
`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator
|
||||
field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for
|
||||
how these drive dispatch.
|
||||
|
||||
### `$ref` resolution and normalization
|
||||
|
||||
| Function | Purpose |
|
||||
|----------|---------|
|
||||
| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/<name>"`. Idempotent. Runs once at `TypedefEngine::compile` time. |
|
||||
| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. |
|
||||
| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). |
|
||||
|
||||
`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s
|
||||
JSON Pointer requirement. The layout engines call `resolve_ref_or_inline`
|
||||
on every `$ref`-bearing node they encounter during traversal.
|
||||
|
||||
## jsonschema Custom Keyword Integration
|
||||
|
||||
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
|
||||
via the `with_keyword` API. Each `TypeDef:*` kind is registered as a
|
||||
custom keyword:
|
||||
|
||||
```rust
|
||||
let validator = jsonschema::options()
|
||||
.with_keyword("TypeDef:Float32", factory)
|
||||
.with_keyword("TypeDef:Int32", factory)
|
||||
.with_keyword("TypeDef:Struct", factory)
|
||||
// ... all 17 kinds
|
||||
.build(&schema)?;
|
||||
```
|
||||
|
||||
The factory closure receives the parent schema object, the keyword's
|
||||
value, and the schema path — enabling cross-keyword awareness. The
|
||||
`TypeDef:Struct` validator, for example, inspects the parent's
|
||||
`properties` to validate each field against its declared `TypeDef:*` kind.
|
||||
|
||||
Each custom keyword implementation is ~10 lines. The `jsonschema` crate
|
||||
handles all structural validation (object properties, required fields,
|
||||
array items, enum values) — the custom keywords only need to validate
|
||||
the leaf type constraints. See [validation.md](validation.md) for the
|
||||
validator implementations.
|
||||
|
||||
This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side.
|
||||
Same semantics, different language, same JSON Schema wire format. A
|
||||
TypeBox schema serialized to JSON feeds into the typedef engine after a
|
||||
single pre-processing step: normalizing `$ref` values (see below).
|
||||
|
||||
## TypeBox Interop
|
||||
|
||||
TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox
|
||||
schema like:
|
||||
|
||||
```typescript
|
||||
const TensorRef = Type.Object({
|
||||
dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]),
|
||||
shape: Type.Array(Type.Number()),
|
||||
data_offsets: Type.Tuple([Type.Number(), Type.Number()])
|
||||
});
|
||||
```
|
||||
|
||||
serialized to JSON is a standard JSON Schema with `type: "object"`,
|
||||
`properties`, and `required`. That JSON feeds into the typedef engine
|
||||
after `$ref` normalization. The `TypeDef:*` custom keywords are added by
|
||||
TypeBox's `TypeRegistry.Set` — they appear in the serialized JSON as
|
||||
additional properties on the schema object.
|
||||
|
||||
### `$ref` normalization
|
||||
|
||||
TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`),
|
||||
referencing sibling definitions within the same `$defs` block. The
|
||||
`jsonschema` crate requires full JSON Pointer paths (e.g.,
|
||||
`"$ref": "#/$defs/Read"`). The typedef engine normalizes TypeBox-style
|
||||
refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization)
|
||||
— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to
|
||||
`"#/$defs/<name>"`. The normalization is idempotent — full JSON Pointer
|
||||
refs pass through unchanged. It runs once at `TypedefEngine::compile`
|
||||
time, before the schema is passed to `jsonschema` or the offset
|
||||
computation.
|
||||
|
||||
**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs
|
||||
with `Resource 'Read' is not present in a registry`. Full JSON Pointer
|
||||
refs (`#/$defs/Read`) resolve correctly. The normalization step bridges
|
||||
the gap between TypeBox's output and jsonschema's input.
|
||||
|
||||
The typedef engine does not depend on TypeBox or any JS toolchain. It
|
||||
consumes JSON — whether that JSON was authored in TypeBox, generated by
|
||||
a ujsx component, or hand-written. The schema is the interface.
|
||||
|
||||
## Schema Annotations
|
||||
|
||||
Schema-level annotations control binary layout behavior. These are
|
||||
decided in [ADR-097](decisions/097-schema-annotations.md).
|
||||
|
||||
### Endianness
|
||||
|
||||
Schema-level annotation with a default of little-endian:
|
||||
|
||||
```json
|
||||
{ "TypeDef:Struct": true, "endian": "big", "properties": { ... } }
|
||||
```
|
||||
|
||||
- `"endian": "little"` (default) — read/write in little-endian byte order.
|
||||
- `"endian": "big"` — read/write in big-endian byte order.
|
||||
- Applies to the entire schema and all nested types.
|
||||
|
||||
### Alignment
|
||||
|
||||
Both struct-level and field-level, with field-level overriding:
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Struct": true,
|
||||
"align": 256,
|
||||
"properties": {
|
||||
"weight": { "TypeDef:Float32": true, "align": 16 }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- Struct-level `"align"` sets the default for all fields.
|
||||
- Field-level `"align"` overrides the struct default.
|
||||
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/
|
||||
enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix),
|
||||
1 for struct/union/array.
|
||||
- Only meaningful in aligned static mode (ADR-096). Ignored in packed
|
||||
sequential mode.
|
||||
|
||||
### Variable-length encoding
|
||||
|
||||
The typedef engine supports three strategies for variable-length types
|
||||
(see §Variable-length types above for full details). The strategy is
|
||||
selected by the `encoding` annotation and the standard JSON Schema
|
||||
`maxLength` keyword:
|
||||
|
||||
```json
|
||||
// Strategy 1: Inline length-prefixing (default, shorthand)
|
||||
{ "TypeDef:String": true }
|
||||
|
||||
// Strategy 1: Explicit inline length-prefixing
|
||||
{ "TypeDef:String": { "encoding": "length-prefixed" } }
|
||||
|
||||
// Strategy 2: Fixed-size reservation (uses standard maxLength)
|
||||
{ "TypeDef:String": true, "maxLength": 256 }
|
||||
|
||||
// Strategy 3: Offset indirection (opt-in)
|
||||
{ "TypeDef:String": { "encoding": "offset-indirect" } }
|
||||
```
|
||||
|
||||
- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at
|
||||
computed offset, variable data follows immediately. Used by protocol
|
||||
wire formats.
|
||||
- `maxLength` (standard JSON Schema keyword) — in aligned static mode,
|
||||
reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the
|
||||
field fixed-size from the layout perspective. In packed sequential
|
||||
mode, `maxLength` is a validation constraint only.
|
||||
- `"encoding": "offset-indirect"` — field is a struct
|
||||
`{offset: u32, length: u32}` pointing into a separate data region.
|
||||
The consumer provides the data region separately. Used by metatensor
|
||||
blob tensors.
|
||||
- Applies to all variable-length types: `TypeDef:String`, `TypeDef:Bytes`,
|
||||
`TypeDef:Array`, `TypeDef:Record`, `TypeDef:Timestamp`.
|
||||
|
||||
### TUnion discriminators
|
||||
|
||||
Two discriminator kinds: byte-offset (protocol dispatch) and field-name
|
||||
(typedef.ts pattern).
|
||||
|
||||
**Byte-offset discriminator** (SFTP type bytes, call protocol event types):
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "TypeDef:Uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" },
|
||||
"101": { "$ref": "#/$defs/Status" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `"offset"` — byte position of the discriminator.
|
||||
- `"type"` — the `TypeDef:*` kind of the discriminator (typically
|
||||
`TypeDef:Uint8`).
|
||||
- Mapping keys are stringified integers. The variant struct starts at
|
||||
`offset + discriminator_size`.
|
||||
|
||||
**Field-name discriminator** (typedef.ts pattern):
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "field",
|
||||
"name": "type"
|
||||
},
|
||||
"mapping": {
|
||||
"read": { "$ref": "#/$defs/Read" },
|
||||
"write": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `"name"` — the field name holding the discriminator value.
|
||||
- Mapping keys are string values matching the discriminator field's value.
|
||||
- The discriminator field is just another field in the struct.
|
||||
|
||||
Mapping values may be either inline schemas or `$ref` pointers. Both work.
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Schema annotations | [ADR-097](decisions/097-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators |
|
||||
| Int64/Uint64 kinds | [ADR-099](decisions/099-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (required by SFTP offsets and metatensor data_offsets) |
|
||||
| Purpose and scope | [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle |
|
||||
|
||||
## Open Questions
|
||||
|
||||
See [open-questions.md](open-questions.md) for full details.
|
||||
|
||||
- **OQ-071** (deferred(scope)): Builder API for schema construction.
|
||||
|
||||
## References
|
||||
|
||||
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||
schema kinds (619 lines)
|
||||
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||
- [ADR-097](decisions/097-schema-annotations.md) — schema
|
||||
annotation shapes
|
||||
- [validation.md](validation.md) — custom keyword validator implementations
|
||||
@@ -0,0 +1,335 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-07-22
|
||||
---
|
||||
|
||||
# alknet-typedef — Validation
|
||||
|
||||
The validation layer: custom keyword validators for all 19 `TypeDef:*`
|
||||
kinds, the `TypedefError` enum, load-time vs access-time validation
|
||||
strategy, and the `TypedefEngine` as the compiled form of a schema.
|
||||
|
||||
## Validation Strategy
|
||||
|
||||
Validation is delegated to the `jsonschema` crate (v0.46.5, Draft
|
||||
2020-12). The typedef engine does not implement its own validation —
|
||||
it registers custom keyword validators for each `TypeDef:*` kind and
|
||||
lets `jsonschema` handle the structural validation (object properties,
|
||||
required fields, array items, enum values).
|
||||
|
||||
The strategy is decided in [ADR-098](decisions/098-error-handling-validation-strategy.md):
|
||||
|
||||
1. **Load time:** Parse the schema JSON, build the layout engine, build the
|
||||
jsonschema validator. This is the `TypedefEngine::compile(schema)` constructor.
|
||||
2. **Access time:** Use the compiled engine for repeated read/write
|
||||
operations. Validation is opt-in per operation.
|
||||
|
||||
### What validation validates
|
||||
|
||||
The jsonschema validator operates on `serde_json::Value` instances — it
|
||||
validates JSON representations of data, not raw byte buffers. This is
|
||||
the correct separation of concerns:
|
||||
|
||||
- **JSON validation** (jsonschema): validates that a JSON document
|
||||
conforms to the schema. Used for validating hand-written schemas,
|
||||
TypeBox output, JSON payloads, or the JSON representation of a binary
|
||||
struct after deserialization.
|
||||
- **Binary access validation** (data access layer): the read/write
|
||||
functions perform type-level validation at access time — range checks
|
||||
for integers, UTF-8 validity for strings, buffer bounds checking.
|
||||
These return `TypedefError::Access` with field paths.
|
||||
|
||||
The "schema is the format" principle means the same schema describes
|
||||
both the JSON shape and the binary layout. The jsonschema validator
|
||||
checks the JSON shape; the data access layer checks the binary layout.
|
||||
A consumer that wants to validate a binary buffer end-to-end reads the
|
||||
buffer into a `Value` tree via the data access layer, then validates
|
||||
that `Value` against the jsonschema validator. This is a two-step
|
||||
process, not a single `validate(buffer)` call.
|
||||
|
||||
### The `TypedefEngine` struct
|
||||
|
||||
The `TypedefEngine` is the compiled form of a schema. It supports both
|
||||
layout modes (ADR-096) via an internal `Layout` enum:
|
||||
|
||||
```rust
|
||||
pub struct TypedefEngine {
|
||||
layout: Layout, // packed or aligned (private enum)
|
||||
validator: jsonschema::Validator, // compiled once at load time
|
||||
endian: Endian, // parsed from the schema's "endian" annotation
|
||||
schema: Value, // the normalized schema (refs resolved)
|
||||
}
|
||||
|
||||
// Private — the consumer selects via LayoutMode at compile time.
|
||||
enum Layout {
|
||||
Packed { builder: LayoutBuilder },
|
||||
Aligned { offset_map: OffsetMap },
|
||||
}
|
||||
```
|
||||
|
||||
The consumer selects the mode at construction time via `LayoutMode`
|
||||
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout`
|
||||
enum is private — the engine exposes mode-appropriate accessors instead:
|
||||
|
||||
```rust
|
||||
impl TypedefEngine {
|
||||
pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result<Self, TypedefError>;
|
||||
pub fn mode(&self) -> LayoutMode;
|
||||
pub fn endian(&self) -> Endian;
|
||||
pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode
|
||||
pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode
|
||||
pub fn sequential_reader(&self) -> Option<SequentialReader>; // owned fresh reader (ADR-101)
|
||||
}
|
||||
```
|
||||
|
||||
`compile` takes `&mut Value` because it normalizes `$ref` values in place
|
||||
(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization))
|
||||
before computing the layout and building the validator. The `schema`
|
||||
field retains the normalized schema for `read_field`'s kind lookup and
|
||||
for `sequential_reader()`'s factory construction. The validator is
|
||||
mode-agnostic (it operates on `Value`, not raw bytes).
|
||||
|
||||
The `Layout::Packed` variant stores only the `LayoutBuilder` (write-side).
|
||||
The `SequentialReader` (read-side) is not stored — it has mutable cursor
|
||||
state that the consumer owns, so `sequential_reader()` constructs a fresh
|
||||
reader on each call (ADR-101).
|
||||
|
||||
The `read_field`/`write_field` methods on `TypedefEngine` are the
|
||||
aligned-mode data-access API — see [data-access.md](data-access.md)
|
||||
§"Higher-level read/write".
|
||||
|
||||
## Custom Keyword Validators
|
||||
|
||||
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
|
||||
`jsonschema::options().with_keyword(...)`. The validators check leaf
|
||||
type constraints; `jsonschema` handles all structural validation.
|
||||
|
||||
### Numeric type validators
|
||||
|
||||
**`TypeDef:Float32` / `TypeDef:Float64`:**
|
||||
- Value must be a finite number.
|
||||
- For `Float32`: value must be representable as `f32` (no precision loss
|
||||
beyond `f32`'s mantissa).
|
||||
|
||||
**`TypeDef:Int8` / `TypeDef:Int16` / `TypeDef:Int32`:**
|
||||
- Value must be an integer within the type's range.
|
||||
- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
|
||||
|
||||
**`TypeDef:Uint8` / `TypeDef:Uint16` / `TypeDef:Uint32`:**
|
||||
- Value must be a non-negative integer within the type's range.
|
||||
- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295.
|
||||
|
||||
### String and binary validators
|
||||
|
||||
**`TypeDef:String`:**
|
||||
- Value must be a valid UTF-8 string.
|
||||
- If `maxLength` is specified in the schema, the string's byte length
|
||||
must not exceed it.
|
||||
|
||||
**`TypeDef:Bytes`:**
|
||||
- Value must be a string (JSON represents binary data as a string — JSON
|
||||
has no native byte type).
|
||||
- If `maxLength` is specified, the byte length must not exceed it.
|
||||
- **Binary representation:** In the binary layout, `TBytes` is raw bytes
|
||||
with no encoding (not base64, not hex). The JSON representation (for
|
||||
validation) uses a string; the binary representation (for data access)
|
||||
uses `&[u8]` directly.
|
||||
|
||||
**`TypeDef:Enum`:**
|
||||
- The `TypeDef:Enum` custom keyword signals that the type is an enum for
|
||||
*layout* purposes (the engine needs to know it's a fixed-size u32 index,
|
||||
not a variable-length string). The built-in `enum` keyword provides the
|
||||
value list and handles value-membership validation. The custom keyword
|
||||
validator is a no-op beyond the built-in check — it exists solely for
|
||||
the layout engine to recognize the type.
|
||||
|
||||
**`TypeDef:Timestamp`:**
|
||||
- Value must be a valid RFC 3339 timestamp string (the internet profile
|
||||
of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
|
||||
|
||||
### Composite type validators
|
||||
|
||||
**`TypeDef:Struct`:**
|
||||
- Value must be an object.
|
||||
- Each property must match its declared `TypeDef:*` kind.
|
||||
- Required fields must be present.
|
||||
- The `jsonschema` crate's built-in `properties` and `required` keywords
|
||||
handle the structural checks — the custom keyword only needs to
|
||||
validate that each field's value matches its `TypeDef:*` kind.
|
||||
|
||||
**`TypeDef:Union`:**
|
||||
- The discriminator value must be one of the mapping keys.
|
||||
- The variant struct must match the declared schema for that discriminator
|
||||
value.
|
||||
|
||||
**`TypeDef:Array`:**
|
||||
- Value must be an array.
|
||||
- Each element must match the array's declared element type.
|
||||
- If `minItems`/`maxItems` is specified, the array length must be within
|
||||
bounds.
|
||||
|
||||
### Other validators
|
||||
|
||||
**`TypeDef:Boolean`:**
|
||||
- Value must be `true` or `false`.
|
||||
|
||||
**`TypeDef:Record`:**
|
||||
- Value must be an object.
|
||||
- All values must match the record's declared value type (specified via
|
||||
the `"values"` property in the schema, e.g.,
|
||||
`"values": { "TypeDef:Float32": true }`).
|
||||
|
||||
### Validator implementation pattern
|
||||
|
||||
Each custom keyword implementation is ~10 lines. Example for
|
||||
`TypeDef:Float32`:
|
||||
|
||||
```rust
|
||||
struct Float32Validator;
|
||||
|
||||
impl Keyword for Float32Validator {
|
||||
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
|
||||
match instance {
|
||||
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
|
||||
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
|
||||
}
|
||||
}
|
||||
fn is_valid(&self, instance: &Value) -> bool {
|
||||
instance.as_f64().map_or(false, |f| f.is_finite())
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Registration:
|
||||
|
||||
```rust
|
||||
let validator = jsonschema::options()
|
||||
.with_keyword("TypeDef:Float32", |parent, value, path| {
|
||||
Ok(Box::new(Float32Validator))
|
||||
})
|
||||
.build(&schema)?;
|
||||
```
|
||||
|
||||
The factory closure receives the parent schema object, the keyword's
|
||||
value, and the schema path. This enables cross-keyword awareness — for
|
||||
example, a `TypeDef:Struct` validator can inspect the parent's
|
||||
`properties` to validate each field against its declared `TypeDef:*` kind.
|
||||
|
||||
## TypedefError
|
||||
|
||||
A single `TypedefError` enum covers all error conditions across the
|
||||
engine's three phases (schema parsing, offset computation, read/write)
|
||||
plus validation. Decided in [ADR-098](decisions/098-error-handling-validation-strategy.md).
|
||||
|
||||
```rust
|
||||
pub enum TypedefError {
|
||||
/// Schema parsing errors (invalid JSON, missing keywords, unknown TypeDef kinds).
|
||||
Schema(String),
|
||||
/// Offset computation errors (field not found, unsupported type).
|
||||
Offset { field_path: String, reason: String },
|
||||
/// Read/write errors (buffer too short, invalid UTF-8, value out of range).
|
||||
Access { field_path: String, reason: String },
|
||||
/// Validation errors (delegated to jsonschema).
|
||||
Validation(ValidationError<'static>),
|
||||
}
|
||||
```
|
||||
|
||||
- **`Schema`** — for errors during `TypedefEngine::compile()`. Invalid
|
||||
JSON, missing required keywords, unknown `TypeDef:*` kinds.
|
||||
- **`Offset`** — for errors during offset computation. Field not found
|
||||
in the schema, type not supported for offset computation, recursive
|
||||
depth exceeded. Carries the field path.
|
||||
- **`Access`** — for errors during read/write. Buffer too short, invalid
|
||||
UTF-8 in a string field, value out of range for the target type.
|
||||
Carries the field path.
|
||||
- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The
|
||||
`'static` lifetime is correct — the validator owns its schema reference
|
||||
and lives for the lifetime of the `TypedefEngine`.
|
||||
|
||||
### Field-path-carrying errors
|
||||
|
||||
Read/write and offset errors include the field path for debugging:
|
||||
|
||||
```rust
|
||||
Err(TypedefError::Access {
|
||||
field_path: "header.version".to_string(),
|
||||
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
|
||||
})
|
||||
```
|
||||
|
||||
This makes debugging binary format issues tractable — the error tells
|
||||
you exactly which field failed and why.
|
||||
|
||||
## Validation Timing
|
||||
|
||||
### Load time: `TypedefEngine::compile()`
|
||||
|
||||
The expensive work happens once at schema load time:
|
||||
1. Normalize `$ref` values in the schema (`normalize_refs`).
|
||||
2. Parse the schema's `"endian"` annotation.
|
||||
3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned).
|
||||
4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
|
||||
|
||||
The result is a `TypedefEngine` that can be used for repeated operations.
|
||||
|
||||
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
|
||||
|
||||
Validation is opt-in per operation. The consumer calls
|
||||
`engine.validate_json(instance)` when validation is desired, or
|
||||
`engine.is_valid_json(instance)` for a boolean check. The jsonschema
|
||||
validator is already compiled — these are fast checks against the
|
||||
compiled validator.
|
||||
|
||||
```rust
|
||||
pub fn validate_json(&self, instance: &Value) -> Result<(), TypedefError>;
|
||||
pub fn is_valid_json(&self, instance: &Value) -> bool;
|
||||
```
|
||||
|
||||
The argument is a `serde_json::Value` (the JSON representation of the
|
||||
data), not a raw byte buffer — see §"What validation validates" above.
|
||||
To validate a binary buffer end-to-end, the consumer reads it into a
|
||||
`Value` tree via the data access layer, then validates that `Value`.
|
||||
|
||||
High-throughput paths can skip validation. Security-sensitive paths
|
||||
(parsing incoming frames from untrusted peers) can validate every frame.
|
||||
The choice is the consumer's.
|
||||
|
||||
## Relationship to Read/Write
|
||||
|
||||
Validation and data access are independent operations on the same data.
|
||||
The consumer can:
|
||||
|
||||
1. Validate the JSON representation of a buffer to ensure it conforms to
|
||||
the schema.
|
||||
2. Read fields from the binary buffer at computed offsets.
|
||||
3. Both — validate the JSON representation first, then read the binary
|
||||
buffer (defense in depth).
|
||||
|
||||
The engine does not couple validation and access. A consumer that trusts
|
||||
its data source can skip validation and go straight to read/write. A
|
||||
consumer that parses untrusted input can validate the JSON
|
||||
representation first, then access the binary buffer.
|
||||
|
||||
## Design Decisions
|
||||
|
||||
| Decision | ADR | Summary |
|
||||
|----------|-----|---------|
|
||||
| Error handling and validation | [ADR-098](decisions/098-error-handling-validation-strategy.md) | `TypedefError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
|
||||
| Purpose and scope | [ADR-095](decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine |
|
||||
|
||||
## Open Questions
|
||||
|
||||
None specific to validation. The three typedef OQs (OQ-069, OQ-070,
|
||||
OQ-071) are about layout, platform support, and schema construction —
|
||||
not validation.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"Validation" — the POC's
|
||||
custom keyword validators for all 17 kinds
|
||||
- [ADR-098](decisions/098-error-handling-validation-strategy.md) —
|
||||
error handling and validation strategy
|
||||
- [schema-layer.md](schema-layer.md) — the 17 TypeDef kinds that the
|
||||
validators check
|
||||
- [data-access.md](data-access.md) — read/write functions that operate
|
||||
on the same buffers
|
||||
Reference in new issue
Block a user