feat(typedef): add implementation task decomposition (12 tasks, 7 generations)

Break the alknet-typedef architecture specs into atomic, dependency-ordered
implementation tasks covering crate init, error types, schema layer, data
access, both layout modes (aligned static + packed sequential), TUnion
dispatch, jsonschema custom keyword validators, TypedefEngine integration,
comprehensive tests, and a final review checkpoint.

Validated: 12 tasks, 0 cycles, 7 parallel generations.
This commit is contained in:
deepseek-v4-pro committed 2026-07-21 09:36:30 +00:00
1 parent dd232c3d47
commit db100f9849
12 files changed
+1996

No files matched your search

+131
View File
@@ -0,0 +1,131 @@
---
id: typedef/crate-init
name: Initialize alknet-typedef crate with Cargo.toml, dependencies, and module skeleton
status: pending
depends_on: []
scope: moderate
risk: low
impact: project
level: implementation
---
## Description
Initialize the `alknet-typedef` crate from scratch. This is a greenfield crate — the
binary struct engine that takes a JSON Schema with `TypeDef:*` custom keywords and
produces an offset map, read/write functions, and validation. The schema is the format
definition; the engine is generic.
Per [ADR-095](../../docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md),
the crate depends on `jsonschema` (v0.46.5, Draft 2020-12) for validation and
`serde_json` (with `preserve_order`) for schema parsing. No tokio, no platform deps.
Compiles to `wasm32-unknown-unknown`.
### Crate setup
Create `crates/alknet-typedef/` with:
- `Cargo.toml` — package metadata, dependencies
- `src/lib.rs` — crate root with module declarations and re-exports
- Module skeleton files for:
- `src/error.rs` — `TypedefError` enum (ADR-098)
- `src/schema.rs` — TypeDef kind detection, annotation parsing, `$ref` normalization, `Endian` enum (ADR-097)
- `src/data_access.rs` — primitive read/write functions for all 17 TypeDef kinds
- `src/offset_map.rs` — aligned static `OffsetMap` computation (ADR-096 Mode 2)
- `src/layout_builder.rs` — packed sequential `LayoutBuilder` (ADR-096 Mode 1, write side)
- `src/sequential_reader.rs` — packed sequential `SequentialReader` (ADR-096 Mode 1, read side)
- `src/tunion.rs` — TUnion discriminator dispatch (byte-offset and field-name, ADR-097 §4)
- `src/validation.rs` — custom keyword validators for all 17 `TypeDef:*` kinds
- `src/engine.rs` — `TypedefEngine` struct combining layout + validator
### Dependencies
Per the architecture spec ([overview.md](../../docs/architecture/crates/typedef/overview.md)):
| Crate | Purpose |
|-------|---------|
| `jsonschema` 0.46 | Validation engine, custom keyword support (workspace path: `../../jsonschema`) |
| `serde_json` (preserve_order) | Schema parsing; field order is load-bearing for binary layouts |
No other dependencies. No tokio, no platform deps.
### Workspace Cargo.toml
Add `crates/alknet-typedef` to the workspace `members` list in the root `Cargo.toml`.
### Module skeleton
```rust
// src/lib.rs
//! alknet-typedef: The binary struct engine.
//!
//! Takes a JSON Schema with `TypeDef:*` custom keywords and produces
//! an offset map, read/write functions, and validation — all driven
//! by the schema. The schema is the format definition; the engine is
//! generic.
//!
//! ## Architecture
//!
//! - **Schema layer** ([`schema`]): TypeDef kind detection, annotation
//! parsing, `$ref` normalization, endianness.
//! - **Layout engine** ([`offset_map`], [`layout_builder`],
//! [`sequential_reader`]): Two layout modes — aligned static for
//! mmap-friendly formats, packed sequential for protocol wire formats.
//! - **Data access** ([`data_access`]): Typed read/write at computed
//! offsets, zero-copy for fixed-size types.
//! - **TUnion dispatch** ([`tunion`]): Byte-offset and field-name
//! discriminator dispatch.
//! - **Validation** ([`validation`]): Custom keyword validators for all
//! 17 `TypeDef:*` kinds, delegated to the `jsonschema` crate.
//! - **Engine** ([`engine`]): `TypedefEngine` — the compiled form of a
//! schema, combining layout and validation.
pub mod data_access;
pub mod engine;
pub mod error;
pub mod layout_builder;
pub mod offset_map;
pub mod schema;
pub mod sequential_reader;
pub mod tunion;
pub mod validation;
// Re-exports (filled in by subsequent tasks)
```
Each module file gets a doc comment and `// TODO: implement` marker.
## Acceptance Criteria
- [ ] `crates/alknet-typedef/Cargo.toml` exists with `jsonschema` and `serde_json` (preserve_order) dependencies
- [ ] `jsonschema` dependency uses workspace path (`path = "../../jsonschema"`) or git/crates.io as appropriate
- [ ] `crates/alknet-typedef/src/lib.rs` exists with module declarations for all 9 modules
- [ ] Module skeleton files exist: `error.rs`, `schema.rs`, `data_access.rs`, `offset_map.rs`, `layout_builder.rs`, `sequential_reader.rs`, `tunion.rs`, `validation.rs`, `engine.rs`
- [ ] Root `Cargo.toml` `members` list includes `crates/alknet-typedef`
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] Dual licensing: `MIT OR Apache-2.0` (workspace-inherited)
- [ ] No tokio dependency (WASM-clean by construction)
- [ ] `serde_json` has `preserve_order` feature enabled
- [ ] `cargo build --workspace` still succeeds (old code untouched)
## References
- docs/architecture/crates/typedef/README.md — crate overview and design principles
- docs/architecture/crates/typedef/overview.md — purpose, dependencies, scope boundaries
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
- docs/research/alknet-typedef/findings.md — POC results
- /workspace/jsonschema/ — the jsonschema crate (v0.46.5)
- /workspace/alknet-typedef-poc/ — POC code (disposable reference)
## Notes
> This is the foundational setup task for alknet-typedef. All subsequent typedef/*
> tasks depend on this one. The crate is dependency-light: `jsonschema` + `serde_json`
> only. No tokio, no platform deps — WASM-clean by construction. The `jsonschema` crate
> is already in the workspace at `/workspace/jsonschema/` but not yet used by any
> alknet crate — typedef is the first consumer.
## Summary
> To be filled on completion
+174
View File
@@ -0,0 +1,174 @@
---
id: typedef/data-access
name: Implement primitive read/write functions for all 17 TypeDef kinds with endianness support
status: pending
depends_on: [typedef/schema-types, typedef/error-type]
scope: moderate
risk: medium
impact: component
level: implementation
---
## Description
Implement the data access layer in `crates/alknet-typedef/src/data_access.rs`. This
module provides the primitive typed read/write functions that operate on raw byte
buffers at given offsets. These are the building blocks used by the layout types
(`OffsetMap`, `SequentialReader`) and the `TypedefEngine`.
Per [data-access.md](../../docs/architecture/crates/typedef/data-access.md).
### Fixed-size type read/write
All fixed-size read/write functions take a buffer, an offset, and an `Endian`, and
return `Result<T, TypedefError>` (or `Result<(), TypedefError>` for writes). They
perform bounds checking and return `TypedefError::Access` with the field path on
failure.
```rust
// Signed integers
pub fn read_i8(buffer: &[u8], offset: usize, field_path: &str) -> Result<i8, TypedefError>;
pub fn read_i16(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<i16, TypedefError>;
pub fn read_i32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<i32, TypedefError>;
// Unsigned integers
pub fn read_u8(buffer: &[u8], offset: usize, field_path: &str) -> Result<u8, TypedefError>;
pub fn read_u16(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u16, TypedefError>;
pub fn read_u32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u32, TypedefError>;
pub fn read_u64(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u64, TypedefError>;
// Floats
pub fn read_f32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<f32, TypedefError>;
pub fn read_f64(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<f64, TypedefError>;
// Boolean (0x00 = false, 0x01 = true; other values are errors)
pub fn read_bool(buffer: &[u8], offset: usize, field_path: &str) -> Result<bool, TypedefError>;
// TEnum (u32 index, endian-aware)
pub fn read_enum(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u32, TypedefError>;
```
Write counterparts:
```rust
pub fn write_i8(buffer: &mut [u8], offset: usize, value: i8, field_path: &str) -> Result<(), TypedefError>;
pub fn write_i16(buffer: &mut [u8], offset: usize, value: i16, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
pub fn write_i32(buffer: &mut [u8], offset: usize, value: i32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
pub fn write_u8(buffer: &mut [u8], offset: usize, value: u8, field_path: &str) -> Result<(), TypedefError>;
pub fn write_u16(buffer: &mut [u8], offset: usize, value: u16, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
pub fn write_u32(buffer: &mut [u8], offset: usize, value: u32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
pub fn write_u64(buffer: &mut [u8], offset: usize, value: u64, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
pub fn write_f32(buffer: &mut [u8], offset: usize, value: f32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
pub fn write_f64(buffer: &mut [u8], offset: usize, value: f64, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
pub fn write_bool(buffer: &mut [u8], offset: usize, value: bool, field_path: &str) -> Result<(), TypedefError>;
pub fn write_enum(buffer: &mut [u8], offset: usize, value: u32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
```
### Variable-length type read/write (inline length-prefixing)
For variable-length types with inline length-prefixing (the default):
```rust
/// Read a length-prefixed UTF-8 string. Returns a slice borrowing from the buffer.
/// Format: [length: u32][UTF-8 bytes]
pub fn read_string<'a>(buffer: &'a [u8], offset: usize, field_path: &str, endian: Endian) -> Result<&'a str, TypedefError>;
/// Read length-prefixed raw bytes. Returns a slice borrowing from the buffer.
pub fn read_bytes<'a>(buffer: &'a [u8], offset: usize, field_path: &str, endian: Endian) -> Result<&'a [u8], TypedefError>;
/// Write a length-prefixed UTF-8 string.
pub fn write_string(buffer: &mut [u8], offset: usize, value: &str, field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
// Returns the number of bytes written (4 + value.len()) so the caller can advance.
/// Write length-prefixed raw bytes.
pub fn write_bytes(buffer: &mut [u8], offset: usize, value: &[u8], field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
```
### Variable-length type read (offset indirection)
For offset-indirect types (opt-in, metatensor blob tensor pattern):
```rust
/// Read an offset-indirect string. The field at `offset` is a struct
/// {data_offset: u32, data_length: u32}. The consumer provides the
/// separate data region.
pub fn read_string_indirect<'a>(
buffer: &'a [u8], // the index struct buffer
offset: usize, // position of {data_offset, data_length}
data_region: &'a [u8], // the separate data region
field_path: &str,
endian: Endian,
) -> Result<&'a str, TypedefError>;
/// Read offset-indirect raw bytes.
pub fn read_bytes_indirect<'a>(
buffer: &'a [u8],
offset: usize,
data_region: &'a [u8],
field_path: &str,
endian: Endian,
) -> Result<&'a [u8], TypedefError>;
```
### Design notes
- **Zero-copy**: Read functions for variable-length types return slices borrowing
from the input buffer — no allocation.
- **Bounds checking**: Every function checks that the buffer is large enough for
the requested read/write at the given offset. Returns `TypedefError::Access`
with the field path on failure.
- **Endianness**: Applied at access time based on the schema's `"endian"` annotation.
The offset computation is endian-agnostic.
- **No `unwrap`**: All fallible operations use proper `Result` returns. The spec
pseudocode uses `unwrap` for brevity; production code must not.
- **`TEnum`**: Reads/writes a `u32` index. The consumer maps the index to string
values using the schema's `"enum"` array. The engine does not perform this mapping.
- **`TBoolean`**: `0x00` = false, `0x01` = true. Other values produce
`TypedefError::Access`.
### What this does NOT include
- The offset computation (that's `offset_map.rs` and `layout_builder.rs`)
- TUnion discriminator dispatch (that's `tunion.rs`)
- The `TypedefEngine` struct (that's `engine.rs`)
- `TRecord` read/write (deferred — requires count-prefixed sequence walking; can be
added when a consumer needs it, or implemented here if straightforward)
## Acceptance Criteria
- [ ] All fixed-size read functions implemented: `read_i8`, `read_i16`, `read_i32`, `read_u8`, `read_u16`, `read_u32`, `read_u64`, `read_f32`, `read_f64`, `read_bool`, `read_enum`
- [ ] All fixed-size write functions implemented: `write_i8`, `write_i16`, `write_i32`, `write_u8`, `write_u16`, `write_u32`, `write_u64`, `write_f32`, `write_f64`, `write_bool`, `write_enum`
- [ ] `read_string` and `read_bytes` (inline length-prefixed) implemented
- [ ] `write_string` and `write_bytes` (inline length-prefixed) implemented, returning bytes written
- [ ] `read_string_indirect` and `read_bytes_indirect` (offset-indirect) implemented
- [ ] All functions respect `Endian` parameter (little-endian vs big-endian byte order)
- [ ] All functions perform bounds checking and return `TypedefError::Access` with field path on failure
- [ ] `read_bool` rejects values other than `0x00` and `0x01`
- [ ] `read_string` validates UTF-8 and returns `TypedefError::Access` on invalid UTF-8
- [ ] Zero-copy: read functions for variable-length types return slices, not owned data
- [ ] No `unwrap()` or `expect()` on error paths — all fallible operations use `Result`
- [ ] All public functions have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/data-access.md — read/write model, TEnum access, variable-length handling
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds and their byte sizes
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (endianness, encoding)
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098 (error handling)
- /workspace/alknet-typedef-poc/src/lib.rs — POC reference for read/write functions
## Notes
> This module provides the primitive read/write operations. The layout types
> (`OffsetMap`, `SequentialReader`) use these to access fields at computed positions.
> The functions are endian-aware — the caller passes the schema's `Endian` and the
> functions byte-swap accordingly. All functions use proper `Result` returns with
> field paths for debugging — no `unwrap()` in production code. The `TRecord`
> read/write is deferred unless it proves straightforward to implement here.
## Summary
> To be filled on completion
+202
View File
@@ -0,0 +1,202 @@
---
id: typedef/engine
name: Implement TypedefEngine struct combining layout and validation, with compile() constructor
status: pending
depends_on: [typedef/offset-map, typedef/layout-builder, typedef/sequential-reader, typedef/tunion, typedef/validation]
scope: moderate
risk: medium
impact: component
level: implementation
---
## Description
Implement the `TypedefEngine` struct in `crates/alknet-typedef/src/engine.rs`. This is
the compiled form of a schema — the main entry point for consumers. It combines the
layout engine (both modes) and the jsonschema validator into a single struct.
Per [validation.md](../../docs/architecture/crates/typedef/validation.md) §"The TypedefEngine struct"
and [overview.md](../../docs/architecture/crates/typedef/overview.md).
### Target shape
```rust
/// The compiled form of a typedef schema. Combines the layout engine
/// (both packed and aligned modes) and the jsonschema validator.
///
/// Built once at schema load time via [`TypedefEngine::compile()`].
/// Used for repeated read/write/validate operations at access time.
pub struct TypedefEngine {
/// The layout strategy — packed sequential or aligned static.
layout: Layout,
/// The compiled jsonschema validator (built once at load time).
validator: jsonschema::Validator,
/// The schema's endianness.
endian: Endian,
/// The original schema (for reference, debugging, and TUnion dispatch).
schema: serde_json::Value,
}
/// The layout strategy selected by the consumer.
enum Layout {
/// Packed sequential layout for protocol wire formats.
Packed {
builder: LayoutBuilder,
reader: SequentialReader,
},
/// Aligned static layout for mmap-friendly formats.
Aligned {
offset_map: OffsetMap,
},
}
/// The layout mode selected at engine construction time.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum LayoutMode {
/// Packed sequential — for protocol wire formats (SFTP, channels, TTY).
Packed,
/// Aligned static — for mmap-friendly formats (metatensor, safetensors).
Aligned,
}
impl TypedefEngine {
/// Compile a schema into a TypedefEngine.
///
/// This is the expensive operation — it parses the schema, normalizes
/// `$ref` values, computes the layout, and builds the jsonschema
/// validator. Call once at load time; use the returned engine for
/// repeated operations.
///
/// The `mode` parameter selects the layout strategy. The same schema
/// can be compiled in either mode.
pub fn compile(
schema: &mut serde_json::Value,
mode: LayoutMode,
) -> Result<Self, TypedefError>;
/// The schema's endianness.
pub fn endian(&self) -> Endian;
/// The layout mode this engine was compiled with.
pub fn mode(&self) -> LayoutMode;
/// Access the aligned offset map. Returns `None` if compiled in
/// packed mode.
pub fn offset_map(&self) -> Option<&OffsetMap>;
/// Access the layout builder (write-side of packed mode).
/// Returns `None` if compiled in aligned mode.
pub fn layout_builder(&self) -> Option<&LayoutBuilder>;
/// Access the sequential reader (read-side of packed mode).
/// Returns `None` if compiled in aligned mode.
pub fn sequential_reader(&self) -> Option<&SequentialReader>;
/// Validate a JSON value against the schema. The jsonschema validator
/// is already compiled — this is a fast check.
///
/// Returns `Ok(())` if valid, `Err(TypedefError::Validation(...))` if invalid.
pub fn validate_json(&self, instance: &serde_json::Value) -> Result<(), TypedefError>;
/// Check if a JSON value is valid against the schema.
pub fn is_valid_json(&self, instance: &serde_json::Value) -> bool;
/// Read a field from a buffer at its computed offset (aligned mode).
/// Returns `None` if compiled in packed mode — use `sequential_reader()` instead.
pub fn read_field<'a>(
&self,
buffer: &'a [u8],
field_path: &str,
) -> Result<FieldValue<'a>, TypedefError>;
/// Write a field to a buffer at its computed offset (aligned mode).
/// Returns `None` if compiled in packed mode — use `layout_builder()` instead.
pub fn write_field(
&self,
buffer: &mut [u8],
field_path: &str,
value: &FieldValue<'_>,
) -> Result<(), TypedefError>;
}
```
### `compile()` constructor
The `compile()` method performs these steps in order:
1. **Normalize `$ref` values** — call `schema::normalize_refs()` to rewrite bare-name
refs to full JSON Pointer paths.
2. **Parse endianness** — extract `"endian"` annotation (defaults to `Little`).
3. **Build the layout** — depending on `mode`:
- `Packed`: build `LayoutBuilder` and `SequentialReader` from the schema.
- `Aligned`: compute `OffsetMap` from the schema.
4. **Build the validator** — call `validation::build_validator()` to register all 17
custom keywords and compile the jsonschema validator.
5. **Return the engine** — all components ready for repeated use.
### Aligned mode convenience methods
When compiled in `Aligned` mode, `read_field()` and `write_field()` provide
convenient access using the `OffsetMap`:
- `read_field(buffer, "header.version")` → looks up the field's byte range in the
`OffsetMap`, reads the appropriate type using `data_access` functions.
- `write_field(buffer, "header.version", &FieldValue::U32(1))` → looks up the byte
range, writes using `data_access` functions.
### What this does NOT include
- A builder API for schema construction (deferred, OQ-071)
- Schema evolution / Value system (out of scope for v1)
- Code generation (out of scope)
## Acceptance Criteria
- [ ] `TypedefEngine` struct with `layout`, `validator`, `endian`, `schema` fields
- [ ] `Layout` enum with `Packed { builder, reader }` and `Aligned { offset_map }` variants
- [ ] `LayoutMode` enum with `Packed` and `Aligned` variants
- [ ] `TypedefEngine::compile(&mut schema, mode)` performs all load-time work
- [ ] `compile()` normalizes `$ref` values before building
- [ ] `compile()` builds the correct layout for the selected mode
- [ ] `compile()` builds the jsonschema validator with all 17 custom keywords
- [ ] `compile()` returns `TypedefError::Schema` for invalid schemas
- [ ] `endian()` returns the schema's endianness
- [ ] `mode()` returns the layout mode
- [ ] `offset_map()` returns `Some(&OffsetMap)` in aligned mode, `None` in packed mode
- [ ] `layout_builder()` returns `Some(&LayoutBuilder)` in packed mode, `None` in aligned mode
- [ ] `sequential_reader()` returns `Some(&SequentialReader)` in packed mode, `None` in aligned mode
- [ ] `validate_json()` delegates to the compiled jsonschema validator
- [ ] `is_valid_json()` returns true/false without error details
- [ ] `read_field()` works in aligned mode for all fixed-size and variable-length types
- [ ] `write_field()` works in aligned mode for all fixed-size and variable-length types
- [ ] `read_field()` and `write_field()` return errors in packed mode (use layout-specific APIs)
- [ ] No `unwrap()` or `expect()` on error paths
- [ ] All public types and functions have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/validation.md — TypedefEngine struct, compile() constructor
- docs/architecture/crates/typedef/overview.md — architecture, consumers
- docs/architecture/crates/typedef/layout-engine.md — the two layout modes
- docs/architecture/crates/typedef/data-access.md — read/write functions
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
## Notes
> This is the integration task — it wires together all the components built in the
> preceding tasks. The `TypedefEngine` is the main entry point for consumers. The
> `compile()` constructor does all the expensive work once at load time. The engine
> supports both layout modes via the `Layout` enum — the consumer selects the mode
> at construction time. The aligned-mode convenience methods (`read_field`,
> `write_field`) provide a simple API for the common mmap use case. Protocol
> consumers use the layout-specific APIs (`layout_builder()`, `sequential_reader()`)
> directly.
## Summary
> To be filled on completion
+110
View File
@@ -0,0 +1,110 @@
---
id: typedef/error-type
name: Implement TypedefError enum with Schema, Offset, Access, and Validation variants
status: pending
depends_on: [typedef/crate-init]
scope: narrow
risk: low
impact: component
level: implementation
---
## Description
Implement the `TypedefError` enum in `crates/alknet-typedef/src/error.rs`. This is the
single error type for all three engine phases (schema parsing, offset computation,
read/write) plus validation. Decided in
[ADR-098](../../docs/architecture/decisions/098-error-handling-validation-strategy.md).
### Target shape (per ADR-098)
```rust
use std::fmt;
/// Errors produced by the typedef engine across all phases.
#[derive(Debug)]
pub enum TypedefError {
/// Schema parsing errors — invalid JSON, missing required keywords,
/// unknown `TypeDef:*` kinds, malformed annotations.
Schema(String),
/// Offset computation errors — field not found, type not supported
/// for offset computation, recursive depth exceeded.
Offset {
field_path: String,
reason: String,
},
/// Read/write errors — buffer too short, invalid UTF-8, value out
/// of range for the target type.
Access {
field_path: String,
reason: String,
},
/// Validation errors — delegated to the `jsonschema` crate.
/// The `'static` lifetime is correct: the validator owns its schema
/// reference and lives for the lifetime of the `TypedefEngine`.
Validation(jsonschema::ValidationError<'static>),
}
```
### Design rationale
- **`Schema(String)`** — for errors during `TypedefEngine::compile()`. Invalid JSON,
missing required keywords, unknown `TypeDef:*` kinds. The error message describes
the problem.
- **`Offset { field_path, reason }`** — for errors during offset computation. Field
not found in the schema, type not supported for offset computation. Carries the
field path for debugging.
- **`Access { field_path, reason }`** — for errors during read/write. Buffer too
short, invalid UTF-8 in a string field, value out of range. Carries the field
path for debugging.
- **`Validation(ValidationError<'static>)`** — wraps `jsonschema`'s `ValidationError`.
The `'static` lifetime is correct — the validator is built once at schema load time
and lives for the lifetime of the `TypedefEngine`.
### Trait implementations
- `Display` — human-readable error messages including field paths where applicable
- `Error` (std::error::Error) — for `?` propagation
- `Debug` — derived
The `Validation` variant requires `jsonschema::ValidationError` to be in scope.
Since the `jsonschema` crate is a dependency, this is straightforward.
### What this does NOT include
- No `From` impls for other error types (those are added as needed by subsequent tasks)
- No `PartialEq` — `ValidationError` may not implement it
- No `Clone` — errors are typically consumed, not cloned
## Acceptance Criteria
- [ ] `TypedefError` enum defined in `crates/alknet-typedef/src/error.rs`
- [ ] Four variants: `Schema(String)`, `Offset { field_path, reason }`, `Access { field_path, reason }`, `Validation(ValidationError<'static>)`
- [ ] `Display` impl with descriptive messages including field paths for `Offset` and `Access`
- [ ] `std::error::Error` impl
- [ ] `Debug` derived
- [ ] Re-exported from `lib.rs`
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
- docs/architecture/crates/typedef/validation.md — TypedefError section
- docs/architecture/crates/typedef/data-access.md — error handling in read/write
## Notes
> This is a small, self-contained task. The error type is used by every other module
> in the crate. `Schema` covers load-time errors, `Offset` covers layout computation
> errors, `Access` covers read/write errors, and `Validation` wraps jsonschema's
> error type. The `'static` lifetime on `ValidationError` is correct because the
> validator is built once and lives for the engine's lifetime.
## Summary
> To be filled on completion
+173
View File
@@ -0,0 +1,173 @@
---
id: typedef/layout-builder
name: Implement packed sequential LayoutBuilder for protocol write-side
status: pending
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
scope: moderate
risk: medium
impact: component
level: implementation
---
## Description
Implement the packed sequential `LayoutBuilder` in `crates/alknet-typedef/src/layout_builder.rs`.
This is the write-side of Mode 1 (ADR-096): fields are packed with no alignment padding.
Variable-length fields shift all subsequent fields. The consumer provides actual data
sizes for variable-length fields; the builder computes byte positions for each field.
Per [layout-engine.md](../../docs/architecture/crates/typedef/layout-engine.md) §"Mode 1: Packed sequential".
### Target shape
```rust
/// Builds a packed sequential layout for protocol wire formats.
/// Fields are packed with no alignment padding. Variable-length fields
/// shift all subsequent fields. The consumer provides actual data sizes
/// for variable-length fields to compute correct positions.
///
/// Used at write time when the consumer knows the data sizes upfront.
#[derive(Debug)]
pub struct LayoutBuilder {
/// The schema being laid out.
schema: serde_json::Value,
/// The endianness for the layout.
endian: Endian,
}
/// A field position computed by the LayoutBuilder.
#[derive(Debug, Clone)]
pub struct FieldPosition {
/// Byte offset of the field within the buffer.
pub offset: usize,
/// Byte size of the field (4 for length prefix of variable-length fields,
/// actual size for fixed-size fields).
pub size: usize,
/// The TypeDef kind of the field.
pub kind: String,
}
/// The result of building a layout: a map of field_path → FieldPosition
/// and the total buffer size needed.
#[derive(Debug)]
pub struct PackedLayout {
fields: Vec<(String, FieldPosition)>,
total_size: usize,
}
impl LayoutBuilder {
/// Create a new LayoutBuilder from a schema.
pub fn new(schema: &serde_json::Value) -> Result<Self, TypedefError>;
/// Build the packed layout given actual data sizes for variable-length fields.
/// `var_sizes` maps field paths to their actual byte sizes (not including
/// the 4-byte length prefix — the builder adds that).
///
/// For fixed-size fields, the size is known from the schema.
/// For variable-length fields, the size comes from `var_sizes`.
/// For TUnion, the consumer provides the discriminator value to select
/// the variant, and the variant's field sizes.
pub fn build(
&self,
var_sizes: &HashMap<String, usize>,
) -> Result<PackedLayout, TypedefError>;
}
impl PackedLayout {
/// Look up a field's position by dotted path.
pub fn get(&self, field_path: &str) -> Option<&FieldPosition>;
/// The total buffer size needed to hold all fields.
pub fn total_size(&self) -> usize;
/// Iterate over all (field_path, position) pairs in layout order.
pub fn iter(&self) -> impl Iterator<Item = &(String, FieldPosition)>;
}
```
### How it works
For a struct with fields `[u8, u32, string]` where the string is 10 bytes:
```
LayoutBuilder::build(var_sizes: {"payload": 10}):
field[0] u8: offset 0, size 1
field[1] u32: offset 1, size 4
field[2] string: offset 5, size 4 (length prefix) + 10 (data) = 14
total: 19
```
There is no alignment padding. The `u32` at offset 1 is unaligned — this is correct
for protocol wire formats, which pack fields tightly.
### Variable-length fields in packed mode
The `LayoutBuilder` takes actual data sizes for variable-length fields to compute
correct positions for subsequent fields. The consumer must know the data sizes before
writing — this is inherent to packed layouts.
For each variable-length field:
1. The builder records the position of the 4-byte length prefix.
2. The builder adds `4 + data_size` to the current offset.
3. Subsequent fields start after the variable data.
### TUnion in packed mode
For `TUnion`, the consumer provides the discriminator value and the variant's field
sizes. The builder:
1. Computes the discriminator's position and size.
2. Looks up the variant schema from the mapping.
3. Computes the variant's field positions starting at `offset + discriminator_size`.
4. The union's total size is `discriminator_size + variant_size`.
### What this does NOT include
- The read-side of packed mode (that's `sequential_reader.rs`)
- Aligned static layout (that's `offset_map.rs`)
- TUnion discriminator dispatch (that's `tunion.rs`)
- The `TypedefEngine` struct (that's `engine.rs`)
## Acceptance Criteria
- [ ] `LayoutBuilder` struct with `new(schema)` constructor
- [ ] `LayoutBuilder::build(var_sizes)` computes packed field positions
- [ ] Fixed-size fields get correct offsets with no alignment padding
- [ ] `u8` at offset 0, `u32` at offset 1 (no padding) — packed sequential
- [ ] Variable-length fields: 4-byte length prefix at computed offset, data follows
- [ ] Variable-length field sizes come from `var_sizes` map
- [ ] Subsequent fields shift based on actual variable-length data sizes
- [ ] Nested structs produce dotted field paths
- [ ] `TArray` of fixed-size elements: correct stride (element size, no padding)
- [ ] `TArray` with variable count: 4-byte count prefix at computed offset
- [ ] `TUnion` with byte-offset discriminator: discriminator at `offset`, variant at `offset + disc_size`
- [ ] `TUnion` with field-name discriminator: discriminator is a regular field
- [ ] `PackedLayout::get("field_name")` returns correct `FieldPosition`
- [ ] `PackedLayout::total_size()` returns correct total buffer size
- [ ] `PackedLayout::iter()` iterates fields in layout order
- [ ] Returns `TypedefError::Schema` for malformed schemas
- [ ] Returns `TypedefError::Offset` for missing variable-length field sizes
- [ ] No `unwrap()` or `expect()` on error paths
- [ ] All public types and functions have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/layout-engine.md — Mode 1: Packed sequential, LayoutBuilder
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds and their byte sizes
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (encoding)
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference (LayoutBuilder in POC 2)
## Notes
> This is the write-side of the packed sequential mode. The consumer knows the data
> sizes upfront (e.g., when constructing an SFTP response packet) and uses the
> `LayoutBuilder` to compute where each field goes. The builder does not write data —
> it only computes positions. The consumer uses the `data_access` module's write
> functions at the computed positions. The POC 2's `LayoutBuilder` is a good reference.
## Summary
> To be filled on completion
+163
View File
@@ -0,0 +1,163 @@
---
id: typedef/offset-map
name: Implement aligned static OffsetMap for mmap-friendly formats
status: pending
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
scope: moderate
risk: medium
impact: component
level: implementation
---
## Description
Implement the aligned static `OffsetMap` in `crates/alknet-typedef/src/offset_map.rs`.
This is Mode 2 of the two layout modes (ADR-096): fields have fixed positions with
natural alignment padding. Variable-length fields get a 4-byte length prefix at a
known offset; the variable data is not included in the static layout.
Per [layout-engine.md](../../docs/architecture/crates/typedef/layout-engine.md) §"Mode 2: Aligned static".
### Target shape
```rust
/// A byte range within a buffer.
#[derive(Debug, Clone, Copy)]
pub struct ByteRange {
pub start: usize,
pub end: usize,
}
/// A flat table of (field_path, byte_range) pairs computed from a schema.
/// Fields have fixed positions with natural alignment padding.
/// Used for mmap-friendly formats (metatensor, safetensors).
#[derive(Debug)]
pub struct OffsetMap {
fields: Vec<(String, ByteRange)>,
total_size: usize,
}
impl OffsetMap {
/// Compute the offset map from a schema JSON value.
/// Walks the schema recursively, computing byte positions for each field
/// based on type sizes, field order, and alignment.
pub fn compute(schema: &serde_json::Value) -> Result<Self, TypedefError>;
/// Look up a field's byte range by dotted path (e.g., "header.version").
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
/// The total size of the struct in bytes (including alignment padding).
pub fn total_size(&self) -> usize;
/// Iterate over all (field_path, byte_range) pairs.
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
}
```
### Offset computation algorithm
The algorithm walks the schema recursively:
1. **Fixed-size types**: Determine the type's byte size from the `TypeDef:*` kind.
Insert alignment padding to satisfy the type's alignment (or the field's `align`
annotation, or the struct's `align` default). Record the field's `(start, end)`
range. Advance the current offset by the type's size.
2. **`TStruct`**: Recurse into the struct's `properties`. Inner fields are computed
relative to the struct's start offset. The struct's total size is the sum of its
fields' sizes plus alignment padding. The struct itself may have an `align`
annotation that rounds up its total size.
3. **`TUnion`**: The discriminator occupies `offset..offset + discriminator_size`
bytes. For byte-offset discriminators, the variant struct starts at
`offset + discriminator_size`. For field-name discriminators, the discriminator
is just another field. The union's total size is `discriminator_size +
max(variant_sizes)`.
4. **`TArray` of fixed-size elements**: Element stride = element size plus alignment
padding. Element `i` starts at `array_offset + i × stride`. The array's total
size is `count × stride`. Count is determined from `minItems`/`maxItems` (when
equal, fixed count; otherwise variable — uses length-prefixed encoding).
5. **Variable-length types (inline length-prefixing)**: Record the position of the
4-byte length prefix. The variable data is not included in the static layout.
The length prefix is aligned to 4 bytes.
6. **Variable-length types (fixed-size reservation, `maxLength`)**: Reserve
`maxLength` bytes at a fixed offset. Data shorter than `maxLength` is zero-padded.
Subsequent fields have known, unchanging offsets.
7. **Variable-length types (offset indirection)**: The field is a struct
`{offset: u32, length: u32}` (8 bytes total). Record its position. The consumer
provides the data region separately.
### Nested structs and field paths
Nested structs produce dotted field paths: `"header.version"`, `"header.magic"`.
The offset computation propagates the field path prefix during recursion. The
`OffsetMap` stores fully-qualified paths.
### Alignment rules
- Default alignment: 1 for u8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for
u64/i64/f64, max field alignment for structs.
- Struct-level `"align"` sets the default for all fields in that struct.
- Field-level `"align"` overrides the struct default.
- The struct's total size is rounded up to its alignment.
- Alignment padding is inserted before each field to satisfy its alignment.
### What this does NOT include
- Packed sequential layout (that's `layout_builder.rs` and `sequential_reader.rs`)
- TUnion discriminator dispatch (that's `tunion.rs`)
- The `TypedefEngine` struct (that's `engine.rs`)
- Arrays of variable-length-element structs (deferred, OQ-069)
## Acceptance Criteria
- [ ] `OffsetMap` struct with `fields: Vec<(String, ByteRange)>` and `total_size: usize`
- [ ] `OffsetMap::compute(schema)` walks the schema and computes byte positions
- [ ] Fixed-size types get correct byte ranges with natural alignment padding
- [ ] `u8` at offset 0, `u32` at offset 4 (3 bytes padding) — natural alignment
- [ ] Nested structs produce dotted field paths (`"header.version"`)
- [ ] `TArray` of fixed-size elements: correct stride and element offsets
- [ ] `TArray` with `minItems == maxItems`: fixed count, known at schema time
- [ ] `TArray` with variable count: length-prefixed encoding (4-byte count prefix)
- [ ] Variable-length types (inline length-prefixing): 4-byte length prefix at known offset
- [ ] Variable-length types (`maxLength`): reserved `maxLength` bytes at fixed offset
- [ ] Variable-length types (offset-indirect): 8-byte `{offset, length}` struct at known offset
- [ ] `TUnion` with byte-offset discriminator: discriminator at `offset`, variant at `offset + disc_size`
- [ ] `TUnion` with field-name discriminator: discriminator is a regular field
- [ ] Struct-level `"align"` annotation: rounds up struct total size
- [ ] Field-level `"align"` annotation: overrides struct default for that field
- [ ] `OffsetMap::get("header.version")` returns the correct `ByteRange`
- [ ] `OffsetMap::total_size()` returns the correct total size
- [ ] `OffsetMap::iter()` iterates all field paths
- [ ] Returns `TypedefError::Schema` for malformed schemas
- [ ] Returns `TypedefError::Offset` for unsupported type combinations
- [ ] No `unwrap()` or `expect()` on error paths
- [ ] All public types and functions have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/layout-engine.md — Mode 2: Aligned static, offset computation algorithm
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds and their byte sizes
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (alignment, encoding)
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference for offset computation
## Notes
> This is the aligned static layout mode — the simpler of the two modes. Fields have
> fixed positions; the consumer can read field N without reading fields 0..N-1 first.
> Used by metatensor and safetensors. The offset computation is a recursive walk of
> the schema JSON. Nested structs propagate field path prefixes. Alignment padding
> is inserted between fields based on type sizes and annotations. The POC's
> `offset.rs` is a good reference — the algorithm is correct and can be adapted.
## Summary
> To be filled on completion
+162
View File
@@ -0,0 +1,162 @@
---
id: typedef/review-typedef
name: Review alknet-typedef implementation for spec conformance, API shape, and test coverage
status: pending
depends_on: [typedef/tests]
scope: moderate
risk: low
impact: phase
level: review
---
## Description
Review checkpoint for the `alknet-typedef` crate. Verify the implementation is
spec-conformant, self-contained, and ready for downstream consumption by metatensor,
SFTP, binary call frames, and TTY negotiation.
### Review Checklist
#### 1. Crate structure
- Module layout matches spec: `error.rs`, `schema.rs`, `data_access.rs`, `offset_map.rs`,
`layout_builder.rs`, `sequential_reader.rs`, `tunion.rs`, `validation.rs`, `engine.rs`
- Public API types: `TypedefEngine`, `TypedefError`, `OffsetMap`, `LayoutBuilder`,
`SequentialReader`, `LayoutMode`, `Endian`, `ByteRange`, `FieldValue`, `UnionDispatch`
- Re-exports in `lib.rs` are correct and minimal
- No tokio dependency (WASM-clean by construction)
- `serde_json` has `preserve_order` feature enabled
#### 2. Schema layer (ADR-097)
- All 17 `TypeDef:*` kinds correctly identified by `get_typedef_kind()`
- `type_size()` returns correct sizes for all fixed-size types
- `Endian` enum with `Little` (default) and `Big`
- `parse_encoding()` handles both `true` (shorthand) and `{ "encoding": "..." }` (object)
- `parse_discriminator()` handles byte-offset and field-name discriminators
- `normalize_refs()` rewrites bare-name refs to full JSON Pointer paths
- `TEnum` uses `u32` index (not variable-length string) — deliberate deviation from TypeBox
#### 3. Data access layer
- All fixed-size read/write functions implemented with endianness support
- `read_bool`: `0x00` = false, `0x01` = true, other values → error
- `read_enum`: reads `u32` index with endianness
- Variable-length types: inline length-prefixing (default) and offset indirection (opt-in)
- Zero-copy: read functions return slices, not owned data
- All functions perform bounds checking and return `TypedefError::Access` with field path
- No `unwrap()` or `expect()` on error paths — all fallible operations use `Result`
#### 4. Layout engine (ADR-096)
- **Aligned static mode** (`OffsetMap`): fields have fixed positions with natural alignment
padding. Variable-length fields get a 4-byte length prefix at known offset.
- **Packed sequential mode** (`LayoutBuilder` + `SequentialReader`): fields packed with
no alignment padding. Variable-length fields shift subsequent fields.
- Nested structs produce dotted field paths (`"header.version"`)
- `TArray` with fixed count (`minItems == maxItems`) and variable count (length-prefixed)
- `TUnion` with byte-offset and field-name discriminators
- Alignment annotations: struct-level and field-level, field-level overrides
- `maxLength` annotation: fixed-size reservation in aligned mode, validation constraint in packed mode
- Endianness: per-schema, default little-endian, applied at access time
#### 5. TUnion dispatch (ADR-097 §4)
- Byte-offset discriminator: reads fixed-size integer at known offset, returns mapping key
- Field-name discriminator: reads named field, returns mapping key
- `resolve_variant()` resolves `$ref` pointers to `$defs`
- Supports `TypeDef:Uint8`, `TypeDef:Uint16`, `TypeDef:Uint32` discriminator types
#### 6. Validation (ADR-098)
- `build_validator()` registers all 17 custom keywords with `jsonschema`
- Each validator is ~10 lines (not hundreds)
- `StructValidator` inspects parent's `properties` for cross-keyword awareness
- `EnumValidator` is a no-op (built-in `enum` keyword handles validation)
- `TypedefError::Validation` wraps `jsonschema::ValidationError<'static>`
#### 7. TypedefEngine
- `compile(&mut schema, mode)` performs all load-time work: normalize refs, build layout, build validator
- `Layout` enum with `Packed { builder, reader }` and `Aligned { offset_map }` variants
- `LayoutMode` enum with `Packed` and `Aligned` variants
- Accessor methods: `endian()`, `mode()`, `offset_map()`, `layout_builder()`, `sequential_reader()`
- `validate_json()` and `is_valid_json()` delegate to compiled validator
- `read_field()` and `write_field()` convenience methods for aligned mode
#### 8. Error handling (ADR-098)
- `TypedefError` enum with four variants: `Schema`, `Offset`, `Access`, `Validation`
- `Offset` and `Access` variants carry field paths for debugging
- `Display` and `Error` trait implementations
- No `unwrap()` or `expect()` on error paths anywhere in the crate
#### 9. Test coverage
- Schema layer tests: all public functions tested
- Data access tests: all read/write functions with round-trip, endianness, error paths
- OffsetMap tests: aligned static layout with alignment, nesting, arrays, unions
- LayoutBuilder tests: packed sequential layout with variable-length shifting
- SequentialReader tests: sequential field reading with position tracking
- TUnion tests: both discriminator kinds, variant resolution
- Validation tests: all 17 custom keyword validators
- Engine tests: compile, accessors, convenience methods
- POC round-trip tests: fixed-size, string, nested struct, endianness
- Error path tests: buffer-too-short, invalid UTF-8, malformed schemas
#### 10. Cross-cutting checks
- `cargo build -p alknet-typedef` succeeds
- `cargo test -p alknet-typedef` succeeds (all tests pass)
- `cargo clippy -p alknet-typedef --all-targets` succeeds with no warnings
- `cargo fmt --check -p alknet-typedef` passes
- `cargo build --workspace` still succeeds (old code untouched)
- `cargo test --workspace` still succeeds (old tests untouched)
- No `unwrap()` or `expect()` in production code (spec pseudocode uses `unwrap` for brevity only)
- `TEnum` uses `u32` index (not variable-length string) — deliberate deviation from TypeBox
## Acceptance Criteria
- [ ] Crate structure matches spec (9 source files, correct module layout)
- [ ] All 17 `TypeDef:*` kinds correctly identified and sized
- [ ] Both layout modes work correctly (aligned static and packed sequential)
- [ ] TUnion dispatch supports both byte-offset and field-name discriminators
- [ ] All 17 custom keyword validators registered and working
- [ ] `TypedefEngine::compile()` correctly wires all components
- [ ] `TypedefError` has correct variants with field-path-carrying errors
- [ ] No `unwrap()` or `expect()` in production code
- [ ] All tests pass (unit + integration)
- [ ] `cargo build -p alknet-typedef` succeeds
- [ ] `cargo test -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef --all-targets` succeeds with no warnings
- [ ] `cargo fmt --check -p alknet-typedef` passes
- [ ] Workspace still green: `cargo build --workspace` + `cargo test --workspace` pass
## References
- docs/architecture/crates/typedef/README.md — crate overview and design principles
- docs/architecture/crates/typedef/overview.md — purpose, dependencies, scope boundaries
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds, annotations
- docs/architecture/crates/typedef/layout-engine.md — the two layout modes
- docs/architecture/crates/typedef/data-access.md — read/write functions
- docs/architecture/crates/typedef/validation.md — custom keyword validators
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
- docs/architecture/decisions/097-schema-annotations.md — ADR-097
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
- docs/research/alknet-typedef/findings.md — POC results
- All task files in tasks/typedef/
## Notes
> This review gates the alknet-typedef implementation. The crate must be self-contained
> and spec-conformant before downstream consumers (metatensor, SFTP, binary call frames,
> TTY negotiation) can depend on it. The POC validated the approach with 26 passing
> tests; the production implementation should match or exceed that coverage. Key things
> to verify: no `unwrap()` in production code (the spec pseudocode uses it for brevity),
> `TEnum` uses `u32` index (not variable-length string), and both layout modes produce
> correct offsets for their respective use cases.
## Summary
> To be filled on completion
+183
View File
@@ -0,0 +1,183 @@
---
id: typedef/schema-types
name: Implement TypeDef kind detection, schema annotation parsing, $ref normalization, and Endian enum
status: pending
depends_on: [typedef/crate-init]
scope: moderate
risk: medium
impact: component
level: implementation
---
## Description
Implement the schema layer in `crates/alknet-typedef/src/schema.rs`. This module
provides the foundational types and functions that every other module depends on:
TypeDef kind detection, type size constants, schema annotation parsing, `$ref`
normalization, and the `Endian` enum.
Per [schema-layer.md](../../docs/architecture/crates/typedef/schema-layer.md) and
[ADR-097](../../docs/architecture/decisions/097-schema-annotations.md).
### TypeDef kind detection
The engine needs to identify which `TypeDef:*` kind a schema node declares. The
17 kinds (16 from TypeBox's `typedef.ts` + `TypeDef:Bytes` added by alknet-typedef):
| Kind | Keyword | Category | Byte size |
|------|---------|----------|-----------|
| Float32 | `TypeDef:Float32` | fixed | 4 |
| Float64 | `TypeDef:Float64` | fixed | 8 |
| Int8 | `TypeDef:Int8` | fixed | 1 |
| Int16 | `TypeDef:Int16` | fixed | 2 |
| Int32 | `TypeDef:Int32` | fixed | 4 |
| Uint8 | `TypeDef:Uint8` | fixed | 1 |
| Uint16 | `TypeDef:Uint16` | fixed | 2 |
| Uint32 | `TypeDef:Uint32` | fixed | 4 |
| Boolean | `TypeDef:Boolean` | fixed | 1 |
| Enum | `TypeDef:Enum` | fixed | 4 (u32 index) |
| String | `TypeDef:String` | variable | — |
| Bytes | `TypeDef:Bytes` | variable | — |
| Struct | `TypeDef:Struct` | composite | sum of fields |
| Union | `TypeDef:Union` | composite | discriminator + variant |
| Array | `TypeDef:Array` | composite | count × element |
| Record | `TypeDef:Record` | variable | — |
| Timestamp | `TypeDef:Timestamp` | variable | — |
Implement:
```rust
/// Returns the `TypeDef:*` kind string if the schema node declares one.
/// Returns `None` if the node has no `TypeDef:*` keyword.
pub fn get_typedef_kind(node: &serde_json::Value) -> Option<&str>;
/// Returns the fixed byte size for a TypeDef kind, or `None` if variable-size.
pub fn type_size(kind: &str) -> Option<usize>;
/// Returns the natural alignment for a TypeDef kind.
pub fn natural_alignment(kind: &str) -> usize;
/// Returns true if the kind is a fixed-size type.
pub fn is_fixed_size(kind: &str) -> bool;
```
### Endian enum
```rust
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Endian {
Little,
Big,
}
impl Endian {
/// Parse from the schema's "endian" annotation. Defaults to Little.
pub fn from_schema(schema: &serde_json::Value) -> Self;
}
```
### Schema annotation parsing
Parse the annotations decided in ADR-097:
```rust
/// Parse the "encoding" annotation from a variable-length type's keyword value.
/// The keyword value may be `true` (shorthand for length-prefixed) or an object
/// with an "encoding" field.
pub fn parse_encoding(keyword_value: &serde_json::Value) -> VariableEncoding;
pub enum VariableEncoding {
LengthPrefixed,
OffsetIndirect,
}
/// Parse the "align" annotation from a schema node. Returns None if not specified.
pub fn parse_align(node: &serde_json::Value) -> Option<usize>;
/// Parse the "maxLength" annotation (standard JSON Schema keyword).
pub fn parse_max_length(node: &serde_json::Value) -> Option<usize>;
/// Parse the "endian" annotation. Defaults to Little if absent or unrecognized.
pub fn parse_endian(node: &serde_json::Value) -> Endian;
```
### TUnion discriminator parsing
Parse the discriminator shape from a `TypeDef:Union` schema node:
```rust
pub enum DiscriminatorKind {
Byte { offset: usize, disc_type: String },
Field { name: String },
}
/// Parse the "discriminator" annotation from a TUnion schema node.
pub fn parse_discriminator(node: &serde_json::Value) -> Result<DiscriminatorKind, TypedefError>;
```
### `$ref` normalization
TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`). The `jsonschema`
crate requires full JSON Pointer paths (e.g., `"$ref": "#/$defs/Read"`). Normalize
at schema load time:
```rust
/// Walk the schema tree. For every "$ref" whose value is a bare name
/// (no "#" prefix), rewrite it to "#/$defs/<name>".
/// Full JSON Pointer refs pass through unchanged. Idempotent.
pub fn normalize_refs(schema: &mut serde_json::Value);
```
This is a ~20-line recursive walk. It runs once at load time, before the schema is
passed to `jsonschema::validator_for` or the offset computation.
### What this does NOT include
- The actual offset computation (that's `offset_map.rs` and `layout_builder.rs`)
- The read/write functions (that's `data_access.rs`)
- The custom keyword validators (that's `validation.rs`)
- The `TypedefEngine` struct (that's `engine.rs`)
## Acceptance Criteria
- [ ] `get_typedef_kind()` correctly identifies all 17 `TypeDef:*` kinds
- [ ] `type_size()` returns correct byte sizes for all fixed-size types
- [ ] `type_size()` returns `None` for variable-length and composite types
- [ ] `natural_alignment()` returns correct alignment for each kind
- [ ] `Endian` enum with `Little` and `Big` variants
- [ ] `Endian::from_schema()` defaults to `Little` when annotation is absent
- [ ] `Endian::from_schema()` returns `Big` when `"endian": "big"`
- [ ] `parse_encoding()` handles both `true` (shorthand) and `{ "encoding": "..." }` (object)
- [ ] `parse_encoding()` defaults to `LengthPrefixed` when annotation is absent
- [ ] `parse_align()` returns `None` when no `"align"` annotation
- [ ] `parse_max_length()` returns `None` when no `"maxLength"` annotation
- [ ] `parse_discriminator()` handles byte-offset (`kind: "byte"`) and field-name (`kind: "field"`)
- [ ] `parse_discriminator()` returns `Err(TypedefError::Schema(...))` for malformed discriminators
- [ ] `normalize_refs()` rewrites `"$ref": "Read"` → `"$ref": "#/$defs/Read"`
- [ ] `normalize_refs()` leaves `"$ref": "#/$defs/Read"` unchanged (idempotent)
- [ ] `normalize_refs()` handles nested objects and arrays recursively
- [ ] All public functions have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds, annotations, $ref normalization
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (concrete JSON shapes)
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
- docs/research/alknet-typedef/findings.md — POC results, open questions
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference for kind detection
## Notes
> This is the foundational types module. Every other module depends on it for
> TypeDef kind identification, size lookups, and annotation parsing. The `$ref`
> normalization bridges TypeBox output to jsonschema input — without it, bare-name
> refs from TypeBox would fail to resolve. The `TEnum` kind uses a `u32` index
> (not a variable-length string) — a deliberate deviation from TypeBox fidelity
> for binary efficiency, per the schema-layer spec.
## Summary
> To be filled on completion
+181
View File
@@ -0,0 +1,181 @@
---
id: typedef/sequential-reader
name: Implement packed sequential SequentialReader for protocol read-side
status: pending
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
scope: moderate
risk: medium
impact: component
level: implementation
---
## Description
Implement the packed sequential `SequentialReader` in `crates/alknet-typedef/src/sequential_reader.rs`.
This is the read-side of Mode 1 (ADR-096): walks a buffer field-by-field according to
the schema, reading length prefixes to determine variable-length data positions. Used
at read time when the consumer is parsing an incoming frame.
Per [layout-engine.md](../../docs/architecture/crates/typedef/layout-engine.md) §"Mode 1: Packed sequential".
### Target shape
```rust
/// Walks a buffer field-by-field according to a schema, reading length
/// prefixes to determine variable-length data positions. Used at read time
/// when parsing incoming protocol frames.
///
/// The reader is sequential — it cannot jump to field N without reading
/// fields 0..N-1 first. This is inherent to packed layouts where
/// variable-length fields shift subsequent fields.
#[derive(Debug)]
pub struct SequentialReader {
/// The schema being read.
schema: serde_json::Value,
/// The endianness for the layout.
endian: Endian,
}
/// A value read from a field during sequential traversal.
#[derive(Debug)]
pub enum FieldValue<'a> {
I8(i8),
I16(i16),
I32(i32),
U8(u8),
U16(u16),
U32(u32),
U64(u64),
F32(f32),
F64(f64),
Bool(bool),
Enum(u32),
String(&'a str),
Bytes(&'a [u8]),
/// A nested struct — the consumer recurses with a new SequentialReader
/// scoped to the struct's byte range.
Struct { start: usize, end: usize },
/// A union — the consumer reads the discriminator, then recurses
/// with the variant schema.
Union { discriminator: String, variant_start: usize },
/// An array — the consumer iterates elements.
Array { count: u32, element_start: usize, element_stride: usize },
}
impl SequentialReader {
/// Create a new SequentialReader from a schema.
pub fn new(schema: &serde_json::Value) -> Result<Self, TypedefError>;
/// Read the next field from the buffer at the current position.
/// Returns the field name, the value, and advances the internal position.
/// Returns `None` when all fields have been read.
pub fn read_next<'a>(
&mut self,
buffer: &'a [u8],
) -> Result<Option<(String, FieldValue<'a>)>, TypedefError>;
/// Read a specific field by path. This walks through all preceding fields
/// to reach the target (sequential access is inherent to packed layouts).
pub fn read_field<'a>(
&mut self,
buffer: &'a [u8],
field_path: &str,
) -> Result<FieldValue<'a>, TypedefError>;
/// Reset the reader to the beginning of the buffer.
pub fn reset(&mut self);
/// The current byte position in the buffer.
pub fn position(&self) -> usize;
}
```
### How it works
For a struct with fields `[u8, u32, string]`:
```
SequentialReader:
read_next() → ("field_0", FieldValue::U8(42)), position = 1
read_next() → ("field_1", FieldValue::U32(1234)), position = 5
read_next() → reads u32 length prefix at offset 5 → data_len = 10
("field_2", FieldValue::String("hello worl")), position = 19
read_next() → None (no more fields)
```
The reader walks the buffer sequentially. It reads each field's type from the schema,
reads the appropriate number of bytes at the current position, and advances. For
variable-length fields, it reads the 4-byte length prefix to determine the data
extent, then advances past the data.
### Nested structs
When the reader encounters a `TStruct` field, it returns `FieldValue::Struct { start, end }`.
The consumer creates a new `SequentialReader` scoped to that byte range and reads the
inner fields.
### TUnion
When the reader encounters a `TUnion` field:
1. For byte-offset discriminators: reads the discriminator value at the known offset,
looks up the variant schema, returns `FieldValue::Union { discriminator, variant_start }`.
2. For field-name discriminators: reads the discriminator field like any other field,
then returns the union value.
### TArray
When the reader encounters a `TArray` field:
1. Reads the count (from `minItems`/`maxItems` if fixed, or from a 4-byte count prefix
if variable).
2. Returns `FieldValue::Array { count, element_start, element_stride }`.
3. The consumer iterates elements using the stride.
### What this does NOT include
- The write-side of packed mode (that's `layout_builder.rs`)
- Aligned static layout (that's `offset_map.rs`)
- TUnion discriminator dispatch (that's `tunion.rs`)
- The `TypedefEngine` struct (that's `engine.rs`)
## Acceptance Criteria
- [ ] `SequentialReader` struct with `new(schema)` constructor
- [ ] `read_next(buffer)` reads the next field and advances position
- [ ] Fixed-size fields read correct number of bytes at current position
- [ ] Variable-length fields: reads 4-byte length prefix, then skips data
- [ ] Position advances correctly through all fields
- [ ] `read_next()` returns `None` when all fields have been read
- [ ] `read_field(field_path)` walks through preceding fields to reach target
- [ ] `reset()` resets position to beginning
- [ ] `position()` returns current byte offset
- [ ] Nested structs: returns `FieldValue::Struct { start, end }` for consumer recursion
- [ ] `TUnion`: reads discriminator, returns `FieldValue::Union { ... }`
- [ ] `TArray`: reads count, returns `FieldValue::Array { count, element_start, element_stride }`
- [ ] Endianness respected for all multi-byte reads
- [ ] Returns `TypedefError::Access` for buffer-too-short
- [ ] Returns `TypedefError::Schema` for malformed schemas
- [ ] No `unwrap()` or `expect()` on error paths
- [ ] All public types and functions have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/layout-engine.md — Mode 1: Packed sequential, SequentialReader
- docs/architecture/crates/typedef/data-access.md — read/write functions used by the reader
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (encoding, discriminators)
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference (SequentialReader in POC 2)
## Notes
> This is the read-side of the packed sequential mode. The reader walks the buffer
> sequentially — it cannot jump to field N without reading fields 0..N-1 first.
> This is inherent to packed layouts where variable-length fields shift subsequent
> fields. The POC 2's `SequentialReader` is a good reference. The reader uses the
> `data_access` module's read functions internally.
## Summary
> To be filled on completion
+192
View File
@@ -0,0 +1,192 @@
---
id: typedef/tests
name: "Write comprehensive tests for alknet-typedef: unit tests, integration tests, and POC round-trip tests"
status: pending
depends_on: [typedef/engine]
scope: moderate
risk: low
impact: component
level: implementation
---
## Description
Write comprehensive tests for the `alknet-typedef` crate. Tests should cover all
17 `TypeDef:*` kinds, both layout modes, endianness, TUnion dispatch, validation,
and error paths. Include the POC-verified round-trip tests that validate
byte-identical output against known-good serialization.
### Test categories
#### 1. Schema layer tests (`tests/schema_tests.rs` or `#[cfg(test)] mod tests` in `schema.rs`)
- `get_typedef_kind()` correctly identifies all 17 kinds
- `type_size()` returns correct sizes for all fixed-size types
- `type_size()` returns `None` for variable-length types
- `natural_alignment()` returns correct alignment
- `Endian::from_schema()` defaults to Little, parses "big" correctly
- `parse_encoding()` handles `true`, `{ "encoding": "length-prefixed" }`, `{ "encoding": "offset-indirect" }`
- `parse_align()` returns `None` when absent, correct value when present
- `parse_max_length()` returns `None` when absent, correct value when present
- `parse_discriminator()` handles byte-offset and field-name discriminators
- `normalize_refs()` rewrites bare-name refs, leaves full paths unchanged, handles nested objects
#### 2. Data access tests (`tests/data_access_tests.rs`)
- Fixed-size read/write round-trip for all types: `i8`, `i16`, `i32`, `u8`, `u16`, `u32`, `u64`, `f32`, `f64`, `bool`, `enum`
- Endianness: little-endian and big-endian produce correct byte order
- `read_bool`: `0x00` → false, `0x01` → true, other values → error
- `read_string`: correct length-prefixed read, UTF-8 validation, buffer-too-short error
- `read_bytes`: correct length-prefixed read, buffer-too-short error
- `write_string` / `write_bytes`: correct length prefix + data written, returns bytes written
- `read_string_indirect` / `read_bytes_indirect`: correct offset-indirect read
- Buffer bounds checking: all functions return `TypedefError::Access` for too-short buffers
- Field path in error messages is correct
#### 3. OffsetMap tests (aligned static mode) (`tests/offset_map_tests.rs`)
- Simple struct: `{ u8, u32, f32 }` → correct offsets with natural alignment (0, 4, 8)
- Nested struct: `{ header: { version: u32, magic: u32 }, payload: bytes }` → correct dotted paths
- `TArray` of fixed-size elements: correct stride and element offsets
- `TArray` with fixed count (`minItems == maxItems`): correct total size
- `TArray` with variable count: length-prefixed encoding
- Variable-length types (inline length-prefixing): 4-byte length prefix at known offset
- Variable-length types (`maxLength`): reserved bytes at fixed offset
- Variable-length types (offset-indirect): 8-byte `{offset, length}` struct
- `TUnion` with byte-offset discriminator: discriminator at offset, variant at offset + disc_size
- `TUnion` with field-name discriminator: discriminator is a regular field
- Struct-level `"align"` annotation: rounds up total size
- Field-level `"align"` annotation: overrides struct default
- `OffsetMap::get()` returns correct `ByteRange` for dotted paths
- `OffsetMap::total_size()` is correct
- `OffsetMap::iter()` returns all fields
#### 4. LayoutBuilder tests (packed sequential mode, write-side) (`tests/layout_builder_tests.rs`)
- Simple struct: `{ u8, u32, string }` → packed offsets (0, 1, 5) with no padding
- Variable-length fields shift subsequent fields based on actual sizes
- Nested struct: correct dotted paths with packed offsets
- `TArray` of fixed-size elements: correct stride (element size, no padding)
- `TUnion` with byte-offset discriminator: correct discriminator and variant positions
- Missing variable-length field size → `TypedefError::Offset`
#### 5. SequentialReader tests (packed sequential mode, read-side) (`tests/sequential_reader_tests.rs`)
- Read fields sequentially: correct values and position advancement
- Variable-length fields: reads length prefix, skips data, advances correctly
- `read_next()` returns `None` when all fields read
- `read_field()` walks through preceding fields
- `reset()` resets position
- Nested struct: returns `FieldValue::Struct { start, end }`
- `TUnion`: returns `FieldValue::Union { ... }`
- `TArray`: returns `FieldValue::Array { count, element_start, element_stride }`
- Endianness respected for multi-byte reads
#### 6. TUnion dispatch tests (`tests/tunion_tests.rs`)
- Byte-offset discriminator: reads u8 at offset 0, returns correct key and variant offset
- Byte-offset discriminator with u16 and u32 types
- Field-name discriminator: reads string field, returns correct key
- Field-name discriminator with u8 and enum field types
- `resolve_variant()`: resolves `$ref` pointers, returns error for unknown keys
- `discriminator_size()`: correct for each discriminator type
#### 7. Validation tests (`tests/validation_tests.rs`)
- `build_validator()` returns a working validator
- Each `TypeDef:*` kind's validator rejects invalid values
- `Float32Validator`: rejects NaN, Infinity
- `Int8Validator`: rejects 128, -129
- `Uint8Validator`: rejects -1, 256
- `StringValidator`: rejects non-string, respects `maxLength`
- `TimestampValidator`: rejects non-RFC 3339 strings
- `StructValidator`: validates nested fields
- `UnionValidator`: validates discriminator membership
- `ArrayValidator`: validates element types and length bounds
#### 8. TypedefEngine integration tests (`tests/engine_tests.rs`)
- `compile()` in aligned mode produces a working engine
- `compile()` in packed mode produces a working engine
- `compile()` normalizes `$ref` values
- `compile()` returns error for invalid schemas
- `endian()` and `mode()` accessors
- `offset_map()`, `layout_builder()`, `sequential_reader()` accessors
- `validate_json()` and `is_valid_json()` work correctly
- `read_field()` and `write_field()` in aligned mode
#### 9. POC round-trip tests (`tests/poc_roundtrip_tests.rs`)
Replicate the key POC tests from `/workspace/alknet-typedef-poc/`:
- **Fixed-size round-trip**: Write u8, u16, u32, u64, f32, f64, bool, enum to a buffer
at computed offsets; read back; verify values match.
- **String round-trip**: Write a length-prefixed string; read back; verify.
- **Nested struct round-trip**: Write a struct with nested fields; read back; verify.
- **Endianness round-trip**: Write in big-endian; read back in big-endian; verify.
- **SFTP packet round-trip** (if SFTP schemas are available): Write an SFTP Read/Write/Status
packet; verify byte-identical output against known-good serialization.
#### 10. Error path tests
- Buffer too short for every read function
- Invalid UTF-8 in `read_string`
- Invalid boolean byte value
- Missing required schema keywords
- Unknown `TypeDef:*` kind
- Malformed discriminator annotation
- Unknown discriminator value
- Missing variable-length field size in `LayoutBuilder::build()`
### Test organization
Tests can be organized as:
- Unit tests in `#[cfg(test)] mod tests` blocks within each source file (for
module-internal functions).
- Integration tests in `tests/` directory (for public API tests that exercise
multiple modules).
### What this does NOT include
- Tests for arrays of variable-length-element structs (deferred, OQ-069)
- WASM-specific tests (deferred, OQ-070)
- Performance benchmarks
## Acceptance Criteria
- [ ] Schema layer tests: all `schema.rs` public functions tested
- [ ] Data access tests: all read/write functions tested with round-trip, endianness, and error paths
- [ ] OffsetMap tests: aligned static layout with alignment, nesting, arrays, unions, variable-length types
- [ ] LayoutBuilder tests: packed sequential layout with variable-length shifting
- [ ] SequentialReader tests: sequential field reading with position tracking
- [ ] TUnion tests: both discriminator kinds, variant resolution
- [ ] Validation tests: all 17 custom keyword validators tested
- [ ] Engine tests: compile, accessors, convenience methods
- [ ] POC round-trip tests: fixed-size, string, nested struct, endianness
- [ ] Error path tests: buffer-too-short, invalid UTF-8, malformed schemas
- [ ] All tests pass: `cargo test -p alknet-typedef`
- [ ] No `unwrap()` in test code that would mask failures (use `?` or `assert!(matches!(...))`)
- [ ] `cargo clippy -p alknet-typedef --all-targets` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
- [ ] `cargo test --workspace` still succeeds (old tests untouched)
## References
- docs/research/alknet-typedef/findings.md — POC results (26 tests passing)
- /workspace/alknet-typedef-poc/src/lib.rs — POC test reference
- docs/architecture/crates/typedef/README.md — design principles
- All preceding task files in tasks/typedef/
## Notes
> This is the testing task — it validates the entire crate. The POC had 26 tests
> passing; this task should cover at least that many scenarios plus additional
> error-path tests. Tests should be organized by module/concern. The SFTP round-trip
> test is a stretch goal — it requires SFTP packet schemas which may not be in the
> repo yet. If SFTP schemas aren't available, test with hand-crafted schemas that
> exercise the same patterns (byte-offset TUnion, big-endian, mixed fixed/variable
> fields).
## Summary
> To be filled on completion
+145
View File
@@ -0,0 +1,145 @@
---
id: typedef/tunion
name: Implement TUnion discriminator dispatch for byte-offset and field-name discriminators
status: pending
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
scope: narrow
risk: medium
impact: component
level: implementation
---
## Description
Implement TUnion discriminator dispatch in `crates/alknet-typedef/src/tunion.rs`.
TUnion supports two discriminator kinds (ADR-097 §4): byte-offset (protocol dispatch,
e.g., SFTP type bytes) and field-name (typedef.ts string pattern).
Per [data-access.md](../../docs/architecture/crates/typedef/data-access.md) §"TUnion Dispatch"
and [schema-layer.md](../../docs/architecture/crates/typedef/schema-layer.md) §"TUnion discriminators".
### Target shape
```rust
/// The result of reading a TUnion discriminator.
#[derive(Debug, Clone)]
pub struct UnionDispatch {
/// The mapping key (stringified discriminator value).
pub key: String,
/// The byte offset where the variant struct starts.
pub variant_offset: usize,
/// The size of the discriminator in bytes.
pub discriminator_size: usize,
}
/// Read the discriminator value from a byte-offset TUnion.
/// The discriminator is a fixed-size integer at a known byte offset.
/// Returns the mapping key (as a string) and the variant struct offset.
///
/// This is the SFTP `Packet` enum pattern — byte 0 is the type byte,
/// bytes 1..N are the variant struct. The call protocol's event type
/// dispatch uses the same pattern.
pub fn read_byte_discriminator(
buffer: &[u8],
union_schema: &serde_json::Value,
endian: Endian,
) -> Result<UnionDispatch, TypedefError>;
/// Read the discriminator value from a field-name TUnion.
/// The discriminator is a named field within the struct — its offset
/// is computed like any other field. The consumer provides the
/// discriminator field's offset (from the OffsetMap or LayoutBuilder).
///
/// This is the typedef.ts `TUnion` pattern — the discriminator is a
/// field like any other, and the mapping keys are string values.
pub fn read_field_discriminator(
buffer: &[u8],
union_schema: &serde_json::Value,
disc_field_offset: usize,
endian: Endian,
) -> Result<UnionDispatch, TypedefError>;
/// Look up a variant schema from the union's mapping.
/// Returns the variant schema (resolving `$ref` if needed).
pub fn resolve_variant<'a>(
union_schema: &'a serde_json::Value,
key: &str,
) -> Result<&'a serde_json::Value, TypedefError>;
/// Get the discriminator size in bytes for a byte-offset discriminator.
pub fn discriminator_size(union_schema: &serde_json::Value) -> Result<usize, TypedefError>;
```
### Byte-offset discriminator
The discriminator is a fixed-size integer at a known byte offset. The mapping keys
are stringified integers (`"5"`, `"6"`, `"101"`).
Supported discriminator types: `TypeDef:Uint8` (1 byte), `TypeDef:Uint16` (2 bytes),
`TypeDef:Uint32` (4 bytes). The discriminator value is read using the appropriate
endian-aware read function from `data_access`.
The variant struct starts at `offset + discriminator_size`.
### Field-name discriminator
The discriminator is a named field within the struct. Its offset is computed like any
other field (by the `OffsetMap` or `LayoutBuilder`). The mapping keys are string
values matching the discriminator field's value.
The discriminator field's `TypeDef:*` kind determines how to read it:
- `TypeDef:String` → read a length-prefixed string
- `TypeDef:Uint8` → read a u8, stringify
- `TypeDef:Enum` → read a u32 index, map to string value from `"enum"` array
### Variant resolution
`resolve_variant()` looks up the mapping key in the union's `"mapping"` object.
Mapping values may be inline schemas or `$ref` pointers. `$ref` pointers are resolved
against the schema's `$defs` (the `$ref` normalization in `schema.rs` ensures they
are full JSON Pointer paths).
### What this does NOT include
- The offset computation for TUnion (that's in `offset_map.rs` and `layout_builder.rs`)
- The `TypedefEngine` struct (that's `engine.rs`)
## Acceptance Criteria
- [ ] `read_byte_discriminator()` reads discriminator at known byte offset
- [ ] Supports `TypeDef:Uint8`, `TypeDef:Uint16`, `TypeDef:Uint32` discriminator types
- [ ] Discriminator value read with correct endianness
- [ ] Returns mapping key as string (e.g., `"5"` for SFTP Read packet)
- [ ] Returns correct `variant_offset` (offset + discriminator_size)
- [ ] `read_field_discriminator()` reads discriminator from named field at given offset
- [ ] Supports `TypeDef:String`, `TypeDef:Uint8`, `TypeDef:Enum` discriminator field types
- [ ] `resolve_variant()` looks up mapping key and returns variant schema
- [ ] `resolve_variant()` resolves `$ref` pointers to `$defs`
- [ ] `resolve_variant()` returns `TypedefError::Schema` for unknown mapping keys
- [ ] `discriminator_size()` returns correct size for each discriminator type
- [ ] Returns `TypedefError::Access` for buffer-too-short
- [ ] Returns `TypedefError::Schema` for malformed discriminator annotations
- [ ] No `unwrap()` or `expect()` on error paths
- [ ] All public functions have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/data-access.md — TUnion dispatch section
- docs/architecture/crates/typedef/schema-layer.md — TUnion discriminators section
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 §4 (TUnion discriminators)
- /workspace/alknet-typedef-poc/src/lib.rs — POC reference (parse_union_discriminator, read_union_discriminator)
## Notes
> This is a focused module for TUnion discriminator dispatch. The two discriminator
> kinds cover both protocol dispatch (SFTP type bytes, call protocol event types) and
> the typedef.ts string pattern. The mapping keys are always strings — integer
> discriminator values are stringified. The POC 2's `parse_union_discriminator` and
> `read_union_discriminator` functions are good references.
## Summary
> To be filled on completion
+180
View File
@@ -0,0 +1,180 @@
---
id: typedef/validation
name: Implement custom keyword validators for all 17 TypeDef kinds via jsonschema with_keyword API
status: pending
depends_on: [typedef/schema-types, typedef/error-type]
scope: moderate
risk: medium
impact: component
level: implementation
---
## Description
Implement custom keyword validators for all 17 `TypeDef:*` kinds in
`crates/alknet-typedef/src/validation.rs`. Each kind gets a `Keyword` implementation
registered via `jsonschema::options().with_keyword(...)`. The validators check leaf
type constraints; `jsonschema` handles all structural validation.
Per [validation.md](../../docs/architecture/crates/typedef/validation.md) and
[ADR-098](../../docs/architecture/decisions/098-error-handling-validation-strategy.md).
### Target shape
```rust
use jsonschema::{Keyword, ValidationError};
use serde_json::Value;
/// Build a jsonschema validator with all 17 TypeDef:* custom keywords registered.
/// The returned validator can validate JSON representations of data against
/// the schema's type constraints.
pub fn build_validator(schema: &Value) -> Result<jsonschema::Validator, TypedefError>;
```
### Custom keyword validators
Each `TypeDef:*` kind gets a struct implementing `Keyword`:
**Numeric validators:**
- `Float32Validator` / `Float64Validator` — value must be a finite number.
For `Float32`: value must be representable as `f32` (no precision loss beyond
`f32`'s mantissa).
- `Int8Validator` / `Int16Validator` / `Int32Validator` — value must be an integer
within the type's range. Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
- `Uint8Validator` / `Uint16Validator` / `Uint32Validator` — value must be a
non-negative integer within the type's range. Uint8: 0..255, Uint16: 0..65535,
Uint32: 0..4294967295.
**String and binary validators:**
- `StringValidator` — value must be a valid UTF-8 string. If `maxLength` is
specified in the parent schema, the string's byte length must not exceed it.
- `BytesValidator` — value must be a string (JSON represents binary data as a
string). If `maxLength` is specified, the byte length must not exceed it.
- `EnumValidator` — the `TypeDef:Enum` custom keyword signals that the type is an
enum for *layout* purposes. The built-in `enum` keyword handles value-membership
validation. The custom keyword validator is a no-op beyond the built-in check —
it exists solely for the layout engine to recognize the type.
- `TimestampValidator` — value must be a valid RFC 3339 timestamp string (the
internet profile of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
**Composite validators:**
- `StructValidator` — value must be an object. Each property must match its
declared `TypeDef:*` kind. The `jsonschema` crate's built-in `properties` and
`required` keywords handle structural checks — the custom keyword only validates
that each field's value matches its `TypeDef:*` kind.
- `UnionValidator` — the discriminator value must be one of the mapping keys.
The variant struct must match the declared schema for that discriminator value.
- `ArrayValidator` — value must be an array. Each element must match the array's
declared element type. If `minItems`/`maxItems` is specified, the array length
must be within bounds.
**Other validators:**
- `BooleanValidator` — value must be `true` or `false`.
- `RecordValidator` — value must be an object. All values must match the record's
declared value type (specified via the `"values"` property).
### Validator implementation pattern
Each validator is ~10 lines. Example:
```rust
struct Float32Validator;
impl Keyword for Float32Validator {
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
match instance {
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
}
}
fn is_valid(&self, instance: &Value) -> bool {
instance.as_f64().map_or(false, |f| f.is_finite())
}
}
```
### Registration
```rust
pub fn build_validator(schema: &Value) -> Result<jsonschema::Validator, TypedefError> {
jsonschema::options()
.with_keyword("TypeDef:Float32", |_parent, _value, _path| {
Ok(Box::new(Float32Validator))
})
.with_keyword("TypeDef:Float64", |_parent, _value, _path| {
Ok(Box::new(Float64Validator))
})
.with_keyword("TypeDef:Int8", |_parent, _value, _path| {
Ok(Box::new(Int8Validator))
})
// ... all 17 kinds
.with_keyword("TypeDef:Struct", |parent, _value, _path| {
Ok(Box::new(StructValidator::from_schema(parent)?))
})
.build(schema)
.map_err(|e| TypedefError::Schema(format!("validator build failed: {e}")))
}
```
The factory closure receives the parent schema object, the keyword's value, and the
schema path. This enables cross-keyword awareness — for example, `StructValidator`
inspects the parent's `properties` to validate each field against its declared
`TypeDef:*` kind.
### What this does NOT include
- The `TypedefEngine` struct (that's `engine.rs`) — though `build_validator()` is
called by the engine
- Schema parsing or annotation extraction (that's `schema.rs`)
- Binary-level validation (that's in `data_access.rs` — bounds checking, UTF-8
validation at read time)
## Acceptance Criteria
- [ ] `build_validator(schema)` returns a `jsonschema::Validator` with all 17 custom keywords registered
- [ ] `Float32Validator` / `Float64Validator`: rejects non-finite numbers
- [ ] `Int8Validator` / `Int16Validator` / `Int32Validator`: rejects out-of-range values
- [ ] `Uint8Validator` / `Uint16Validator` / `Uint32Validator`: rejects negative and out-of-range values
- [ ] `StringValidator`: rejects non-string values; respects `maxLength` from parent schema
- [ ] `BytesValidator`: rejects non-string values; respects `maxLength`
- [ ] `EnumValidator`: no-op (built-in `enum` keyword handles validation)
- [ ] `TimestampValidator`: rejects non-RFC 3339 strings
- [ ] `StructValidator`: validates each property against its declared `TypeDef:*` kind
- [ ] `UnionValidator`: validates discriminator membership and variant conformance
- [ ] `ArrayValidator`: validates element type conformance and length bounds
- [ ] `BooleanValidator`: rejects non-boolean values
- [ ] `RecordValidator`: validates all values match declared value type
- [ ] Each validator is ~10 lines (not hundreds)
- [ ] Factory closures handle errors gracefully (return `TypedefError::Schema`)
- [ ] No `unwrap()` or `expect()` in factory closures
- [ ] All validators have doc comments
- [ ] `cargo check -p alknet-typedef` succeeds
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
- [ ] `cargo build --workspace` still succeeds
## References
- docs/architecture/crates/typedef/validation.md — custom keyword validators, TypedefError, validation timing
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
- docs/research/alknet-typedef/findings.md — POC validation results
- /workspace/alknet-typedef-poc/src/validate.rs — POC reference for validators
- /workspace/jsonschema/ — the jsonschema crate API
## Notes
> This module registers custom keyword validators with the `jsonschema` crate. Each
> validator is small (~10 lines) because `jsonschema` handles all structural
> validation. The `StructValidator` is the most complex — it needs to inspect the
> parent schema's `properties` to validate each field. The `EnumValidator` is a
> no-op because the built-in `enum` keyword already handles value-membership
> validation; the custom keyword exists solely for the layout engine to recognize
> the type as a fixed-size u32 index. The POC's `validate.rs` is a good reference.
## Summary
> To be filled on completion