feat(typedef): add implementation task decomposition (12 tasks, 7 generations)
Break the alknet-typedef architecture specs into atomic, dependency-ordered implementation tasks covering crate init, error types, schema layer, data access, both layout modes (aligned static + packed sequential), TUnion dispatch, jsonschema custom keyword validators, TypedefEngine integration, comprehensive tests, and a final review checkpoint. Validated: 12 tasks, 0 cycles, 7 parallel generations.
This commit is contained in:
1 parent
dd232c3d47
commit
db100f9849
12 files changed
+1996
No files matched your search
@@ -0,0 +1,131 @@
|
||||
---
|
||||
id: typedef/crate-init
|
||||
name: Initialize alknet-typedef crate with Cargo.toml, dependencies, and module skeleton
|
||||
status: pending
|
||||
depends_on: []
|
||||
scope: moderate
|
||||
risk: low
|
||||
impact: project
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Initialize the `alknet-typedef` crate from scratch. This is a greenfield crate — the
|
||||
binary struct engine that takes a JSON Schema with `TypeDef:*` custom keywords and
|
||||
produces an offset map, read/write functions, and validation. The schema is the format
|
||||
definition; the engine is generic.
|
||||
|
||||
Per [ADR-095](../../docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md),
|
||||
the crate depends on `jsonschema` (v0.46.5, Draft 2020-12) for validation and
|
||||
`serde_json` (with `preserve_order`) for schema parsing. No tokio, no platform deps.
|
||||
Compiles to `wasm32-unknown-unknown`.
|
||||
|
||||
### Crate setup
|
||||
|
||||
Create `crates/alknet-typedef/` with:
|
||||
|
||||
- `Cargo.toml` — package metadata, dependencies
|
||||
- `src/lib.rs` — crate root with module declarations and re-exports
|
||||
- Module skeleton files for:
|
||||
- `src/error.rs` — `TypedefError` enum (ADR-098)
|
||||
- `src/schema.rs` — TypeDef kind detection, annotation parsing, `$ref` normalization, `Endian` enum (ADR-097)
|
||||
- `src/data_access.rs` — primitive read/write functions for all 17 TypeDef kinds
|
||||
- `src/offset_map.rs` — aligned static `OffsetMap` computation (ADR-096 Mode 2)
|
||||
- `src/layout_builder.rs` — packed sequential `LayoutBuilder` (ADR-096 Mode 1, write side)
|
||||
- `src/sequential_reader.rs` — packed sequential `SequentialReader` (ADR-096 Mode 1, read side)
|
||||
- `src/tunion.rs` — TUnion discriminator dispatch (byte-offset and field-name, ADR-097 §4)
|
||||
- `src/validation.rs` — custom keyword validators for all 17 `TypeDef:*` kinds
|
||||
- `src/engine.rs` — `TypedefEngine` struct combining layout + validator
|
||||
|
||||
### Dependencies
|
||||
|
||||
Per the architecture spec ([overview.md](../../docs/architecture/crates/typedef/overview.md)):
|
||||
|
||||
| Crate | Purpose |
|
||||
|-------|---------|
|
||||
| `jsonschema` 0.46 | Validation engine, custom keyword support (workspace path: `../../jsonschema`) |
|
||||
| `serde_json` (preserve_order) | Schema parsing; field order is load-bearing for binary layouts |
|
||||
|
||||
No other dependencies. No tokio, no platform deps.
|
||||
|
||||
### Workspace Cargo.toml
|
||||
|
||||
Add `crates/alknet-typedef` to the workspace `members` list in the root `Cargo.toml`.
|
||||
|
||||
### Module skeleton
|
||||
|
||||
```rust
|
||||
// src/lib.rs
|
||||
//! alknet-typedef: The binary struct engine.
|
||||
//!
|
||||
//! Takes a JSON Schema with `TypeDef:*` custom keywords and produces
|
||||
//! an offset map, read/write functions, and validation — all driven
|
||||
//! by the schema. The schema is the format definition; the engine is
|
||||
//! generic.
|
||||
//!
|
||||
//! ## Architecture
|
||||
//!
|
||||
//! - **Schema layer** ([`schema`]): TypeDef kind detection, annotation
|
||||
//! parsing, `$ref` normalization, endianness.
|
||||
//! - **Layout engine** ([`offset_map`], [`layout_builder`],
|
||||
//! [`sequential_reader`]): Two layout modes — aligned static for
|
||||
//! mmap-friendly formats, packed sequential for protocol wire formats.
|
||||
//! - **Data access** ([`data_access`]): Typed read/write at computed
|
||||
//! offsets, zero-copy for fixed-size types.
|
||||
//! - **TUnion dispatch** ([`tunion`]): Byte-offset and field-name
|
||||
//! discriminator dispatch.
|
||||
//! - **Validation** ([`validation`]): Custom keyword validators for all
|
||||
//! 17 `TypeDef:*` kinds, delegated to the `jsonschema` crate.
|
||||
//! - **Engine** ([`engine`]): `TypedefEngine` — the compiled form of a
|
||||
//! schema, combining layout and validation.
|
||||
|
||||
pub mod data_access;
|
||||
pub mod engine;
|
||||
pub mod error;
|
||||
pub mod layout_builder;
|
||||
pub mod offset_map;
|
||||
pub mod schema;
|
||||
pub mod sequential_reader;
|
||||
pub mod tunion;
|
||||
pub mod validation;
|
||||
|
||||
// Re-exports (filled in by subsequent tasks)
|
||||
```
|
||||
|
||||
Each module file gets a doc comment and `// TODO: implement` marker.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `crates/alknet-typedef/Cargo.toml` exists with `jsonschema` and `serde_json` (preserve_order) dependencies
|
||||
- [ ] `jsonschema` dependency uses workspace path (`path = "../../jsonschema"`) or git/crates.io as appropriate
|
||||
- [ ] `crates/alknet-typedef/src/lib.rs` exists with module declarations for all 9 modules
|
||||
- [ ] Module skeleton files exist: `error.rs`, `schema.rs`, `data_access.rs`, `offset_map.rs`, `layout_builder.rs`, `sequential_reader.rs`, `tunion.rs`, `validation.rs`, `engine.rs`
|
||||
- [ ] Root `Cargo.toml` `members` list includes `crates/alknet-typedef`
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] Dual licensing: `MIT OR Apache-2.0` (workspace-inherited)
|
||||
- [ ] No tokio dependency (WASM-clean by construction)
|
||||
- [ ] `serde_json` has `preserve_order` feature enabled
|
||||
- [ ] `cargo build --workspace` still succeeds (old code untouched)
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/README.md — crate overview and design principles
|
||||
- docs/architecture/crates/typedef/overview.md — purpose, dependencies, scope boundaries
|
||||
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
|
||||
- docs/research/alknet-typedef/findings.md — POC results
|
||||
- /workspace/jsonschema/ — the jsonschema crate (v0.46.5)
|
||||
- /workspace/alknet-typedef-poc/ — POC code (disposable reference)
|
||||
|
||||
## Notes
|
||||
|
||||
> This is the foundational setup task for alknet-typedef. All subsequent typedef/*
|
||||
> tasks depend on this one. The crate is dependency-light: `jsonschema` + `serde_json`
|
||||
> only. No tokio, no platform deps — WASM-clean by construction. The `jsonschema` crate
|
||||
> is already in the workspace at `/workspace/jsonschema/` but not yet used by any
|
||||
> alknet crate — typedef is the first consumer.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,174 @@
|
||||
---
|
||||
id: typedef/data-access
|
||||
name: Implement primitive read/write functions for all 17 TypeDef kinds with endianness support
|
||||
status: pending
|
||||
depends_on: [typedef/schema-types, typedef/error-type]
|
||||
scope: moderate
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement the data access layer in `crates/alknet-typedef/src/data_access.rs`. This
|
||||
module provides the primitive typed read/write functions that operate on raw byte
|
||||
buffers at given offsets. These are the building blocks used by the layout types
|
||||
(`OffsetMap`, `SequentialReader`) and the `TypedefEngine`.
|
||||
|
||||
Per [data-access.md](../../docs/architecture/crates/typedef/data-access.md).
|
||||
|
||||
### Fixed-size type read/write
|
||||
|
||||
All fixed-size read/write functions take a buffer, an offset, and an `Endian`, and
|
||||
return `Result<T, TypedefError>` (or `Result<(), TypedefError>` for writes). They
|
||||
perform bounds checking and return `TypedefError::Access` with the field path on
|
||||
failure.
|
||||
|
||||
```rust
|
||||
// Signed integers
|
||||
pub fn read_i8(buffer: &[u8], offset: usize, field_path: &str) -> Result<i8, TypedefError>;
|
||||
pub fn read_i16(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<i16, TypedefError>;
|
||||
pub fn read_i32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<i32, TypedefError>;
|
||||
|
||||
// Unsigned integers
|
||||
pub fn read_u8(buffer: &[u8], offset: usize, field_path: &str) -> Result<u8, TypedefError>;
|
||||
pub fn read_u16(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u16, TypedefError>;
|
||||
pub fn read_u32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u32, TypedefError>;
|
||||
pub fn read_u64(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u64, TypedefError>;
|
||||
|
||||
// Floats
|
||||
pub fn read_f32(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<f32, TypedefError>;
|
||||
pub fn read_f64(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<f64, TypedefError>;
|
||||
|
||||
// Boolean (0x00 = false, 0x01 = true; other values are errors)
|
||||
pub fn read_bool(buffer: &[u8], offset: usize, field_path: &str) -> Result<bool, TypedefError>;
|
||||
|
||||
// TEnum (u32 index, endian-aware)
|
||||
pub fn read_enum(buffer: &[u8], offset: usize, field_path: &str, endian: Endian) -> Result<u32, TypedefError>;
|
||||
```
|
||||
|
||||
Write counterparts:
|
||||
|
||||
```rust
|
||||
pub fn write_i8(buffer: &mut [u8], offset: usize, value: i8, field_path: &str) -> Result<(), TypedefError>;
|
||||
pub fn write_i16(buffer: &mut [u8], offset: usize, value: i16, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
pub fn write_i32(buffer: &mut [u8], offset: usize, value: i32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
pub fn write_u8(buffer: &mut [u8], offset: usize, value: u8, field_path: &str) -> Result<(), TypedefError>;
|
||||
pub fn write_u16(buffer: &mut [u8], offset: usize, value: u16, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
pub fn write_u32(buffer: &mut [u8], offset: usize, value: u32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
pub fn write_u64(buffer: &mut [u8], offset: usize, value: u64, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
pub fn write_f32(buffer: &mut [u8], offset: usize, value: f32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
pub fn write_f64(buffer: &mut [u8], offset: usize, value: f64, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
pub fn write_bool(buffer: &mut [u8], offset: usize, value: bool, field_path: &str) -> Result<(), TypedefError>;
|
||||
pub fn write_enum(buffer: &mut [u8], offset: usize, value: u32, field_path: &str, endian: Endian) -> Result<(), TypedefError>;
|
||||
```
|
||||
|
||||
### Variable-length type read/write (inline length-prefixing)
|
||||
|
||||
For variable-length types with inline length-prefixing (the default):
|
||||
|
||||
```rust
|
||||
/// Read a length-prefixed UTF-8 string. Returns a slice borrowing from the buffer.
|
||||
/// Format: [length: u32][UTF-8 bytes]
|
||||
pub fn read_string<'a>(buffer: &'a [u8], offset: usize, field_path: &str, endian: Endian) -> Result<&'a str, TypedefError>;
|
||||
|
||||
/// Read length-prefixed raw bytes. Returns a slice borrowing from the buffer.
|
||||
pub fn read_bytes<'a>(buffer: &'a [u8], offset: usize, field_path: &str, endian: Endian) -> Result<&'a [u8], TypedefError>;
|
||||
|
||||
/// Write a length-prefixed UTF-8 string.
|
||||
pub fn write_string(buffer: &mut [u8], offset: usize, value: &str, field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
|
||||
// Returns the number of bytes written (4 + value.len()) so the caller can advance.
|
||||
|
||||
/// Write length-prefixed raw bytes.
|
||||
pub fn write_bytes(buffer: &mut [u8], offset: usize, value: &[u8], field_path: &str, endian: Endian) -> Result<usize, TypedefError>;
|
||||
```
|
||||
|
||||
### Variable-length type read (offset indirection)
|
||||
|
||||
For offset-indirect types (opt-in, metatensor blob tensor pattern):
|
||||
|
||||
```rust
|
||||
/// Read an offset-indirect string. The field at `offset` is a struct
|
||||
/// {data_offset: u32, data_length: u32}. The consumer provides the
|
||||
/// separate data region.
|
||||
pub fn read_string_indirect<'a>(
|
||||
buffer: &'a [u8], // the index struct buffer
|
||||
offset: usize, // position of {data_offset, data_length}
|
||||
data_region: &'a [u8], // the separate data region
|
||||
field_path: &str,
|
||||
endian: Endian,
|
||||
) -> Result<&'a str, TypedefError>;
|
||||
|
||||
/// Read offset-indirect raw bytes.
|
||||
pub fn read_bytes_indirect<'a>(
|
||||
buffer: &'a [u8],
|
||||
offset: usize,
|
||||
data_region: &'a [u8],
|
||||
field_path: &str,
|
||||
endian: Endian,
|
||||
) -> Result<&'a [u8], TypedefError>;
|
||||
```
|
||||
|
||||
### Design notes
|
||||
|
||||
- **Zero-copy**: Read functions for variable-length types return slices borrowing
|
||||
from the input buffer — no allocation.
|
||||
- **Bounds checking**: Every function checks that the buffer is large enough for
|
||||
the requested read/write at the given offset. Returns `TypedefError::Access`
|
||||
with the field path on failure.
|
||||
- **Endianness**: Applied at access time based on the schema's `"endian"` annotation.
|
||||
The offset computation is endian-agnostic.
|
||||
- **No `unwrap`**: All fallible operations use proper `Result` returns. The spec
|
||||
pseudocode uses `unwrap` for brevity; production code must not.
|
||||
- **`TEnum`**: Reads/writes a `u32` index. The consumer maps the index to string
|
||||
values using the schema's `"enum"` array. The engine does not perform this mapping.
|
||||
- **`TBoolean`**: `0x00` = false, `0x01` = true. Other values produce
|
||||
`TypedefError::Access`.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- The offset computation (that's `offset_map.rs` and `layout_builder.rs`)
|
||||
- TUnion discriminator dispatch (that's `tunion.rs`)
|
||||
- The `TypedefEngine` struct (that's `engine.rs`)
|
||||
- `TRecord` read/write (deferred — requires count-prefixed sequence walking; can be
|
||||
added when a consumer needs it, or implemented here if straightforward)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] All fixed-size read functions implemented: `read_i8`, `read_i16`, `read_i32`, `read_u8`, `read_u16`, `read_u32`, `read_u64`, `read_f32`, `read_f64`, `read_bool`, `read_enum`
|
||||
- [ ] All fixed-size write functions implemented: `write_i8`, `write_i16`, `write_i32`, `write_u8`, `write_u16`, `write_u32`, `write_u64`, `write_f32`, `write_f64`, `write_bool`, `write_enum`
|
||||
- [ ] `read_string` and `read_bytes` (inline length-prefixed) implemented
|
||||
- [ ] `write_string` and `write_bytes` (inline length-prefixed) implemented, returning bytes written
|
||||
- [ ] `read_string_indirect` and `read_bytes_indirect` (offset-indirect) implemented
|
||||
- [ ] All functions respect `Endian` parameter (little-endian vs big-endian byte order)
|
||||
- [ ] All functions perform bounds checking and return `TypedefError::Access` with field path on failure
|
||||
- [ ] `read_bool` rejects values other than `0x00` and `0x01`
|
||||
- [ ] `read_string` validates UTF-8 and returns `TypedefError::Access` on invalid UTF-8
|
||||
- [ ] Zero-copy: read functions for variable-length types return slices, not owned data
|
||||
- [ ] No `unwrap()` or `expect()` on error paths — all fallible operations use `Result`
|
||||
- [ ] All public functions have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/data-access.md — read/write model, TEnum access, variable-length handling
|
||||
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds and their byte sizes
|
||||
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (endianness, encoding)
|
||||
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098 (error handling)
|
||||
- /workspace/alknet-typedef-poc/src/lib.rs — POC reference for read/write functions
|
||||
|
||||
## Notes
|
||||
|
||||
> This module provides the primitive read/write operations. The layout types
|
||||
> (`OffsetMap`, `SequentialReader`) use these to access fields at computed positions.
|
||||
> The functions are endian-aware — the caller passes the schema's `Endian` and the
|
||||
> functions byte-swap accordingly. All functions use proper `Result` returns with
|
||||
> field paths for debugging — no `unwrap()` in production code. The `TRecord`
|
||||
> read/write is deferred unless it proves straightforward to implement here.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,202 @@
|
||||
---
|
||||
id: typedef/engine
|
||||
name: Implement TypedefEngine struct combining layout and validation, with compile() constructor
|
||||
status: pending
|
||||
depends_on: [typedef/offset-map, typedef/layout-builder, typedef/sequential-reader, typedef/tunion, typedef/validation]
|
||||
scope: moderate
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement the `TypedefEngine` struct in `crates/alknet-typedef/src/engine.rs`. This is
|
||||
the compiled form of a schema — the main entry point for consumers. It combines the
|
||||
layout engine (both modes) and the jsonschema validator into a single struct.
|
||||
|
||||
Per [validation.md](../../docs/architecture/crates/typedef/validation.md) §"The TypedefEngine struct"
|
||||
and [overview.md](../../docs/architecture/crates/typedef/overview.md).
|
||||
|
||||
### Target shape
|
||||
|
||||
```rust
|
||||
/// The compiled form of a typedef schema. Combines the layout engine
|
||||
/// (both packed and aligned modes) and the jsonschema validator.
|
||||
///
|
||||
/// Built once at schema load time via [`TypedefEngine::compile()`].
|
||||
/// Used for repeated read/write/validate operations at access time.
|
||||
pub struct TypedefEngine {
|
||||
/// The layout strategy — packed sequential or aligned static.
|
||||
layout: Layout,
|
||||
/// The compiled jsonschema validator (built once at load time).
|
||||
validator: jsonschema::Validator,
|
||||
/// The schema's endianness.
|
||||
endian: Endian,
|
||||
/// The original schema (for reference, debugging, and TUnion dispatch).
|
||||
schema: serde_json::Value,
|
||||
}
|
||||
|
||||
/// The layout strategy selected by the consumer.
|
||||
enum Layout {
|
||||
/// Packed sequential layout for protocol wire formats.
|
||||
Packed {
|
||||
builder: LayoutBuilder,
|
||||
reader: SequentialReader,
|
||||
},
|
||||
/// Aligned static layout for mmap-friendly formats.
|
||||
Aligned {
|
||||
offset_map: OffsetMap,
|
||||
},
|
||||
}
|
||||
|
||||
/// The layout mode selected at engine construction time.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum LayoutMode {
|
||||
/// Packed sequential — for protocol wire formats (SFTP, channels, TTY).
|
||||
Packed,
|
||||
/// Aligned static — for mmap-friendly formats (metatensor, safetensors).
|
||||
Aligned,
|
||||
}
|
||||
|
||||
impl TypedefEngine {
|
||||
/// Compile a schema into a TypedefEngine.
|
||||
///
|
||||
/// This is the expensive operation — it parses the schema, normalizes
|
||||
/// `$ref` values, computes the layout, and builds the jsonschema
|
||||
/// validator. Call once at load time; use the returned engine for
|
||||
/// repeated operations.
|
||||
///
|
||||
/// The `mode` parameter selects the layout strategy. The same schema
|
||||
/// can be compiled in either mode.
|
||||
pub fn compile(
|
||||
schema: &mut serde_json::Value,
|
||||
mode: LayoutMode,
|
||||
) -> Result<Self, TypedefError>;
|
||||
|
||||
/// The schema's endianness.
|
||||
pub fn endian(&self) -> Endian;
|
||||
|
||||
/// The layout mode this engine was compiled with.
|
||||
pub fn mode(&self) -> LayoutMode;
|
||||
|
||||
/// Access the aligned offset map. Returns `None` if compiled in
|
||||
/// packed mode.
|
||||
pub fn offset_map(&self) -> Option<&OffsetMap>;
|
||||
|
||||
/// Access the layout builder (write-side of packed mode).
|
||||
/// Returns `None` if compiled in aligned mode.
|
||||
pub fn layout_builder(&self) -> Option<&LayoutBuilder>;
|
||||
|
||||
/// Access the sequential reader (read-side of packed mode).
|
||||
/// Returns `None` if compiled in aligned mode.
|
||||
pub fn sequential_reader(&self) -> Option<&SequentialReader>;
|
||||
|
||||
/// Validate a JSON value against the schema. The jsonschema validator
|
||||
/// is already compiled — this is a fast check.
|
||||
///
|
||||
/// Returns `Ok(())` if valid, `Err(TypedefError::Validation(...))` if invalid.
|
||||
pub fn validate_json(&self, instance: &serde_json::Value) -> Result<(), TypedefError>;
|
||||
|
||||
/// Check if a JSON value is valid against the schema.
|
||||
pub fn is_valid_json(&self, instance: &serde_json::Value) -> bool;
|
||||
|
||||
/// Read a field from a buffer at its computed offset (aligned mode).
|
||||
/// Returns `None` if compiled in packed mode — use `sequential_reader()` instead.
|
||||
pub fn read_field<'a>(
|
||||
&self,
|
||||
buffer: &'a [u8],
|
||||
field_path: &str,
|
||||
) -> Result<FieldValue<'a>, TypedefError>;
|
||||
|
||||
/// Write a field to a buffer at its computed offset (aligned mode).
|
||||
/// Returns `None` if compiled in packed mode — use `layout_builder()` instead.
|
||||
pub fn write_field(
|
||||
&self,
|
||||
buffer: &mut [u8],
|
||||
field_path: &str,
|
||||
value: &FieldValue<'_>,
|
||||
) -> Result<(), TypedefError>;
|
||||
}
|
||||
```
|
||||
|
||||
### `compile()` constructor
|
||||
|
||||
The `compile()` method performs these steps in order:
|
||||
|
||||
1. **Normalize `$ref` values** — call `schema::normalize_refs()` to rewrite bare-name
|
||||
refs to full JSON Pointer paths.
|
||||
2. **Parse endianness** — extract `"endian"` annotation (defaults to `Little`).
|
||||
3. **Build the layout** — depending on `mode`:
|
||||
- `Packed`: build `LayoutBuilder` and `SequentialReader` from the schema.
|
||||
- `Aligned`: compute `OffsetMap` from the schema.
|
||||
4. **Build the validator** — call `validation::build_validator()` to register all 17
|
||||
custom keywords and compile the jsonschema validator.
|
||||
5. **Return the engine** — all components ready for repeated use.
|
||||
|
||||
### Aligned mode convenience methods
|
||||
|
||||
When compiled in `Aligned` mode, `read_field()` and `write_field()` provide
|
||||
convenient access using the `OffsetMap`:
|
||||
|
||||
- `read_field(buffer, "header.version")` → looks up the field's byte range in the
|
||||
`OffsetMap`, reads the appropriate type using `data_access` functions.
|
||||
- `write_field(buffer, "header.version", &FieldValue::U32(1))` → looks up the byte
|
||||
range, writes using `data_access` functions.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- A builder API for schema construction (deferred, OQ-071)
|
||||
- Schema evolution / Value system (out of scope for v1)
|
||||
- Code generation (out of scope)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `TypedefEngine` struct with `layout`, `validator`, `endian`, `schema` fields
|
||||
- [ ] `Layout` enum with `Packed { builder, reader }` and `Aligned { offset_map }` variants
|
||||
- [ ] `LayoutMode` enum with `Packed` and `Aligned` variants
|
||||
- [ ] `TypedefEngine::compile(&mut schema, mode)` performs all load-time work
|
||||
- [ ] `compile()` normalizes `$ref` values before building
|
||||
- [ ] `compile()` builds the correct layout for the selected mode
|
||||
- [ ] `compile()` builds the jsonschema validator with all 17 custom keywords
|
||||
- [ ] `compile()` returns `TypedefError::Schema` for invalid schemas
|
||||
- [ ] `endian()` returns the schema's endianness
|
||||
- [ ] `mode()` returns the layout mode
|
||||
- [ ] `offset_map()` returns `Some(&OffsetMap)` in aligned mode, `None` in packed mode
|
||||
- [ ] `layout_builder()` returns `Some(&LayoutBuilder)` in packed mode, `None` in aligned mode
|
||||
- [ ] `sequential_reader()` returns `Some(&SequentialReader)` in packed mode, `None` in aligned mode
|
||||
- [ ] `validate_json()` delegates to the compiled jsonschema validator
|
||||
- [ ] `is_valid_json()` returns true/false without error details
|
||||
- [ ] `read_field()` works in aligned mode for all fixed-size and variable-length types
|
||||
- [ ] `write_field()` works in aligned mode for all fixed-size and variable-length types
|
||||
- [ ] `read_field()` and `write_field()` return errors in packed mode (use layout-specific APIs)
|
||||
- [ ] No `unwrap()` or `expect()` on error paths
|
||||
- [ ] All public types and functions have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/validation.md — TypedefEngine struct, compile() constructor
|
||||
- docs/architecture/crates/typedef/overview.md — architecture, consumers
|
||||
- docs/architecture/crates/typedef/layout-engine.md — the two layout modes
|
||||
- docs/architecture/crates/typedef/data-access.md — read/write functions
|
||||
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
|
||||
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
|
||||
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
|
||||
|
||||
## Notes
|
||||
|
||||
> This is the integration task — it wires together all the components built in the
|
||||
> preceding tasks. The `TypedefEngine` is the main entry point for consumers. The
|
||||
> `compile()` constructor does all the expensive work once at load time. The engine
|
||||
> supports both layout modes via the `Layout` enum — the consumer selects the mode
|
||||
> at construction time. The aligned-mode convenience methods (`read_field`,
|
||||
> `write_field`) provide a simple API for the common mmap use case. Protocol
|
||||
> consumers use the layout-specific APIs (`layout_builder()`, `sequential_reader()`)
|
||||
> directly.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,110 @@
|
||||
---
|
||||
id: typedef/error-type
|
||||
name: Implement TypedefError enum with Schema, Offset, Access, and Validation variants
|
||||
status: pending
|
||||
depends_on: [typedef/crate-init]
|
||||
scope: narrow
|
||||
risk: low
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement the `TypedefError` enum in `crates/alknet-typedef/src/error.rs`. This is the
|
||||
single error type for all three engine phases (schema parsing, offset computation,
|
||||
read/write) plus validation. Decided in
|
||||
[ADR-098](../../docs/architecture/decisions/098-error-handling-validation-strategy.md).
|
||||
|
||||
### Target shape (per ADR-098)
|
||||
|
||||
```rust
|
||||
use std::fmt;
|
||||
|
||||
/// Errors produced by the typedef engine across all phases.
|
||||
#[derive(Debug)]
|
||||
pub enum TypedefError {
|
||||
/// Schema parsing errors — invalid JSON, missing required keywords,
|
||||
/// unknown `TypeDef:*` kinds, malformed annotations.
|
||||
Schema(String),
|
||||
|
||||
/// Offset computation errors — field not found, type not supported
|
||||
/// for offset computation, recursive depth exceeded.
|
||||
Offset {
|
||||
field_path: String,
|
||||
reason: String,
|
||||
},
|
||||
|
||||
/// Read/write errors — buffer too short, invalid UTF-8, value out
|
||||
/// of range for the target type.
|
||||
Access {
|
||||
field_path: String,
|
||||
reason: String,
|
||||
},
|
||||
|
||||
/// Validation errors — delegated to the `jsonschema` crate.
|
||||
/// The `'static` lifetime is correct: the validator owns its schema
|
||||
/// reference and lives for the lifetime of the `TypedefEngine`.
|
||||
Validation(jsonschema::ValidationError<'static>),
|
||||
}
|
||||
```
|
||||
|
||||
### Design rationale
|
||||
|
||||
- **`Schema(String)`** — for errors during `TypedefEngine::compile()`. Invalid JSON,
|
||||
missing required keywords, unknown `TypeDef:*` kinds. The error message describes
|
||||
the problem.
|
||||
- **`Offset { field_path, reason }`** — for errors during offset computation. Field
|
||||
not found in the schema, type not supported for offset computation. Carries the
|
||||
field path for debugging.
|
||||
- **`Access { field_path, reason }`** — for errors during read/write. Buffer too
|
||||
short, invalid UTF-8 in a string field, value out of range. Carries the field
|
||||
path for debugging.
|
||||
- **`Validation(ValidationError<'static>)`** — wraps `jsonschema`'s `ValidationError`.
|
||||
The `'static` lifetime is correct — the validator is built once at schema load time
|
||||
and lives for the lifetime of the `TypedefEngine`.
|
||||
|
||||
### Trait implementations
|
||||
|
||||
- `Display` — human-readable error messages including field paths where applicable
|
||||
- `Error` (std::error::Error) — for `?` propagation
|
||||
- `Debug` — derived
|
||||
|
||||
The `Validation` variant requires `jsonschema::ValidationError` to be in scope.
|
||||
Since the `jsonschema` crate is a dependency, this is straightforward.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- No `From` impls for other error types (those are added as needed by subsequent tasks)
|
||||
- No `PartialEq` — `ValidationError` may not implement it
|
||||
- No `Clone` — errors are typically consumed, not cloned
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `TypedefError` enum defined in `crates/alknet-typedef/src/error.rs`
|
||||
- [ ] Four variants: `Schema(String)`, `Offset { field_path, reason }`, `Access { field_path, reason }`, `Validation(ValidationError<'static>)`
|
||||
- [ ] `Display` impl with descriptive messages including field paths for `Offset` and `Access`
|
||||
- [ ] `std::error::Error` impl
|
||||
- [ ] `Debug` derived
|
||||
- [ ] Re-exported from `lib.rs`
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
|
||||
- docs/architecture/crates/typedef/validation.md — TypedefError section
|
||||
- docs/architecture/crates/typedef/data-access.md — error handling in read/write
|
||||
|
||||
## Notes
|
||||
|
||||
> This is a small, self-contained task. The error type is used by every other module
|
||||
> in the crate. `Schema` covers load-time errors, `Offset` covers layout computation
|
||||
> errors, `Access` covers read/write errors, and `Validation` wraps jsonschema's
|
||||
> error type. The `'static` lifetime on `ValidationError` is correct because the
|
||||
> validator is built once and lives for the engine's lifetime.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,173 @@
|
||||
---
|
||||
id: typedef/layout-builder
|
||||
name: Implement packed sequential LayoutBuilder for protocol write-side
|
||||
status: pending
|
||||
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
|
||||
scope: moderate
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement the packed sequential `LayoutBuilder` in `crates/alknet-typedef/src/layout_builder.rs`.
|
||||
This is the write-side of Mode 1 (ADR-096): fields are packed with no alignment padding.
|
||||
Variable-length fields shift all subsequent fields. The consumer provides actual data
|
||||
sizes for variable-length fields; the builder computes byte positions for each field.
|
||||
|
||||
Per [layout-engine.md](../../docs/architecture/crates/typedef/layout-engine.md) §"Mode 1: Packed sequential".
|
||||
|
||||
### Target shape
|
||||
|
||||
```rust
|
||||
/// Builds a packed sequential layout for protocol wire formats.
|
||||
/// Fields are packed with no alignment padding. Variable-length fields
|
||||
/// shift all subsequent fields. The consumer provides actual data sizes
|
||||
/// for variable-length fields to compute correct positions.
|
||||
///
|
||||
/// Used at write time when the consumer knows the data sizes upfront.
|
||||
#[derive(Debug)]
|
||||
pub struct LayoutBuilder {
|
||||
/// The schema being laid out.
|
||||
schema: serde_json::Value,
|
||||
/// The endianness for the layout.
|
||||
endian: Endian,
|
||||
}
|
||||
|
||||
/// A field position computed by the LayoutBuilder.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct FieldPosition {
|
||||
/// Byte offset of the field within the buffer.
|
||||
pub offset: usize,
|
||||
/// Byte size of the field (4 for length prefix of variable-length fields,
|
||||
/// actual size for fixed-size fields).
|
||||
pub size: usize,
|
||||
/// The TypeDef kind of the field.
|
||||
pub kind: String,
|
||||
}
|
||||
|
||||
/// The result of building a layout: a map of field_path → FieldPosition
|
||||
/// and the total buffer size needed.
|
||||
#[derive(Debug)]
|
||||
pub struct PackedLayout {
|
||||
fields: Vec<(String, FieldPosition)>,
|
||||
total_size: usize,
|
||||
}
|
||||
|
||||
impl LayoutBuilder {
|
||||
/// Create a new LayoutBuilder from a schema.
|
||||
pub fn new(schema: &serde_json::Value) -> Result<Self, TypedefError>;
|
||||
|
||||
/// Build the packed layout given actual data sizes for variable-length fields.
|
||||
/// `var_sizes` maps field paths to their actual byte sizes (not including
|
||||
/// the 4-byte length prefix — the builder adds that).
|
||||
///
|
||||
/// For fixed-size fields, the size is known from the schema.
|
||||
/// For variable-length fields, the size comes from `var_sizes`.
|
||||
/// For TUnion, the consumer provides the discriminator value to select
|
||||
/// the variant, and the variant's field sizes.
|
||||
pub fn build(
|
||||
&self,
|
||||
var_sizes: &HashMap<String, usize>,
|
||||
) -> Result<PackedLayout, TypedefError>;
|
||||
}
|
||||
|
||||
impl PackedLayout {
|
||||
/// Look up a field's position by dotted path.
|
||||
pub fn get(&self, field_path: &str) -> Option<&FieldPosition>;
|
||||
|
||||
/// The total buffer size needed to hold all fields.
|
||||
pub fn total_size(&self) -> usize;
|
||||
|
||||
/// Iterate over all (field_path, position) pairs in layout order.
|
||||
pub fn iter(&self) -> impl Iterator<Item = &(String, FieldPosition)>;
|
||||
}
|
||||
```
|
||||
|
||||
### How it works
|
||||
|
||||
For a struct with fields `[u8, u32, string]` where the string is 10 bytes:
|
||||
|
||||
```
|
||||
LayoutBuilder::build(var_sizes: {"payload": 10}):
|
||||
field[0] u8: offset 0, size 1
|
||||
field[1] u32: offset 1, size 4
|
||||
field[2] string: offset 5, size 4 (length prefix) + 10 (data) = 14
|
||||
total: 19
|
||||
```
|
||||
|
||||
There is no alignment padding. The `u32` at offset 1 is unaligned — this is correct
|
||||
for protocol wire formats, which pack fields tightly.
|
||||
|
||||
### Variable-length fields in packed mode
|
||||
|
||||
The `LayoutBuilder` takes actual data sizes for variable-length fields to compute
|
||||
correct positions for subsequent fields. The consumer must know the data sizes before
|
||||
writing — this is inherent to packed layouts.
|
||||
|
||||
For each variable-length field:
|
||||
1. The builder records the position of the 4-byte length prefix.
|
||||
2. The builder adds `4 + data_size` to the current offset.
|
||||
3. Subsequent fields start after the variable data.
|
||||
|
||||
### TUnion in packed mode
|
||||
|
||||
For `TUnion`, the consumer provides the discriminator value and the variant's field
|
||||
sizes. The builder:
|
||||
1. Computes the discriminator's position and size.
|
||||
2. Looks up the variant schema from the mapping.
|
||||
3. Computes the variant's field positions starting at `offset + discriminator_size`.
|
||||
4. The union's total size is `discriminator_size + variant_size`.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- The read-side of packed mode (that's `sequential_reader.rs`)
|
||||
- Aligned static layout (that's `offset_map.rs`)
|
||||
- TUnion discriminator dispatch (that's `tunion.rs`)
|
||||
- The `TypedefEngine` struct (that's `engine.rs`)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `LayoutBuilder` struct with `new(schema)` constructor
|
||||
- [ ] `LayoutBuilder::build(var_sizes)` computes packed field positions
|
||||
- [ ] Fixed-size fields get correct offsets with no alignment padding
|
||||
- [ ] `u8` at offset 0, `u32` at offset 1 (no padding) — packed sequential
|
||||
- [ ] Variable-length fields: 4-byte length prefix at computed offset, data follows
|
||||
- [ ] Variable-length field sizes come from `var_sizes` map
|
||||
- [ ] Subsequent fields shift based on actual variable-length data sizes
|
||||
- [ ] Nested structs produce dotted field paths
|
||||
- [ ] `TArray` of fixed-size elements: correct stride (element size, no padding)
|
||||
- [ ] `TArray` with variable count: 4-byte count prefix at computed offset
|
||||
- [ ] `TUnion` with byte-offset discriminator: discriminator at `offset`, variant at `offset + disc_size`
|
||||
- [ ] `TUnion` with field-name discriminator: discriminator is a regular field
|
||||
- [ ] `PackedLayout::get("field_name")` returns correct `FieldPosition`
|
||||
- [ ] `PackedLayout::total_size()` returns correct total buffer size
|
||||
- [ ] `PackedLayout::iter()` iterates fields in layout order
|
||||
- [ ] Returns `TypedefError::Schema` for malformed schemas
|
||||
- [ ] Returns `TypedefError::Offset` for missing variable-length field sizes
|
||||
- [ ] No `unwrap()` or `expect()` on error paths
|
||||
- [ ] All public types and functions have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/layout-engine.md — Mode 1: Packed sequential, LayoutBuilder
|
||||
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds and their byte sizes
|
||||
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
|
||||
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (encoding)
|
||||
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference (LayoutBuilder in POC 2)
|
||||
|
||||
## Notes
|
||||
|
||||
> This is the write-side of the packed sequential mode. The consumer knows the data
|
||||
> sizes upfront (e.g., when constructing an SFTP response packet) and uses the
|
||||
> `LayoutBuilder` to compute where each field goes. The builder does not write data —
|
||||
> it only computes positions. The consumer uses the `data_access` module's write
|
||||
> functions at the computed positions. The POC 2's `LayoutBuilder` is a good reference.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,163 @@
|
||||
---
|
||||
id: typedef/offset-map
|
||||
name: Implement aligned static OffsetMap for mmap-friendly formats
|
||||
status: pending
|
||||
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
|
||||
scope: moderate
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement the aligned static `OffsetMap` in `crates/alknet-typedef/src/offset_map.rs`.
|
||||
This is Mode 2 of the two layout modes (ADR-096): fields have fixed positions with
|
||||
natural alignment padding. Variable-length fields get a 4-byte length prefix at a
|
||||
known offset; the variable data is not included in the static layout.
|
||||
|
||||
Per [layout-engine.md](../../docs/architecture/crates/typedef/layout-engine.md) §"Mode 2: Aligned static".
|
||||
|
||||
### Target shape
|
||||
|
||||
```rust
|
||||
/// A byte range within a buffer.
|
||||
#[derive(Debug, Clone, Copy)]
|
||||
pub struct ByteRange {
|
||||
pub start: usize,
|
||||
pub end: usize,
|
||||
}
|
||||
|
||||
/// A flat table of (field_path, byte_range) pairs computed from a schema.
|
||||
/// Fields have fixed positions with natural alignment padding.
|
||||
/// Used for mmap-friendly formats (metatensor, safetensors).
|
||||
#[derive(Debug)]
|
||||
pub struct OffsetMap {
|
||||
fields: Vec<(String, ByteRange)>,
|
||||
total_size: usize,
|
||||
}
|
||||
|
||||
impl OffsetMap {
|
||||
/// Compute the offset map from a schema JSON value.
|
||||
/// Walks the schema recursively, computing byte positions for each field
|
||||
/// based on type sizes, field order, and alignment.
|
||||
pub fn compute(schema: &serde_json::Value) -> Result<Self, TypedefError>;
|
||||
|
||||
/// Look up a field's byte range by dotted path (e.g., "header.version").
|
||||
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
|
||||
|
||||
/// The total size of the struct in bytes (including alignment padding).
|
||||
pub fn total_size(&self) -> usize;
|
||||
|
||||
/// Iterate over all (field_path, byte_range) pairs.
|
||||
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
|
||||
}
|
||||
```
|
||||
|
||||
### Offset computation algorithm
|
||||
|
||||
The algorithm walks the schema recursively:
|
||||
|
||||
1. **Fixed-size types**: Determine the type's byte size from the `TypeDef:*` kind.
|
||||
Insert alignment padding to satisfy the type's alignment (or the field's `align`
|
||||
annotation, or the struct's `align` default). Record the field's `(start, end)`
|
||||
range. Advance the current offset by the type's size.
|
||||
|
||||
2. **`TStruct`**: Recurse into the struct's `properties`. Inner fields are computed
|
||||
relative to the struct's start offset. The struct's total size is the sum of its
|
||||
fields' sizes plus alignment padding. The struct itself may have an `align`
|
||||
annotation that rounds up its total size.
|
||||
|
||||
3. **`TUnion`**: The discriminator occupies `offset..offset + discriminator_size`
|
||||
bytes. For byte-offset discriminators, the variant struct starts at
|
||||
`offset + discriminator_size`. For field-name discriminators, the discriminator
|
||||
is just another field. The union's total size is `discriminator_size +
|
||||
max(variant_sizes)`.
|
||||
|
||||
4. **`TArray` of fixed-size elements**: Element stride = element size plus alignment
|
||||
padding. Element `i` starts at `array_offset + i × stride`. The array's total
|
||||
size is `count × stride`. Count is determined from `minItems`/`maxItems` (when
|
||||
equal, fixed count; otherwise variable — uses length-prefixed encoding).
|
||||
|
||||
5. **Variable-length types (inline length-prefixing)**: Record the position of the
|
||||
4-byte length prefix. The variable data is not included in the static layout.
|
||||
The length prefix is aligned to 4 bytes.
|
||||
|
||||
6. **Variable-length types (fixed-size reservation, `maxLength`)**: Reserve
|
||||
`maxLength` bytes at a fixed offset. Data shorter than `maxLength` is zero-padded.
|
||||
Subsequent fields have known, unchanging offsets.
|
||||
|
||||
7. **Variable-length types (offset indirection)**: The field is a struct
|
||||
`{offset: u32, length: u32}` (8 bytes total). Record its position. The consumer
|
||||
provides the data region separately.
|
||||
|
||||
### Nested structs and field paths
|
||||
|
||||
Nested structs produce dotted field paths: `"header.version"`, `"header.magic"`.
|
||||
The offset computation propagates the field path prefix during recursion. The
|
||||
`OffsetMap` stores fully-qualified paths.
|
||||
|
||||
### Alignment rules
|
||||
|
||||
- Default alignment: 1 for u8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for
|
||||
u64/i64/f64, max field alignment for structs.
|
||||
- Struct-level `"align"` sets the default for all fields in that struct.
|
||||
- Field-level `"align"` overrides the struct default.
|
||||
- The struct's total size is rounded up to its alignment.
|
||||
- Alignment padding is inserted before each field to satisfy its alignment.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- Packed sequential layout (that's `layout_builder.rs` and `sequential_reader.rs`)
|
||||
- TUnion discriminator dispatch (that's `tunion.rs`)
|
||||
- The `TypedefEngine` struct (that's `engine.rs`)
|
||||
- Arrays of variable-length-element structs (deferred, OQ-069)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `OffsetMap` struct with `fields: Vec<(String, ByteRange)>` and `total_size: usize`
|
||||
- [ ] `OffsetMap::compute(schema)` walks the schema and computes byte positions
|
||||
- [ ] Fixed-size types get correct byte ranges with natural alignment padding
|
||||
- [ ] `u8` at offset 0, `u32` at offset 4 (3 bytes padding) — natural alignment
|
||||
- [ ] Nested structs produce dotted field paths (`"header.version"`)
|
||||
- [ ] `TArray` of fixed-size elements: correct stride and element offsets
|
||||
- [ ] `TArray` with `minItems == maxItems`: fixed count, known at schema time
|
||||
- [ ] `TArray` with variable count: length-prefixed encoding (4-byte count prefix)
|
||||
- [ ] Variable-length types (inline length-prefixing): 4-byte length prefix at known offset
|
||||
- [ ] Variable-length types (`maxLength`): reserved `maxLength` bytes at fixed offset
|
||||
- [ ] Variable-length types (offset-indirect): 8-byte `{offset, length}` struct at known offset
|
||||
- [ ] `TUnion` with byte-offset discriminator: discriminator at `offset`, variant at `offset + disc_size`
|
||||
- [ ] `TUnion` with field-name discriminator: discriminator is a regular field
|
||||
- [ ] Struct-level `"align"` annotation: rounds up struct total size
|
||||
- [ ] Field-level `"align"` annotation: overrides struct default for that field
|
||||
- [ ] `OffsetMap::get("header.version")` returns the correct `ByteRange`
|
||||
- [ ] `OffsetMap::total_size()` returns the correct total size
|
||||
- [ ] `OffsetMap::iter()` iterates all field paths
|
||||
- [ ] Returns `TypedefError::Schema` for malformed schemas
|
||||
- [ ] Returns `TypedefError::Offset` for unsupported type combinations
|
||||
- [ ] No `unwrap()` or `expect()` on error paths
|
||||
- [ ] All public types and functions have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/layout-engine.md — Mode 2: Aligned static, offset computation algorithm
|
||||
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds and their byte sizes
|
||||
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
|
||||
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (alignment, encoding)
|
||||
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference for offset computation
|
||||
|
||||
## Notes
|
||||
|
||||
> This is the aligned static layout mode — the simpler of the two modes. Fields have
|
||||
> fixed positions; the consumer can read field N without reading fields 0..N-1 first.
|
||||
> Used by metatensor and safetensors. The offset computation is a recursive walk of
|
||||
> the schema JSON. Nested structs propagate field path prefixes. Alignment padding
|
||||
> is inserted between fields based on type sizes and annotations. The POC's
|
||||
> `offset.rs` is a good reference — the algorithm is correct and can be adapted.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,162 @@
|
||||
---
|
||||
id: typedef/review-typedef
|
||||
name: Review alknet-typedef implementation for spec conformance, API shape, and test coverage
|
||||
status: pending
|
||||
depends_on: [typedef/tests]
|
||||
scope: moderate
|
||||
risk: low
|
||||
impact: phase
|
||||
level: review
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Review checkpoint for the `alknet-typedef` crate. Verify the implementation is
|
||||
spec-conformant, self-contained, and ready for downstream consumption by metatensor,
|
||||
SFTP, binary call frames, and TTY negotiation.
|
||||
|
||||
### Review Checklist
|
||||
|
||||
#### 1. Crate structure
|
||||
|
||||
- Module layout matches spec: `error.rs`, `schema.rs`, `data_access.rs`, `offset_map.rs`,
|
||||
`layout_builder.rs`, `sequential_reader.rs`, `tunion.rs`, `validation.rs`, `engine.rs`
|
||||
- Public API types: `TypedefEngine`, `TypedefError`, `OffsetMap`, `LayoutBuilder`,
|
||||
`SequentialReader`, `LayoutMode`, `Endian`, `ByteRange`, `FieldValue`, `UnionDispatch`
|
||||
- Re-exports in `lib.rs` are correct and minimal
|
||||
- No tokio dependency (WASM-clean by construction)
|
||||
- `serde_json` has `preserve_order` feature enabled
|
||||
|
||||
#### 2. Schema layer (ADR-097)
|
||||
|
||||
- All 17 `TypeDef:*` kinds correctly identified by `get_typedef_kind()`
|
||||
- `type_size()` returns correct sizes for all fixed-size types
|
||||
- `Endian` enum with `Little` (default) and `Big`
|
||||
- `parse_encoding()` handles both `true` (shorthand) and `{ "encoding": "..." }` (object)
|
||||
- `parse_discriminator()` handles byte-offset and field-name discriminators
|
||||
- `normalize_refs()` rewrites bare-name refs to full JSON Pointer paths
|
||||
- `TEnum` uses `u32` index (not variable-length string) — deliberate deviation from TypeBox
|
||||
|
||||
#### 3. Data access layer
|
||||
|
||||
- All fixed-size read/write functions implemented with endianness support
|
||||
- `read_bool`: `0x00` = false, `0x01` = true, other values → error
|
||||
- `read_enum`: reads `u32` index with endianness
|
||||
- Variable-length types: inline length-prefixing (default) and offset indirection (opt-in)
|
||||
- Zero-copy: read functions return slices, not owned data
|
||||
- All functions perform bounds checking and return `TypedefError::Access` with field path
|
||||
- No `unwrap()` or `expect()` on error paths — all fallible operations use `Result`
|
||||
|
||||
#### 4. Layout engine (ADR-096)
|
||||
|
||||
- **Aligned static mode** (`OffsetMap`): fields have fixed positions with natural alignment
|
||||
padding. Variable-length fields get a 4-byte length prefix at known offset.
|
||||
- **Packed sequential mode** (`LayoutBuilder` + `SequentialReader`): fields packed with
|
||||
no alignment padding. Variable-length fields shift subsequent fields.
|
||||
- Nested structs produce dotted field paths (`"header.version"`)
|
||||
- `TArray` with fixed count (`minItems == maxItems`) and variable count (length-prefixed)
|
||||
- `TUnion` with byte-offset and field-name discriminators
|
||||
- Alignment annotations: struct-level and field-level, field-level overrides
|
||||
- `maxLength` annotation: fixed-size reservation in aligned mode, validation constraint in packed mode
|
||||
- Endianness: per-schema, default little-endian, applied at access time
|
||||
|
||||
#### 5. TUnion dispatch (ADR-097 §4)
|
||||
|
||||
- Byte-offset discriminator: reads fixed-size integer at known offset, returns mapping key
|
||||
- Field-name discriminator: reads named field, returns mapping key
|
||||
- `resolve_variant()` resolves `$ref` pointers to `$defs`
|
||||
- Supports `TypeDef:Uint8`, `TypeDef:Uint16`, `TypeDef:Uint32` discriminator types
|
||||
|
||||
#### 6. Validation (ADR-098)
|
||||
|
||||
- `build_validator()` registers all 17 custom keywords with `jsonschema`
|
||||
- Each validator is ~10 lines (not hundreds)
|
||||
- `StructValidator` inspects parent's `properties` for cross-keyword awareness
|
||||
- `EnumValidator` is a no-op (built-in `enum` keyword handles validation)
|
||||
- `TypedefError::Validation` wraps `jsonschema::ValidationError<'static>`
|
||||
|
||||
#### 7. TypedefEngine
|
||||
|
||||
- `compile(&mut schema, mode)` performs all load-time work: normalize refs, build layout, build validator
|
||||
- `Layout` enum with `Packed { builder, reader }` and `Aligned { offset_map }` variants
|
||||
- `LayoutMode` enum with `Packed` and `Aligned` variants
|
||||
- Accessor methods: `endian()`, `mode()`, `offset_map()`, `layout_builder()`, `sequential_reader()`
|
||||
- `validate_json()` and `is_valid_json()` delegate to compiled validator
|
||||
- `read_field()` and `write_field()` convenience methods for aligned mode
|
||||
|
||||
#### 8. Error handling (ADR-098)
|
||||
|
||||
- `TypedefError` enum with four variants: `Schema`, `Offset`, `Access`, `Validation`
|
||||
- `Offset` and `Access` variants carry field paths for debugging
|
||||
- `Display` and `Error` trait implementations
|
||||
- No `unwrap()` or `expect()` on error paths anywhere in the crate
|
||||
|
||||
#### 9. Test coverage
|
||||
|
||||
- Schema layer tests: all public functions tested
|
||||
- Data access tests: all read/write functions with round-trip, endianness, error paths
|
||||
- OffsetMap tests: aligned static layout with alignment, nesting, arrays, unions
|
||||
- LayoutBuilder tests: packed sequential layout with variable-length shifting
|
||||
- SequentialReader tests: sequential field reading with position tracking
|
||||
- TUnion tests: both discriminator kinds, variant resolution
|
||||
- Validation tests: all 17 custom keyword validators
|
||||
- Engine tests: compile, accessors, convenience methods
|
||||
- POC round-trip tests: fixed-size, string, nested struct, endianness
|
||||
- Error path tests: buffer-too-short, invalid UTF-8, malformed schemas
|
||||
|
||||
#### 10. Cross-cutting checks
|
||||
|
||||
- `cargo build -p alknet-typedef` succeeds
|
||||
- `cargo test -p alknet-typedef` succeeds (all tests pass)
|
||||
- `cargo clippy -p alknet-typedef --all-targets` succeeds with no warnings
|
||||
- `cargo fmt --check -p alknet-typedef` passes
|
||||
- `cargo build --workspace` still succeeds (old code untouched)
|
||||
- `cargo test --workspace` still succeeds (old tests untouched)
|
||||
- No `unwrap()` or `expect()` in production code (spec pseudocode uses `unwrap` for brevity only)
|
||||
- `TEnum` uses `u32` index (not variable-length string) — deliberate deviation from TypeBox
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] Crate structure matches spec (9 source files, correct module layout)
|
||||
- [ ] All 17 `TypeDef:*` kinds correctly identified and sized
|
||||
- [ ] Both layout modes work correctly (aligned static and packed sequential)
|
||||
- [ ] TUnion dispatch supports both byte-offset and field-name discriminators
|
||||
- [ ] All 17 custom keyword validators registered and working
|
||||
- [ ] `TypedefEngine::compile()` correctly wires all components
|
||||
- [ ] `TypedefError` has correct variants with field-path-carrying errors
|
||||
- [ ] No `unwrap()` or `expect()` in production code
|
||||
- [ ] All tests pass (unit + integration)
|
||||
- [ ] `cargo build -p alknet-typedef` succeeds
|
||||
- [ ] `cargo test -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef --all-targets` succeeds with no warnings
|
||||
- [ ] `cargo fmt --check -p alknet-typedef` passes
|
||||
- [ ] Workspace still green: `cargo build --workspace` + `cargo test --workspace` pass
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/README.md — crate overview and design principles
|
||||
- docs/architecture/crates/typedef/overview.md — purpose, dependencies, scope boundaries
|
||||
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds, annotations
|
||||
- docs/architecture/crates/typedef/layout-engine.md — the two layout modes
|
||||
- docs/architecture/crates/typedef/data-access.md — read/write functions
|
||||
- docs/architecture/crates/typedef/validation.md — custom keyword validators
|
||||
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
|
||||
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
|
||||
- docs/architecture/decisions/097-schema-annotations.md — ADR-097
|
||||
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
|
||||
- docs/research/alknet-typedef/findings.md — POC results
|
||||
- All task files in tasks/typedef/
|
||||
|
||||
## Notes
|
||||
|
||||
> This review gates the alknet-typedef implementation. The crate must be self-contained
|
||||
> and spec-conformant before downstream consumers (metatensor, SFTP, binary call frames,
|
||||
> TTY negotiation) can depend on it. The POC validated the approach with 26 passing
|
||||
> tests; the production implementation should match or exceed that coverage. Key things
|
||||
> to verify: no `unwrap()` in production code (the spec pseudocode uses it for brevity),
|
||||
> `TEnum` uses `u32` index (not variable-length string), and both layout modes produce
|
||||
> correct offsets for their respective use cases.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,183 @@
|
||||
---
|
||||
id: typedef/schema-types
|
||||
name: Implement TypeDef kind detection, schema annotation parsing, $ref normalization, and Endian enum
|
||||
status: pending
|
||||
depends_on: [typedef/crate-init]
|
||||
scope: moderate
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement the schema layer in `crates/alknet-typedef/src/schema.rs`. This module
|
||||
provides the foundational types and functions that every other module depends on:
|
||||
TypeDef kind detection, type size constants, schema annotation parsing, `$ref`
|
||||
normalization, and the `Endian` enum.
|
||||
|
||||
Per [schema-layer.md](../../docs/architecture/crates/typedef/schema-layer.md) and
|
||||
[ADR-097](../../docs/architecture/decisions/097-schema-annotations.md).
|
||||
|
||||
### TypeDef kind detection
|
||||
|
||||
The engine needs to identify which `TypeDef:*` kind a schema node declares. The
|
||||
17 kinds (16 from TypeBox's `typedef.ts` + `TypeDef:Bytes` added by alknet-typedef):
|
||||
|
||||
| Kind | Keyword | Category | Byte size |
|
||||
|------|---------|----------|-----------|
|
||||
| Float32 | `TypeDef:Float32` | fixed | 4 |
|
||||
| Float64 | `TypeDef:Float64` | fixed | 8 |
|
||||
| Int8 | `TypeDef:Int8` | fixed | 1 |
|
||||
| Int16 | `TypeDef:Int16` | fixed | 2 |
|
||||
| Int32 | `TypeDef:Int32` | fixed | 4 |
|
||||
| Uint8 | `TypeDef:Uint8` | fixed | 1 |
|
||||
| Uint16 | `TypeDef:Uint16` | fixed | 2 |
|
||||
| Uint32 | `TypeDef:Uint32` | fixed | 4 |
|
||||
| Boolean | `TypeDef:Boolean` | fixed | 1 |
|
||||
| Enum | `TypeDef:Enum` | fixed | 4 (u32 index) |
|
||||
| String | `TypeDef:String` | variable | — |
|
||||
| Bytes | `TypeDef:Bytes` | variable | — |
|
||||
| Struct | `TypeDef:Struct` | composite | sum of fields |
|
||||
| Union | `TypeDef:Union` | composite | discriminator + variant |
|
||||
| Array | `TypeDef:Array` | composite | count × element |
|
||||
| Record | `TypeDef:Record` | variable | — |
|
||||
| Timestamp | `TypeDef:Timestamp` | variable | — |
|
||||
|
||||
Implement:
|
||||
|
||||
```rust
|
||||
/// Returns the `TypeDef:*` kind string if the schema node declares one.
|
||||
/// Returns `None` if the node has no `TypeDef:*` keyword.
|
||||
pub fn get_typedef_kind(node: &serde_json::Value) -> Option<&str>;
|
||||
|
||||
/// Returns the fixed byte size for a TypeDef kind, or `None` if variable-size.
|
||||
pub fn type_size(kind: &str) -> Option<usize>;
|
||||
|
||||
/// Returns the natural alignment for a TypeDef kind.
|
||||
pub fn natural_alignment(kind: &str) -> usize;
|
||||
|
||||
/// Returns true if the kind is a fixed-size type.
|
||||
pub fn is_fixed_size(kind: &str) -> bool;
|
||||
```
|
||||
|
||||
### Endian enum
|
||||
|
||||
```rust
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum Endian {
|
||||
Little,
|
||||
Big,
|
||||
}
|
||||
|
||||
impl Endian {
|
||||
/// Parse from the schema's "endian" annotation. Defaults to Little.
|
||||
pub fn from_schema(schema: &serde_json::Value) -> Self;
|
||||
}
|
||||
```
|
||||
|
||||
### Schema annotation parsing
|
||||
|
||||
Parse the annotations decided in ADR-097:
|
||||
|
||||
```rust
|
||||
/// Parse the "encoding" annotation from a variable-length type's keyword value.
|
||||
/// The keyword value may be `true` (shorthand for length-prefixed) or an object
|
||||
/// with an "encoding" field.
|
||||
pub fn parse_encoding(keyword_value: &serde_json::Value) -> VariableEncoding;
|
||||
|
||||
pub enum VariableEncoding {
|
||||
LengthPrefixed,
|
||||
OffsetIndirect,
|
||||
}
|
||||
|
||||
/// Parse the "align" annotation from a schema node. Returns None if not specified.
|
||||
pub fn parse_align(node: &serde_json::Value) -> Option<usize>;
|
||||
|
||||
/// Parse the "maxLength" annotation (standard JSON Schema keyword).
|
||||
pub fn parse_max_length(node: &serde_json::Value) -> Option<usize>;
|
||||
|
||||
/// Parse the "endian" annotation. Defaults to Little if absent or unrecognized.
|
||||
pub fn parse_endian(node: &serde_json::Value) -> Endian;
|
||||
```
|
||||
|
||||
### TUnion discriminator parsing
|
||||
|
||||
Parse the discriminator shape from a `TypeDef:Union` schema node:
|
||||
|
||||
```rust
|
||||
pub enum DiscriminatorKind {
|
||||
Byte { offset: usize, disc_type: String },
|
||||
Field { name: String },
|
||||
}
|
||||
|
||||
/// Parse the "discriminator" annotation from a TUnion schema node.
|
||||
pub fn parse_discriminator(node: &serde_json::Value) -> Result<DiscriminatorKind, TypedefError>;
|
||||
```
|
||||
|
||||
### `$ref` normalization
|
||||
|
||||
TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`). The `jsonschema`
|
||||
crate requires full JSON Pointer paths (e.g., `"$ref": "#/$defs/Read"`). Normalize
|
||||
at schema load time:
|
||||
|
||||
```rust
|
||||
/// Walk the schema tree. For every "$ref" whose value is a bare name
|
||||
/// (no "#" prefix), rewrite it to "#/$defs/<name>".
|
||||
/// Full JSON Pointer refs pass through unchanged. Idempotent.
|
||||
pub fn normalize_refs(schema: &mut serde_json::Value);
|
||||
```
|
||||
|
||||
This is a ~20-line recursive walk. It runs once at load time, before the schema is
|
||||
passed to `jsonschema::validator_for` or the offset computation.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- The actual offset computation (that's `offset_map.rs` and `layout_builder.rs`)
|
||||
- The read/write functions (that's `data_access.rs`)
|
||||
- The custom keyword validators (that's `validation.rs`)
|
||||
- The `TypedefEngine` struct (that's `engine.rs`)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `get_typedef_kind()` correctly identifies all 17 `TypeDef:*` kinds
|
||||
- [ ] `type_size()` returns correct byte sizes for all fixed-size types
|
||||
- [ ] `type_size()` returns `None` for variable-length and composite types
|
||||
- [ ] `natural_alignment()` returns correct alignment for each kind
|
||||
- [ ] `Endian` enum with `Little` and `Big` variants
|
||||
- [ ] `Endian::from_schema()` defaults to `Little` when annotation is absent
|
||||
- [ ] `Endian::from_schema()` returns `Big` when `"endian": "big"`
|
||||
- [ ] `parse_encoding()` handles both `true` (shorthand) and `{ "encoding": "..." }` (object)
|
||||
- [ ] `parse_encoding()` defaults to `LengthPrefixed` when annotation is absent
|
||||
- [ ] `parse_align()` returns `None` when no `"align"` annotation
|
||||
- [ ] `parse_max_length()` returns `None` when no `"maxLength"` annotation
|
||||
- [ ] `parse_discriminator()` handles byte-offset (`kind: "byte"`) and field-name (`kind: "field"`)
|
||||
- [ ] `parse_discriminator()` returns `Err(TypedefError::Schema(...))` for malformed discriminators
|
||||
- [ ] `normalize_refs()` rewrites `"$ref": "Read"` → `"$ref": "#/$defs/Read"`
|
||||
- [ ] `normalize_refs()` leaves `"$ref": "#/$defs/Read"` unchanged (idempotent)
|
||||
- [ ] `normalize_refs()` handles nested objects and arrays recursively
|
||||
- [ ] All public functions have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds, annotations, $ref normalization
|
||||
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (concrete JSON shapes)
|
||||
- docs/architecture/decisions/095-alknet-typedef-purpose-scope-jsonschema-engine.md — ADR-095
|
||||
- docs/research/alknet-typedef/findings.md — POC results, open questions
|
||||
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference for kind detection
|
||||
|
||||
## Notes
|
||||
|
||||
> This is the foundational types module. Every other module depends on it for
|
||||
> TypeDef kind identification, size lookups, and annotation parsing. The `$ref`
|
||||
> normalization bridges TypeBox output to jsonschema input — without it, bare-name
|
||||
> refs from TypeBox would fail to resolve. The `TEnum` kind uses a `u32` index
|
||||
> (not a variable-length string) — a deliberate deviation from TypeBox fidelity
|
||||
> for binary efficiency, per the schema-layer spec.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,181 @@
|
||||
---
|
||||
id: typedef/sequential-reader
|
||||
name: Implement packed sequential SequentialReader for protocol read-side
|
||||
status: pending
|
||||
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
|
||||
scope: moderate
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement the packed sequential `SequentialReader` in `crates/alknet-typedef/src/sequential_reader.rs`.
|
||||
This is the read-side of Mode 1 (ADR-096): walks a buffer field-by-field according to
|
||||
the schema, reading length prefixes to determine variable-length data positions. Used
|
||||
at read time when the consumer is parsing an incoming frame.
|
||||
|
||||
Per [layout-engine.md](../../docs/architecture/crates/typedef/layout-engine.md) §"Mode 1: Packed sequential".
|
||||
|
||||
### Target shape
|
||||
|
||||
```rust
|
||||
/// Walks a buffer field-by-field according to a schema, reading length
|
||||
/// prefixes to determine variable-length data positions. Used at read time
|
||||
/// when parsing incoming protocol frames.
|
||||
///
|
||||
/// The reader is sequential — it cannot jump to field N without reading
|
||||
/// fields 0..N-1 first. This is inherent to packed layouts where
|
||||
/// variable-length fields shift subsequent fields.
|
||||
#[derive(Debug)]
|
||||
pub struct SequentialReader {
|
||||
/// The schema being read.
|
||||
schema: serde_json::Value,
|
||||
/// The endianness for the layout.
|
||||
endian: Endian,
|
||||
}
|
||||
|
||||
/// A value read from a field during sequential traversal.
|
||||
#[derive(Debug)]
|
||||
pub enum FieldValue<'a> {
|
||||
I8(i8),
|
||||
I16(i16),
|
||||
I32(i32),
|
||||
U8(u8),
|
||||
U16(u16),
|
||||
U32(u32),
|
||||
U64(u64),
|
||||
F32(f32),
|
||||
F64(f64),
|
||||
Bool(bool),
|
||||
Enum(u32),
|
||||
String(&'a str),
|
||||
Bytes(&'a [u8]),
|
||||
/// A nested struct — the consumer recurses with a new SequentialReader
|
||||
/// scoped to the struct's byte range.
|
||||
Struct { start: usize, end: usize },
|
||||
/// A union — the consumer reads the discriminator, then recurses
|
||||
/// with the variant schema.
|
||||
Union { discriminator: String, variant_start: usize },
|
||||
/// An array — the consumer iterates elements.
|
||||
Array { count: u32, element_start: usize, element_stride: usize },
|
||||
}
|
||||
|
||||
impl SequentialReader {
|
||||
/// Create a new SequentialReader from a schema.
|
||||
pub fn new(schema: &serde_json::Value) -> Result<Self, TypedefError>;
|
||||
|
||||
/// Read the next field from the buffer at the current position.
|
||||
/// Returns the field name, the value, and advances the internal position.
|
||||
/// Returns `None` when all fields have been read.
|
||||
pub fn read_next<'a>(
|
||||
&mut self,
|
||||
buffer: &'a [u8],
|
||||
) -> Result<Option<(String, FieldValue<'a>)>, TypedefError>;
|
||||
|
||||
/// Read a specific field by path. This walks through all preceding fields
|
||||
/// to reach the target (sequential access is inherent to packed layouts).
|
||||
pub fn read_field<'a>(
|
||||
&mut self,
|
||||
buffer: &'a [u8],
|
||||
field_path: &str,
|
||||
) -> Result<FieldValue<'a>, TypedefError>;
|
||||
|
||||
/// Reset the reader to the beginning of the buffer.
|
||||
pub fn reset(&mut self);
|
||||
|
||||
/// The current byte position in the buffer.
|
||||
pub fn position(&self) -> usize;
|
||||
}
|
||||
```
|
||||
|
||||
### How it works
|
||||
|
||||
For a struct with fields `[u8, u32, string]`:
|
||||
|
||||
```
|
||||
SequentialReader:
|
||||
read_next() → ("field_0", FieldValue::U8(42)), position = 1
|
||||
read_next() → ("field_1", FieldValue::U32(1234)), position = 5
|
||||
read_next() → reads u32 length prefix at offset 5 → data_len = 10
|
||||
("field_2", FieldValue::String("hello worl")), position = 19
|
||||
read_next() → None (no more fields)
|
||||
```
|
||||
|
||||
The reader walks the buffer sequentially. It reads each field's type from the schema,
|
||||
reads the appropriate number of bytes at the current position, and advances. For
|
||||
variable-length fields, it reads the 4-byte length prefix to determine the data
|
||||
extent, then advances past the data.
|
||||
|
||||
### Nested structs
|
||||
|
||||
When the reader encounters a `TStruct` field, it returns `FieldValue::Struct { start, end }`.
|
||||
The consumer creates a new `SequentialReader` scoped to that byte range and reads the
|
||||
inner fields.
|
||||
|
||||
### TUnion
|
||||
|
||||
When the reader encounters a `TUnion` field:
|
||||
1. For byte-offset discriminators: reads the discriminator value at the known offset,
|
||||
looks up the variant schema, returns `FieldValue::Union { discriminator, variant_start }`.
|
||||
2. For field-name discriminators: reads the discriminator field like any other field,
|
||||
then returns the union value.
|
||||
|
||||
### TArray
|
||||
|
||||
When the reader encounters a `TArray` field:
|
||||
1. Reads the count (from `minItems`/`maxItems` if fixed, or from a 4-byte count prefix
|
||||
if variable).
|
||||
2. Returns `FieldValue::Array { count, element_start, element_stride }`.
|
||||
3. The consumer iterates elements using the stride.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- The write-side of packed mode (that's `layout_builder.rs`)
|
||||
- Aligned static layout (that's `offset_map.rs`)
|
||||
- TUnion discriminator dispatch (that's `tunion.rs`)
|
||||
- The `TypedefEngine` struct (that's `engine.rs`)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `SequentialReader` struct with `new(schema)` constructor
|
||||
- [ ] `read_next(buffer)` reads the next field and advances position
|
||||
- [ ] Fixed-size fields read correct number of bytes at current position
|
||||
- [ ] Variable-length fields: reads 4-byte length prefix, then skips data
|
||||
- [ ] Position advances correctly through all fields
|
||||
- [ ] `read_next()` returns `None` when all fields have been read
|
||||
- [ ] `read_field(field_path)` walks through preceding fields to reach target
|
||||
- [ ] `reset()` resets position to beginning
|
||||
- [ ] `position()` returns current byte offset
|
||||
- [ ] Nested structs: returns `FieldValue::Struct { start, end }` for consumer recursion
|
||||
- [ ] `TUnion`: reads discriminator, returns `FieldValue::Union { ... }`
|
||||
- [ ] `TArray`: reads count, returns `FieldValue::Array { count, element_start, element_stride }`
|
||||
- [ ] Endianness respected for all multi-byte reads
|
||||
- [ ] Returns `TypedefError::Access` for buffer-too-short
|
||||
- [ ] Returns `TypedefError::Schema` for malformed schemas
|
||||
- [ ] No `unwrap()` or `expect()` on error paths
|
||||
- [ ] All public types and functions have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/layout-engine.md — Mode 1: Packed sequential, SequentialReader
|
||||
- docs/architecture/crates/typedef/data-access.md — read/write functions used by the reader
|
||||
- docs/architecture/decisions/096-two-layout-modes-packed-vs-aligned.md — ADR-096
|
||||
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 (encoding, discriminators)
|
||||
- /workspace/alknet-typedef-poc/src/offset.rs — POC reference (SequentialReader in POC 2)
|
||||
|
||||
## Notes
|
||||
|
||||
> This is the read-side of the packed sequential mode. The reader walks the buffer
|
||||
> sequentially — it cannot jump to field N without reading fields 0..N-1 first.
|
||||
> This is inherent to packed layouts where variable-length fields shift subsequent
|
||||
> fields. The POC 2's `SequentialReader` is a good reference. The reader uses the
|
||||
> `data_access` module's read functions internally.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,192 @@
|
||||
---
|
||||
id: typedef/tests
|
||||
name: "Write comprehensive tests for alknet-typedef: unit tests, integration tests, and POC round-trip tests"
|
||||
status: pending
|
||||
depends_on: [typedef/engine]
|
||||
scope: moderate
|
||||
risk: low
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Write comprehensive tests for the `alknet-typedef` crate. Tests should cover all
|
||||
17 `TypeDef:*` kinds, both layout modes, endianness, TUnion dispatch, validation,
|
||||
and error paths. Include the POC-verified round-trip tests that validate
|
||||
byte-identical output against known-good serialization.
|
||||
|
||||
### Test categories
|
||||
|
||||
#### 1. Schema layer tests (`tests/schema_tests.rs` or `#[cfg(test)] mod tests` in `schema.rs`)
|
||||
|
||||
- `get_typedef_kind()` correctly identifies all 17 kinds
|
||||
- `type_size()` returns correct sizes for all fixed-size types
|
||||
- `type_size()` returns `None` for variable-length types
|
||||
- `natural_alignment()` returns correct alignment
|
||||
- `Endian::from_schema()` defaults to Little, parses "big" correctly
|
||||
- `parse_encoding()` handles `true`, `{ "encoding": "length-prefixed" }`, `{ "encoding": "offset-indirect" }`
|
||||
- `parse_align()` returns `None` when absent, correct value when present
|
||||
- `parse_max_length()` returns `None` when absent, correct value when present
|
||||
- `parse_discriminator()` handles byte-offset and field-name discriminators
|
||||
- `normalize_refs()` rewrites bare-name refs, leaves full paths unchanged, handles nested objects
|
||||
|
||||
#### 2. Data access tests (`tests/data_access_tests.rs`)
|
||||
|
||||
- Fixed-size read/write round-trip for all types: `i8`, `i16`, `i32`, `u8`, `u16`, `u32`, `u64`, `f32`, `f64`, `bool`, `enum`
|
||||
- Endianness: little-endian and big-endian produce correct byte order
|
||||
- `read_bool`: `0x00` → false, `0x01` → true, other values → error
|
||||
- `read_string`: correct length-prefixed read, UTF-8 validation, buffer-too-short error
|
||||
- `read_bytes`: correct length-prefixed read, buffer-too-short error
|
||||
- `write_string` / `write_bytes`: correct length prefix + data written, returns bytes written
|
||||
- `read_string_indirect` / `read_bytes_indirect`: correct offset-indirect read
|
||||
- Buffer bounds checking: all functions return `TypedefError::Access` for too-short buffers
|
||||
- Field path in error messages is correct
|
||||
|
||||
#### 3. OffsetMap tests (aligned static mode) (`tests/offset_map_tests.rs`)
|
||||
|
||||
- Simple struct: `{ u8, u32, f32 }` → correct offsets with natural alignment (0, 4, 8)
|
||||
- Nested struct: `{ header: { version: u32, magic: u32 }, payload: bytes }` → correct dotted paths
|
||||
- `TArray` of fixed-size elements: correct stride and element offsets
|
||||
- `TArray` with fixed count (`minItems == maxItems`): correct total size
|
||||
- `TArray` with variable count: length-prefixed encoding
|
||||
- Variable-length types (inline length-prefixing): 4-byte length prefix at known offset
|
||||
- Variable-length types (`maxLength`): reserved bytes at fixed offset
|
||||
- Variable-length types (offset-indirect): 8-byte `{offset, length}` struct
|
||||
- `TUnion` with byte-offset discriminator: discriminator at offset, variant at offset + disc_size
|
||||
- `TUnion` with field-name discriminator: discriminator is a regular field
|
||||
- Struct-level `"align"` annotation: rounds up total size
|
||||
- Field-level `"align"` annotation: overrides struct default
|
||||
- `OffsetMap::get()` returns correct `ByteRange` for dotted paths
|
||||
- `OffsetMap::total_size()` is correct
|
||||
- `OffsetMap::iter()` returns all fields
|
||||
|
||||
#### 4. LayoutBuilder tests (packed sequential mode, write-side) (`tests/layout_builder_tests.rs`)
|
||||
|
||||
- Simple struct: `{ u8, u32, string }` → packed offsets (0, 1, 5) with no padding
|
||||
- Variable-length fields shift subsequent fields based on actual sizes
|
||||
- Nested struct: correct dotted paths with packed offsets
|
||||
- `TArray` of fixed-size elements: correct stride (element size, no padding)
|
||||
- `TUnion` with byte-offset discriminator: correct discriminator and variant positions
|
||||
- Missing variable-length field size → `TypedefError::Offset`
|
||||
|
||||
#### 5. SequentialReader tests (packed sequential mode, read-side) (`tests/sequential_reader_tests.rs`)
|
||||
|
||||
- Read fields sequentially: correct values and position advancement
|
||||
- Variable-length fields: reads length prefix, skips data, advances correctly
|
||||
- `read_next()` returns `None` when all fields read
|
||||
- `read_field()` walks through preceding fields
|
||||
- `reset()` resets position
|
||||
- Nested struct: returns `FieldValue::Struct { start, end }`
|
||||
- `TUnion`: returns `FieldValue::Union { ... }`
|
||||
- `TArray`: returns `FieldValue::Array { count, element_start, element_stride }`
|
||||
- Endianness respected for multi-byte reads
|
||||
|
||||
#### 6. TUnion dispatch tests (`tests/tunion_tests.rs`)
|
||||
|
||||
- Byte-offset discriminator: reads u8 at offset 0, returns correct key and variant offset
|
||||
- Byte-offset discriminator with u16 and u32 types
|
||||
- Field-name discriminator: reads string field, returns correct key
|
||||
- Field-name discriminator with u8 and enum field types
|
||||
- `resolve_variant()`: resolves `$ref` pointers, returns error for unknown keys
|
||||
- `discriminator_size()`: correct for each discriminator type
|
||||
|
||||
#### 7. Validation tests (`tests/validation_tests.rs`)
|
||||
|
||||
- `build_validator()` returns a working validator
|
||||
- Each `TypeDef:*` kind's validator rejects invalid values
|
||||
- `Float32Validator`: rejects NaN, Infinity
|
||||
- `Int8Validator`: rejects 128, -129
|
||||
- `Uint8Validator`: rejects -1, 256
|
||||
- `StringValidator`: rejects non-string, respects `maxLength`
|
||||
- `TimestampValidator`: rejects non-RFC 3339 strings
|
||||
- `StructValidator`: validates nested fields
|
||||
- `UnionValidator`: validates discriminator membership
|
||||
- `ArrayValidator`: validates element types and length bounds
|
||||
|
||||
#### 8. TypedefEngine integration tests (`tests/engine_tests.rs`)
|
||||
|
||||
- `compile()` in aligned mode produces a working engine
|
||||
- `compile()` in packed mode produces a working engine
|
||||
- `compile()` normalizes `$ref` values
|
||||
- `compile()` returns error for invalid schemas
|
||||
- `endian()` and `mode()` accessors
|
||||
- `offset_map()`, `layout_builder()`, `sequential_reader()` accessors
|
||||
- `validate_json()` and `is_valid_json()` work correctly
|
||||
- `read_field()` and `write_field()` in aligned mode
|
||||
|
||||
#### 9. POC round-trip tests (`tests/poc_roundtrip_tests.rs`)
|
||||
|
||||
Replicate the key POC tests from `/workspace/alknet-typedef-poc/`:
|
||||
|
||||
- **Fixed-size round-trip**: Write u8, u16, u32, u64, f32, f64, bool, enum to a buffer
|
||||
at computed offsets; read back; verify values match.
|
||||
- **String round-trip**: Write a length-prefixed string; read back; verify.
|
||||
- **Nested struct round-trip**: Write a struct with nested fields; read back; verify.
|
||||
- **Endianness round-trip**: Write in big-endian; read back in big-endian; verify.
|
||||
- **SFTP packet round-trip** (if SFTP schemas are available): Write an SFTP Read/Write/Status
|
||||
packet; verify byte-identical output against known-good serialization.
|
||||
|
||||
#### 10. Error path tests
|
||||
|
||||
- Buffer too short for every read function
|
||||
- Invalid UTF-8 in `read_string`
|
||||
- Invalid boolean byte value
|
||||
- Missing required schema keywords
|
||||
- Unknown `TypeDef:*` kind
|
||||
- Malformed discriminator annotation
|
||||
- Unknown discriminator value
|
||||
- Missing variable-length field size in `LayoutBuilder::build()`
|
||||
|
||||
### Test organization
|
||||
|
||||
Tests can be organized as:
|
||||
- Unit tests in `#[cfg(test)] mod tests` blocks within each source file (for
|
||||
module-internal functions).
|
||||
- Integration tests in `tests/` directory (for public API tests that exercise
|
||||
multiple modules).
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- Tests for arrays of variable-length-element structs (deferred, OQ-069)
|
||||
- WASM-specific tests (deferred, OQ-070)
|
||||
- Performance benchmarks
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] Schema layer tests: all `schema.rs` public functions tested
|
||||
- [ ] Data access tests: all read/write functions tested with round-trip, endianness, and error paths
|
||||
- [ ] OffsetMap tests: aligned static layout with alignment, nesting, arrays, unions, variable-length types
|
||||
- [ ] LayoutBuilder tests: packed sequential layout with variable-length shifting
|
||||
- [ ] SequentialReader tests: sequential field reading with position tracking
|
||||
- [ ] TUnion tests: both discriminator kinds, variant resolution
|
||||
- [ ] Validation tests: all 17 custom keyword validators tested
|
||||
- [ ] Engine tests: compile, accessors, convenience methods
|
||||
- [ ] POC round-trip tests: fixed-size, string, nested struct, endianness
|
||||
- [ ] Error path tests: buffer-too-short, invalid UTF-8, malformed schemas
|
||||
- [ ] All tests pass: `cargo test -p alknet-typedef`
|
||||
- [ ] No `unwrap()` in test code that would mask failures (use `?` or `assert!(matches!(...))`)
|
||||
- [ ] `cargo clippy -p alknet-typedef --all-targets` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
- [ ] `cargo test --workspace` still succeeds (old tests untouched)
|
||||
|
||||
## References
|
||||
|
||||
- docs/research/alknet-typedef/findings.md — POC results (26 tests passing)
|
||||
- /workspace/alknet-typedef-poc/src/lib.rs — POC test reference
|
||||
- docs/architecture/crates/typedef/README.md — design principles
|
||||
- All preceding task files in tasks/typedef/
|
||||
|
||||
## Notes
|
||||
|
||||
> This is the testing task — it validates the entire crate. The POC had 26 tests
|
||||
> passing; this task should cover at least that many scenarios plus additional
|
||||
> error-path tests. Tests should be organized by module/concern. The SFTP round-trip
|
||||
> test is a stretch goal — it requires SFTP packet schemas which may not be in the
|
||||
> repo yet. If SFTP schemas aren't available, test with hand-crafted schemas that
|
||||
> exercise the same patterns (byte-offset TUnion, big-endian, mixed fixed/variable
|
||||
> fields).
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,145 @@
|
||||
---
|
||||
id: typedef/tunion
|
||||
name: Implement TUnion discriminator dispatch for byte-offset and field-name discriminators
|
||||
status: pending
|
||||
depends_on: [typedef/schema-types, typedef/error-type, typedef/data-access]
|
||||
scope: narrow
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement TUnion discriminator dispatch in `crates/alknet-typedef/src/tunion.rs`.
|
||||
TUnion supports two discriminator kinds (ADR-097 §4): byte-offset (protocol dispatch,
|
||||
e.g., SFTP type bytes) and field-name (typedef.ts string pattern).
|
||||
|
||||
Per [data-access.md](../../docs/architecture/crates/typedef/data-access.md) §"TUnion Dispatch"
|
||||
and [schema-layer.md](../../docs/architecture/crates/typedef/schema-layer.md) §"TUnion discriminators".
|
||||
|
||||
### Target shape
|
||||
|
||||
```rust
|
||||
/// The result of reading a TUnion discriminator.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct UnionDispatch {
|
||||
/// The mapping key (stringified discriminator value).
|
||||
pub key: String,
|
||||
/// The byte offset where the variant struct starts.
|
||||
pub variant_offset: usize,
|
||||
/// The size of the discriminator in bytes.
|
||||
pub discriminator_size: usize,
|
||||
}
|
||||
|
||||
/// Read the discriminator value from a byte-offset TUnion.
|
||||
/// The discriminator is a fixed-size integer at a known byte offset.
|
||||
/// Returns the mapping key (as a string) and the variant struct offset.
|
||||
///
|
||||
/// This is the SFTP `Packet` enum pattern — byte 0 is the type byte,
|
||||
/// bytes 1..N are the variant struct. The call protocol's event type
|
||||
/// dispatch uses the same pattern.
|
||||
pub fn read_byte_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &serde_json::Value,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, TypedefError>;
|
||||
|
||||
/// Read the discriminator value from a field-name TUnion.
|
||||
/// The discriminator is a named field within the struct — its offset
|
||||
/// is computed like any other field. The consumer provides the
|
||||
/// discriminator field's offset (from the OffsetMap or LayoutBuilder).
|
||||
///
|
||||
/// This is the typedef.ts `TUnion` pattern — the discriminator is a
|
||||
/// field like any other, and the mapping keys are string values.
|
||||
pub fn read_field_discriminator(
|
||||
buffer: &[u8],
|
||||
union_schema: &serde_json::Value,
|
||||
disc_field_offset: usize,
|
||||
endian: Endian,
|
||||
) -> Result<UnionDispatch, TypedefError>;
|
||||
|
||||
/// Look up a variant schema from the union's mapping.
|
||||
/// Returns the variant schema (resolving `$ref` if needed).
|
||||
pub fn resolve_variant<'a>(
|
||||
union_schema: &'a serde_json::Value,
|
||||
key: &str,
|
||||
) -> Result<&'a serde_json::Value, TypedefError>;
|
||||
|
||||
/// Get the discriminator size in bytes for a byte-offset discriminator.
|
||||
pub fn discriminator_size(union_schema: &serde_json::Value) -> Result<usize, TypedefError>;
|
||||
```
|
||||
|
||||
### Byte-offset discriminator
|
||||
|
||||
The discriminator is a fixed-size integer at a known byte offset. The mapping keys
|
||||
are stringified integers (`"5"`, `"6"`, `"101"`).
|
||||
|
||||
Supported discriminator types: `TypeDef:Uint8` (1 byte), `TypeDef:Uint16` (2 bytes),
|
||||
`TypeDef:Uint32` (4 bytes). The discriminator value is read using the appropriate
|
||||
endian-aware read function from `data_access`.
|
||||
|
||||
The variant struct starts at `offset + discriminator_size`.
|
||||
|
||||
### Field-name discriminator
|
||||
|
||||
The discriminator is a named field within the struct. Its offset is computed like any
|
||||
other field (by the `OffsetMap` or `LayoutBuilder`). The mapping keys are string
|
||||
values matching the discriminator field's value.
|
||||
|
||||
The discriminator field's `TypeDef:*` kind determines how to read it:
|
||||
- `TypeDef:String` → read a length-prefixed string
|
||||
- `TypeDef:Uint8` → read a u8, stringify
|
||||
- `TypeDef:Enum` → read a u32 index, map to string value from `"enum"` array
|
||||
|
||||
### Variant resolution
|
||||
|
||||
`resolve_variant()` looks up the mapping key in the union's `"mapping"` object.
|
||||
Mapping values may be inline schemas or `$ref` pointers. `$ref` pointers are resolved
|
||||
against the schema's `$defs` (the `$ref` normalization in `schema.rs` ensures they
|
||||
are full JSON Pointer paths).
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- The offset computation for TUnion (that's in `offset_map.rs` and `layout_builder.rs`)
|
||||
- The `TypedefEngine` struct (that's `engine.rs`)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `read_byte_discriminator()` reads discriminator at known byte offset
|
||||
- [ ] Supports `TypeDef:Uint8`, `TypeDef:Uint16`, `TypeDef:Uint32` discriminator types
|
||||
- [ ] Discriminator value read with correct endianness
|
||||
- [ ] Returns mapping key as string (e.g., `"5"` for SFTP Read packet)
|
||||
- [ ] Returns correct `variant_offset` (offset + discriminator_size)
|
||||
- [ ] `read_field_discriminator()` reads discriminator from named field at given offset
|
||||
- [ ] Supports `TypeDef:String`, `TypeDef:Uint8`, `TypeDef:Enum` discriminator field types
|
||||
- [ ] `resolve_variant()` looks up mapping key and returns variant schema
|
||||
- [ ] `resolve_variant()` resolves `$ref` pointers to `$defs`
|
||||
- [ ] `resolve_variant()` returns `TypedefError::Schema` for unknown mapping keys
|
||||
- [ ] `discriminator_size()` returns correct size for each discriminator type
|
||||
- [ ] Returns `TypedefError::Access` for buffer-too-short
|
||||
- [ ] Returns `TypedefError::Schema` for malformed discriminator annotations
|
||||
- [ ] No `unwrap()` or `expect()` on error paths
|
||||
- [ ] All public functions have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/data-access.md — TUnion dispatch section
|
||||
- docs/architecture/crates/typedef/schema-layer.md — TUnion discriminators section
|
||||
- docs/architecture/decisions/097-schema-annotations.md — ADR-097 §4 (TUnion discriminators)
|
||||
- /workspace/alknet-typedef-poc/src/lib.rs — POC reference (parse_union_discriminator, read_union_discriminator)
|
||||
|
||||
## Notes
|
||||
|
||||
> This is a focused module for TUnion discriminator dispatch. The two discriminator
|
||||
> kinds cover both protocol dispatch (SFTP type bytes, call protocol event types) and
|
||||
> the typedef.ts string pattern. The mapping keys are always strings — integer
|
||||
> discriminator values are stringified. The POC 2's `parse_union_discriminator` and
|
||||
> `read_union_discriminator` functions are good references.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
@@ -0,0 +1,180 @@
|
||||
---
|
||||
id: typedef/validation
|
||||
name: Implement custom keyword validators for all 17 TypeDef kinds via jsonschema with_keyword API
|
||||
status: pending
|
||||
depends_on: [typedef/schema-types, typedef/error-type]
|
||||
scope: moderate
|
||||
risk: medium
|
||||
impact: component
|
||||
level: implementation
|
||||
---
|
||||
|
||||
## Description
|
||||
|
||||
Implement custom keyword validators for all 17 `TypeDef:*` kinds in
|
||||
`crates/alknet-typedef/src/validation.rs`. Each kind gets a `Keyword` implementation
|
||||
registered via `jsonschema::options().with_keyword(...)`. The validators check leaf
|
||||
type constraints; `jsonschema` handles all structural validation.
|
||||
|
||||
Per [validation.md](../../docs/architecture/crates/typedef/validation.md) and
|
||||
[ADR-098](../../docs/architecture/decisions/098-error-handling-validation-strategy.md).
|
||||
|
||||
### Target shape
|
||||
|
||||
```rust
|
||||
use jsonschema::{Keyword, ValidationError};
|
||||
use serde_json::Value;
|
||||
|
||||
/// Build a jsonschema validator with all 17 TypeDef:* custom keywords registered.
|
||||
/// The returned validator can validate JSON representations of data against
|
||||
/// the schema's type constraints.
|
||||
pub fn build_validator(schema: &Value) -> Result<jsonschema::Validator, TypedefError>;
|
||||
```
|
||||
|
||||
### Custom keyword validators
|
||||
|
||||
Each `TypeDef:*` kind gets a struct implementing `Keyword`:
|
||||
|
||||
**Numeric validators:**
|
||||
|
||||
- `Float32Validator` / `Float64Validator` — value must be a finite number.
|
||||
For `Float32`: value must be representable as `f32` (no precision loss beyond
|
||||
`f32`'s mantissa).
|
||||
- `Int8Validator` / `Int16Validator` / `Int32Validator` — value must be an integer
|
||||
within the type's range. Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
|
||||
- `Uint8Validator` / `Uint16Validator` / `Uint32Validator` — value must be a
|
||||
non-negative integer within the type's range. Uint8: 0..255, Uint16: 0..65535,
|
||||
Uint32: 0..4294967295.
|
||||
|
||||
**String and binary validators:**
|
||||
|
||||
- `StringValidator` — value must be a valid UTF-8 string. If `maxLength` is
|
||||
specified in the parent schema, the string's byte length must not exceed it.
|
||||
- `BytesValidator` — value must be a string (JSON represents binary data as a
|
||||
string). If `maxLength` is specified, the byte length must not exceed it.
|
||||
- `EnumValidator` — the `TypeDef:Enum` custom keyword signals that the type is an
|
||||
enum for *layout* purposes. The built-in `enum` keyword handles value-membership
|
||||
validation. The custom keyword validator is a no-op beyond the built-in check —
|
||||
it exists solely for the layout engine to recognize the type.
|
||||
- `TimestampValidator` — value must be a valid RFC 3339 timestamp string (the
|
||||
internet profile of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
|
||||
|
||||
**Composite validators:**
|
||||
|
||||
- `StructValidator` — value must be an object. Each property must match its
|
||||
declared `TypeDef:*` kind. The `jsonschema` crate's built-in `properties` and
|
||||
`required` keywords handle structural checks — the custom keyword only validates
|
||||
that each field's value matches its `TypeDef:*` kind.
|
||||
- `UnionValidator` — the discriminator value must be one of the mapping keys.
|
||||
The variant struct must match the declared schema for that discriminator value.
|
||||
- `ArrayValidator` — value must be an array. Each element must match the array's
|
||||
declared element type. If `minItems`/`maxItems` is specified, the array length
|
||||
must be within bounds.
|
||||
|
||||
**Other validators:**
|
||||
|
||||
- `BooleanValidator` — value must be `true` or `false`.
|
||||
- `RecordValidator` — value must be an object. All values must match the record's
|
||||
declared value type (specified via the `"values"` property).
|
||||
|
||||
### Validator implementation pattern
|
||||
|
||||
Each validator is ~10 lines. Example:
|
||||
|
||||
```rust
|
||||
struct Float32Validator;
|
||||
|
||||
impl Keyword for Float32Validator {
|
||||
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
|
||||
match instance {
|
||||
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
|
||||
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
|
||||
}
|
||||
}
|
||||
fn is_valid(&self, instance: &Value) -> bool {
|
||||
instance.as_f64().map_or(false, |f| f.is_finite())
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Registration
|
||||
|
||||
```rust
|
||||
pub fn build_validator(schema: &Value) -> Result<jsonschema::Validator, TypedefError> {
|
||||
jsonschema::options()
|
||||
.with_keyword("TypeDef:Float32", |_parent, _value, _path| {
|
||||
Ok(Box::new(Float32Validator))
|
||||
})
|
||||
.with_keyword("TypeDef:Float64", |_parent, _value, _path| {
|
||||
Ok(Box::new(Float64Validator))
|
||||
})
|
||||
.with_keyword("TypeDef:Int8", |_parent, _value, _path| {
|
||||
Ok(Box::new(Int8Validator))
|
||||
})
|
||||
// ... all 17 kinds
|
||||
.with_keyword("TypeDef:Struct", |parent, _value, _path| {
|
||||
Ok(Box::new(StructValidator::from_schema(parent)?))
|
||||
})
|
||||
.build(schema)
|
||||
.map_err(|e| TypedefError::Schema(format!("validator build failed: {e}")))
|
||||
}
|
||||
```
|
||||
|
||||
The factory closure receives the parent schema object, the keyword's value, and the
|
||||
schema path. This enables cross-keyword awareness — for example, `StructValidator`
|
||||
inspects the parent's `properties` to validate each field against its declared
|
||||
`TypeDef:*` kind.
|
||||
|
||||
### What this does NOT include
|
||||
|
||||
- The `TypedefEngine` struct (that's `engine.rs`) — though `build_validator()` is
|
||||
called by the engine
|
||||
- Schema parsing or annotation extraction (that's `schema.rs`)
|
||||
- Binary-level validation (that's in `data_access.rs` — bounds checking, UTF-8
|
||||
validation at read time)
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
- [ ] `build_validator(schema)` returns a `jsonschema::Validator` with all 17 custom keywords registered
|
||||
- [ ] `Float32Validator` / `Float64Validator`: rejects non-finite numbers
|
||||
- [ ] `Int8Validator` / `Int16Validator` / `Int32Validator`: rejects out-of-range values
|
||||
- [ ] `Uint8Validator` / `Uint16Validator` / `Uint32Validator`: rejects negative and out-of-range values
|
||||
- [ ] `StringValidator`: rejects non-string values; respects `maxLength` from parent schema
|
||||
- [ ] `BytesValidator`: rejects non-string values; respects `maxLength`
|
||||
- [ ] `EnumValidator`: no-op (built-in `enum` keyword handles validation)
|
||||
- [ ] `TimestampValidator`: rejects non-RFC 3339 strings
|
||||
- [ ] `StructValidator`: validates each property against its declared `TypeDef:*` kind
|
||||
- [ ] `UnionValidator`: validates discriminator membership and variant conformance
|
||||
- [ ] `ArrayValidator`: validates element type conformance and length bounds
|
||||
- [ ] `BooleanValidator`: rejects non-boolean values
|
||||
- [ ] `RecordValidator`: validates all values match declared value type
|
||||
- [ ] Each validator is ~10 lines (not hundreds)
|
||||
- [ ] Factory closures handle errors gracefully (return `TypedefError::Schema`)
|
||||
- [ ] No `unwrap()` or `expect()` in factory closures
|
||||
- [ ] All validators have doc comments
|
||||
- [ ] `cargo check -p alknet-typedef` succeeds
|
||||
- [ ] `cargo clippy -p alknet-typedef` succeeds with no warnings
|
||||
- [ ] `cargo build --workspace` still succeeds
|
||||
|
||||
## References
|
||||
|
||||
- docs/architecture/crates/typedef/validation.md — custom keyword validators, TypedefError, validation timing
|
||||
- docs/architecture/crates/typedef/schema-layer.md — the 17 TypeDef kinds
|
||||
- docs/architecture/decisions/098-error-handling-validation-strategy.md — ADR-098
|
||||
- docs/research/alknet-typedef/findings.md — POC validation results
|
||||
- /workspace/alknet-typedef-poc/src/validate.rs — POC reference for validators
|
||||
- /workspace/jsonschema/ — the jsonschema crate API
|
||||
|
||||
## Notes
|
||||
|
||||
> This module registers custom keyword validators with the `jsonschema` crate. Each
|
||||
> validator is small (~10 lines) because `jsonschema` handles all structural
|
||||
> validation. The `StructValidator` is the most complex — it needs to inspect the
|
||||
> parent schema's `properties` to validate each field. The `EnumValidator` is a
|
||||
> no-op because the built-in `enum` keyword already handles value-membership
|
||||
> validation; the custom keyword exists solely for the layout engine to recognize
|
||||
> the type as a fixed-size u32 index. The POC's `validate.rs` is a good reference.
|
||||
|
||||
## Summary
|
||||
|
||||
> To be filled on completion
|
||||
Reference in new issue
Block a user