Port alknet-typedef crate from alknet

Copy the binary struct engine (src/, tests/) verbatim from
alknet/crates/alknet-typedef and create a standalone Cargo.toml
(workspace-inherited fields inlined). Port the architecture docs
(specs, ADRs 095-102, OQs 069-071) from alknet's nested multi-crate
layout to a flat single-crate layout, fixing relative link paths.

Build, 295 tests, and clippy all pass clean.
This commit is contained in:
glm-5.2 committed 2026-08-02 05:59:12 +00:00
1 parent eac7ad88b3
commit 2c4a4994dc
36 files changed
+13805

No files matched your search

@@ -0,0 +1,173 @@
# ADR-095: alknet-typedef — Purpose, Scope, and the jsonschema Engine
## Status
Accepted
## Context
Three threads in the codebase converge on the same pattern: a JSON Schema
describes the shape of binary data, and the binary data is the struct's
bytes at computed offsets.
1. **typedef.ts** (`/workspace/@alkdev/typebox/example/typedef/typedef.ts`,
619 lines) defines custom TypeBox schema kinds (`TFloat32`, `TStruct`,
`TUnion`, etc.) that carry binary layout semantics. These are registered
via `TypeRegistry.Set` with custom validators.
2. **russh-sftp** has 29 packet types, each a struct with typed fields
(`Read { id: u32, handle: String, offset: u64, len: u32 }`). The wire
format is `[length: u32][type: u8][payload]` where payload is the
struct's serde bytes. The `Packet` enum dispatches on the type byte —
a tagged union of structs. Under the typedef lens, each packet is a
`TStruct`; the `Packet` enum is a `TUnion` with a byte-offset
discriminator.
3. **metatensor** needs an offset map for mmap-friendly tensor access —
given a schema describing a model layout (ConvNet struct, tensor refs),
compute byte offsets for each field so the consumer can read tensor
data at known positions without parsing.
The common pattern: **a JSON Schema with `TypeDef:*` custom keywords
describes the shape of binary data; the binary data is the struct's bytes
at computed offsets.** The schema is the format definition; the engine is
generic.
Two prior attempts built their own jsonschema engines — the fatal flaw:
- **typebox-rs** (`/workspace/@alkimiadev/typebox-rs/`, ~8,400 lines):
a full 26-variant `SchemaKind` enum, a custom `Value` type with typed
arrays, and a 912-line hand-written validator.
- **alktype** (`/workspace/@alkimiadev/alktype/`, ~5,600 lines): a
handler-registry pattern that also implements its own validation for
each type.
The `jsonschema` crate (v0.46.5, Draft 2020-12) is already in the
workspace at `/workspace/jsonschema/`. It handles validation with custom
keyword support — the novel code is the offset computation, not the
validation.
The call-channels-unification research
(`docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine") identified the convergence and
bumped typedef up in the timeline. The POC
(`docs/research/alknet-typedef/findings.md`, 26 tests passing) validated
the approach: a ~1,900-line Rust crate that takes a JSON Schema with
`TypeDef:*` custom keywords and produces an offset map, read/write
functions, and validation — all driven by the schema.
## Decision
**alknet-typedef is a small Rust crate that takes a JSON Schema with
`TypeDef:*` custom keywords and produces three capabilities:**
1. **An offset map** — walks the schema, computes byte offsets for each
field based on type sizes, field order, and alignment.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that data
conforms to the schema's type constraints. The jsonschema validator
operates on `serde_json::Value` instances (JSON representations), not
raw byte buffers directly. A consumer that wants to validate a binary
buffer reads it into a `Value` tree via the data access layer, then
validates that `Value` against the jsonschema validator.
**The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing).** The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are ~10 lines each.
**The schema is the format.** A JSON Schema with `TypeDef:Float32`,
`TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
the layout spec. No separate format definition, no separate parser, no
separate validator. One schema, three uses: validate, compute offsets,
access data.
**The crate depends on `jsonschema` and `serde_json` (with
`preserve_order`).** No tokio, no platform deps. Compiles to
`wasm32-unknown-unknown` for browser use. The `jsonschema` crate's
`with_keyword("TypeDef:Float32", factory)` API is the integration point
for custom type kinds — each `TypeDef:*` kind maps to a custom keyword
validator in Rust. Same semantics as TypeBox's `TypeRegistry.Set`, same
JSON Schema wire format.
**The crate targets `std` for v1.** The WASM target has `std` available
via `wasm-bindgen`. If embedded use cases emerge, `no_std` + `alloc` can
be added as a feature gate later — the engine's core (offset computation,
read/write) is already allocation-free. See OQ-070.
## Consequences
### Positive
- **Eliminates ~14,000 lines of hand-rolled schema engines.** typebox-rs
and alktype are replaced by `jsonschema` + an offset map + ~50 lines of
custom keyword implementations. The codebase drops from "a port of
TypeBox" to "jsonschema + an offset map."
- **One schema, three uses.** The same JSON Schema validates, computes
offsets, and drives data access. No separate format definition, parser,
or validator per protocol.
- **Schema-driven, not code-driven.** Adding a new SFTP packet type is
adding a variant to the schema JSON, not writing a new Rust struct +
serde impl. The engine is generic; the schema is the configuration.
- **WASM-clean.** `serde_json` + `jsonschema` + byte manipulation. No
tokio, no platform deps. The same typedef schemas work in browser,
Node, Python (via `wasmtime-py`), Go (via `wazero`), and any other
WASM host.
- **TypeBox interop.** TypeBox modules render to standard JSON Schema
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
on the Rust side. Zero translation. The same schema validates in both
ecosystems.
- **Defense in depth.** Schema validation via jsonschema custom keywords —
a malformed binary payload can be read into a `Value` tree via the data
access layer and validated against the schema before any consumer
touches it. The `jsonschema` crate's compiled validators are fast enough
to run on every incoming frame.
### Negative
- **New dependency on `jsonschema`.** The crate is already in the
workspace but not yet used by any alknet crate. This is the first
consumer.
- **`serde_json` with `preserve_order` is required.** Field order is
load-bearing for binary layouts. The `preserve_order` feature adds a
small compile-time cost.
- **Schema authoring is external.** Schemas are authored in TypeBox (JS)
or hand-written JSON. The typedef engine consumes schemas; it does not
generate them. A builder API is deferred (OQ-071).
## Scope Boundaries (What This Is Not)
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
is a *format* (8-byte header + JSON header + binary data) that uses the
typedef engine for its offset computation and tensor access.
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
`Value.Convert` — schema evolution — is out of scope for v1. The engine
should not do anything that explicitly blocks adding a Value system
later.
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
concern. The typedef engine consumes schemas; it does not generate them.
- **Not a schema builder.** The typedef engine does not provide a fluent
API for constructing schemas. Schemas are plain JSON.
- **Not a serialization framework.** The typedef engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use typedef.
## References
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
passing, two layout modes, TUnion dispatch, endianness)
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
JSON Schema as the binary struct engine" — the origin of this research
thread
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes decision
- [ADR-097](097-schema-annotations.md) — schema annotation shapes
- [ADR-098](098-error-handling-validation-strategy.md) — error handling
and validation strategy
@@ -0,0 +1,137 @@
# ADR-096: Two Layout Modes — Packed Sequential vs Aligned Static
## Status
Accepted
## Context
The POCs surfaced that protocols and mmap-friendly formats need different
layout strategies. POC 1 built an aligned `OffsetMap` with natural
alignment padding — correct for mmap-friendly formats (metatensor) but
wrong for protocol wire formats (SFTP, channels, TTY). POC 2 built a
`LayoutBuilder` and `SequentialReader` for packed sequential layouts —
correct for protocol wire formats but wrong for mmap-friendly formats.
This is the most important architectural finding from the POCs. The
engine must support both modes; a single layout strategy cannot serve
both use cases.
### Packed sequential layout (protocol wire formats)
Protocols pack fields sequentially with no alignment padding.
Variable-length fields shift all subsequent fields. Writing requires
knowing actual data sizes upfront; reading walks the buffer sequentially,
reading length prefixes to determine positions.
This is the layout used by SFTP (all strings and byte arrays are
length-prefixed inline), channels (`[channel_id: u32][size: u32][payload]`),
TTY (`[stream_type: u8][length: u32][payload]`), and most binary protocols.
### Aligned static layout (mmap-friendly formats)
Fields have fixed positions with natural alignment padding.
Variable-length fields get a 4-byte length prefix at a known offset; the
variable data is not included in the static layout. This enables
mmap-friendly random access — the consumer can read field N at a known
offset without parsing the fields before it.
This is the layout used by metatensor (blob tensor pattern: index struct
in one region, blob data in another) and safetensors (header + aligned
tensor data).
## Decision
**The typedef engine supports two layout modes, selected by the consumer
at engine construction time:**
### Mode 1: Packed sequential (`LayoutBuilder` / `SequentialReader`)
For protocol wire formats. Fields are packed with no alignment padding.
Variable-length fields shift all subsequent fields.
- **LayoutBuilder** — takes a schema and actual data sizes for
variable-length fields, computes byte positions for each field in a
packed layout. Used at write time when the consumer knows the data
sizes upfront.
- **SequentialReader** — walks a buffer field-by-field according to the
schema, reading length prefixes to determine variable-length data
positions. Used at read time when the consumer is parsing an incoming
frame.
The `LayoutBuilder` and `SequentialReader` are the primary interface for
protocol consumers (SFTP, binary call frames, TTY negotiation).
### Mode 2: Aligned static (`OffsetMap`)
For mmap-friendly formats. Fields have fixed positions with natural
alignment padding. Variable-length fields get a 4-byte length prefix at
a known offset; the variable data is not included in the static layout.
- **OffsetMap** — walks the schema once, computes fixed byte positions
for each field based on type sizes and alignment. The output is a flat
table of `(field_path, byte_range)` pairs. Used for both read and write
at known offsets.
The `OffsetMap` is the primary interface for mmap consumers (metatensor).
### Variable-length handling in each mode
**Packed sequential mode:** Variable-length fields are inline
length-prefixed by default (`[length: u32][data]`). The `LayoutBuilder`
takes the actual data size to compute the length prefix value and the
position of subsequent fields. The `SequentialReader` reads the length
prefix to determine the data extent and the position of the next field.
**Aligned static mode:** Variable-length fields get a 4-byte length
prefix at a known offset. The variable data lives outside the static
layout — either immediately after the fixed fields (inline
length-prefixing) or in a separate data region (offset indirection, the
metatensor blob tensor pattern). The `OffsetMap` records the position of
the length prefix (or the `{offset, length}` pair for offset-indirect
fields).
### Default for variable-length types
Inline length-prefixing (`[length: u32][data]`) is the default for all
variable-length types in both modes. This is the universal pattern used
by channels, SFTP, TTY, and most binary protocols. Offset indirection is
opt-in via the `encoding` annotation (see ADR-097).
## Consequences
### Positive
- **One engine, two modes.** The same schema can be used in either mode.
A schema describing an SFTP packet can be consumed by a `SequentialReader`
(for parsing incoming frames) and a `LayoutBuilder` (for constructing
outgoing frames). A schema describing a metatensor layout can be
consumed by an `OffsetMap` (for mmap access).
- **Correct for both use cases.** Packed sequential mode produces
byte-identical output to hand-written protocol serialization (validated
by POC 2's russh-sftp round-trip tests). Aligned static mode produces
correct offsets for mmap-friendly access (validated by POC 1's
alignment tests).
- **No mode confusion.** The consumer explicitly selects the mode at
engine construction time. A protocol consumer never accidentally gets
alignment padding; an mmap consumer never accidentally gets
variable-length field shifting.
### Negative
- **Two APIs to learn.** Consumers must choose between
`LayoutBuilder`/`SequentialReader` and `OffsetMap`. The choice is
determined by the use case (protocol vs mmap), not by the schema.
- **Variable-length fields in packed mode require size foreknowledge.**
The `LayoutBuilder` needs actual data sizes for variable-length fields
to compute correct positions for subsequent fields. This is inherent
to packed layouts — the consumer must know the data sizes before
writing.
## References
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-097](097-schema-annotations.md) — schema annotations including
the `encoding` field for variable-length types
@@ -0,0 +1,259 @@
# ADR-097: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
## Status
Accepted
## Context
The typedef engine needs concrete JSON shapes for schema-level
annotations that control binary layout behavior. The POCs validated the
semantics; this ADR pins the shapes.
Four annotation categories need concrete shapes:
1. **Endianness** — safetensors is little-endian, SFTP is big-endian.
The engine needs to know which to use.
2. **Alignment** — different backends have different alignment
requirements (wgpu: 256-byte, protocols: natural, mmap: page).
3. **Variable-length encoding** — inline length-prefixing vs offset
indirection for strings, byte arrays, and other variable-length types.
4. **TUnion discriminators** — byte-offset (protocol dispatch) vs
field-name (typedef.ts pattern).
## Decision
### 1. Endianness
**Schema-level annotation with a default of little-endian.**
```json
{
"TypeDef:Struct": true,
"endian": "big",
"properties": { ... }
}
```
- `"endian": "little"` (default) — read/write in little-endian byte order.
- `"endian": "big"` — read/write in big-endian byte order.
- The annotation applies to the entire schema and all nested types.
- Mixed endianness within one schema is not supported (pathological; no
known protocol requires it).
- The default is little-endian, matching safetensors, wgpu, and most
modern formats. SFTP consumers specify `"endian": "big"`.
### 2. Alignment
**Both struct-level and field-level, with field-level overriding
struct-level.**
```json
{
"TypeDef:Struct": true,
"align": 256,
"properties": {
"header": { "TypeDef:Struct": true, "properties": { ... } },
"weight": { "TypeDef:Float32": true, "align": 16 }
}
}
```
- Struct-level `"align"` sets the default alignment for all fields in
that struct. The struct's total size is rounded up to this alignment.
- Field-level `"align"` overrides the struct default for that specific
field.
- Default alignment (when no annotation is present): 1 for u8/bool, 2
for u16/i16, 4 for u32/i32/f32, 8 for u64/i64/f64, max field alignment
for structs.
- Alignment is only meaningful in aligned static mode (ADR-096). In
packed sequential mode, alignment annotations are ignored — fields are
packed with no padding.
### 3. Variable-length encoding
**Three strategies for variable-length types, selected by the `encoding`
annotation and the standard JSON Schema `maxLength` keyword.**
```json
// Strategy 1: Inline length-prefixing (default, shorthand)
{ "TypeDef:String": true }
// Strategy 1: Explicit inline length-prefixing
{ "TypeDef:String": { "encoding": "length-prefixed" } }
// Strategy 2: Fixed-size reservation (uses standard maxLength)
{ "TypeDef:String": true, "maxLength": 256 }
// Strategy 3: Offset indirection (opt-in)
{ "TypeDef:String": { "encoding": "offset-indirect" } }
```
**Strategy 1: Inline length-prefixing (default).** The field's fixed
portion is a 4-byte length prefix at a computed offset. The variable data
follows immediately after. In packed sequential mode, the length prefix
determines the position of subsequent fields. In aligned static mode, the
length prefix is at a known offset; the variable data is not included in
the static layout. This is the universal pattern used by channels, SFTP,
TTY, and most binary protocols.
**Strategy 2: Fixed-size reservation.** When a variable-length field
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
than `maxLength` is zero-padded; data longer than `maxLength` is a
validation error. This makes the field fixed-size from the layout
perspective — subsequent fields have known, unchanging offsets. This is
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
pattern for fields with known maximum sizes. In packed sequential mode,
`maxLength` is a validation constraint only — the engine still uses
inline length-prefixing (strategy 1).
**Strategy 3: Offset indirection.** The field is a struct
`{offset: u32, length: u32}` that points into a separate data region.
This is the metatensor blob tensor pattern — the index struct lives in
one region, the blob data lives in another. The consumer provides the
data region separately. Enables mmap-friendly random access to
variable-length data without parsing length prefixes and without
reserving worst-case space.
**Default strategy selection:**
- In packed sequential mode: always strategy 1 (inline length-prefixing).
`maxLength` is a validation constraint only.
- In aligned static mode: strategy 2 (fixed-size reservation) if
`maxLength` is declared; strategy 3 (offset indirection) if
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
length-prefixing) otherwise.
- `true` is a shorthand for the default (length-prefixed). This keeps
the common case concise and the override explicit.
- The `encoding` annotation and `maxLength` apply to all variable-length
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
`TypeDef:Record`, `TypeDef:Timestamp`.
### 3a. TRecord value type
`TypeDef:Record` is a string-keyed map. The value type is declared via
the `"values"` property in the schema:
```json
{
"TypeDef:Record": true,
"values": { "TypeDef:Float32": true }
}
```
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
values in the record. All values share the same type.
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count`
times. Each key is a length-prefixed UTF-8 string. Each value is
encoded according to its declared `TypeDef:*` kind — a `Record<Uint32>`
value is 4 raw bytes; a `Record<String>` value is itself a
length-prefixed string; a `Record<Struct>` value is the struct's
fields laid out inline. There is **no separate `value_len` prefix** —
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix).
- The count and key-length prefixes respect the schema's endianness.
- In aligned static mode with `maxLength`, the entire record is reserved
at `maxLength` bytes (zero-padded).
### 4. TUnion discriminators
**Two discriminator kinds: byte-offset (protocol dispatch) and
field-name (typedef.ts pattern).**
#### Kind A: Byte-offset discriminator
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "TypeDef:Uint8"
},
"mapping": {
"1": { "$ref": "#/$defs/Init" },
"3": { "$ref": "#/$defs/Open" },
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" },
"101": { "$ref": "#/$defs/Status" }
}
}
```
- The discriminator is a fixed-size integer at a known byte offset.
- `"offset"` is the byte position of the discriminator within the union's
buffer.
- `"type"` is the `TypeDef:*` kind of the discriminator (typically
`TypeDef:Uint8` for protocol type bytes).
- The mapping keys are stringified integers (`"1"`, `"5"`, `"101"`).
The engine parses the key to match the discriminator value.
- The variant struct starts at `offset + discriminator_size`.
- This is the SFTP `Packet` enum pattern and the call protocol's event
type dispatch.
#### Kind B: Field-name discriminator
```json
{
"TypeDef:Union": true,
"discriminator": {
"kind": "field",
"name": "type"
},
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
}
```
- The discriminator is a named field within the struct.
- `"name"` is the field name that holds the discriminator value.
- The mapping keys are string values matching the discriminator field's
value.
- The discriminator field is just another field in the struct — its
offset is computed like any other field.
- This is the typedef.ts `TUnion` pattern.
#### Mapping values
Mapping values may be either inline schemas or `$ref` pointers. `$ref`
is cleaner for large unions (29 SFTP variants) but requires a `$defs`
section. Inline schemas are simpler for small unions (5 call protocol
event types). Both work.
## Consequences
### Positive
- **Concrete, validated shapes.** All four annotation categories have
concrete JSON shapes that were validated by the POCs.
- **Sensible defaults.** Little-endian, natural alignment, inline
length-prefixing — the common case requires no annotations.
- **Explicit overrides.** Big-endian, custom alignment, offset
indirection — the uncommon case is explicit and self-documenting.
- **TUnion covers both protocol and typedef.ts patterns.** The
byte-offset discriminator handles SFTP type bytes and call protocol
event types. The field-name discriminator handles the typedef.ts string
pattern. No separate union type needed.
### Negative
- **Keyword value shape change.** `"TypeDef:String": true` (boolean) and
`"TypeDef:String": { "encoding": "length-prefixed" }` (object) are both
valid. The engine must handle both shapes. This is a minor parsing
concern — the POC already handles it.
- **Alignment annotations are mode-specific.** Alignment is only
meaningful in aligned static mode. In packed sequential mode, alignment
annotations are ignored. This is documented, not enforced — a consumer
that specifies alignment in packed mode gets no error, just no effect.
## References
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — the
annotation shape questions this ADR resolves
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes (alignment only meaningful in aligned static mode)
@@ -0,0 +1,157 @@
# ADR-098: Error Handling and Validation Strategy
## Status
Accepted
## Context
The typedef engine operates in three phases, each with distinct error
conditions:
1. **Schema parsing** — invalid JSON, missing required keywords, unknown
`TypeDef:*` kinds, malformed annotations.
2. **Offset computation** — field not found, type not supported for
offset computation, recursive schema depth exceeded.
3. **Read/write** — buffer too short, invalid UTF-8, value out of range
for the target type.
4. **Validation** — type constraint violations (range, UTF-8, field
presence, discriminator membership).
The engine also needs a clear strategy for *when* validation happens:
once at schema load time (build the validator) vs repeatedly at access
time (validate each buffer).
## Decision
### Error type: `TypedefError`
A single `TypedefError` enum with variants for each error category:
```rust
pub enum TypedefError {
/// Schema parsing errors.
Schema(String),
/// Offset computation errors.
Offset { field_path: String, reason: String },
/// Read/write errors.
Access { field_path: String, reason: String },
/// Validation errors (delegated to jsonschema).
Validation(ValidationError<'static>),
}
```
- `Schema` — for invalid JSON, missing required keywords, unknown
`TypeDef:*` kinds. The error message describes the problem.
- `Offset` — for field-not-found, unsupported type for offset
computation, etc. Carries the field path for debugging.
- `Access` — for buffer-too-short, invalid UTF-8, value out of range.
Carries the field path for debugging.
- `Validation` — wraps `jsonschema`'s `ValidationError`. The
`jsonschema` crate already provides rich error messages with schema
paths; the typedef engine does not re-wrap or re-interpret them.
The `Validation` variant uses `ValidationError<'static>` because the
validator is built once at schema load time and lives for the lifetime
of the `TypedefEngine`. The `'static` lifetime is correct — the validator
owns its schema reference.
### Validation timing: load-time build, access-time check
The jsonschema validator is built once at schema load time
(`validator_for(&schema)?`) and then called repeatedly
(`validator.is_valid(&instance)`). The typedef engine follows the same
pattern:
1. **Load time:** Parse the schema JSON, build the offset map (or
`LayoutBuilder`/`SequentialReader`), build the jsonschema validator.
This is the `TypedefEngine::compile(schema: &Value) -> Result<Self,
TypedefError>` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation — the consumer calls
`engine.validate(buffer)` when validation is desired.
The `TypedefEngine` struct is the compiled form of a schema:
```rust
pub struct TypedefEngine {
offset_map: OffsetMap, // or LayoutBuilder/SequentialReader
validator: jsonschema::Validator, // compiled once at load time
}
```
### Custom keyword validators
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
`jsonschema::options().with_keyword(...)`. The validators check:
- **Numeric types** (`TypeDef:Float32`, `TypeDef:Int8`, etc.): range
constraints (Int8: -128..127, Uint8: 0..255, etc.), finiteness for
floats.
- **`TypeDef:String`**: UTF-8 validity.
- **`TypeDef:Struct`**: field presence and types (delegated to
jsonschema's structural validation — the custom keyword only needs to
validate that the struct's fields match their declared `TypeDef:*`
kinds).
- **`TypeDef:Union`**: discriminator value membership in the mapping.
- **`TypeDef:Array`**: element type conformance.
- **`TypeDef:Boolean`**: value is `true` or `false`.
- **`TypeDef:Timestamp`**: RFC 3339 string format (the internet profile of ISO 8601).
The `jsonschema` crate handles all the structural validation (object
properties, required fields, array items, enum values) — the custom
keywords only need to validate the leaf type constraints. Each custom
keyword implementation is ~10 lines.
### Read/write errors carry field paths
Read/write errors include the field path for debugging:
```rust
// Example: reading a u32 from a buffer that's too short
Err(TypedefError::Access {
field_path: "header.version".to_string(),
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
})
```
This makes debugging binary format issues tractable — the error tells
you exactly which field failed and why.
## Consequences
### Positive
- **Single error type.** Consumers handle one `TypedefError` enum, not
multiple error types from different engine phases.
- **Field-path-carrying errors.** Read/write errors include the field
path, making binary format debugging tractable.
- **Validation is opt-in.** The consumer decides when to validate.
High-throughput paths can skip validation; security-sensitive paths
can validate every frame.
- **jsonschema integration is clean.** The `ValidationError` is wrapped
as-is — no re-interpretation, no information loss.
- **Load-time build, access-time use.** The expensive work (schema
parsing, validator compilation, offset computation) happens once at
load time. Access-time operations are cheap (pointer casts, slice
operations, length-prefix reads).
### Negative
- **`ValidationError<'static>` lifetime.** The `'static` lifetime on the
`Validation` variant means the error cannot borrow from the buffer
being validated. This is correct (the validator owns its schema
reference) but may surprise readers who expect a shorter lifetime.
- **No error recovery.** The engine does not attempt to recover from
partial reads or writes. A buffer-too-short error on field N means
fields N+1.. are also unreadable. This is inherent to binary formats
— there is no "skip to next field" without a schema-driven parser.
## References
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — error
handling strategy question (OQ 8)
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes
- [ADR-097](097-schema-annotations.md) — schema annotations
@@ -0,0 +1,114 @@
# ADR-099: Int64/Uint64 as First-Class Kinds
## Status
Accepted
## Context
The typedef engine's kind set (ADR-095, ADR-097) tops out at 32-bit
integers. The POC included `u64` read/write primitives, and the
call-channels-unification research's own SFTP schema example uses
`"TypeDef:Uint64"` for the `offset` field (`Read`/`Write` packets have
`offset: u64`). Metatensor/safetensors `data_offsets` are also `u64`.
A `TypeDef:Uint64` variant was added to the `TypeDefKind` enum during
implementation (the task decomposition correctly identified the gap),
but without an ADR the addition was half-finished: `type_size()`
returned `None`, the layout engines couldn't compute offsets for it, and
the validator didn't register a `TypeDef:Uint64` keyword. The variant
was then removed (commit `14d9cf2`) on the grounds that it was
unintended and a latent panic — but the underlying gap is real: SFTP and
metatensor, the two primary POC targets, both require 64-bit integers.
The presumed reason 64-bit integers were left out of the original
specification is a JSON-level concern: `serde_json::Number` loses
precision past 2^53 when parsing from JSON text. This is a
*validation-layer* caveat, not a *layout-layer* one — the binary layout
is 8 raw bytes, and `from_le_bytes`/`from_be_bytes` work correctly for
the full `u64`/`i64` range. The validation concern is handled by
accepting integer-form JSON values (the `jsonschema` crate's
`as_i64`/`as_u64` methods handle the common range; values past 2^53 are
a JSON representation limitation, not a typedef limitation).
## Decision
**Add `TypeDef:Int64` and `TypeDef:Uint64` as first-class kinds.**
Both are fixed-size (8 bytes), with natural alignment 8. They follow
the schema's endianness annotation like all other fixed-size types.
Read/write is via `data_access::read_i64`/`write_i64`/`read_u64`/
`write_u64` (endian-aware, 8 bytes).
### Kind table additions
| Kind | TypeBox key | Rust type | Size | Alignment |
|------|-------------|-----------|------|-----------|
| `TInt64` | `TypeDef:Int64` | `i64` | 8 | 8 |
| `TUint64` | `TypeDef:Uint64` | `u64` | 8 | 8 |
### Validation
The custom keyword validators check:
- `TypeDef:Int64`: value must be an integer in `i64::MIN..=i64::MAX`
(`-9223372036854775808` to `9223372036854775807`).
- `TypeDef:Uint64`: value must be a non-negative integer in
`0..=u64::MAX` (`0` to `18446744073709551615`).
The `jsonschema` crate's `as_i64`/`as_u64` handle the common range.
JSON numbers past 2^53 lose precision in the JSON representation —
this is a JSON limitation, not a typedef limitation. The binary
representation (8 raw bytes) is always exact. A consumer that needs
to validate the full 64-bit range from JSON should provide the value
as a JSON integer (which `serde_json` preserves for values up to
`u64::MAX`/`i64::MIN` when the `arbitrary_precision` feature is
enabled, or when the value fits in `i64`/`u64` without the feature).
### `FieldValue` additions
`FieldValue::I64(i64)` and `FieldValue::U64(u64)` are added to the
unified return type. The `SequentialReader`, `TypedefEngine::read_field`,
and `TypedefEngine::write_field` dispatch on the new kinds.
### Kind count
The engine now has **19** first-class kinds (17 + Int64 + Uint64).
`TypeDefKind::is_fixed_size()` returns `true` for both new kinds.
`type_size()` returns `Some(8)`. `natural_alignment()` returns `8`.
`needs_endian()` returns `true`.
## Consequences
### Positive
- **Unblocks the two primary POC targets.** SFTP `Read`/`Write` packets
(`offset: u64`) and metatensor `data_offsets` (`u64`) are now
expressible in typedef schemas.
- **Completes the half-finished addition.** The `TypeDefKind` enum,
`data_access` primitives, and `FieldValue` variants for 64-bit
integers now have matching layout, validator, and engine support.
- **No new design surface.** Int64/Uint64 are fixed-size types that
follow all existing patterns (endianness, alignment, zero-copy
read/write). They are mechanical additions.
### Negative
- **JSON precision caveat.** Values past 2^53 lose precision in the
JSON representation (not in the binary representation). This is a
JSON limitation, not a typedef limitation, but it means the
validation layer cannot perfectly round-trip the full 64-bit range
through JSON `Number` without `arbitrary_precision`. In practice,
SFTP offsets and tensor data offsets are well within 2^53.
- **Two more kinds to maintain.** The kind table, validator
registration, `FieldValue` enum, and dispatch arms all grow by two
variants. This is the cost of completeness.
## References
- `docs/research/call-channels-unification/findings.md` §"russh-sftp" —
the SFTP schema with `"offset": { "TypeDef:Uint64": true }`
- `docs/research/alknet-typedef/findings.md` §"POC 1" — the POC included
u64 read/write
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
purpose and scope (the kind set)
- [ADR-097](097-schema-annotations.md) — schema annotations
(endianness applies to the new kinds)
@@ -0,0 +1,112 @@
# ADR-100: Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode
## Status
Accepted
## Context
The aligned static layout mode (ADR-096) is designed for mmap-friendly
formats: fields have fixed positions with natural alignment padding,
enabling random access by field path without parsing preceding fields.
The spec (layout-engine.md) says variable-length fields in aligned mode
get a 4-byte length prefix at a known offset, and "the variable data
lives outside the static layout — either immediately after the fixed
fields (inline length-prefixing) or in a separate data region (offset
indirection)."
The implementation has a bug: `OffsetMap::compute` reserves only 4 bytes
for an inline length-prefixed variable field (the length prefix), but
`TypedefEngine::write_field` for a `String`/`Bytes` field calls
`data_access::write_string` at `range.start`, which writes
`[4-byte length][data]` inline — clobbering every subsequent field. The
`read_field` path has the mirror behavior (reads inline), so the engine
is self-consistent but only works correctly when the variable field is
the last field in the struct (no subsequent field to clobber).
Concretely, `{name: String, id: Uint32}` in aligned mode maps
`name → 0..4`, `id → 4..8`. Writing `"hello"` to `name` writes
`[5,0,0,0,h,e,l,l,o]` at offset 0, overwriting `id`'s range with
`hello`. All existing tests happen to put the variable field last, so
the bug is latent.
The spec's "data region after fixed fields" model (where variable data
lives after all fixed fields) is the correct design for aligned mode,
but implementing it would require a two-region layout (fixed fields +
variable data region) with the `OffsetMap` tracking both the prefix
position and the data position. This is a significant design addition
for a use case that doesn't exist yet — real aligned-format consumers
(metatensor, safetensors) use `maxLength` reservation or
`offset-indirect` encoding for variable data, not inline
length-prefixing.
## Decision
**Reject non-final inline length-prefixed variable fields in aligned
static mode at `OffsetMap::compute` time.**
A variable-length field (`TypeDef:String`, `TypeDef:Bytes`,
`TypeDef:Timestamp`, `TypeDef:Record`) in aligned static mode that uses
the default inline length-prefixing strategy (no `maxLength`, no
`offset-indirect`) must be the last field in its struct. If a non-final
inline length-prefixed variable field is encountered,
`OffsetMap::compute` returns `TypedefError::Offset` with a message
explaining that non-final variable fields in aligned mode require
`maxLength` (fixed-size reservation) or `"encoding": "offset-indirect"`
(offset indirection).
This is a validation-time rejection (schema load time), not a runtime
check. The consumer learns about the problem when compiling the schema,
not when writing data.
### What is NOT rejected
- Inline length-prefixed variable fields that are the last field in
their struct — these are fine (no subsequent field to clobber).
- `maxLength` reservation and `offset-indirect` encoding in any
position — these make the field fixed-size from the layout
perspective (known size at a known offset), so they don't clobber.
- Inline length-prefixed variable fields in packed sequential mode —
packed mode doesn't have fixed offsets; variable fields shift
subsequent fields by design.
## Consequences
### Positive
- **Eliminates a silent data-corruption bug.** A consumer that writes
a non-final string in aligned mode currently clobbers subsequent
fields with no error. After this fix, the schema is rejected at
compile time.
- **Matches real aligned-format usage.** mmap-friendly formats use
`maxLength` or `offset-indirect` for variable data; inline
length-prefixing in aligned mode is only meaningful as the last
field.
- **Simple to implement.** A single check in `compute_struct` (is this
variable field non-final and using inline length-prefixing? → reject).
No two-region layout needed.
- **Defers the two-region design without blocking consumers.** If a
future consumer needs inline length-prefixing in non-final position
in aligned mode, the two-region layout can be implemented then. The
rejection is reversible (remove the check, add the two-region logic).
### Negative
- **A schema that worked before (silently corrupting data) now fails
at compile time.** This is the correct behavior — the schema was
always broken, it just wasn't caught.
- **The "data region after fixed fields" model from the spec is not
implemented.** A consumer that wants inline variable data in a
non-final position must use packed mode or wait for the two-region
layout. This is acceptable for v1 — no current consumer needs it.
## References
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes (aligned static mode's variable-length handling)
- [ADR-097](097-schema-annotations.md) — the three variable-length
encoding strategies (`maxLength`, `offset-indirect`, inline
length-prefixing)
- `../layout-engine.md` §"Variable-length
fields in aligned mode" — the spec's "data region after fixed fields"
description
@@ -0,0 +1,103 @@
# ADR-101: Packed-Mode Read API — Engine as SequentialReader Factory
## Status
Accepted
## Context
`TypedefEngine` stores a `SequentialReader` inside its `Layout::Packed`
variant. The engine exposes it via
`engine.sequential_reader() -> Option<&SequentialReader>`.
The problem: `SequentialReader`'s read methods (`read_next`,
`read_field`, `reset`) all take `&mut self` — they mutate the reader's
internal cursor (`field_index`, `position`). But the engine hands out
`&SequentialReader` (a shared reference), which cannot be used to call
`&mut self` methods. The accessor can only give the consumer
`position()` and `endian()` (the `&self` methods) — the actual read
API is unreachable.
This makes the engine's packed read-side dead API. A consumer that
wants to read a packed buffer must construct their own
`SequentialReader::new(&schema)` from the schema, bypassing the engine
entirely. The stored reader is dead weight.
Three options were considered:
1. **Factory method** — the engine provides a method that returns an
owned fresh `SequentialReader` (reconstructed from the stored
schema). The consumer owns the reader and drives it with `&mut self`.
2. **Interior mutability** — wrap the reader in `Mutex` or `RefCell`
so `&SequentialReader` can be upgraded to `&mut`. Adds overhead and
complexity for mutable cursor state that the consumer legitimately
wants to own.
3. **`sequential_reader_mut()`** — return `&mut SequentialReader`.
Requires `&mut self` on the engine, which is overly restrictive
(the consumer may share the engine across threads or hold it behind
an `Arc`).
## Decision
**The engine is a `SequentialReader` factory.** Replace
`sequential_reader() -> Option<&SequentialReader>` with
`sequential_reader() -> Option<SequentialReader>` — the method returns
an owned fresh reader, reconstructed from the stored schema.
```rust
impl TypedefEngine {
/// Construct a fresh SequentialReader for packed-mode reads.
/// Returns None if compiled in aligned mode.
pub fn sequential_reader(&self) -> Option<SequentialReader>;
}
```
Each call returns a new reader with the cursor at position 0. The
consumer owns the reader and calls `read_next`/`read_field`/`reset` on
it directly. The engine still stores its own reader (used for schema
validation during construction), but no longer exposes it by
reference.
The same applies to `LayoutBuilder`: `layout_builder()` returns
`Option<&LayoutBuilder>` which is fine — `LayoutBuilder::build` takes
`&self`, so the shared reference is usable. No change needed for the
write-side.
### Cost
`SequentialReader::new` clones the top-level struct's field schemas (a
`Vec<(String, Value)>` of the `properties` entries) and clones the
schema itself. This is cheap — a struct has a small number of fields
(SFTP's largest packet has 5). The construction cost is negligible
compared to the cost of reading a buffer.
## Consequences
### Positive
- **The packed read API is now usable.** A consumer calls
`engine.sequential_reader()` to get an owned reader and drives it
directly. No dead API.
- **No interior mutability overhead.** The reader's mutable cursor
state is owned by the consumer, not shared through a lock.
- **Thread-safe engine.** The engine remains `Send + Sync` (it only
exposes `&self` methods). The reader is owned by the calling thread.
- **Simple.** One method signature change. The stored reader in
`Layout::Packed` can be removed (it was only used for schema
validation during construction, which is done by the time the
consumer calls `sequential_reader()`).
### Negative
- **Each call to `sequential_reader()` allocates a new reader.** The
cost is a `Vec` of field schemas + a schema clone. Acceptable for
the use case (one reader per buffer read).
- **The engine no longer holds a live reader.** If a future use case
needs to share a reader's cursor state across calls, the consumer
must manage that themselves. This is the correct separation — cursor
state is consumer-owned, not engine-owned.
## References
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — packed
sequential mode (`SequentialReader` as the read-side)
- `../data-access.md` §"Higher-level
read/write" — the `SequentialReader` API
@@ -0,0 +1,111 @@
# ADR-102: Reject TUnion in Aligned Mode for v1
## Status
Accepted
## Context
The aligned static layout mode (ADR-096) computes fixed byte positions
for each field, enabling random access by field path. `TUnion` in
aligned mode has three implementation problems:
1. **No variant field offsets.** Only the `__discriminator` byte range
is recorded in the `OffsetMap`. Variant field offsets are not
available anywhere in aligned mode — the consumer must recompute
them by hand. This makes `TypedefEngine::read_field` on a union
variant field impossible.
2. **`find_discriminator_field` takes the first variant's offset.** For
a field-name discriminator, the code probes the first variant that
contains the discriminator field and records that offset globally.
If variants order fields differently, the discriminator sits at
different offsets per variant and the recorded range is silently
wrong. The code should validate that the offset is identical across
all variants (or require the discriminator field to be first).
3. **Byte-discriminator union total misaligns the variant.** The union
total is `disc_off + disc_size + variant_max_size`, but the variant
was probed from offset 0 with alignment. A `u8` discriminator
before a `u32`-bearing variant produces a variant region that
starts at an unaligned offset in a mode whose entire purpose is
alignment.
The real question is whether `TUnion` in aligned mode is even needed.
The two consumer profiles are:
- **Protocol consumers** (SFTP, call protocol event types): use packed
sequential mode. `TUnion` with byte-offset discriminators is the
core dispatch mechanism. This is well-supported.
- **mmap consumers** (metatensor, safetensors): use aligned static
mode. These formats are structs and arrays of structs — they don't
use tagged unions. A tensor file has a header struct with tensor
descriptors, not a "which variant is this?" dispatch.
`TUnion` in aligned mode is a combination that no current or planned
consumer needs. Shipping broken semantics for an unused use case is
worse than rejecting it clearly.
## Decision
**Reject `TUnion` in aligned static mode for v1.**
`OffsetMap::compute` returns `TypedefError::Offset` when it encounters
a `TypeDef:Union` field, with a message explaining that unions are not
supported in aligned mode and the consumer should use packed mode (or
restructure as a struct with an explicit discriminator field).
This is a schema-load-time rejection. The consumer learns about the
problem when compiling the schema, not at runtime.
### What is NOT rejected
- `TUnion` in packed sequential mode — this is the core use case
(SFTP `Packet` dispatch, call protocol event types) and is fully
supported by `LayoutBuilder` and `SequentialReader`.
- `TStruct`, `TArray`, and all primitive kinds in aligned mode — these
are the mmap-format primitives and are fully supported.
### Reversal
This is a two-way door. If a future mmap-format consumer needs tagged
unions in aligned mode, the rejection can be lifted and the three
implementation problems fixed. The fix would require:
- Recording per-variant field offsets in the `OffsetMap` (which
variant's offsets to record when variants have different layouts?).
- Validating that field-name discriminators have identical offsets
across all variants.
- Aligning the variant region correctly after the byte discriminator.
These are design questions that should be answered when the use case
arrives, not speculatively now.
## Consequences
### Positive
- **No broken semantics shipped.** The three implementation problems
are removed from the API surface rather than silently producing
wrong offsets.
- **Clear scope boundary.** Aligned mode is for structs and arrays;
packed mode is for protocols (including union dispatch). The
consumer chooses the mode based on the use case.
- **Reversible.** When a real consumer needs aligned-mode unions, the
rejection is lifted and the design questions are worked through with
a concrete use case.
### Negative
- **A schema with a `TUnion` field cannot be compiled in aligned
mode.** A consumer that wants both aligned layout and union dispatch
must use packed mode or restructure. No current consumer needs this.
- **The aligned-mode union code in `offset_map.rs` is dead.** It can
be removed or left as a reference for when the rejection is lifted.
Removing it is cleaner.
## References
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
modes
- [ADR-097](097-schema-annotations.md) §4 — TUnion discriminators
- `../layout-engine.md` §"TUnion" —
aligned-mode union sizing