Port alknet-typedef crate from alknet
Copy the binary struct engine (src/, tests/) verbatim from alknet/crates/alknet-typedef and create a standalone Cargo.toml (workspace-inherited fields inlined). Port the architecture docs (specs, ADRs 095-102, OQs 069-071) from alknet's nested multi-crate layout to a flat single-crate layout, fixing relative link paths. Build, 295 tests, and clippy all pass clean.
This commit is contained in:
1 parent
eac7ad88b3
commit
2c4a4994dc
36 files changed
+13805
No files matched your search
@@ -0,0 +1,173 @@
|
||||
# ADR-095: alknet-typedef — Purpose, Scope, and the jsonschema Engine
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Three threads in the codebase converge on the same pattern: a JSON Schema
|
||||
describes the shape of binary data, and the binary data is the struct's
|
||||
bytes at computed offsets.
|
||||
|
||||
1. **typedef.ts** (`/workspace/@alkdev/typebox/example/typedef/typedef.ts`,
|
||||
619 lines) defines custom TypeBox schema kinds (`TFloat32`, `TStruct`,
|
||||
`TUnion`, etc.) that carry binary layout semantics. These are registered
|
||||
via `TypeRegistry.Set` with custom validators.
|
||||
|
||||
2. **russh-sftp** has 29 packet types, each a struct with typed fields
|
||||
(`Read { id: u32, handle: String, offset: u64, len: u32 }`). The wire
|
||||
format is `[length: u32][type: u8][payload]` where payload is the
|
||||
struct's serde bytes. The `Packet` enum dispatches on the type byte —
|
||||
a tagged union of structs. Under the typedef lens, each packet is a
|
||||
`TStruct`; the `Packet` enum is a `TUnion` with a byte-offset
|
||||
discriminator.
|
||||
|
||||
3. **metatensor** needs an offset map for mmap-friendly tensor access —
|
||||
given a schema describing a model layout (ConvNet struct, tensor refs),
|
||||
compute byte offsets for each field so the consumer can read tensor
|
||||
data at known positions without parsing.
|
||||
|
||||
The common pattern: **a JSON Schema with `TypeDef:*` custom keywords
|
||||
describes the shape of binary data; the binary data is the struct's bytes
|
||||
at computed offsets.** The schema is the format definition; the engine is
|
||||
generic.
|
||||
|
||||
Two prior attempts built their own jsonschema engines — the fatal flaw:
|
||||
|
||||
- **typebox-rs** (`/workspace/@alkimiadev/typebox-rs/`, ~8,400 lines):
|
||||
a full 26-variant `SchemaKind` enum, a custom `Value` type with typed
|
||||
arrays, and a 912-line hand-written validator.
|
||||
- **alktype** (`/workspace/@alkimiadev/alktype/`, ~5,600 lines): a
|
||||
handler-registry pattern that also implements its own validation for
|
||||
each type.
|
||||
|
||||
The `jsonschema` crate (v0.46.5, Draft 2020-12) is already in the
|
||||
workspace at `/workspace/jsonschema/`. It handles validation with custom
|
||||
keyword support — the novel code is the offset computation, not the
|
||||
validation.
|
||||
|
||||
The call-channels-unification research
|
||||
(`docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||
JSON Schema as the binary struct engine") identified the convergence and
|
||||
bumped typedef up in the timeline. The POC
|
||||
(`docs/research/alknet-typedef/findings.md`, 26 tests passing) validated
|
||||
the approach: a ~1,900-line Rust crate that takes a JSON Schema with
|
||||
`TypeDef:*` custom keywords and produces an offset map, read/write
|
||||
functions, and validation — all driven by the schema.
|
||||
|
||||
## Decision
|
||||
|
||||
**alknet-typedef is a small Rust crate that takes a JSON Schema with
|
||||
`TypeDef:*` custom keywords and produces three capabilities:**
|
||||
|
||||
1. **An offset map** — walks the schema, computes byte offsets for each
|
||||
field based on type sizes, field order, and alignment.
|
||||
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
|
||||
read the field's bytes at its offset (zero-copy for fixed-size types).
|
||||
Given a `&mut [u8]` buffer, write a value at its offset.
|
||||
3. **Validation** — via `jsonschema` custom keywords, validates that data
|
||||
conforms to the schema's type constraints. The jsonschema validator
|
||||
operates on `serde_json::Value` instances (JSON representations), not
|
||||
raw byte buffers directly. A consumer that wants to validate a binary
|
||||
buffer reads it into a `Value` tree via the data access layer, then
|
||||
validates that `Value` against the jsonschema validator.
|
||||
|
||||
**The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||
`serde_json` (schema parsing).** The novel code is the offset computation
|
||||
— a recursive walk of the schema JSON that computes byte positions for
|
||||
each field. The custom keyword implementations are ~10 lines each.
|
||||
|
||||
**The schema is the format.** A JSON Schema with `TypeDef:Float32`,
|
||||
`TypeDef:Struct`, `TypeDef:Union` etc. is both the validation spec and
|
||||
the layout spec. No separate format definition, no separate parser, no
|
||||
separate validator. One schema, three uses: validate, compute offsets,
|
||||
access data.
|
||||
|
||||
**The crate depends on `jsonschema` and `serde_json` (with
|
||||
`preserve_order`).** No tokio, no platform deps. Compiles to
|
||||
`wasm32-unknown-unknown` for browser use. The `jsonschema` crate's
|
||||
`with_keyword("TypeDef:Float32", factory)` API is the integration point
|
||||
for custom type kinds — each `TypeDef:*` kind maps to a custom keyword
|
||||
validator in Rust. Same semantics as TypeBox's `TypeRegistry.Set`, same
|
||||
JSON Schema wire format.
|
||||
|
||||
**The crate targets `std` for v1.** The WASM target has `std` available
|
||||
via `wasm-bindgen`. If embedded use cases emerge, `no_std` + `alloc` can
|
||||
be added as a feature gate later — the engine's core (offset computation,
|
||||
read/write) is already allocation-free. See OQ-070.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Eliminates ~14,000 lines of hand-rolled schema engines.** typebox-rs
|
||||
and alktype are replaced by `jsonschema` + an offset map + ~50 lines of
|
||||
custom keyword implementations. The codebase drops from "a port of
|
||||
TypeBox" to "jsonschema + an offset map."
|
||||
- **One schema, three uses.** The same JSON Schema validates, computes
|
||||
offsets, and drives data access. No separate format definition, parser,
|
||||
or validator per protocol.
|
||||
- **Schema-driven, not code-driven.** Adding a new SFTP packet type is
|
||||
adding a variant to the schema JSON, not writing a new Rust struct +
|
||||
serde impl. The engine is generic; the schema is the configuration.
|
||||
- **WASM-clean.** `serde_json` + `jsonschema` + byte manipulation. No
|
||||
tokio, no platform deps. The same typedef schemas work in browser,
|
||||
Node, Python (via `wasmtime-py`), Go (via `wazero`), and any other
|
||||
WASM host.
|
||||
- **TypeBox interop.** TypeBox modules render to standard JSON Schema
|
||||
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
|
||||
on the Rust side. Zero translation. The same schema validates in both
|
||||
ecosystems.
|
||||
- **Defense in depth.** Schema validation via jsonschema custom keywords —
|
||||
a malformed binary payload can be read into a `Value` tree via the data
|
||||
access layer and validated against the schema before any consumer
|
||||
touches it. The `jsonschema` crate's compiled validators are fast enough
|
||||
to run on every incoming frame.
|
||||
|
||||
### Negative
|
||||
|
||||
- **New dependency on `jsonschema`.** The crate is already in the
|
||||
workspace but not yet used by any alknet crate. This is the first
|
||||
consumer.
|
||||
- **`serde_json` with `preserve_order` is required.** Field order is
|
||||
load-bearing for binary layouts. The `preserve_order` feature adds a
|
||||
small compile-time cost.
|
||||
- **Schema authoring is external.** Schemas are authored in TypeBox (JS)
|
||||
or hand-written JSON. The typedef engine consumes schemas; it does not
|
||||
generate them. A builder API is deferred (OQ-071).
|
||||
|
||||
## Scope Boundaries (What This Is Not)
|
||||
|
||||
- **Not metatensor.** typedef is the binary struct *engine*. Metatensor
|
||||
is a *format* (8-byte header + JSON header + binary data) that uses the
|
||||
typedef engine for its offset computation and tensor access.
|
||||
- **Not a Value system.** TypeBox's `Value.Diff`, `Value.Migrate`,
|
||||
`Value.Convert` — schema evolution — is out of scope for v1. The engine
|
||||
should not do anything that explicitly blocks adding a Value system
|
||||
later.
|
||||
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
|
||||
concern. The typedef engine consumes schemas; it does not generate them.
|
||||
- **Not a schema builder.** The typedef engine does not provide a fluent
|
||||
API for constructing schemas. Schemas are plain JSON.
|
||||
- **Not a serialization framework.** The typedef engine is not a
|
||||
general-purpose serde replacement. It operates on raw byte buffers at
|
||||
computed offsets — no intermediate `Value` tree, no reflection, no
|
||||
dynamic dispatch per field. For JSON data, use serde. For binary data
|
||||
with a known schema, use typedef.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` — POC results (26 tests
|
||||
passing, two layout modes, TUnion dispatch, endianness)
|
||||
- `docs/research/call-channels-unification/findings.md` §"alknet-typedef:
|
||||
JSON Schema as the binary struct engine" — the origin of this research
|
||||
thread
|
||||
- `/workspace/@alkdev/typebox/example/typedef/typedef.ts` — the TypeBox
|
||||
schema kinds (619 lines)
|
||||
- `/workspace/jsonschema/` — the jsonschema crate (v0.46.5, Draft 2020-12)
|
||||
- `/workspace/alknet-typedef-poc/` — the POC code (disposable)
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes decision
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotation shapes
|
||||
- [ADR-098](098-error-handling-validation-strategy.md) — error handling
|
||||
and validation strategy
|
||||
@@ -0,0 +1,137 @@
|
||||
# ADR-096: Two Layout Modes — Packed Sequential vs Aligned Static
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The POCs surfaced that protocols and mmap-friendly formats need different
|
||||
layout strategies. POC 1 built an aligned `OffsetMap` with natural
|
||||
alignment padding — correct for mmap-friendly formats (metatensor) but
|
||||
wrong for protocol wire formats (SFTP, channels, TTY). POC 2 built a
|
||||
`LayoutBuilder` and `SequentialReader` for packed sequential layouts —
|
||||
correct for protocol wire formats but wrong for mmap-friendly formats.
|
||||
|
||||
This is the most important architectural finding from the POCs. The
|
||||
engine must support both modes; a single layout strategy cannot serve
|
||||
both use cases.
|
||||
|
||||
### Packed sequential layout (protocol wire formats)
|
||||
|
||||
Protocols pack fields sequentially with no alignment padding.
|
||||
Variable-length fields shift all subsequent fields. Writing requires
|
||||
knowing actual data sizes upfront; reading walks the buffer sequentially,
|
||||
reading length prefixes to determine positions.
|
||||
|
||||
This is the layout used by SFTP (all strings and byte arrays are
|
||||
length-prefixed inline), channels (`[channel_id: u32][size: u32][payload]`),
|
||||
TTY (`[stream_type: u8][length: u32][payload]`), and most binary protocols.
|
||||
|
||||
### Aligned static layout (mmap-friendly formats)
|
||||
|
||||
Fields have fixed positions with natural alignment padding.
|
||||
Variable-length fields get a 4-byte length prefix at a known offset; the
|
||||
variable data is not included in the static layout. This enables
|
||||
mmap-friendly random access — the consumer can read field N at a known
|
||||
offset without parsing the fields before it.
|
||||
|
||||
This is the layout used by metatensor (blob tensor pattern: index struct
|
||||
in one region, blob data in another) and safetensors (header + aligned
|
||||
tensor data).
|
||||
|
||||
## Decision
|
||||
|
||||
**The typedef engine supports two layout modes, selected by the consumer
|
||||
at engine construction time:**
|
||||
|
||||
### Mode 1: Packed sequential (`LayoutBuilder` / `SequentialReader`)
|
||||
|
||||
For protocol wire formats. Fields are packed with no alignment padding.
|
||||
Variable-length fields shift all subsequent fields.
|
||||
|
||||
- **LayoutBuilder** — takes a schema and actual data sizes for
|
||||
variable-length fields, computes byte positions for each field in a
|
||||
packed layout. Used at write time when the consumer knows the data
|
||||
sizes upfront.
|
||||
- **SequentialReader** — walks a buffer field-by-field according to the
|
||||
schema, reading length prefixes to determine variable-length data
|
||||
positions. Used at read time when the consumer is parsing an incoming
|
||||
frame.
|
||||
|
||||
The `LayoutBuilder` and `SequentialReader` are the primary interface for
|
||||
protocol consumers (SFTP, binary call frames, TTY negotiation).
|
||||
|
||||
### Mode 2: Aligned static (`OffsetMap`)
|
||||
|
||||
For mmap-friendly formats. Fields have fixed positions with natural
|
||||
alignment padding. Variable-length fields get a 4-byte length prefix at
|
||||
a known offset; the variable data is not included in the static layout.
|
||||
|
||||
- **OffsetMap** — walks the schema once, computes fixed byte positions
|
||||
for each field based on type sizes and alignment. The output is a flat
|
||||
table of `(field_path, byte_range)` pairs. Used for both read and write
|
||||
at known offsets.
|
||||
|
||||
The `OffsetMap` is the primary interface for mmap consumers (metatensor).
|
||||
|
||||
### Variable-length handling in each mode
|
||||
|
||||
**Packed sequential mode:** Variable-length fields are inline
|
||||
length-prefixed by default (`[length: u32][data]`). The `LayoutBuilder`
|
||||
takes the actual data size to compute the length prefix value and the
|
||||
position of subsequent fields. The `SequentialReader` reads the length
|
||||
prefix to determine the data extent and the position of the next field.
|
||||
|
||||
**Aligned static mode:** Variable-length fields get a 4-byte length
|
||||
prefix at a known offset. The variable data lives outside the static
|
||||
layout — either immediately after the fixed fields (inline
|
||||
length-prefixing) or in a separate data region (offset indirection, the
|
||||
metatensor blob tensor pattern). The `OffsetMap` records the position of
|
||||
the length prefix (or the `{offset, length}` pair for offset-indirect
|
||||
fields).
|
||||
|
||||
### Default for variable-length types
|
||||
|
||||
Inline length-prefixing (`[length: u32][data]`) is the default for all
|
||||
variable-length types in both modes. This is the universal pattern used
|
||||
by channels, SFTP, TTY, and most binary protocols. Offset indirection is
|
||||
opt-in via the `encoding` annotation (see ADR-097).
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **One engine, two modes.** The same schema can be used in either mode.
|
||||
A schema describing an SFTP packet can be consumed by a `SequentialReader`
|
||||
(for parsing incoming frames) and a `LayoutBuilder` (for constructing
|
||||
outgoing frames). A schema describing a metatensor layout can be
|
||||
consumed by an `OffsetMap` (for mmap access).
|
||||
- **Correct for both use cases.** Packed sequential mode produces
|
||||
byte-identical output to hand-written protocol serialization (validated
|
||||
by POC 2's russh-sftp round-trip tests). Aligned static mode produces
|
||||
correct offsets for mmap-friendly access (validated by POC 1's
|
||||
alignment tests).
|
||||
- **No mode confusion.** The consumer explicitly selects the mode at
|
||||
engine construction time. A protocol consumer never accidentally gets
|
||||
alignment padding; an mmap consumer never accidentally gets
|
||||
variable-length field shifting.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Two APIs to learn.** Consumers must choose between
|
||||
`LayoutBuilder`/`SequentialReader` and `OffsetMap`. The choice is
|
||||
determined by the use case (protocol vs mmap), not by the schema.
|
||||
- **Variable-length fields in packed mode require size foreknowledge.**
|
||||
The `LayoutBuilder` needs actual data sizes for variable-length fields
|
||||
to compute correct positions for subsequent fields. This is inherent
|
||||
to packed layouts — the consumer must know the data sizes before
|
||||
writing.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"POC Results" — POC 1
|
||||
(aligned OffsetMap) and POC 2 (packed LayoutBuilder/SequentialReader)
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotations including
|
||||
the `encoding` field for variable-length types
|
||||
@@ -0,0 +1,259 @@
|
||||
# ADR-097: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The typedef engine needs concrete JSON shapes for schema-level
|
||||
annotations that control binary layout behavior. The POCs validated the
|
||||
semantics; this ADR pins the shapes.
|
||||
|
||||
Four annotation categories need concrete shapes:
|
||||
|
||||
1. **Endianness** — safetensors is little-endian, SFTP is big-endian.
|
||||
The engine needs to know which to use.
|
||||
2. **Alignment** — different backends have different alignment
|
||||
requirements (wgpu: 256-byte, protocols: natural, mmap: page).
|
||||
3. **Variable-length encoding** — inline length-prefixing vs offset
|
||||
indirection for strings, byte arrays, and other variable-length types.
|
||||
4. **TUnion discriminators** — byte-offset (protocol dispatch) vs
|
||||
field-name (typedef.ts pattern).
|
||||
|
||||
## Decision
|
||||
|
||||
### 1. Endianness
|
||||
|
||||
**Schema-level annotation with a default of little-endian.**
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Struct": true,
|
||||
"endian": "big",
|
||||
"properties": { ... }
|
||||
}
|
||||
```
|
||||
|
||||
- `"endian": "little"` (default) — read/write in little-endian byte order.
|
||||
- `"endian": "big"` — read/write in big-endian byte order.
|
||||
- The annotation applies to the entire schema and all nested types.
|
||||
- Mixed endianness within one schema is not supported (pathological; no
|
||||
known protocol requires it).
|
||||
- The default is little-endian, matching safetensors, wgpu, and most
|
||||
modern formats. SFTP consumers specify `"endian": "big"`.
|
||||
|
||||
### 2. Alignment
|
||||
|
||||
**Both struct-level and field-level, with field-level overriding
|
||||
struct-level.**
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Struct": true,
|
||||
"align": 256,
|
||||
"properties": {
|
||||
"header": { "TypeDef:Struct": true, "properties": { ... } },
|
||||
"weight": { "TypeDef:Float32": true, "align": 16 }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- Struct-level `"align"` sets the default alignment for all fields in
|
||||
that struct. The struct's total size is rounded up to this alignment.
|
||||
- Field-level `"align"` overrides the struct default for that specific
|
||||
field.
|
||||
- Default alignment (when no annotation is present): 1 for u8/bool, 2
|
||||
for u16/i16, 4 for u32/i32/f32, 8 for u64/i64/f64, max field alignment
|
||||
for structs.
|
||||
- Alignment is only meaningful in aligned static mode (ADR-096). In
|
||||
packed sequential mode, alignment annotations are ignored — fields are
|
||||
packed with no padding.
|
||||
|
||||
### 3. Variable-length encoding
|
||||
|
||||
**Three strategies for variable-length types, selected by the `encoding`
|
||||
annotation and the standard JSON Schema `maxLength` keyword.**
|
||||
|
||||
```json
|
||||
// Strategy 1: Inline length-prefixing (default, shorthand)
|
||||
{ "TypeDef:String": true }
|
||||
|
||||
// Strategy 1: Explicit inline length-prefixing
|
||||
{ "TypeDef:String": { "encoding": "length-prefixed" } }
|
||||
|
||||
// Strategy 2: Fixed-size reservation (uses standard maxLength)
|
||||
{ "TypeDef:String": true, "maxLength": 256 }
|
||||
|
||||
// Strategy 3: Offset indirection (opt-in)
|
||||
{ "TypeDef:String": { "encoding": "offset-indirect" } }
|
||||
```
|
||||
|
||||
**Strategy 1: Inline length-prefixing (default).** The field's fixed
|
||||
portion is a 4-byte length prefix at a computed offset. The variable data
|
||||
follows immediately after. In packed sequential mode, the length prefix
|
||||
determines the position of subsequent fields. In aligned static mode, the
|
||||
length prefix is at a known offset; the variable data is not included in
|
||||
the static layout. This is the universal pattern used by channels, SFTP,
|
||||
TTY, and most binary protocols.
|
||||
|
||||
**Strategy 2: Fixed-size reservation.** When a variable-length field
|
||||
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
|
||||
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
|
||||
than `maxLength` is zero-padded; data longer than `maxLength` is a
|
||||
validation error. This makes the field fixed-size from the layout
|
||||
perspective — subsequent fields have known, unchanging offsets. This is
|
||||
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
|
||||
pattern for fields with known maximum sizes. In packed sequential mode,
|
||||
`maxLength` is a validation constraint only — the engine still uses
|
||||
inline length-prefixing (strategy 1).
|
||||
|
||||
**Strategy 3: Offset indirection.** The field is a struct
|
||||
`{offset: u32, length: u32}` that points into a separate data region.
|
||||
This is the metatensor blob tensor pattern — the index struct lives in
|
||||
one region, the blob data lives in another. The consumer provides the
|
||||
data region separately. Enables mmap-friendly random access to
|
||||
variable-length data without parsing length prefixes and without
|
||||
reserving worst-case space.
|
||||
|
||||
**Default strategy selection:**
|
||||
- In packed sequential mode: always strategy 1 (inline length-prefixing).
|
||||
`maxLength` is a validation constraint only.
|
||||
- In aligned static mode: strategy 2 (fixed-size reservation) if
|
||||
`maxLength` is declared; strategy 3 (offset indirection) if
|
||||
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
|
||||
length-prefixing) otherwise.
|
||||
|
||||
- `true` is a shorthand for the default (length-prefixed). This keeps
|
||||
the common case concise and the override explicit.
|
||||
- The `encoding` annotation and `maxLength` apply to all variable-length
|
||||
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
|
||||
`TypeDef:Record`, `TypeDef:Timestamp`.
|
||||
|
||||
### 3a. TRecord value type
|
||||
|
||||
`TypeDef:Record` is a string-keyed map. The value type is declared via
|
||||
the `"values"` property in the schema:
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Record": true,
|
||||
"values": { "TypeDef:Float32": true }
|
||||
}
|
||||
```
|
||||
|
||||
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
|
||||
values in the record. All values share the same type.
|
||||
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
|
||||
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count`
|
||||
times. Each key is a length-prefixed UTF-8 string. Each value is
|
||||
encoded according to its declared `TypeDef:*` kind — a `Record<Uint32>`
|
||||
value is 4 raw bytes; a `Record<String>` value is itself a
|
||||
length-prefixed string; a `Record<Struct>` value is the struct's
|
||||
fields laid out inline. There is **no separate `value_len` prefix** —
|
||||
the value's size is determined by its kind (fixed-size kinds have a
|
||||
known size; variable-length kinds carry their own length prefix).
|
||||
- The count and key-length prefixes respect the schema's endianness.
|
||||
- In aligned static mode with `maxLength`, the entire record is reserved
|
||||
at `maxLength` bytes (zero-padded).
|
||||
|
||||
### 4. TUnion discriminators
|
||||
|
||||
**Two discriminator kinds: byte-offset (protocol dispatch) and
|
||||
field-name (typedef.ts pattern).**
|
||||
|
||||
#### Kind A: Byte-offset discriminator
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "byte",
|
||||
"offset": 0,
|
||||
"type": "TypeDef:Uint8"
|
||||
},
|
||||
"mapping": {
|
||||
"1": { "$ref": "#/$defs/Init" },
|
||||
"3": { "$ref": "#/$defs/Open" },
|
||||
"5": { "$ref": "#/$defs/Read" },
|
||||
"6": { "$ref": "#/$defs/Write" },
|
||||
"101": { "$ref": "#/$defs/Status" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- The discriminator is a fixed-size integer at a known byte offset.
|
||||
- `"offset"` is the byte position of the discriminator within the union's
|
||||
buffer.
|
||||
- `"type"` is the `TypeDef:*` kind of the discriminator (typically
|
||||
`TypeDef:Uint8` for protocol type bytes).
|
||||
- The mapping keys are stringified integers (`"1"`, `"5"`, `"101"`).
|
||||
The engine parses the key to match the discriminator value.
|
||||
- The variant struct starts at `offset + discriminator_size`.
|
||||
- This is the SFTP `Packet` enum pattern and the call protocol's event
|
||||
type dispatch.
|
||||
|
||||
#### Kind B: Field-name discriminator
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Union": true,
|
||||
"discriminator": {
|
||||
"kind": "field",
|
||||
"name": "type"
|
||||
},
|
||||
"mapping": {
|
||||
"read": { "$ref": "#/$defs/Read" },
|
||||
"write": { "$ref": "#/$defs/Write" }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- The discriminator is a named field within the struct.
|
||||
- `"name"` is the field name that holds the discriminator value.
|
||||
- The mapping keys are string values matching the discriminator field's
|
||||
value.
|
||||
- The discriminator field is just another field in the struct — its
|
||||
offset is computed like any other field.
|
||||
- This is the typedef.ts `TUnion` pattern.
|
||||
|
||||
#### Mapping values
|
||||
|
||||
Mapping values may be either inline schemas or `$ref` pointers. `$ref`
|
||||
is cleaner for large unions (29 SFTP variants) but requires a `$defs`
|
||||
section. Inline schemas are simpler for small unions (5 call protocol
|
||||
event types). Both work.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Concrete, validated shapes.** All four annotation categories have
|
||||
concrete JSON shapes that were validated by the POCs.
|
||||
- **Sensible defaults.** Little-endian, natural alignment, inline
|
||||
length-prefixing — the common case requires no annotations.
|
||||
- **Explicit overrides.** Big-endian, custom alignment, offset
|
||||
indirection — the uncommon case is explicit and self-documenting.
|
||||
- **TUnion covers both protocol and typedef.ts patterns.** The
|
||||
byte-offset discriminator handles SFTP type bytes and call protocol
|
||||
event types. The field-name discriminator handles the typedef.ts string
|
||||
pattern. No separate union type needed.
|
||||
|
||||
### Negative
|
||||
|
||||
- **Keyword value shape change.** `"TypeDef:String": true` (boolean) and
|
||||
`"TypeDef:String": { "encoding": "length-prefixed" }` (object) are both
|
||||
valid. The engine must handle both shapes. This is a minor parsing
|
||||
concern — the POC already handles it.
|
||||
- **Alignment annotations are mode-specific.** Alignment is only
|
||||
meaningful in aligned static mode. In packed sequential mode, alignment
|
||||
annotations are ignored. This is documented, not enforced — a consumer
|
||||
that specifies alignment in packed mode gets no error, just no effect.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — the
|
||||
annotation shape questions this ADR resolves
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes (alignment only meaningful in aligned static mode)
|
||||
@@ -0,0 +1,157 @@
|
||||
# ADR-098: Error Handling and Validation Strategy
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The typedef engine operates in three phases, each with distinct error
|
||||
conditions:
|
||||
|
||||
1. **Schema parsing** — invalid JSON, missing required keywords, unknown
|
||||
`TypeDef:*` kinds, malformed annotations.
|
||||
2. **Offset computation** — field not found, type not supported for
|
||||
offset computation, recursive schema depth exceeded.
|
||||
3. **Read/write** — buffer too short, invalid UTF-8, value out of range
|
||||
for the target type.
|
||||
4. **Validation** — type constraint violations (range, UTF-8, field
|
||||
presence, discriminator membership).
|
||||
|
||||
The engine also needs a clear strategy for *when* validation happens:
|
||||
once at schema load time (build the validator) vs repeatedly at access
|
||||
time (validate each buffer).
|
||||
|
||||
## Decision
|
||||
|
||||
### Error type: `TypedefError`
|
||||
|
||||
A single `TypedefError` enum with variants for each error category:
|
||||
|
||||
```rust
|
||||
pub enum TypedefError {
|
||||
/// Schema parsing errors.
|
||||
Schema(String),
|
||||
/// Offset computation errors.
|
||||
Offset { field_path: String, reason: String },
|
||||
/// Read/write errors.
|
||||
Access { field_path: String, reason: String },
|
||||
/// Validation errors (delegated to jsonschema).
|
||||
Validation(ValidationError<'static>),
|
||||
}
|
||||
```
|
||||
|
||||
- `Schema` — for invalid JSON, missing required keywords, unknown
|
||||
`TypeDef:*` kinds. The error message describes the problem.
|
||||
- `Offset` — for field-not-found, unsupported type for offset
|
||||
computation, etc. Carries the field path for debugging.
|
||||
- `Access` — for buffer-too-short, invalid UTF-8, value out of range.
|
||||
Carries the field path for debugging.
|
||||
- `Validation` — wraps `jsonschema`'s `ValidationError`. The
|
||||
`jsonschema` crate already provides rich error messages with schema
|
||||
paths; the typedef engine does not re-wrap or re-interpret them.
|
||||
|
||||
The `Validation` variant uses `ValidationError<'static>` because the
|
||||
validator is built once at schema load time and lives for the lifetime
|
||||
of the `TypedefEngine`. The `'static` lifetime is correct — the validator
|
||||
owns its schema reference.
|
||||
|
||||
### Validation timing: load-time build, access-time check
|
||||
|
||||
The jsonschema validator is built once at schema load time
|
||||
(`validator_for(&schema)?`) and then called repeatedly
|
||||
(`validator.is_valid(&instance)`). The typedef engine follows the same
|
||||
pattern:
|
||||
|
||||
1. **Load time:** Parse the schema JSON, build the offset map (or
|
||||
`LayoutBuilder`/`SequentialReader`), build the jsonschema validator.
|
||||
This is the `TypedefEngine::compile(schema: &Value) -> Result<Self,
|
||||
TypedefError>` constructor.
|
||||
2. **Access time:** Use the compiled engine for repeated read/write
|
||||
operations. Validation is opt-in per operation — the consumer calls
|
||||
`engine.validate(buffer)` when validation is desired.
|
||||
|
||||
The `TypedefEngine` struct is the compiled form of a schema:
|
||||
|
||||
```rust
|
||||
pub struct TypedefEngine {
|
||||
offset_map: OffsetMap, // or LayoutBuilder/SequentialReader
|
||||
validator: jsonschema::Validator, // compiled once at load time
|
||||
}
|
||||
```
|
||||
|
||||
### Custom keyword validators
|
||||
|
||||
Each `TypeDef:*` kind gets a `Keyword` implementation registered via
|
||||
`jsonschema::options().with_keyword(...)`. The validators check:
|
||||
|
||||
- **Numeric types** (`TypeDef:Float32`, `TypeDef:Int8`, etc.): range
|
||||
constraints (Int8: -128..127, Uint8: 0..255, etc.), finiteness for
|
||||
floats.
|
||||
- **`TypeDef:String`**: UTF-8 validity.
|
||||
- **`TypeDef:Struct`**: field presence and types (delegated to
|
||||
jsonschema's structural validation — the custom keyword only needs to
|
||||
validate that the struct's fields match their declared `TypeDef:*`
|
||||
kinds).
|
||||
- **`TypeDef:Union`**: discriminator value membership in the mapping.
|
||||
- **`TypeDef:Array`**: element type conformance.
|
||||
- **`TypeDef:Boolean`**: value is `true` or `false`.
|
||||
- **`TypeDef:Timestamp`**: RFC 3339 string format (the internet profile of ISO 8601).
|
||||
|
||||
The `jsonschema` crate handles all the structural validation (object
|
||||
properties, required fields, array items, enum values) — the custom
|
||||
keywords only need to validate the leaf type constraints. Each custom
|
||||
keyword implementation is ~10 lines.
|
||||
|
||||
### Read/write errors carry field paths
|
||||
|
||||
Read/write errors include the field path for debugging:
|
||||
|
||||
```rust
|
||||
// Example: reading a u32 from a buffer that's too short
|
||||
Err(TypedefError::Access {
|
||||
field_path: "header.version".to_string(),
|
||||
reason: "buffer too short: need 4 bytes at offset 12, have 2".to_string(),
|
||||
})
|
||||
```
|
||||
|
||||
This makes debugging binary format issues tractable — the error tells
|
||||
you exactly which field failed and why.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Single error type.** Consumers handle one `TypedefError` enum, not
|
||||
multiple error types from different engine phases.
|
||||
- **Field-path-carrying errors.** Read/write errors include the field
|
||||
path, making binary format debugging tractable.
|
||||
- **Validation is opt-in.** The consumer decides when to validate.
|
||||
High-throughput paths can skip validation; security-sensitive paths
|
||||
can validate every frame.
|
||||
- **jsonschema integration is clean.** The `ValidationError` is wrapped
|
||||
as-is — no re-interpretation, no information loss.
|
||||
- **Load-time build, access-time use.** The expensive work (schema
|
||||
parsing, validator compilation, offset computation) happens once at
|
||||
load time. Access-time operations are cheap (pointer casts, slice
|
||||
operations, length-prefix reads).
|
||||
|
||||
### Negative
|
||||
|
||||
- **`ValidationError<'static>` lifetime.** The `'static` lifetime on the
|
||||
`Validation` variant means the error cannot borrow from the buffer
|
||||
being validated. This is correct (the validator owns its schema
|
||||
reference) but may surprise readers who expect a shorter lifetime.
|
||||
- **No error recovery.** The engine does not attempt to recover from
|
||||
partial reads or writes. A buffer-too-short error on field N means
|
||||
fields N+1.. are also unreadable. This is inherent to binary formats
|
||||
— there is no "skip to next field" without a schema-driven parser.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/alknet-typedef/findings.md` §"Open Questions" — error
|
||||
handling strategy question (OQ 8)
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotations
|
||||
@@ -0,0 +1,114 @@
|
||||
# ADR-099: Int64/Uint64 as First-Class Kinds
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The typedef engine's kind set (ADR-095, ADR-097) tops out at 32-bit
|
||||
integers. The POC included `u64` read/write primitives, and the
|
||||
call-channels-unification research's own SFTP schema example uses
|
||||
`"TypeDef:Uint64"` for the `offset` field (`Read`/`Write` packets have
|
||||
`offset: u64`). Metatensor/safetensors `data_offsets` are also `u64`.
|
||||
|
||||
A `TypeDef:Uint64` variant was added to the `TypeDefKind` enum during
|
||||
implementation (the task decomposition correctly identified the gap),
|
||||
but without an ADR the addition was half-finished: `type_size()`
|
||||
returned `None`, the layout engines couldn't compute offsets for it, and
|
||||
the validator didn't register a `TypeDef:Uint64` keyword. The variant
|
||||
was then removed (commit `14d9cf2`) on the grounds that it was
|
||||
unintended and a latent panic — but the underlying gap is real: SFTP and
|
||||
metatensor, the two primary POC targets, both require 64-bit integers.
|
||||
|
||||
The presumed reason 64-bit integers were left out of the original
|
||||
specification is a JSON-level concern: `serde_json::Number` loses
|
||||
precision past 2^53 when parsing from JSON text. This is a
|
||||
*validation-layer* caveat, not a *layout-layer* one — the binary layout
|
||||
is 8 raw bytes, and `from_le_bytes`/`from_be_bytes` work correctly for
|
||||
the full `u64`/`i64` range. The validation concern is handled by
|
||||
accepting integer-form JSON values (the `jsonschema` crate's
|
||||
`as_i64`/`as_u64` methods handle the common range; values past 2^53 are
|
||||
a JSON representation limitation, not a typedef limitation).
|
||||
|
||||
## Decision
|
||||
|
||||
**Add `TypeDef:Int64` and `TypeDef:Uint64` as first-class kinds.**
|
||||
|
||||
Both are fixed-size (8 bytes), with natural alignment 8. They follow
|
||||
the schema's endianness annotation like all other fixed-size types.
|
||||
Read/write is via `data_access::read_i64`/`write_i64`/`read_u64`/
|
||||
`write_u64` (endian-aware, 8 bytes).
|
||||
|
||||
### Kind table additions
|
||||
|
||||
| Kind | TypeBox key | Rust type | Size | Alignment |
|
||||
|------|-------------|-----------|------|-----------|
|
||||
| `TInt64` | `TypeDef:Int64` | `i64` | 8 | 8 |
|
||||
| `TUint64` | `TypeDef:Uint64` | `u64` | 8 | 8 |
|
||||
|
||||
### Validation
|
||||
|
||||
The custom keyword validators check:
|
||||
- `TypeDef:Int64`: value must be an integer in `i64::MIN..=i64::MAX`
|
||||
(`-9223372036854775808` to `9223372036854775807`).
|
||||
- `TypeDef:Uint64`: value must be a non-negative integer in
|
||||
`0..=u64::MAX` (`0` to `18446744073709551615`).
|
||||
|
||||
The `jsonschema` crate's `as_i64`/`as_u64` handle the common range.
|
||||
JSON numbers past 2^53 lose precision in the JSON representation —
|
||||
this is a JSON limitation, not a typedef limitation. The binary
|
||||
representation (8 raw bytes) is always exact. A consumer that needs
|
||||
to validate the full 64-bit range from JSON should provide the value
|
||||
as a JSON integer (which `serde_json` preserves for values up to
|
||||
`u64::MAX`/`i64::MIN` when the `arbitrary_precision` feature is
|
||||
enabled, or when the value fits in `i64`/`u64` without the feature).
|
||||
|
||||
### `FieldValue` additions
|
||||
|
||||
`FieldValue::I64(i64)` and `FieldValue::U64(u64)` are added to the
|
||||
unified return type. The `SequentialReader`, `TypedefEngine::read_field`,
|
||||
and `TypedefEngine::write_field` dispatch on the new kinds.
|
||||
|
||||
### Kind count
|
||||
|
||||
The engine now has **19** first-class kinds (17 + Int64 + Uint64).
|
||||
`TypeDefKind::is_fixed_size()` returns `true` for both new kinds.
|
||||
`type_size()` returns `Some(8)`. `natural_alignment()` returns `8`.
|
||||
`needs_endian()` returns `true`.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Unblocks the two primary POC targets.** SFTP `Read`/`Write` packets
|
||||
(`offset: u64`) and metatensor `data_offsets` (`u64`) are now
|
||||
expressible in typedef schemas.
|
||||
- **Completes the half-finished addition.** The `TypeDefKind` enum,
|
||||
`data_access` primitives, and `FieldValue` variants for 64-bit
|
||||
integers now have matching layout, validator, and engine support.
|
||||
- **No new design surface.** Int64/Uint64 are fixed-size types that
|
||||
follow all existing patterns (endianness, alignment, zero-copy
|
||||
read/write). They are mechanical additions.
|
||||
|
||||
### Negative
|
||||
|
||||
- **JSON precision caveat.** Values past 2^53 lose precision in the
|
||||
JSON representation (not in the binary representation). This is a
|
||||
JSON limitation, not a typedef limitation, but it means the
|
||||
validation layer cannot perfectly round-trip the full 64-bit range
|
||||
through JSON `Number` without `arbitrary_precision`. In practice,
|
||||
SFTP offsets and tensor data offsets are well within 2^53.
|
||||
- **Two more kinds to maintain.** The kind table, validator
|
||||
registration, `FieldValue` enum, and dispatch arms all grow by two
|
||||
variants. This is the cost of completeness.
|
||||
|
||||
## References
|
||||
|
||||
- `docs/research/call-channels-unification/findings.md` §"russh-sftp" —
|
||||
the SFTP schema with `"offset": { "TypeDef:Uint64": true }`
|
||||
- `docs/research/alknet-typedef/findings.md` §"POC 1" — the POC included
|
||||
u64 read/write
|
||||
- [ADR-095](095-alknet-typedef-purpose-scope-jsonschema-engine.md) —
|
||||
purpose and scope (the kind set)
|
||||
- [ADR-097](097-schema-annotations.md) — schema annotations
|
||||
(endianness applies to the new kinds)
|
||||
+112
@@ -0,0 +1,112 @@
|
||||
# ADR-100: Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The aligned static layout mode (ADR-096) is designed for mmap-friendly
|
||||
formats: fields have fixed positions with natural alignment padding,
|
||||
enabling random access by field path without parsing preceding fields.
|
||||
|
||||
The spec (layout-engine.md) says variable-length fields in aligned mode
|
||||
get a 4-byte length prefix at a known offset, and "the variable data
|
||||
lives outside the static layout — either immediately after the fixed
|
||||
fields (inline length-prefixing) or in a separate data region (offset
|
||||
indirection)."
|
||||
|
||||
The implementation has a bug: `OffsetMap::compute` reserves only 4 bytes
|
||||
for an inline length-prefixed variable field (the length prefix), but
|
||||
`TypedefEngine::write_field` for a `String`/`Bytes` field calls
|
||||
`data_access::write_string` at `range.start`, which writes
|
||||
`[4-byte length][data]` inline — clobbering every subsequent field. The
|
||||
`read_field` path has the mirror behavior (reads inline), so the engine
|
||||
is self-consistent but only works correctly when the variable field is
|
||||
the last field in the struct (no subsequent field to clobber).
|
||||
|
||||
Concretely, `{name: String, id: Uint32}` in aligned mode maps
|
||||
`name → 0..4`, `id → 4..8`. Writing `"hello"` to `name` writes
|
||||
`[5,0,0,0,h,e,l,l,o]` at offset 0, overwriting `id`'s range with
|
||||
`hello`. All existing tests happen to put the variable field last, so
|
||||
the bug is latent.
|
||||
|
||||
The spec's "data region after fixed fields" model (where variable data
|
||||
lives after all fixed fields) is the correct design for aligned mode,
|
||||
but implementing it would require a two-region layout (fixed fields +
|
||||
variable data region) with the `OffsetMap` tracking both the prefix
|
||||
position and the data position. This is a significant design addition
|
||||
for a use case that doesn't exist yet — real aligned-format consumers
|
||||
(metatensor, safetensors) use `maxLength` reservation or
|
||||
`offset-indirect` encoding for variable data, not inline
|
||||
length-prefixing.
|
||||
|
||||
## Decision
|
||||
|
||||
**Reject non-final inline length-prefixed variable fields in aligned
|
||||
static mode at `OffsetMap::compute` time.**
|
||||
|
||||
A variable-length field (`TypeDef:String`, `TypeDef:Bytes`,
|
||||
`TypeDef:Timestamp`, `TypeDef:Record`) in aligned static mode that uses
|
||||
the default inline length-prefixing strategy (no `maxLength`, no
|
||||
`offset-indirect`) must be the last field in its struct. If a non-final
|
||||
inline length-prefixed variable field is encountered,
|
||||
`OffsetMap::compute` returns `TypedefError::Offset` with a message
|
||||
explaining that non-final variable fields in aligned mode require
|
||||
`maxLength` (fixed-size reservation) or `"encoding": "offset-indirect"`
|
||||
(offset indirection).
|
||||
|
||||
This is a validation-time rejection (schema load time), not a runtime
|
||||
check. The consumer learns about the problem when compiling the schema,
|
||||
not when writing data.
|
||||
|
||||
### What is NOT rejected
|
||||
|
||||
- Inline length-prefixed variable fields that are the last field in
|
||||
their struct — these are fine (no subsequent field to clobber).
|
||||
- `maxLength` reservation and `offset-indirect` encoding in any
|
||||
position — these make the field fixed-size from the layout
|
||||
perspective (known size at a known offset), so they don't clobber.
|
||||
- Inline length-prefixed variable fields in packed sequential mode —
|
||||
packed mode doesn't have fixed offsets; variable fields shift
|
||||
subsequent fields by design.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **Eliminates a silent data-corruption bug.** A consumer that writes
|
||||
a non-final string in aligned mode currently clobbers subsequent
|
||||
fields with no error. After this fix, the schema is rejected at
|
||||
compile time.
|
||||
- **Matches real aligned-format usage.** mmap-friendly formats use
|
||||
`maxLength` or `offset-indirect` for variable data; inline
|
||||
length-prefixing in aligned mode is only meaningful as the last
|
||||
field.
|
||||
- **Simple to implement.** A single check in `compute_struct` (is this
|
||||
variable field non-final and using inline length-prefixing? → reject).
|
||||
No two-region layout needed.
|
||||
- **Defers the two-region design without blocking consumers.** If a
|
||||
future consumer needs inline length-prefixing in non-final position
|
||||
in aligned mode, the two-region layout can be implemented then. The
|
||||
rejection is reversible (remove the check, add the two-region logic).
|
||||
|
||||
### Negative
|
||||
|
||||
- **A schema that worked before (silently corrupting data) now fails
|
||||
at compile time.** This is the correct behavior — the schema was
|
||||
always broken, it just wasn't caught.
|
||||
- **The "data region after fixed fields" model from the spec is not
|
||||
implemented.** A consumer that wants inline variable data in a
|
||||
non-final position must use packed mode or wait for the two-region
|
||||
layout. This is acceptable for v1 — no current consumer needs it.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes (aligned static mode's variable-length handling)
|
||||
- [ADR-097](097-schema-annotations.md) — the three variable-length
|
||||
encoding strategies (`maxLength`, `offset-indirect`, inline
|
||||
length-prefixing)
|
||||
- `../layout-engine.md` §"Variable-length
|
||||
fields in aligned mode" — the spec's "data region after fixed fields"
|
||||
description
|
||||
@@ -0,0 +1,103 @@
|
||||
# ADR-101: Packed-Mode Read API — Engine as SequentialReader Factory
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
`TypedefEngine` stores a `SequentialReader` inside its `Layout::Packed`
|
||||
variant. The engine exposes it via
|
||||
`engine.sequential_reader() -> Option<&SequentialReader>`.
|
||||
|
||||
The problem: `SequentialReader`'s read methods (`read_next`,
|
||||
`read_field`, `reset`) all take `&mut self` — they mutate the reader's
|
||||
internal cursor (`field_index`, `position`). But the engine hands out
|
||||
`&SequentialReader` (a shared reference), which cannot be used to call
|
||||
`&mut self` methods. The accessor can only give the consumer
|
||||
`position()` and `endian()` (the `&self` methods) — the actual read
|
||||
API is unreachable.
|
||||
|
||||
This makes the engine's packed read-side dead API. A consumer that
|
||||
wants to read a packed buffer must construct their own
|
||||
`SequentialReader::new(&schema)` from the schema, bypassing the engine
|
||||
entirely. The stored reader is dead weight.
|
||||
|
||||
Three options were considered:
|
||||
1. **Factory method** — the engine provides a method that returns an
|
||||
owned fresh `SequentialReader` (reconstructed from the stored
|
||||
schema). The consumer owns the reader and drives it with `&mut self`.
|
||||
2. **Interior mutability** — wrap the reader in `Mutex` or `RefCell`
|
||||
so `&SequentialReader` can be upgraded to `&mut`. Adds overhead and
|
||||
complexity for mutable cursor state that the consumer legitimately
|
||||
wants to own.
|
||||
3. **`sequential_reader_mut()`** — return `&mut SequentialReader`.
|
||||
Requires `&mut self` on the engine, which is overly restrictive
|
||||
(the consumer may share the engine across threads or hold it behind
|
||||
an `Arc`).
|
||||
|
||||
## Decision
|
||||
|
||||
**The engine is a `SequentialReader` factory.** Replace
|
||||
`sequential_reader() -> Option<&SequentialReader>` with
|
||||
`sequential_reader() -> Option<SequentialReader>` — the method returns
|
||||
an owned fresh reader, reconstructed from the stored schema.
|
||||
|
||||
```rust
|
||||
impl TypedefEngine {
|
||||
/// Construct a fresh SequentialReader for packed-mode reads.
|
||||
/// Returns None if compiled in aligned mode.
|
||||
pub fn sequential_reader(&self) -> Option<SequentialReader>;
|
||||
}
|
||||
```
|
||||
|
||||
Each call returns a new reader with the cursor at position 0. The
|
||||
consumer owns the reader and calls `read_next`/`read_field`/`reset` on
|
||||
it directly. The engine still stores its own reader (used for schema
|
||||
validation during construction), but no longer exposes it by
|
||||
reference.
|
||||
|
||||
The same applies to `LayoutBuilder`: `layout_builder()` returns
|
||||
`Option<&LayoutBuilder>` which is fine — `LayoutBuilder::build` takes
|
||||
`&self`, so the shared reference is usable. No change needed for the
|
||||
write-side.
|
||||
|
||||
### Cost
|
||||
|
||||
`SequentialReader::new` clones the top-level struct's field schemas (a
|
||||
`Vec<(String, Value)>` of the `properties` entries) and clones the
|
||||
schema itself. This is cheap — a struct has a small number of fields
|
||||
(SFTP's largest packet has 5). The construction cost is negligible
|
||||
compared to the cost of reading a buffer.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **The packed read API is now usable.** A consumer calls
|
||||
`engine.sequential_reader()` to get an owned reader and drives it
|
||||
directly. No dead API.
|
||||
- **No interior mutability overhead.** The reader's mutable cursor
|
||||
state is owned by the consumer, not shared through a lock.
|
||||
- **Thread-safe engine.** The engine remains `Send + Sync` (it only
|
||||
exposes `&self` methods). The reader is owned by the calling thread.
|
||||
- **Simple.** One method signature change. The stored reader in
|
||||
`Layout::Packed` can be removed (it was only used for schema
|
||||
validation during construction, which is done by the time the
|
||||
consumer calls `sequential_reader()`).
|
||||
|
||||
### Negative
|
||||
|
||||
- **Each call to `sequential_reader()` allocates a new reader.** The
|
||||
cost is a `Vec` of field schemas + a schema clone. Acceptable for
|
||||
the use case (one reader per buffer read).
|
||||
- **The engine no longer holds a live reader.** If a future use case
|
||||
needs to share a reader's cursor state across calls, the consumer
|
||||
must manage that themselves. This is the correct separation — cursor
|
||||
state is consumer-owned, not engine-owned.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — packed
|
||||
sequential mode (`SequentialReader` as the read-side)
|
||||
- `../data-access.md` §"Higher-level
|
||||
read/write" — the `SequentialReader` API
|
||||
@@ -0,0 +1,111 @@
|
||||
# ADR-102: Reject TUnion in Aligned Mode for v1
|
||||
|
||||
## Status
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The aligned static layout mode (ADR-096) computes fixed byte positions
|
||||
for each field, enabling random access by field path. `TUnion` in
|
||||
aligned mode has three implementation problems:
|
||||
|
||||
1. **No variant field offsets.** Only the `__discriminator` byte range
|
||||
is recorded in the `OffsetMap`. Variant field offsets are not
|
||||
available anywhere in aligned mode — the consumer must recompute
|
||||
them by hand. This makes `TypedefEngine::read_field` on a union
|
||||
variant field impossible.
|
||||
|
||||
2. **`find_discriminator_field` takes the first variant's offset.** For
|
||||
a field-name discriminator, the code probes the first variant that
|
||||
contains the discriminator field and records that offset globally.
|
||||
If variants order fields differently, the discriminator sits at
|
||||
different offsets per variant and the recorded range is silently
|
||||
wrong. The code should validate that the offset is identical across
|
||||
all variants (or require the discriminator field to be first).
|
||||
|
||||
3. **Byte-discriminator union total misaligns the variant.** The union
|
||||
total is `disc_off + disc_size + variant_max_size`, but the variant
|
||||
was probed from offset 0 with alignment. A `u8` discriminator
|
||||
before a `u32`-bearing variant produces a variant region that
|
||||
starts at an unaligned offset in a mode whose entire purpose is
|
||||
alignment.
|
||||
|
||||
The real question is whether `TUnion` in aligned mode is even needed.
|
||||
The two consumer profiles are:
|
||||
|
||||
- **Protocol consumers** (SFTP, call protocol event types): use packed
|
||||
sequential mode. `TUnion` with byte-offset discriminators is the
|
||||
core dispatch mechanism. This is well-supported.
|
||||
- **mmap consumers** (metatensor, safetensors): use aligned static
|
||||
mode. These formats are structs and arrays of structs — they don't
|
||||
use tagged unions. A tensor file has a header struct with tensor
|
||||
descriptors, not a "which variant is this?" dispatch.
|
||||
|
||||
`TUnion` in aligned mode is a combination that no current or planned
|
||||
consumer needs. Shipping broken semantics for an unused use case is
|
||||
worse than rejecting it clearly.
|
||||
|
||||
## Decision
|
||||
|
||||
**Reject `TUnion` in aligned static mode for v1.**
|
||||
|
||||
`OffsetMap::compute` returns `TypedefError::Offset` when it encounters
|
||||
a `TypeDef:Union` field, with a message explaining that unions are not
|
||||
supported in aligned mode and the consumer should use packed mode (or
|
||||
restructure as a struct with an explicit discriminator field).
|
||||
|
||||
This is a schema-load-time rejection. The consumer learns about the
|
||||
problem when compiling the schema, not at runtime.
|
||||
|
||||
### What is NOT rejected
|
||||
|
||||
- `TUnion` in packed sequential mode — this is the core use case
|
||||
(SFTP `Packet` dispatch, call protocol event types) and is fully
|
||||
supported by `LayoutBuilder` and `SequentialReader`.
|
||||
- `TStruct`, `TArray`, and all primitive kinds in aligned mode — these
|
||||
are the mmap-format primitives and are fully supported.
|
||||
|
||||
### Reversal
|
||||
|
||||
This is a two-way door. If a future mmap-format consumer needs tagged
|
||||
unions in aligned mode, the rejection can be lifted and the three
|
||||
implementation problems fixed. The fix would require:
|
||||
- Recording per-variant field offsets in the `OffsetMap` (which
|
||||
variant's offsets to record when variants have different layouts?).
|
||||
- Validating that field-name discriminators have identical offsets
|
||||
across all variants.
|
||||
- Aligning the variant region correctly after the byte discriminator.
|
||||
|
||||
These are design questions that should be answered when the use case
|
||||
arrives, not speculatively now.
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- **No broken semantics shipped.** The three implementation problems
|
||||
are removed from the API surface rather than silently producing
|
||||
wrong offsets.
|
||||
- **Clear scope boundary.** Aligned mode is for structs and arrays;
|
||||
packed mode is for protocols (including union dispatch). The
|
||||
consumer chooses the mode based on the use case.
|
||||
- **Reversible.** When a real consumer needs aligned-mode unions, the
|
||||
rejection is lifted and the design questions are worked through with
|
||||
a concrete use case.
|
||||
|
||||
### Negative
|
||||
|
||||
- **A schema with a `TUnion` field cannot be compiled in aligned
|
||||
mode.** A consumer that wants both aligned layout and union dispatch
|
||||
must use packed mode or restructure. No current consumer needs this.
|
||||
- **The aligned-mode union code in `offset_map.rs` is dead.** It can
|
||||
be removed or left as a reference for when the rejection is lifted.
|
||||
Removing it is cleaner.
|
||||
|
||||
## References
|
||||
|
||||
- [ADR-096](096-two-layout-modes-packed-vs-aligned.md) — the two layout
|
||||
modes
|
||||
- [ADR-097](097-schema-annotations.md) §4 — TUnion discriminators
|
||||
- `../layout-engine.md` §"TUnion" —
|
||||
aligned-mode union sizing
|
||||
Reference in new issue
Block a user