fix(typedef): fix code examples and cross-doc inconsistencies from second review

- Fix code examples hardcoding little-endian: read_string, write_string,
  read_string_indirect now take endian parameter and use match on Endian.
- Fix TUnion dispatch examples: remove undefined functions (read_u8,
  read_field, read_struct, read_f32_raw), remove Value returns
  (contradicts 'no intermediate Value tree'), add endian-aware
  discriminator reading for Uint8/Uint16/Uint32.
- Fix read_f32 example: inline the endian-aware conversion instead of
  calling undefined read_f32_raw; document it as aligned-mode only.
- Fix architecture README: 16→17 kinds in schema-layer and validation
  descriptions.
- Fix ADR-095: clarify validation operates on Value instances, not raw
  byte buffers directly. Fix 'defense in depth' paragraph.
- Fix ADR-097: add §3a defining TRecord 'values' property shape.
- Fix ADR-098: TTimestamp format ISO 8601→RFC 3339.
- Tighten OQ-071 impacts field: state what IS blocked, not just what
  isn't.
This commit is contained in:
deepseek-v4-pro committed 2026-07-21 08:42:13 +00:00
1 parent 01cc3a0367
commit dd232c3d47
6 files changed
+172 -50

No files matched your search

+2 -2
View File
@@ -298,10 +298,10 @@ adapter location map is now consistent: all HTTP-backed adapters
| [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved | | [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved |
| [crates/typedef/README.md](crates/typedef/README.md) | draft | alknet-typedef crate — binary struct engine; JSON Schema with `TypeDef:*` custom keywords → offset map + read/write + validation | | [crates/typedef/README.md](crates/typedef/README.md) | draft | alknet-typedef crate — binary struct engine; JSON Schema with `TypeDef:*` custom keywords → offset map + read/write + validation |
| [crates/typedef/overview.md](crates/typedef/overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries | | [crates/typedef/overview.md](crates/typedef/overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 16 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations | | [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 17 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md) | draft | Offset computation, two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling | | [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md) | draft | Offset computation, two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [crates/typedef/data-access.md](crates/typedef/data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading | | [crates/typedef/data-access.md](crates/typedef/data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation | | [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 17 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation |
## ADR Table ## ADR Table
+132 -37
View File
@@ -76,16 +76,33 @@ For variable-length types with inline length-prefixing (the default):
```rust ```rust
// Read a length-prefixed string // Read a length-prefixed string
fn read_string<'a>(buffer: &'a [u8], offset: usize) -> &'a str { fn read_string<'a>(buffer: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
let len = u32::from_le_bytes(buffer[offset..offset+4].try_into().unwrap()) as usize; let len_bytes: [u8; 4] = buffer[offset..offset+4].try_into()
std::str::from_utf8(&buffer[offset+4..offset+4+len]).unwrap() .map_err(|_| TypedefError::Access { /* ... */ })?;
let len = match endian {
Endian::Little => u32::from_le_bytes(len_bytes),
Endian::Big => u32::from_be_bytes(len_bytes),
} as usize;
let data = buffer.get(offset+4..offset+4+len)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
std::str::from_utf8(data)
.map_err(|e| TypedefError::Access { /* ... */ })
} }
// Write a length-prefixed string // Write a length-prefixed string
fn write_string(buffer: &mut [u8], offset: usize, value: &str) { fn write_string(buffer: &mut [u8], offset: usize, value: &str, endian: Endian) -> Result<(), TypedefError> {
let data = value.as_bytes(); let data = value.as_bytes();
buffer[offset..offset+4].copy_from_slice(&(data.len() as u32).to_le_bytes()); let len_bytes = match endian {
buffer[offset+4..offset+4+data.len()].copy_from_slice(data); Endian::Little => (data.len() as u32).to_le_bytes(),
Endian::Big => (data.len() as u32).to_be_bytes(),
};
buffer.get_mut(offset..offset+4)
.ok_or_else(|| TypedefError::Access { /* ... */ })?
.copy_from_slice(&len_bytes);
buffer.get_mut(offset+4..offset+4+data.len())
.ok_or_else(|| TypedefError::Access { /* ... */ })?
.copy_from_slice(data);
Ok(())
} }
``` ```
@@ -104,10 +121,19 @@ For variable-length types with offset indirection (opt-in):
```rust ```rust
// Read an offset-indirect string // Read an offset-indirect string
fn read_string_indirect(data_region: &[u8], offset: usize) -> &str { fn read_string_indirect<'a>(data_region: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
let ptr_offset = u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()) as usize; let ptr_offset = match endian {
let ptr_length = u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()) as usize; Endian::Little => u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()),
std::str::from_utf8(&data_region[ptr_offset..ptr_offset+ptr_length]).unwrap() Endian::Big => u32::from_be_bytes(data_region[offset..offset+4].try_into().unwrap()),
} as usize;
let ptr_length = match endian {
Endian::Little => u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
Endian::Big => u32::from_be_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
} as usize;
let data = data_region.get(ptr_offset..ptr_offset+ptr_length)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
std::str::from_utf8(data)
.map_err(|e| TypedefError::Access { /* ... */ })
} }
``` ```
@@ -124,28 +150,62 @@ differs by discriminator kind (ADR-097).
### Byte-offset discriminator ### Byte-offset discriminator
```rust ```rust
fn read_union(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> { /// Read the discriminator value from a byte-offset TUnion.
let disc = &schema["discriminator"]; /// Returns the mapping key (as a string) so the consumer can look up
let offset = disc["offset"].as_u64().unwrap() as usize; /// the variant schema and read the variant's fields.
let disc_type = disc["type"].as_str().unwrap(); // e.g., "TypeDef:Uint8" fn read_union_discriminator(
buffer: &[u8],
schema: &Value,
endian: Endian,
) -> Result<String, TypedefError> {
let disc = schema["discriminator"].as_object()
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
let offset = disc["offset"].as_u64().unwrap_or(0) as usize;
let disc_type = disc["type"].as_str().unwrap_or("TypeDef:Uint8");
// Read the discriminator value let (disc_value, disc_size) = match disc_type {
let disc_value: u8 = read_u8(buffer, offset); "TypeDef:Uint8" => {
let key = disc_value.to_string(); // "5", "6", "101" let b = *buffer.get(offset)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
(b as u32, 1)
}
"TypeDef:Uint16" => {
let bytes: [u8; 2] = buffer[offset..offset+2].try_into().unwrap();
let v = match endian {
Endian::Little => u16::from_le_bytes(bytes),
Endian::Big => u16::from_be_bytes(bytes),
};
(v as u32, 2)
}
"TypeDef:Uint32" => {
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
let v = match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
};
(v, 4)
}
_ => return Err(TypedefError::Schema(format!("unsupported discriminator type: {disc_type}"))),
};
// Look up the variant schema let key = disc_value.to_string();
let mapping = &schema["mapping"]; if schema["mapping"].as_object().map_or(false, |m| m.contains_key(&key)) {
let variant_schema = &mapping[&key]; Ok(key)
} else {
// Read the variant struct starting at offset + discriminator_size Err(TypedefError::Access {
let variant_offset = offset + 1; // discriminator_size for Uint8 field_path: "__discriminator".into(),
read_struct(buffer, variant_offset, variant_schema) reason: format!("unknown discriminator value: {disc_value}"),
})
}
} }
``` ```
The discriminator is a fixed-size integer at a known byte offset. The The discriminator is a fixed-size integer at a known byte offset. The
mapping keys are stringified integers. The variant struct starts at mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`. `offset + discriminator_size`. After reading the discriminator, the
consumer looks up the variant schema and reads the variant's fields
using the normal typed read functions (e.g., `read_u32`, `read_string`)
at `offset + discriminator_size`.
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
1..N are the variant struct. The call protocol's 5 event types 1..N are the variant struct. The call protocol's 5 event types
@@ -154,24 +214,53 @@ This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
### Field-name discriminator ### Field-name discriminator
```rust ```rust
fn read_union_field(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> { /// Read the discriminator value from a field-name TUnion.
let disc = &schema["discriminator"]; /// The discriminator is a named field within the struct — its offset
let field_name = disc["name"].as_str().unwrap(); // e.g., "type" /// is computed like any other field. The consumer reads the field's
/// value, looks up the variant schema, then reads the variant's fields.
fn read_union_field_discriminator(
buffer: &[u8],
schema: &Value,
offset_map: &OffsetMap,
endian: Endian,
) -> Result<String, TypedefError> {
let disc = schema["discriminator"].as_object()
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
let field_name = disc["name"].as_str()
.ok_or_else(|| TypedefError::Schema("discriminator has no 'name'".into()))?;
// Read the discriminator field like any other field // Read the discriminator field at its computed offset.
let disc_value = read_field(buffer, field_name, schema)?; // The field's TypeDef kind determines how to read it (typically a string).
let field_schema = schema["properties"].get(field_name)
.ok_or_else(|| TypedefError::Schema(format!("discriminator field '{field_name}' not found")))?;
let kind = get_typedef_kind(field_schema)
.ok_or_else(|| TypedefError::Schema("discriminator field has no TypeDef kind".into()))?;
// Look up the variant schema match kind {
let mapping = &schema["mapping"]; "TypeDef:String" => {
let variant_schema = &mapping[disc_value.as_str().unwrap()]; let range = offset_map.get(field_name)
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
// Read the variant struct read_string(buffer, range.start, endian)
read_struct(buffer, variant_offset, variant_schema) .map(|s| s.to_string())
}
"TypeDef:Uint8" => {
let range = offset_map.get(field_name)
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
Ok(buffer[range.start].to_string())
}
_ => Err(TypedefError::Schema(format!(
"unsupported discriminator field type: {kind}"
))),
}
} }
``` ```
The discriminator is a named field within the struct. Its offset is The discriminator is a named field within the struct. Its offset is
computed like any other field. The mapping keys are string values. computed like any other field. The mapping keys are string values.
After reading the discriminator, the consumer looks up the variant
schema and reads the variant's fields starting at the end of the
discriminator field (or at the start of the union buffer if the
discriminator is the first field).
## Field Paths ## Field Paths
@@ -180,6 +269,8 @@ The `OffsetMap` stores fully-qualified paths. The read/write functions
accept a field path and look up the byte range: accept a field path and look up the byte range:
```rust ```rust
/// Read an f32 field by path. This is the aligned-mode path (uses OffsetMap).
/// In packed mode, the consumer uses SequentialReader instead.
fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> { fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> {
let range = self.offset_map.get(field_path) let range = self.offset_map.get(field_path)
.ok_or_else(|| TypedefError::Offset { .ok_or_else(|| TypedefError::Offset {
@@ -192,7 +283,11 @@ fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError>
reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()), reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()),
}); });
} }
Ok(read_f32_raw(buffer, range.start, self.endian)) let bytes: [u8; 4] = buffer[range.start..range.end].try_into().unwrap();
Ok(match self.endian {
Endian::Little => f32::from_le_bytes(bytes),
Endian::Big => f32::from_be_bytes(bytes),
})
} }
``` ```
@@ -65,8 +65,12 @@ functions, and validation — all driven by the schema.
2. **Read/write functions** — given a `&[u8]` buffer and a field path, 2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types). read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset. Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that a 3. **Validation** — via `jsonschema` custom keywords, validates that data
buffer's bytes match the schema's type constraints. conforms to the schema's type constraints. The jsonschema validator
operates on `serde_json::Value` instances (JSON representations), not
raw byte buffers directly. A consumer that wants to validate a binary
buffer reads it into a `Value` tree via the data access layer, then
validates that `Value` against the jsonschema validator.
**The heavy lifting is done by the `jsonschema` crate (validation) and **The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing).** The novel code is the offset computation `serde_json` (schema parsing).** The novel code is the offset computation
@@ -114,10 +118,11 @@ read/write) is already allocation-free. See OQ-070.
under `$defs`. That JSON feeds directly into `jsonschema::validator_for` under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
on the Rust side. Zero translation. The same schema validates in both on the Rust side. Zero translation. The same schema validates in both
ecosystems. ecosystems.
- **Defense in depth.** Schema validation at the byte level — a malformed - **Defense in depth.** Schema validation via jsonschema custom keywords —
binary payload fails validation before any consumer touches it. The a malformed binary payload can be read into a `Value` tree via the data
`jsonschema` crate's compiled validators are fast enough to run on access layer and validated against the schema before any consumer
every incoming frame. touches it. The `jsonschema` crate's compiled validators are fast enough
to run on every incoming frame.
### Negative ### Negative
@@ -129,6 +129,26 @@ reserving worst-case space.
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`, types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
`TypeDef:Record`, `TypeDef:Timestamp`. `TypeDef:Record`, `TypeDef:Timestamp`.
### 3a. TRecord value type
`TypeDef:Record` is a string-keyed map. The value type is declared via
the `"values"` property in the schema:
```json
{
"TypeDef:Record": true,
"values": { "TypeDef:Float32": true }
}
```
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
values in the record. All values share the same type.
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value_len: u32][value_bytes]...`.
- The count prefix respects the schema's endianness.
- In aligned static mode with `maxLength`, the entire record is reserved
at `maxLength` bytes (zero-padded).
### 4. TUnion discriminators ### 4. TUnion discriminators
**Two discriminator kinds: byte-offset (protocol dispatch) and **Two discriminator kinds: byte-offset (protocol dispatch) and
@@ -95,7 +95,7 @@ Each `TypeDef:*` kind gets a `Keyword` implementation registered via
- **`TypeDef:Union`**: discriminator value membership in the mapping. - **`TypeDef:Union`**: discriminator value membership in the mapping.
- **`TypeDef:Array`**: element type conformance. - **`TypeDef:Array`**: element type conformance.
- **`TypeDef:Boolean`**: value is `true` or `false`. - **`TypeDef:Boolean`**: value is `true` or `false`.
- **`TypeDef:Timestamp`**: ISO 8601 string format. - **`TypeDef:Timestamp`**: RFC 3339 string format (the internet profile of ISO 8601).
The `jsonschema` crate handles all the structural validation (object The `jsonschema` crate handles all the structural validation (object
properties, required fields, array items, enum values) — the custom properties, required fields, array items, enum values) — the custom
@@ -8,10 +8,12 @@
- **Door type**: Two-way (additive — a builder API can be added without - **Door type**: Two-way (additive — a builder API can be added without
changing the existing JSON-consumption path) changing the existing JSON-consumption path)
- **Priority**: medium - **Priority**: medium
- **Impacts**: No current consumer. Schemas are authored in TypeBox (JS) - **Impacts**: Blocks programmatic schema construction in Rust without a
or hand-written JSON for v1. A Rust builder API would enable JS toolchain. Any consumer that wants to build typedef schemas at
programmatic schema construction in Rust without depending on a JS runtime from Rust code (rather than loading pre-authored JSON) must
toolchain, but no current consumer needs this. construct the JSON manually or depend on TypeBox. Does NOT block any
current consumer — all v1 consumers (SFTP, metatensor, binary call
frames, TTY negotiation) use pre-authored schemas.
- **Blocked on**: A concrete need for programmatic schema construction - **Blocked on**: A concrete need for programmatic schema construction
in Rust. The current consumers (SFTP, metatensor, binary call frames, in Rust. The current consumers (SFTP, metatensor, binary call frames,
TTY negotiation) all have schemas that can be hand-written or generated TTY negotiation) all have schemas that can be hand-written or generated