fix(typedef): fix code examples and cross-doc inconsistencies from second review

- Fix code examples hardcoding little-endian: read_string, write_string,
  read_string_indirect now take endian parameter and use match on Endian.
- Fix TUnion dispatch examples: remove undefined functions (read_u8,
  read_field, read_struct, read_f32_raw), remove Value returns
  (contradicts 'no intermediate Value tree'), add endian-aware
  discriminator reading for Uint8/Uint16/Uint32.
- Fix read_f32 example: inline the endian-aware conversion instead of
  calling undefined read_f32_raw; document it as aligned-mode only.
- Fix architecture README: 16→17 kinds in schema-layer and validation
  descriptions.
- Fix ADR-095: clarify validation operates on Value instances, not raw
  byte buffers directly. Fix 'defense in depth' paragraph.
- Fix ADR-097: add §3a defining TRecord 'values' property shape.
- Fix ADR-098: TTimestamp format ISO 8601→RFC 3339.
- Tighten OQ-071 impacts field: state what IS blocked, not just what
  isn't.
This commit is contained in:
deepseek-v4-pro committed 2026-07-21 08:42:13 +00:00
1 parent 01cc3a0367
commit dd232c3d47
6 files changed
+172 -50

No files matched your search

+2 -2
View File
@@ -298,10 +298,10 @@ adapter location map is now consistent: all HTTP-backed adapters
| [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved |
| [crates/typedef/README.md](crates/typedef/README.md) | draft | alknet-typedef crate — binary struct engine; JSON Schema with `TypeDef:*` custom keywords → offset map + read/write + validation |
| [crates/typedef/overview.md](crates/typedef/overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 16 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 17 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md) | draft | Offset computation, two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [crates/typedef/data-access.md](crates/typedef/data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation |
| [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 17 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation |
## ADR Table
+132 -37
View File
@@ -76,16 +76,33 @@ For variable-length types with inline length-prefixing (the default):
```rust
// Read a length-prefixed string
fn read_string<'a>(buffer: &'a [u8], offset: usize) -> &'a str {
let len = u32::from_le_bytes(buffer[offset..offset+4].try_into().unwrap()) as usize;
std::str::from_utf8(&buffer[offset+4..offset+4+len]).unwrap()
fn read_string<'a>(buffer: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
let len_bytes: [u8; 4] = buffer[offset..offset+4].try_into()
.map_err(|_| TypedefError::Access { /* ... */ })?;
let len = match endian {
Endian::Little => u32::from_le_bytes(len_bytes),
Endian::Big => u32::from_be_bytes(len_bytes),
} as usize;
let data = buffer.get(offset+4..offset+4+len)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
std::str::from_utf8(data)
.map_err(|e| TypedefError::Access { /* ... */ })
}
// Write a length-prefixed string
fn write_string(buffer: &mut [u8], offset: usize, value: &str) {
fn write_string(buffer: &mut [u8], offset: usize, value: &str, endian: Endian) -> Result<(), TypedefError> {
let data = value.as_bytes();
buffer[offset..offset+4].copy_from_slice(&(data.len() as u32).to_le_bytes());
buffer[offset+4..offset+4+data.len()].copy_from_slice(data);
let len_bytes = match endian {
Endian::Little => (data.len() as u32).to_le_bytes(),
Endian::Big => (data.len() as u32).to_be_bytes(),
};
buffer.get_mut(offset..offset+4)
.ok_or_else(|| TypedefError::Access { /* ... */ })?
.copy_from_slice(&len_bytes);
buffer.get_mut(offset+4..offset+4+data.len())
.ok_or_else(|| TypedefError::Access { /* ... */ })?
.copy_from_slice(data);
Ok(())
}
```
@@ -104,10 +121,19 @@ For variable-length types with offset indirection (opt-in):
```rust
// Read an offset-indirect string
fn read_string_indirect(data_region: &[u8], offset: usize) -> &str {
let ptr_offset = u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()) as usize;
let ptr_length = u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()) as usize;
std::str::from_utf8(&data_region[ptr_offset..ptr_offset+ptr_length]).unwrap()
fn read_string_indirect<'a>(data_region: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
let ptr_offset = match endian {
Endian::Little => u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()),
Endian::Big => u32::from_be_bytes(data_region[offset..offset+4].try_into().unwrap()),
} as usize;
let ptr_length = match endian {
Endian::Little => u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
Endian::Big => u32::from_be_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
} as usize;
let data = data_region.get(ptr_offset..ptr_offset+ptr_length)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
std::str::from_utf8(data)
.map_err(|e| TypedefError::Access { /* ... */ })
}
```
@@ -124,28 +150,62 @@ differs by discriminator kind (ADR-097).
### Byte-offset discriminator
```rust
fn read_union(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
let disc = &schema["discriminator"];
let offset = disc["offset"].as_u64().unwrap() as usize;
let disc_type = disc["type"].as_str().unwrap(); // e.g., "TypeDef:Uint8"
/// Read the discriminator value from a byte-offset TUnion.
/// Returns the mapping key (as a string) so the consumer can look up
/// the variant schema and read the variant's fields.
fn read_union_discriminator(
buffer: &[u8],
schema: &Value,
endian: Endian,
) -> Result<String, TypedefError> {
let disc = schema["discriminator"].as_object()
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
let offset = disc["offset"].as_u64().unwrap_or(0) as usize;
let disc_type = disc["type"].as_str().unwrap_or("TypeDef:Uint8");
// Read the discriminator value
let disc_value: u8 = read_u8(buffer, offset);
let key = disc_value.to_string(); // "5", "6", "101"
let (disc_value, disc_size) = match disc_type {
"TypeDef:Uint8" => {
let b = *buffer.get(offset)
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
(b as u32, 1)
}
"TypeDef:Uint16" => {
let bytes: [u8; 2] = buffer[offset..offset+2].try_into().unwrap();
let v = match endian {
Endian::Little => u16::from_le_bytes(bytes),
Endian::Big => u16::from_be_bytes(bytes),
};
(v as u32, 2)
}
"TypeDef:Uint32" => {
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
let v = match endian {
Endian::Little => u32::from_le_bytes(bytes),
Endian::Big => u32::from_be_bytes(bytes),
};
(v, 4)
}
_ => return Err(TypedefError::Schema(format!("unsupported discriminator type: {disc_type}"))),
};
// Look up the variant schema
let mapping = &schema["mapping"];
let variant_schema = &mapping[&key];
// Read the variant struct starting at offset + discriminator_size
let variant_offset = offset + 1; // discriminator_size for Uint8
read_struct(buffer, variant_offset, variant_schema)
let key = disc_value.to_string();
if schema["mapping"].as_object().map_or(false, |m| m.contains_key(&key)) {
Ok(key)
} else {
Err(TypedefError::Access {
field_path: "__discriminator".into(),
reason: format!("unknown discriminator value: {disc_value}"),
})
}
}
```
The discriminator is a fixed-size integer at a known byte offset. The
mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`.
`offset + discriminator_size`. After reading the discriminator, the
consumer looks up the variant schema and reads the variant's fields
using the normal typed read functions (e.g., `read_u32`, `read_string`)
at `offset + discriminator_size`.
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
1..N are the variant struct. The call protocol's 5 event types
@@ -154,24 +214,53 @@ This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
### Field-name discriminator
```rust
fn read_union_field(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
let disc = &schema["discriminator"];
let field_name = disc["name"].as_str().unwrap(); // e.g., "type"
/// Read the discriminator value from a field-name TUnion.
/// The discriminator is a named field within the struct — its offset
/// is computed like any other field. The consumer reads the field's
/// value, looks up the variant schema, then reads the variant's fields.
fn read_union_field_discriminator(
buffer: &[u8],
schema: &Value,
offset_map: &OffsetMap,
endian: Endian,
) -> Result<String, TypedefError> {
let disc = schema["discriminator"].as_object()
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
let field_name = disc["name"].as_str()
.ok_or_else(|| TypedefError::Schema("discriminator has no 'name'".into()))?;
// Read the discriminator field like any other field
let disc_value = read_field(buffer, field_name, schema)?;
// Read the discriminator field at its computed offset.
// The field's TypeDef kind determines how to read it (typically a string).
let field_schema = schema["properties"].get(field_name)
.ok_or_else(|| TypedefError::Schema(format!("discriminator field '{field_name}' not found")))?;
let kind = get_typedef_kind(field_schema)
.ok_or_else(|| TypedefError::Schema("discriminator field has no TypeDef kind".into()))?;
// Look up the variant schema
let mapping = &schema["mapping"];
let variant_schema = &mapping[disc_value.as_str().unwrap()];
// Read the variant struct
read_struct(buffer, variant_offset, variant_schema)
match kind {
"TypeDef:String" => {
let range = offset_map.get(field_name)
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
read_string(buffer, range.start, endian)
.map(|s| s.to_string())
}
"TypeDef:Uint8" => {
let range = offset_map.get(field_name)
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
Ok(buffer[range.start].to_string())
}
_ => Err(TypedefError::Schema(format!(
"unsupported discriminator field type: {kind}"
))),
}
}
```
The discriminator is a named field within the struct. Its offset is
computed like any other field. The mapping keys are string values.
After reading the discriminator, the consumer looks up the variant
schema and reads the variant's fields starting at the end of the
discriminator field (or at the start of the union buffer if the
discriminator is the first field).
## Field Paths
@@ -180,6 +269,8 @@ The `OffsetMap` stores fully-qualified paths. The read/write functions
accept a field path and look up the byte range:
```rust
/// Read an f32 field by path. This is the aligned-mode path (uses OffsetMap).
/// In packed mode, the consumer uses SequentialReader instead.
fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> {
let range = self.offset_map.get(field_path)
.ok_or_else(|| TypedefError::Offset {
@@ -192,7 +283,11 @@ fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError>
reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()),
});
}
Ok(read_f32_raw(buffer, range.start, self.endian))
let bytes: [u8; 4] = buffer[range.start..range.end].try_into().unwrap();
Ok(match self.endian {
Endian::Little => f32::from_le_bytes(bytes),
Endian::Big => f32::from_be_bytes(bytes),
})
}
```
@@ -65,8 +65,12 @@ functions, and validation — all driven by the schema.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that a
buffer's bytes match the schema's type constraints.
3. **Validation** — via `jsonschema` custom keywords, validates that data
conforms to the schema's type constraints. The jsonschema validator
operates on `serde_json::Value` instances (JSON representations), not
raw byte buffers directly. A consumer that wants to validate a binary
buffer reads it into a `Value` tree via the data access layer, then
validates that `Value` against the jsonschema validator.
**The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing).** The novel code is the offset computation
@@ -114,10 +118,11 @@ read/write) is already allocation-free. See OQ-070.
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
on the Rust side. Zero translation. The same schema validates in both
ecosystems.
- **Defense in depth.** Schema validation at the byte level — a malformed
binary payload fails validation before any consumer touches it. The
`jsonschema` crate's compiled validators are fast enough to run on
every incoming frame.
- **Defense in depth.** Schema validation via jsonschema custom keywords —
a malformed binary payload can be read into a `Value` tree via the data
access layer and validated against the schema before any consumer
touches it. The `jsonschema` crate's compiled validators are fast enough
to run on every incoming frame.
### Negative
@@ -129,6 +129,26 @@ reserving worst-case space.
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
`TypeDef:Record`, `TypeDef:Timestamp`.
### 3a. TRecord value type
`TypeDef:Record` is a string-keyed map. The value type is declared via
the `"values"` property in the schema:
```json
{
"TypeDef:Record": true,
"values": { "TypeDef:Float32": true }
}
```
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
values in the record. All values share the same type.
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value_len: u32][value_bytes]...`.
- The count prefix respects the schema's endianness.
- In aligned static mode with `maxLength`, the entire record is reserved
at `maxLength` bytes (zero-padded).
### 4. TUnion discriminators
**Two discriminator kinds: byte-offset (protocol dispatch) and
@@ -95,7 +95,7 @@ Each `TypeDef:*` kind gets a `Keyword` implementation registered via
- **`TypeDef:Union`**: discriminator value membership in the mapping.
- **`TypeDef:Array`**: element type conformance.
- **`TypeDef:Boolean`**: value is `true` or `false`.
- **`TypeDef:Timestamp`**: ISO 8601 string format.
- **`TypeDef:Timestamp`**: RFC 3339 string format (the internet profile of ISO 8601).
The `jsonschema` crate handles all the structural validation (object
properties, required fields, array items, enum values) — the custom
@@ -8,10 +8,12 @@
- **Door type**: Two-way (additive — a builder API can be added without
changing the existing JSON-consumption path)
- **Priority**: medium
- **Impacts**: No current consumer. Schemas are authored in TypeBox (JS)
or hand-written JSON for v1. A Rust builder API would enable
programmatic schema construction in Rust without depending on a JS
toolchain, but no current consumer needs this.
- **Impacts**: Blocks programmatic schema construction in Rust without a
JS toolchain. Any consumer that wants to build typedef schemas at
runtime from Rust code (rather than loading pre-authored JSON) must
construct the JSON manually or depend on TypeBox. Does NOT block any
current consumer — all v1 consumers (SFTP, metatensor, binary call
frames, TTY negotiation) use pre-authored schemas.
- **Blocked on**: A concrete need for programmatic schema construction
in Rust. The current consumers (SFTP, metatensor, binary call frames,
TTY negotiation) all have schemas that can be hand-written or generated