fix(typedef): fix code examples and cross-doc inconsistencies from second review
- Fix code examples hardcoding little-endian: read_string, write_string, read_string_indirect now take endian parameter and use match on Endian. - Fix TUnion dispatch examples: remove undefined functions (read_u8, read_field, read_struct, read_f32_raw), remove Value returns (contradicts 'no intermediate Value tree'), add endian-aware discriminator reading for Uint8/Uint16/Uint32. - Fix read_f32 example: inline the endian-aware conversion instead of calling undefined read_f32_raw; document it as aligned-mode only. - Fix architecture README: 16→17 kinds in schema-layer and validation descriptions. - Fix ADR-095: clarify validation operates on Value instances, not raw byte buffers directly. Fix 'defense in depth' paragraph. - Fix ADR-097: add §3a defining TRecord 'values' property shape. - Fix ADR-098: TTimestamp format ISO 8601→RFC 3339. - Tighten OQ-071 impacts field: state what IS blocked, not just what isn't.
This commit is contained in:
1 parent
01cc3a0367
commit
dd232c3d47
6 files changed
+172
-50
No files matched your search
@@ -298,10 +298,10 @@ adapter location map is now consistent: all HTTP-backed adapters
|
||||
| [crates/channels/channel-client.md](crates/channels/channel-client.md) | draft | `ChannelClient` — client side of a channels connection, transport-agnostic `from_connection` primary; dial lives in `AlknetClient` (ADR-089); bidirectionality preserved |
|
||||
| [crates/typedef/README.md](crates/typedef/README.md) | draft | alknet-typedef crate — binary struct engine; JSON Schema with `TypeDef:*` custom keywords → offset map + read/write + validation |
|
||||
| [crates/typedef/overview.md](crates/typedef/overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
|
||||
| [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 16 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
|
||||
| [crates/typedef/schema-layer.md](crates/typedef/schema-layer.md) | draft | The 17 `TypeDef:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
|
||||
| [crates/typedef/layout-engine.md](crates/typedef/layout-engine.md) | draft | Offset computation, two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
|
||||
| [crates/typedef/data-access.md](crates/typedef/data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
|
||||
| [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 16 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation |
|
||||
| [crates/typedef/validation.md](crates/typedef/validation.md) | draft | Custom keyword validators for all 17 `TypeDef:*` kinds, `TypedefError`, load-time vs access-time validation |
|
||||
|
||||
## ADR Table
|
||||
|
||||
|
||||
@@ -76,16 +76,33 @@ For variable-length types with inline length-prefixing (the default):
|
||||
|
||||
```rust
|
||||
// Read a length-prefixed string
|
||||
fn read_string<'a>(buffer: &'a [u8], offset: usize) -> &'a str {
|
||||
let len = u32::from_le_bytes(buffer[offset..offset+4].try_into().unwrap()) as usize;
|
||||
std::str::from_utf8(&buffer[offset+4..offset+4+len]).unwrap()
|
||||
fn read_string<'a>(buffer: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
|
||||
let len_bytes: [u8; 4] = buffer[offset..offset+4].try_into()
|
||||
.map_err(|_| TypedefError::Access { /* ... */ })?;
|
||||
let len = match endian {
|
||||
Endian::Little => u32::from_le_bytes(len_bytes),
|
||||
Endian::Big => u32::from_be_bytes(len_bytes),
|
||||
} as usize;
|
||||
let data = buffer.get(offset+4..offset+4+len)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
|
||||
std::str::from_utf8(data)
|
||||
.map_err(|e| TypedefError::Access { /* ... */ })
|
||||
}
|
||||
|
||||
// Write a length-prefixed string
|
||||
fn write_string(buffer: &mut [u8], offset: usize, value: &str) {
|
||||
fn write_string(buffer: &mut [u8], offset: usize, value: &str, endian: Endian) -> Result<(), TypedefError> {
|
||||
let data = value.as_bytes();
|
||||
buffer[offset..offset+4].copy_from_slice(&(data.len() as u32).to_le_bytes());
|
||||
buffer[offset+4..offset+4+data.len()].copy_from_slice(data);
|
||||
let len_bytes = match endian {
|
||||
Endian::Little => (data.len() as u32).to_le_bytes(),
|
||||
Endian::Big => (data.len() as u32).to_be_bytes(),
|
||||
};
|
||||
buffer.get_mut(offset..offset+4)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?
|
||||
.copy_from_slice(&len_bytes);
|
||||
buffer.get_mut(offset+4..offset+4+data.len())
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?
|
||||
.copy_from_slice(data);
|
||||
Ok(())
|
||||
}
|
||||
```
|
||||
|
||||
@@ -104,10 +121,19 @@ For variable-length types with offset indirection (opt-in):
|
||||
|
||||
```rust
|
||||
// Read an offset-indirect string
|
||||
fn read_string_indirect(data_region: &[u8], offset: usize) -> &str {
|
||||
let ptr_offset = u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()) as usize;
|
||||
let ptr_length = u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()) as usize;
|
||||
std::str::from_utf8(&data_region[ptr_offset..ptr_offset+ptr_length]).unwrap()
|
||||
fn read_string_indirect<'a>(data_region: &'a [u8], offset: usize, endian: Endian) -> Result<&'a str, TypedefError> {
|
||||
let ptr_offset = match endian {
|
||||
Endian::Little => u32::from_le_bytes(data_region[offset..offset+4].try_into().unwrap()),
|
||||
Endian::Big => u32::from_be_bytes(data_region[offset..offset+4].try_into().unwrap()),
|
||||
} as usize;
|
||||
let ptr_length = match endian {
|
||||
Endian::Little => u32::from_le_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
|
||||
Endian::Big => u32::from_be_bytes(data_region[offset+4..offset+8].try_into().unwrap()),
|
||||
} as usize;
|
||||
let data = data_region.get(ptr_offset..ptr_offset+ptr_length)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
|
||||
std::str::from_utf8(data)
|
||||
.map_err(|e| TypedefError::Access { /* ... */ })
|
||||
}
|
||||
```
|
||||
|
||||
@@ -124,28 +150,62 @@ differs by discriminator kind (ADR-097).
|
||||
### Byte-offset discriminator
|
||||
|
||||
```rust
|
||||
fn read_union(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
|
||||
let disc = &schema["discriminator"];
|
||||
let offset = disc["offset"].as_u64().unwrap() as usize;
|
||||
let disc_type = disc["type"].as_str().unwrap(); // e.g., "TypeDef:Uint8"
|
||||
/// Read the discriminator value from a byte-offset TUnion.
|
||||
/// Returns the mapping key (as a string) so the consumer can look up
|
||||
/// the variant schema and read the variant's fields.
|
||||
fn read_union_discriminator(
|
||||
buffer: &[u8],
|
||||
schema: &Value,
|
||||
endian: Endian,
|
||||
) -> Result<String, TypedefError> {
|
||||
let disc = schema["discriminator"].as_object()
|
||||
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
|
||||
let offset = disc["offset"].as_u64().unwrap_or(0) as usize;
|
||||
let disc_type = disc["type"].as_str().unwrap_or("TypeDef:Uint8");
|
||||
|
||||
// Read the discriminator value
|
||||
let disc_value: u8 = read_u8(buffer, offset);
|
||||
let key = disc_value.to_string(); // "5", "6", "101"
|
||||
let (disc_value, disc_size) = match disc_type {
|
||||
"TypeDef:Uint8" => {
|
||||
let b = *buffer.get(offset)
|
||||
.ok_or_else(|| TypedefError::Access { /* ... */ })?;
|
||||
(b as u32, 1)
|
||||
}
|
||||
"TypeDef:Uint16" => {
|
||||
let bytes: [u8; 2] = buffer[offset..offset+2].try_into().unwrap();
|
||||
let v = match endian {
|
||||
Endian::Little => u16::from_le_bytes(bytes),
|
||||
Endian::Big => u16::from_be_bytes(bytes),
|
||||
};
|
||||
(v as u32, 2)
|
||||
}
|
||||
"TypeDef:Uint32" => {
|
||||
let bytes: [u8; 4] = buffer[offset..offset+4].try_into().unwrap();
|
||||
let v = match endian {
|
||||
Endian::Little => u32::from_le_bytes(bytes),
|
||||
Endian::Big => u32::from_be_bytes(bytes),
|
||||
};
|
||||
(v, 4)
|
||||
}
|
||||
_ => return Err(TypedefError::Schema(format!("unsupported discriminator type: {disc_type}"))),
|
||||
};
|
||||
|
||||
// Look up the variant schema
|
||||
let mapping = &schema["mapping"];
|
||||
let variant_schema = &mapping[&key];
|
||||
|
||||
// Read the variant struct starting at offset + discriminator_size
|
||||
let variant_offset = offset + 1; // discriminator_size for Uint8
|
||||
read_struct(buffer, variant_offset, variant_schema)
|
||||
let key = disc_value.to_string();
|
||||
if schema["mapping"].as_object().map_or(false, |m| m.contains_key(&key)) {
|
||||
Ok(key)
|
||||
} else {
|
||||
Err(TypedefError::Access {
|
||||
field_path: "__discriminator".into(),
|
||||
reason: format!("unknown discriminator value: {disc_value}"),
|
||||
})
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The discriminator is a fixed-size integer at a known byte offset. The
|
||||
mapping keys are stringified integers. The variant struct starts at
|
||||
`offset + discriminator_size`.
|
||||
`offset + discriminator_size`. After reading the discriminator, the
|
||||
consumer looks up the variant schema and reads the variant's fields
|
||||
using the normal typed read functions (e.g., `read_u32`, `read_string`)
|
||||
at `offset + discriminator_size`.
|
||||
|
||||
This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
|
||||
1..N are the variant struct. The call protocol's 5 event types
|
||||
@@ -154,24 +214,53 @@ This is the SFTP `Packet` enum pattern — byte 0 is the type byte, bytes
|
||||
### Field-name discriminator
|
||||
|
||||
```rust
|
||||
fn read_union_field(buffer: &[u8], schema: &Value) -> Result<Value, TypedefError> {
|
||||
let disc = &schema["discriminator"];
|
||||
let field_name = disc["name"].as_str().unwrap(); // e.g., "type"
|
||||
/// Read the discriminator value from a field-name TUnion.
|
||||
/// The discriminator is a named field within the struct — its offset
|
||||
/// is computed like any other field. The consumer reads the field's
|
||||
/// value, looks up the variant schema, then reads the variant's fields.
|
||||
fn read_union_field_discriminator(
|
||||
buffer: &[u8],
|
||||
schema: &Value,
|
||||
offset_map: &OffsetMap,
|
||||
endian: Endian,
|
||||
) -> Result<String, TypedefError> {
|
||||
let disc = schema["discriminator"].as_object()
|
||||
.ok_or_else(|| TypedefError::Schema("missing discriminator".into()))?;
|
||||
let field_name = disc["name"].as_str()
|
||||
.ok_or_else(|| TypedefError::Schema("discriminator has no 'name'".into()))?;
|
||||
|
||||
// Read the discriminator field like any other field
|
||||
let disc_value = read_field(buffer, field_name, schema)?;
|
||||
// Read the discriminator field at its computed offset.
|
||||
// The field's TypeDef kind determines how to read it (typically a string).
|
||||
let field_schema = schema["properties"].get(field_name)
|
||||
.ok_or_else(|| TypedefError::Schema(format!("discriminator field '{field_name}' not found")))?;
|
||||
let kind = get_typedef_kind(field_schema)
|
||||
.ok_or_else(|| TypedefError::Schema("discriminator field has no TypeDef kind".into()))?;
|
||||
|
||||
// Look up the variant schema
|
||||
let mapping = &schema["mapping"];
|
||||
let variant_schema = &mapping[disc_value.as_str().unwrap()];
|
||||
|
||||
// Read the variant struct
|
||||
read_struct(buffer, variant_offset, variant_schema)
|
||||
match kind {
|
||||
"TypeDef:String" => {
|
||||
let range = offset_map.get(field_name)
|
||||
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
|
||||
read_string(buffer, range.start, endian)
|
||||
.map(|s| s.to_string())
|
||||
}
|
||||
"TypeDef:Uint8" => {
|
||||
let range = offset_map.get(field_name)
|
||||
.ok_or_else(|| TypedefError::Offset { /* ... */ })?;
|
||||
Ok(buffer[range.start].to_string())
|
||||
}
|
||||
_ => Err(TypedefError::Schema(format!(
|
||||
"unsupported discriminator field type: {kind}"
|
||||
))),
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The discriminator is a named field within the struct. Its offset is
|
||||
computed like any other field. The mapping keys are string values.
|
||||
After reading the discriminator, the consumer looks up the variant
|
||||
schema and reads the variant's fields starting at the end of the
|
||||
discriminator field (or at the start of the union buffer if the
|
||||
discriminator is the first field).
|
||||
|
||||
## Field Paths
|
||||
|
||||
@@ -180,6 +269,8 @@ The `OffsetMap` stores fully-qualified paths. The read/write functions
|
||||
accept a field path and look up the byte range:
|
||||
|
||||
```rust
|
||||
/// Read an f32 field by path. This is the aligned-mode path (uses OffsetMap).
|
||||
/// In packed mode, the consumer uses SequentialReader instead.
|
||||
fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError> {
|
||||
let range = self.offset_map.get(field_path)
|
||||
.ok_or_else(|| TypedefError::Offset {
|
||||
@@ -192,7 +283,11 @@ fn read_f32(&self, buffer: &[u8], field_path: &str) -> Result<f32, TypedefError>
|
||||
reason: format!("buffer too short: need {} bytes, have {}", range.end, buffer.len()),
|
||||
});
|
||||
}
|
||||
Ok(read_f32_raw(buffer, range.start, self.endian))
|
||||
let bytes: [u8; 4] = buffer[range.start..range.end].try_into().unwrap();
|
||||
Ok(match self.endian {
|
||||
Endian::Little => f32::from_le_bytes(bytes),
|
||||
Endian::Big => f32::from_be_bytes(bytes),
|
||||
})
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
@@ -65,8 +65,12 @@ functions, and validation — all driven by the schema.
|
||||
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
|
||||
read the field's bytes at its offset (zero-copy for fixed-size types).
|
||||
Given a `&mut [u8]` buffer, write a value at its offset.
|
||||
3. **Validation** — via `jsonschema` custom keywords, validates that a
|
||||
buffer's bytes match the schema's type constraints.
|
||||
3. **Validation** — via `jsonschema` custom keywords, validates that data
|
||||
conforms to the schema's type constraints. The jsonschema validator
|
||||
operates on `serde_json::Value` instances (JSON representations), not
|
||||
raw byte buffers directly. A consumer that wants to validate a binary
|
||||
buffer reads it into a `Value` tree via the data access layer, then
|
||||
validates that `Value` against the jsonschema validator.
|
||||
|
||||
**The heavy lifting is done by the `jsonschema` crate (validation) and
|
||||
`serde_json` (schema parsing).** The novel code is the offset computation
|
||||
@@ -114,10 +118,11 @@ read/write) is already allocation-free. See OQ-070.
|
||||
under `$defs`. That JSON feeds directly into `jsonschema::validator_for`
|
||||
on the Rust side. Zero translation. The same schema validates in both
|
||||
ecosystems.
|
||||
- **Defense in depth.** Schema validation at the byte level — a malformed
|
||||
binary payload fails validation before any consumer touches it. The
|
||||
`jsonschema` crate's compiled validators are fast enough to run on
|
||||
every incoming frame.
|
||||
- **Defense in depth.** Schema validation via jsonschema custom keywords —
|
||||
a malformed binary payload can be read into a `Value` tree via the data
|
||||
access layer and validated against the schema before any consumer
|
||||
touches it. The `jsonschema` crate's compiled validators are fast enough
|
||||
to run on every incoming frame.
|
||||
|
||||
### Negative
|
||||
|
||||
|
||||
@@ -129,6 +129,26 @@ reserving worst-case space.
|
||||
types: `TypeDef:String`, `TypeDef:Bytes`, `TypeDef:Array`,
|
||||
`TypeDef:Record`, `TypeDef:Timestamp`.
|
||||
|
||||
### 3a. TRecord value type
|
||||
|
||||
`TypeDef:Record` is a string-keyed map. The value type is declared via
|
||||
the `"values"` property in the schema:
|
||||
|
||||
```json
|
||||
{
|
||||
"TypeDef:Record": true,
|
||||
"values": { "TypeDef:Float32": true }
|
||||
}
|
||||
```
|
||||
|
||||
- `"values"` is a schema object declaring the `TypeDef:*` kind of all
|
||||
values in the record. All values share the same type.
|
||||
- The binary layout is a count-prefixed sequence of `(key, value)` pairs:
|
||||
`[count: u32][key_len: u32][key_bytes][value_len: u32][value_bytes]...`.
|
||||
- The count prefix respects the schema's endianness.
|
||||
- In aligned static mode with `maxLength`, the entire record is reserved
|
||||
at `maxLength` bytes (zero-padded).
|
||||
|
||||
### 4. TUnion discriminators
|
||||
|
||||
**Two discriminator kinds: byte-offset (protocol dispatch) and
|
||||
|
||||
@@ -95,7 +95,7 @@ Each `TypeDef:*` kind gets a `Keyword` implementation registered via
|
||||
- **`TypeDef:Union`**: discriminator value membership in the mapping.
|
||||
- **`TypeDef:Array`**: element type conformance.
|
||||
- **`TypeDef:Boolean`**: value is `true` or `false`.
|
||||
- **`TypeDef:Timestamp`**: ISO 8601 string format.
|
||||
- **`TypeDef:Timestamp`**: RFC 3339 string format (the internet profile of ISO 8601).
|
||||
|
||||
The `jsonschema` crate handles all the structural validation (object
|
||||
properties, required fields, array items, enum values) — the custom
|
||||
|
||||
@@ -8,10 +8,12 @@
|
||||
- **Door type**: Two-way (additive — a builder API can be added without
|
||||
changing the existing JSON-consumption path)
|
||||
- **Priority**: medium
|
||||
- **Impacts**: No current consumer. Schemas are authored in TypeBox (JS)
|
||||
or hand-written JSON for v1. A Rust builder API would enable
|
||||
programmatic schema construction in Rust without depending on a JS
|
||||
toolchain, but no current consumer needs this.
|
||||
- **Impacts**: Blocks programmatic schema construction in Rust without a
|
||||
JS toolchain. Any consumer that wants to build typedef schemas at
|
||||
runtime from Rust code (rather than loading pre-authored JSON) must
|
||||
construct the JSON manually or depend on TypeBox. Does NOT block any
|
||||
current consumer — all v1 consumers (SFTP, metatensor, binary call
|
||||
frames, TTY negotiation) use pre-authored schemas.
|
||||
- **Blocked on**: A concrete need for programmatic schema construction
|
||||
in Rust. The current consumers (SFTP, metatensor, binary call frames,
|
||||
TTY negotiation) all have schemas that can be hand-written or generated
|
||||
|
||||
Reference in new issue
Block a user