Sync architecture docs and ADRs to BAST pivot (steps 9-10)

Step 9 (convert tests to BAST format) was a no-op: steps 4-8 converted
the tests as they went. The only remaining  reference in
src/tests was the intentional  rejection test at
src/schema.rs:462 (asserting the old keyword form is rejected). Full
suite passes: 389 tests (312 lib + 77 integration).

Step 10 (sync architecture docs and ADRs):

Descriptive docs rewritten/updated for BAST:
- schema-layer.md: rewritten for the BAST parser (BastDoc/BastDef/
  BastType typed tree, AlkTypeKind enum with to_bast_str/from_bast_str,
  what was removed). Points at bast-format.md for the normative format.
- validation.md: rewritten for the two-validator model
  (bast_validation for validate_bytes, standard jsonschema for
  validate_json). Documents the repurposed build_validator, the
  AlkTypeError::Validation uniform payload (D-BAST-009), and what is
  removed.
- builder.md: updated all output examples to BAST JSON
  (struct_() -> { kind: struct, fields: [...] }; object() -> standard
  JSON Schema). Documents build_doc, count(), and the field-name union
  fields requirement (D-BAST-005).
- overview.md: updated for BAST (what/why, schema-is-the-format table,
  dependencies, architecture pointers, design decisions table).
- README.md (architecture index): updated document table, ADR table
  (new ADR-BAST + ADR-VAL-SPLIT, superseded ADR-001), OQ table
  (OQ-007/OQ-008 resolutions updated for BAST-native validator), and
  key design principles (#1, #2, #7, #10 reworded for BAST).
- data-access.md: updated tunion function signatures to BastUnion and
  the variant resolution to return BastType (resolve_typeref for refs).
- layout-engine.md: updated construct signatures
  (LayoutBuilder::new(bast_doc, root_name), OffsetMap::compute(&doc),
  SequentialReader::new(bast_doc, root_name)), the recursive-walk
  description (BAST typed tree), and composite-kind headings
  (TStruct/TUnion/TArray -> struct/union/array). Added D-BAST-004
  note on array count requirement.

New ADRs:
- ADR-BAST (bast-bast-format.md): the BAST format, meta-schema,
  //kind vocabulary, design principles, what is removed, the
  enum index bounds bug fix. Supersedes ADR-001's format-specific
  content; records D-BAST-001..009.
- ADR-VAL-SPLIT (val-split-two-validator-model.md): the two-validator
  model (BAST-native for validate_bytes, standard jsonschema for
  validate_json), the repurposed build_validator, the uniform
  AlkTypeError::Validation payload. Refines ADR-004's validation
  strategy and ADR-010's validation step; records D-BAST-006/007/009.

Amended ADRs (supersession/amendment notes added; original decision
text preserved as historical record):
- ADR-001: format-specific content superseded by ADR-BAST;
  purpose/scope and schema-is-the-format principle retained.
- ADR-002: unchanged under the pivot; one-line note that the input
  format changed but the modes didn't.
- ADR-003: annotation semantics retained; annotation location moved
  to BAST type-level properties (amended by ADR-BAST).
- ADR-004: AlkTypeError enum retained (D-BAST-009); validation
  strategy section refined by ADR-VAL-SPLIT.
- ADR-009: builder API surface retained; build() output format
  amended to BAST / standard JSON Schema by ADR-BAST (D-BAST-008).
- ADR-010: validate_bytes two-step concept retained; validation step
  amended to the BAST-native validator by ADR-VAL-SPLIT.

Other:
- Cargo.toml description: JSON Schema with AlkType:* custom keywords
  -> BAST document.
- bast-pivot.md research record: status draft -> implemented, with a
  pointer to the ADRs that superseded its decisions.
- bast-implementation.md plan: status draft -> complete, with a note
  that step 9 was a no-op and step 10 is this commit.
- open-questions.md: OQ-006/OQ-007/OQ-008 resolutions updated for the
  BAST-native validator.
- questions/008-unionvalidator-variant-dispatch.md: added a
  post-BAST-pivot note pointing to the current bast_validation
  implementation; v0.1.0 resolution text preserved as historical
  record.

Verification:
- cargo test --release: 389 pass (312 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cross-reference check: every relative link in the new/updated docs
  resolves (verified by script).
This commit is contained in:
glm-5.2 committed 2026-08-15 14:03:21 +00:00
1 parent 54fd112fde
commit 62270b03ca
20 files changed
+1642 -1075

No files matched your search

+19 -18
View File
@@ -7,8 +7,8 @@ last_updated: 2026-07-22
The layout engine: offset computation, the two layout modes (packed
sequential vs aligned static), alignment, endianness, and variable-length
field handling. This is the novel code — the recursive walk of the schema
JSON that computes byte positions for each field.
field handling. This is the novel code — the recursive walk of the BAST
typed tree that computes byte positions for each field.
## The Two Layout Modes
@@ -24,8 +24,8 @@ protocols.
**Components:**
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `AlkType:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(bast_doc, root_name)` (requires a `struct` at the root), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `SequentialReader::new(bast_doc, root_name)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
**How it works:**
@@ -69,7 +69,7 @@ and safetensors.
**Component:**
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, AlkTypeError>` (requires `AlkType:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
- **`OffsetMap`** — constructed via `OffsetMap::compute(&doc) -> Result<Self, AlkTypeError>` (requires a `struct` at the root). Walks the BAST typed tree once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
**How it works:**
@@ -118,14 +118,15 @@ fields must use `maxLength` (fixed-size reservation) or
## Offset Computation Algorithm
The offset computation is a recursive walk of the schema JSON. The
The offset computation is a recursive walk of the BAST typed tree
([`BastDoc`](schema-layer.md#the-bast-parser-bast-module)). The
algorithm is the same for both modes; the difference is whether alignment
padding is inserted between fields.
### Fixed-size types
For each fixed-size type, the algorithm:
1. Determines the type's byte size from the `AlkType:*` kind.
1. Determines the type's byte size from the `AlkTypeKind`.
2. In aligned mode: inserts padding to satisfy the type's alignment
(or the field's `align` annotation, or the struct's `align` default).
3. Records the field's `(start, end)` range.
@@ -133,14 +134,14 @@ For each fixed-size type, the algorithm:
### Composite types
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
**`struct`:** Recurse into the struct's `fields` array. The inner fields
are computed relative to the struct's start offset. The struct's total
size is the sum of its fields' sizes (plus alignment padding in aligned
mode). The struct itself may have an `align` annotation that rounds up
its total size.
**`TUnion`:** TUnion is supported in packed sequential mode only. In
aligned static mode, `OffsetMap::compute` rejects `TUnion` fields with
**`union`:** TUnion is supported in packed sequential mode only. In
aligned static mode, `OffsetMap::compute` rejects `union` fields with
`AlkTypeError::Offset` — see
[ADR-008](decisions/008-reject-tunion-in-aligned-mode.md). Unions
are the protocol dispatch pattern (SFTP type bytes, call protocol event
@@ -161,12 +162,12 @@ total size. The `SequentialReader` reads the discriminator first, looks
up the variant schema, then reads the variant struct sequentially — it
doesn't need to know the union's total size upfront.
**`TArray` of fixed-size elements:** Element stride = element size (plus
**`array` of fixed-size elements:** Element stride = element size (plus
alignment padding in aligned mode). Element `i` starts at
`array_offset + i × stride`. The array's total size is `count × stride`.
**`TArray` of variable-length-element structs:** Deferred for v1
(OQ-001).
**`array` of variable-length-element structs:** Deferred for v1
(OQ-001, D-BAST-004 — BAST arrays require `count` in v1).
### Variable-length types
@@ -289,7 +290,7 @@ pub struct FieldPosition {
A field's computed position in a packed layout, produced by
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
length prefix); for fixed-size fields, `size` is the type's byte size.
`kind` records the field's `AlkType:*` kind so the consumer can dispatch
`kind` records the field's `AlkTypeKind` so the consumer can dispatch
to the correct `data_access` read/write function.
### `PackedLayout` (packed mode)
@@ -317,16 +318,16 @@ A flat table of `(field_path, byte_range)` pairs computed from a schema.
```rust
impl OffsetMap {
pub fn compute(schema: &Value) -> Result<Self, AlkTypeError>;
pub fn compute<'a>(doc: &'a BastDoc<'a>) -> Result<Self, AlkTypeError>;
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
}
```
`compute` requires a `AlkType:Struct` at the top level. `total_size`
includes trailing alignment padding. `iter` yields fields in insertion
order (schema `properties` order, nested struct fields appearing inline).
`compute` requires a `struct` at the root. `total_size`
includes trailing alignment padding. `iter` yields fields in the BAST
`fields` array order (nested struct fields appearing inline).
## Design Decisions