Clean up BAST pivot: drop POCs, remove recursion, add spec gaps

- Remove the four proposed POCs: they were implementation smoke tests,
  not de-risking probes. The pivot is a backend swap on a proven layout
  engine; byte-identity is already proven and the layout code is
  unchanged, so there is nothing empirical left to de-risk.
- Remove the recursive TreeNode example and the recursion mention in
  design principle 2: recursion is not a binary-layout concern and the
  engine has no cycle detection.
- Add a Spec Gaps section: validate_bytes semantics after keyword
  validator removal (UnionValidator variant dispatch regression),
  arrays of variable-length elements (engine rejects them), and
  field-name discriminator unions (meta-schema cannot express them).
- Reframe the engine change as an accessor-layer refactor: the
  walkers' (kind, field list, annotations) reads change; everything
  beneath them carries over unchanged.
This commit is contained in:
deepseek-v4-pro committed 2026-08-15 07:46:03 +00:00
1 parent 37bd5d7b7d
commit 82bc8f29c0
1 file changed
+59 -63
+59 -63
View File
@@ -16,9 +16,10 @@ self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser.
The engine's core logic (layout computation, data access, union
dispatch, two layout modes) is unchanged. Only the JSON parsing layer
changes: instead of detecting `AlkType:*` keywords scattered through a
JSON Schema tree, the engine parses a purpose-built `kind`-based format.
dispatch, two layout modes) is unchanged. Only the schema-walking
accessor layer changes: instead of detecting `AlkType:*` keywords
scattered through a JSON Schema tree, the walkers read `kind`/`fields`/
annotation properties from a purpose-built format.
The builder API's public surface stays the same; only the JSON output
format changes internally.
@@ -96,8 +97,8 @@ adoption creates migration cost.
validator can validate a BAST document's structure.
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
top-level `$defs` block. `$ref` handles recursion, cross-references,
and union variant references. This is the same pattern as TypeBox's
top-level `$defs` block. `$ref` handles cross-references and union
variant references. This is the same pattern as TypeBox's
`Type.Module` and JSON Schema's own `$defs` — no custom reference
resolution mechanism needed.
@@ -244,23 +245,6 @@ adoption creates migration cost.
}
```
#### Recursive type (e.g., a tree node)
```json
{
"$defs": {
"TreeNode": {
"kind": "struct",
"fields": [
{ "name": "value", "kind": "uint32" },
{ "name": "left", "kind": { "$ref": "#/$defs/TreeNode" } },
{ "name": "right", "kind": { "$ref": "#/$defs/TreeNode" } }
]
}
}
}
```
#### Array of fixed-size elements with known count
```json
@@ -578,6 +562,14 @@ aligned static mode (same as ADR-003):
| Top-level container | Schema object with `AlkType:Struct` at root | `{ "$defs": { ... } }` with a root type name |
| JSON validation info | `"type": "object"`, `"required"`, etc. on the same object | Separate concern — not in BAST |
The engine change is an accessor-layer refactor, not a rewrite. The
schema walkers currently read (kind, field list, annotations) out of
JSON nodes via custom-keyword accessors; under BAST they read the same
information from `kind`/`fields`/annotation properties. Everything
beneath the accessors — checked offset arithmetic, union dispatch,
materialization, the two layout modes — is format-agnostic and carries
over unchanged.
### Engine internals
| Component | Change |
@@ -846,6 +838,51 @@ through JSON Schema trees do not.
`kind` values
3. External Handlebars templates in `src/codegen/templates/`
## Spec Gaps
No POCs are needed for this pivot. It is a backend swap (custom
keywords → `kind`-based format) on top of a proven layout engine, not
new protocol invention. The layout engine's byte-identity is already
proven (alknet-typedef-poc, alktype-builder-poc) and the layout code is
unchanged, so there is nothing empirical left to de-risk. What remains
are three spec gaps where the draft promises more than the engine
delivers, or expresses less than the engine supports. Each is a
decision, not an unknown.
### Gap 1: `validate_bytes` semantics after keyword validator removal
The current engine's `UnionValidator` dispatches to variant schemas at
validation time (OQ-008), so `validate_bytes` on a union checks variant
field constraints (e.g. `maxLength` on a `Bytes` field inside a
variant). Removing the 19 custom keyword validators deletes this
dispatch. The spec must state what `validate_bytes` validates:
- Structural readability only (bounds, discriminator lookup, UTF-8) —
constraint checks move to a separately-provided JSON Schema
- Full constraint validation via a JSON Schema the consumer provides
(OQ-BAST-006's leaning) — but variant constraints then require the
consumer to hand-author `__discriminator` dispatch, a regression
from today's behavior
### Gap 2: Arrays of variable-length elements
The example in §"The BAST Format" shows
`{ "kind": "array", "element": "string" }` (count-prefixed). The engine
explicitly rejects arrays of variable-length elements ("TArray of
variable-length element kind ... is not supported"). Either the example
is out of scope for v1 (remove it and note the restriction in the
meta-schema), or variable-element arrays are a new feature with their
own layout semantics.
### Gap 3: Field-name discriminator unions
The meta-schema's `UnionDef` has no `fields` array, but the engine's
field-name discriminator reads the discriminator field from the union's
`properties` object. As sketched, BAST cannot express a feature the
engine already supports. The meta-schema needs a field-name-
discriminator shape (e.g. a `fields` array on `UnionDef`), or
field-name discriminators are dropped for v1.
## Open Questions
### OQ-BAST-001: Root type selection
@@ -962,47 +999,6 @@ naming conventions (`string()` vs `string_()`).
| Builder API output format change breaks consumers | No real consumers exist yet (v0.1.0 has zero adoption). The builder's public methods are unchanged; only the JSON output format changes. |
| Meta-schema maintenance burden | The meta-schema is small (~150 lines) and changes rarely. It's embedded in the crate and published at a stable URL. |
## Proposed POCs
Before committing to the full pivot, these focused POCs would de-risk
the key decision points:
### POC 1: BAST meta-schema validation
Prove that a standard JSON Schema validator can validate BAST documents.
- Write the BAST meta-schema as a JSON Schema document
- Validate the example BAST documents (chunk header, SFTP packets,
recursive types) against it using `jsonschema`
- Verify that malformed BAST documents (wrong `kind`, missing `fields`,
invalid `$ref`) are rejected with clear errors
### POC 2: BAST → layout compilation
Prove the engine can compile BAST documents into layouts.
- Implement a minimal BAST parser that extracts type definitions from
a BAST document
- Feed a BAST ChunkHeader schema through the existing `LayoutBuilder`
and `SequentialReader`
- Verify byte-identical output with the current custom-keyword path
### POC 3: BAST round-trip against russh-sftp
Prove byte-identical SFTP packet serialization using BAST schemas.
- Port the SFTP Read/Write/Status schemas from the alknet-typedef-poc
to BAST format
- Run the existing round-trip tests (`sftp_roundtrip_test.rs`) against
BAST-compiled layouts
- Verify byte-identical output with `russh_sftp::protocol` serialization
### POC 4: Builder API output switch
Prove the builder API can produce BAST JSON without changing its public
methods.
- Add a `build_bast()` method (or change `build()` output) to the
existing `Schema` builder
- Construct the ChunkHeader schema via the builder
- Verify the output is valid BAST JSON that passes meta-schema validation
## References
- [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) — current "schema is the format" principle (to be updated)