Contract-suite scheduler rows (task suite-scheduler-rows): the determinism-posture extension in properties.rs's module doc (runner-driving rows admitted — spawned run_schedules(stop) tasks on a core StopToken against state outcomes only, elapsed-boundary-band tolerances, never timing-value assertions, never two-task race windows; ADR-017 §2 class 4) and three version-stamped rows: scheduler_boundary_fires (fired jobs are ordinary claimable work with ScheduleOpts over the plain-queue derived defaults 300/9/5/none + payload exact, clean-stop Ok(()), fired count inside the elapsed-boundary band — no double-fire per boundary while one leader runs; ADR-009 §3/§4, ADR-020 §3, ADR-019 §4), scheduler_bounded_catchup (runner-less downtime proves no fire without a runner, ≥3 elapsed boundaries replay boundary-by-boundary bounded below the 64-cap and inside the band; ADR-009 §4), scheduler_leadership_discipline (two spawned runners on one store: exactly one Ok(())/Err(LeadershipLost) pair by value, no duplicated fires inside the band; ADR-009 §1/§6, ADR-019 §4) — wired into both engines' suite targets (SQLite tests, pg harness_row!s), suite tokio dep added for the runner rows. Dispositions recorded in Notes: the beyond-cap skip-forward leg stays pinned engine-side (both engines' scheduler tests already backdate next_fire_at directly — catch_up_replays_up_to_the_cap_then_skips_forward twins), and the pg fire-wake parity gap is not demanded by these rows' shapes (claim-polling observation only; the one-call wake_tx disposition recorded for a later task). Verified: sqlite suite 16/16, pg suite 16/16 vs harness twice + solo re-runs of the three rows per engine (determinism), workspace build/test green, clippy -D warnings, fmt clean

This commit is contained in:
glm-5.3-flash committed 2026-10-10 06:04:30 +00:00
1 parent 5c9ae6a7fc
commit 7f749ac673
6 files changed
+523 -16

No files matched your search

+88 -3
View File
@@ -1,7 +1,7 @@
---
id: suite-scheduler-rows
name: Contract-suite rows — scheduler boundary/catch-up/leadership (+ determinism-posture extension)
status: pending
status: completed
depends_on: []
scope: moderate
risk: medium
@@ -86,8 +86,93 @@ latency-posture in Notes and leave the engines as they are.
## Notes
> To be filled by implementation agent
- **Beyond-cap skip-forward leg: pinned engine-side, already done in
both engines.** The leg needs >64 elapsed boundaries (~65+ s of
downtime) and a direct schedule-row backdate the public surface does
not express (`next_fire_at` is engine-internal). Disposition: pinned
where `next_fire_at` is directly manipulable —
`catch_up_replays_up_to_the_cap_then_skips_forward` exists in
`alkstore-sqlite/src/store/scheduler_tests.rs` (substrate row
backdate, exactly 64 fires, skip-forward past now, skipped
boundaries never enqueue) and the twin in
`alkstore-postgres/src/store/scheduler_tests.rs` (admin-SQL backdate,
same shape). The suite row (`scheduler_bounded_catchup`) pins the
≤64 in-band catch-up leg — the suite-observable one — and states the
engine-side disposition in its doc text.
- **Fire-wake parity: not demanded by this row set; engines left
as-is.** The suite rows observe fire outcomes through ordinary
claim polling, never through wake latency, so the row shapes demand
no wake. The recorded cross-engine latency-parity gap stands (the pg
tick's in-tx enqueue fires no `pg_notify`; SQLite's watcher wakes on
those commits — `tasks/review-wave-4-fixes.md`'s recorded-for-later
note): correctness is covered by the queues row's pinned re-poll
safety net on both engines; the row set here pins state outcomes
with tolerances bounded well above the boundary resolution, so wake
latency is out of scope. If a later task demands the parity, the
one-call `wake_tx` inside the pg tick's tx is the shape to ride.
- **Runner-driving posture details** (the extension's riders, beyond
what the description pinned): the spawned runners share the row's
single store handle via `Arc<dyn Store>` clones — leadership is
storage-level, not handle-level, so two concurrent runners on one
handle still contend on the leadership lock (the SQLite engine's own
tests open a second handle on the same path; the suite's factory
gives one isolated store per row, and the shared-handle shape
exercises the same storage-level discipline). A leadership-lock
acquire through the writer slot is a short lease on both engines, so
the shared handle neither deadlocks nor serializes the tick away.
The two-concurrent-runners row takes whichever runner wins (no
identity assumption): exactly one `Ok(())` and one
`Err(LeadershipLost)` pair, checked by value.
- **Band tolerances** (the determinism posture's "bounded well above
the boundary resolution"): every fired-count assertion is bounded by
an elapsed-boundary band measured from just before registration
(`unix-second boundary arithmetic` slack + one in-progress tick at
cancel, +2 s — the engine tests' fire-window tolerance), never a
timing value. The catch-up row's lower bound (≥ 3 fires) is the
boundary-by-boundary pin — strictly more than a coalescing
single-fire would leave; its registration's first-fire grace of one
interval is counted in the band.
## Summary
> To be filled on completion
**Landed:** the determinism-posture extension in
`alkstore-contract-suite/src/properties.rs`'s module doc (runner-driving
rows admitted — spawned `run_schedules(stop)` tasks with a core
`StopToken`, state outcomes only, elapsed-boundary-band tolerances,
never timing-value assertions, never two-task race windows; ADR-017 §2
class 4 posture note, cross-connection wake orchestration still
engine-test-side) and the three scheduler rows, version-stamped and
re-exported from the suite lib:
- `scheduler_boundary_fires` — stamps ADR-009 §3/§4, ADR-020 §3,
ADR-019 §4; spawned runner, ≥2 boundary fires observed via ordinary
claim polling, the fire row's stamp source pinned (ScheduleOpts
priority/max_attempts over the plain-queue derived defaults
300/5/none, payload exact), clean-stop `Ok(())` prompt, fired count
inside the elapsed-boundary band (no double-fire per boundary under
one leader).
- `scheduler_bounded_catchup` — stamp ADR-009 §4 (plus ADR-019 §4);
runner-less downtime proves schedules never fire without a runner,
a 4 s pause yields ≥3 replayed fires (boundary-by-boundary, not
coalesced) bounded below the 64-cap and inside the band; the
beyond-cap skip-forward disposition recorded in Notes (pinned
engine-side on both engines, already implemented there).
- `scheduler_leadership_discipline` — stamps ADR-009 §1/§6,
ADR-019 §4; two spawned runners on one store, exactly one
`Ok(())`/`Err(LeadershipLost)` pair by value, fired count inside the
elapsed-boundary band with both alive (no duplicated fires).
Wired into `alkstore-sqlite/tests/contract_suite.rs` (three direct
tests) and `alkstore-postgres/tests/contract_suite.rs` (three
`harness_row!` entries). Suite crate gained `tokio` (rt + time) as a
regular dependency for the runner-driving rows (the unused
dev-dependency dropped).
**Gates:** `cargo build`; `cargo test -p alkstore-sqlite` (189 lib +
16 suite rows green); `cargo test -p alkstore-postgres` serverless
green (pg rows skip cleanly); the full pg column green against the
harness server (`pglo-poc` :15432 — 121 lib incl. the previously
skipped rows + 16 suite rows + 9 schema); the three new rows re-run
twice solo on each engine (determinism check, green both columns both
runs); `cargo clippy --all-targets -- -D warnings`; `cargo fmt --check`
(clean after fmt).