9.9 KiB
id, name, status, depends_on, scope, risk, impact, level, tags
| id | name | status | depends_on | scope | risk | impact | level | tags | |||
|---|---|---|---|---|---|---|---|---|---|---|---|
| suite-scheduler-rows | Contract-suite rows — scheduler boundary/catch-up/leadership (+ determinism-posture extension) | completed | moderate | medium | phase | implementation |
|
Description
Add the scheduler contract rows to
alkstore-contract-suite/src/properties.rs and wire them into both
engines' suite targets, discharging the backlog row "Scheduler + depth
properties on Postgres" (core-contract.md §Verification backlog): the
tick machinery — boundary advance, bounded catch-up, leadership-loss
discipline — pinned engine-uniform (SQLite inherits the substrate's
test-pinned implementation; pg re-derived; the suite pins identical
behavior).
Posture extension first (test-side, ADR-017 §2 class 4): the
suite's determinism statement (properties.rs module doc) currently
says runner spawning stays engine-test-side. The backlog explicitly
demands these rows in the suite, so amend the module doc to admit
runner-driving rows: a spawned run_schedules(stop) task with a
StopToken, assertions on state outcomes (jobs fired, leadership
returned) with tolerances bounded well above the boundary resolution —
never timing-value assertions, never two-task race windows. Document
the extension where the posture is stated.
Rows to add:
scheduler_boundary_fires— register@every 1s, run the runner briefly, stop: boundary fires land as ordinary claimable jobs stamped with theScheduleOpts/derived-default resolution (ADR-009 §3, ADR-020 §3); no double-fire per boundary while one leader runs (fire + boundary-advance commit atomically).scheduler_bounded_catchup— register, delay before starting the runner (several elapsed boundaries), then run: missed boundaries replay boundary-by-boundary (at-least-once per elapsed boundary), the fired count bounded by the 64-cap (a multi-second downtime stays far below it — the ≤64 catch-up leg is the suite-pinnable one). The beyond-cap skip-forward leg needs >64 elapsed boundaries (~65+ s) and row-backdating the public surface does not express — pin it engine-side where the schedule row'snext_fire_atis directly manipulable (both engines' scheduler tests), and record that disposition in Notes for the review gate.scheduler_leadership_discipline— two concurrent runners: exactly one leader (fires not duplicated), the loser surfacesLeadershipLost(ADR-009 §6) rather than silently co-ticking; a stopped runner's clean-stop return shape per ADR-019 §4.
Disposition note from the wave-4 fix-batch gate (recorded, not a
defect): the pg tick's in-tx enqueue fires issue no pg_notify wake
(SQLite's watcher wakes on those commits) — a cross-engine
wake-latency-parity gap the re-poll safety net covers for correctness.
If this row's shape demands the parity, the one-call wake_tx inside
the tick's tx rides this task; otherwise record the row's
latency-posture in Notes and leave the engines as they are. [Retired
2026-10-10] The parity was not demanded here, but the gap itself is
now closed: tasks/pg-fix-scheduler-fire-wake.md landed the tick's
commit-atomic fire wake (one coalesced wake per firing schedule per
tick) — see the updated Notes bullet below.
Acceptance Criteria
- The determinism-posture extension documented in
properties.rs's module doc (runner-driving rows admitted, tolerance-bounded state assertions only) - Three rows exist, version-stamped (ADR-009 §3/§4/§6, ADR-019 §4, ADR-020 §3 as applicable)
- Rows wired into both engines'
contract_suite.rstargets - SQLite column green server-less; pg column green against the harness server
- The skip-forward-leg disposition and the fire-wake-parity disposition recorded in Notes
cargo test -p alkstore-sqlite -p alkstore-postgresgreen (pg rows skip cleanly server-less); clippy-D warnings; fmt clean
References
- docs/architecture/core-contract.md §Verification backlog (scheduler
- depth properties on Postgres)
- docs/architecture/decisions/009-scheduler-collapse.md §3/§4/§6
- docs/architecture/decisions/019-mechanism-handle-surfaces.md §4
- tasks/review-wave-4-fixes.md (the fire-wake-parity recorded-for-later note)
Notes
- Beyond-cap skip-forward leg: pinned engine-side, already done in
both engines. The leg needs >64 elapsed boundaries (~65+ s of
downtime) and a direct schedule-row backdate the public surface does
not express (
next_fire_atis engine-internal). Disposition: pinned wherenext_fire_atis directly manipulable —catch_up_replays_up_to_the_cap_then_skips_forwardexists inalkstore-sqlite/src/store/scheduler_tests.rs(substrate row backdate, exactly 64 fires, skip-forward past now, skipped boundaries never enqueue) and the twin inalkstore-postgres/src/store/scheduler_tests.rs(admin-SQL backdate, same shape). The suite row (scheduler_bounded_catchup) pins the ≤64 in-band catch-up leg — the suite-observable one — and states the engine-side disposition in its doc text. - Fire-wake parity: not demanded by this row set; engines left
as-is. The suite rows observe fire outcomes through ordinary
claim polling, never through wake latency, so the row shapes demand
no wake. The recorded cross-engine latency-parity gap stood (the pg
tick's in-tx enqueue fires no
pg_notify; SQLite's watcher wakes on those commits —tasks/review-wave-4-fixes.md's recorded-for-later note): correctness is covered by the queues row's pinned re-poll safety net on both engines; the row set here pins state outcomes with tolerances bounded well above the boundary resolution, so wake latency is out of scope. [Retired 2026-10-10] The gap itself is closed bytasks/pg-fix-scheduler-fire-wake.md: the pg tick now issues its commit-atomicpg_notifyfire wake inside the tick's tx (one coalesced wake per firing schedule per tick, channel = the fired queue's name, best-effort per ADR-006), so both engines wake on a firing tick's commit and the three-way suite story starts symmetric. The row shapes here stay as landed — claim-polling observation, wake latency out of scope — with no re-shape owed. - Runner-driving posture details (the extension's riders, beyond
what the description pinned): the spawned runners share the row's
single store handle via
Arc<dyn Store>clones — leadership is storage-level, not handle-level, so two concurrent runners on one handle still contend on the leadership lock (the SQLite engine's own tests open a second handle on the same path; the suite's factory gives one isolated store per row, and the shared-handle shape exercises the same storage-level discipline). A leadership-lock acquire through the writer slot is a short lease on both engines, so the shared handle neither deadlocks nor serializes the tick away. The two-concurrent-runners row takes whichever runner wins (no identity assumption): exactly oneOk(())and oneErr(LeadershipLost)pair, checked by value. - Band tolerances (the determinism posture's "bounded well above
the boundary resolution"): every fired-count assertion is bounded by
an elapsed-boundary band measured from just before registration
(
unix-second boundary arithmeticslack + one in-progress tick at cancel, +2 s — the engine tests' fire-window tolerance), never a timing value. The catch-up row's lower bound (≥ 3 fires) is the boundary-by-boundary pin — strictly more than a coalescing single-fire would leave; its registration's first-fire grace of one interval is counted in the band.
Summary
Landed: the determinism-posture extension in
alkstore-contract-suite/src/properties.rs's module doc (runner-driving
rows admitted — spawned run_schedules(stop) tasks with a core
StopToken, state outcomes only, elapsed-boundary-band tolerances,
never timing-value assertions, never two-task race windows; ADR-017 §2
class 4 posture note, cross-connection wake orchestration still
engine-test-side) and the three scheduler rows, version-stamped and
re-exported from the suite lib:
scheduler_boundary_fires— stamps ADR-009 §3/§4, ADR-020 §3, ADR-019 §4; spawned runner, ≥2 boundary fires observed via ordinary claim polling, the fire row's stamp source pinned (ScheduleOpts priority/max_attempts over the plain-queue derived defaults 300/5/none, payload exact), clean-stopOk(())prompt, fired count inside the elapsed-boundary band (no double-fire per boundary under one leader).scheduler_bounded_catchup— stamp ADR-009 §4 (plus ADR-019 §4); runner-less downtime proves schedules never fire without a runner, a 4 s pause yields ≥3 replayed fires (boundary-by-boundary, not coalesced) bounded below the 64-cap and inside the band; the beyond-cap skip-forward disposition recorded in Notes (pinned engine-side on both engines, already implemented there).scheduler_leadership_discipline— stamps ADR-009 §1/§6, ADR-019 §4; two spawned runners on one store, exactly oneOk(())/Err(LeadershipLost)pair by value, fired count inside the elapsed-boundary band with both alive (no duplicated fires).
Wired into alkstore-sqlite/tests/contract_suite.rs (three direct
tests) and alkstore-postgres/tests/contract_suite.rs (three
harness_row! entries). Suite crate gained tokio (rt + time) as a
regular dependency for the runner-driving rows (the unused
dev-dependency dropped).
Gates: cargo build; cargo test -p alkstore-sqlite (189 lib +
16 suite rows green); cargo test -p alkstore-postgres serverless
green (pg rows skip cleanly); the full pg column green against the
harness server (pglo-poc :15432 — 121 lib incl. the previously
skipped rows + 16 suite rows + 9 schema); the three new rows re-run
twice solo on each engine (determinism check, green both columns both
runs); cargo clippy --all-targets -- -D warnings; cargo fmt --check
(clean after fmt).