Files
alkstore/tasks/suite-scheduler-rows.md
glm-5.3-flash 57451cef46 pg-fix-scheduler-fire-wake — the scheduler tick's fire wake lands, closing the recorded cross-engine fire-wake latency-parity gap (the wave-5 gate's one standing item, review-wave-4-fixes' recorded-for-later note): scheduler.rs::tick now issues one coalesced pg_notify per firing schedule per tick after the fire loop, still inside the tick's in_tx frame — channel = the fired queue's name (schedule() rejects reserved names per ADR-021 §3, so the plain queue name is always the right channel), empty payload (Wake { channel } only, ADR-008 §3), best-effort via the shared tx::wake_tx (widened private → pub(crate) per the recorded one-call shape; enqueue_row stays wake-free — no double-wake; auto-commit Queue::enqueue and the tx producer paths untouched). NOTIFY's native transactional delivery makes the wake commit-atomic: a rolled-back tick (crash mid-tick ⇒ boundary refires) discards the wake with its fire rows — the tx/no-ghosts discipline the tx producer paths already pin; coalescing default one-per-firing-schedule-per-tick (not per-fire) matching the SQLite watcher's per-committed-tick cadence, decision recorded in the task's Notes. New four-arm pinning test (scheduler_tests.rs, tx_producer_wakes_are_commit_atomic pattern): pre-commit silence inside the runner's own in_tx frame driving the engine's own tick (rogue-ticks shape), deterministic delivery at/after commit, coalescing (≥2 boundaries → one wake), nothing-due silence (same-tick not-due schedule + a later all-not-due tick, in-frame and post-commit), and the end-to-end runner leg (a live run_schedules leader's wake reaches a registered listener); local listener_hears helper per the per-file helper convention; rollback arm native (NOTIFY transactional delivery — no tick-fault seam built, per the task pin). Docs: tx.rs 'The tx wakes' section + wake_tx doc name the scheduler fire path as a shared caller; scheduler.rs module docs + tick doc pin the coalesced commit-atomic semantics. Record updates: review-wave-4-fixes' recorded-for-later bullet, suite-scheduler-rows' gate disposition note + Notes bullet, review-wave-5 §Notes 4's audit row — all retired with pointers; implementation.md gains the review-rounds line (fires visible to registered listeners ahead of the mem engine's posture being written). Also carried: review-wave-5 §6 flake-ledger update — row_lock_ttl_expiry_and_reacquisition failed once more in a full pg run this session (green on the next two full runs + six focused; unrelated mechanism to this change), meeting the ledger's own 'fails again' trigger with the dedicated investigative session owed by the wave-7 watch carriage. Verified: full pg lib suite 122/122 vs the harness (new test green ×3 solo); pg contract suite 25/25 + schema 9/9; workspace cargo test green server-less (pg skips clean incl. the new test); clippy --all-targets -D warnings clean; fmt clean
2026-10-10 10:27:35 +00:00

9.9 KiB

id, name, status, depends_on, scope, risk, impact, level, tags
id name status depends_on scope risk impact level tags
suite-scheduler-rows Contract-suite rows — scheduler boundary/catch-up/leadership (+ determinism-posture extension) completed
moderate medium phase implementation
wave-5
contract-suite
scheduler

Description

Add the scheduler contract rows to alkstore-contract-suite/src/properties.rs and wire them into both engines' suite targets, discharging the backlog row "Scheduler + depth properties on Postgres" (core-contract.md §Verification backlog): the tick machinery — boundary advance, bounded catch-up, leadership-loss discipline — pinned engine-uniform (SQLite inherits the substrate's test-pinned implementation; pg re-derived; the suite pins identical behavior).

Posture extension first (test-side, ADR-017 §2 class 4): the suite's determinism statement (properties.rs module doc) currently says runner spawning stays engine-test-side. The backlog explicitly demands these rows in the suite, so amend the module doc to admit runner-driving rows: a spawned run_schedules(stop) task with a StopToken, assertions on state outcomes (jobs fired, leadership returned) with tolerances bounded well above the boundary resolution — never timing-value assertions, never two-task race windows. Document the extension where the posture is stated.

Rows to add:

  • scheduler_boundary_fires — register @every 1s, run the runner briefly, stop: boundary fires land as ordinary claimable jobs stamped with the ScheduleOpts/derived-default resolution (ADR-009 §3, ADR-020 §3); no double-fire per boundary while one leader runs (fire + boundary-advance commit atomically).
  • scheduler_bounded_catchup — register, delay before starting the runner (several elapsed boundaries), then run: missed boundaries replay boundary-by-boundary (at-least-once per elapsed boundary), the fired count bounded by the 64-cap (a multi-second downtime stays far below it — the ≤64 catch-up leg is the suite-pinnable one). The beyond-cap skip-forward leg needs >64 elapsed boundaries (~65+ s) and row-backdating the public surface does not express — pin it engine-side where the schedule row's next_fire_at is directly manipulable (both engines' scheduler tests), and record that disposition in Notes for the review gate.
  • scheduler_leadership_discipline — two concurrent runners: exactly one leader (fires not duplicated), the loser surfaces LeadershipLost (ADR-009 §6) rather than silently co-ticking; a stopped runner's clean-stop return shape per ADR-019 §4.

Disposition note from the wave-4 fix-batch gate (recorded, not a defect): the pg tick's in-tx enqueue fires issue no pg_notify wake (SQLite's watcher wakes on those commits) — a cross-engine wake-latency-parity gap the re-poll safety net covers for correctness. If this row's shape demands the parity, the one-call wake_tx inside the tick's tx rides this task; otherwise record the row's latency-posture in Notes and leave the engines as they are. [Retired 2026-10-10] The parity was not demanded here, but the gap itself is now closed: tasks/pg-fix-scheduler-fire-wake.md landed the tick's commit-atomic fire wake (one coalesced wake per firing schedule per tick) — see the updated Notes bullet below.

Acceptance Criteria

  • The determinism-posture extension documented in properties.rs's module doc (runner-driving rows admitted, tolerance-bounded state assertions only)
  • Three rows exist, version-stamped (ADR-009 §3/§4/§6, ADR-019 §4, ADR-020 §3 as applicable)
  • Rows wired into both engines' contract_suite.rs targets
  • SQLite column green server-less; pg column green against the harness server
  • The skip-forward-leg disposition and the fire-wake-parity disposition recorded in Notes
  • cargo test -p alkstore-sqlite -p alkstore-postgres green (pg rows skip cleanly server-less); clippy -D warnings; fmt clean

References

  • docs/architecture/core-contract.md §Verification backlog (scheduler
    • depth properties on Postgres)
  • docs/architecture/decisions/009-scheduler-collapse.md §3/§4/§6
  • docs/architecture/decisions/019-mechanism-handle-surfaces.md §4
  • tasks/review-wave-4-fixes.md (the fire-wake-parity recorded-for-later note)

Notes

  • Beyond-cap skip-forward leg: pinned engine-side, already done in both engines. The leg needs >64 elapsed boundaries (~65+ s of downtime) and a direct schedule-row backdate the public surface does not express (next_fire_at is engine-internal). Disposition: pinned where next_fire_at is directly manipulable — catch_up_replays_up_to_the_cap_then_skips_forward exists in alkstore-sqlite/src/store/scheduler_tests.rs (substrate row backdate, exactly 64 fires, skip-forward past now, skipped boundaries never enqueue) and the twin in alkstore-postgres/src/store/scheduler_tests.rs (admin-SQL backdate, same shape). The suite row (scheduler_bounded_catchup) pins the ≤64 in-band catch-up leg — the suite-observable one — and states the engine-side disposition in its doc text.
  • Fire-wake parity: not demanded by this row set; engines left as-is. The suite rows observe fire outcomes through ordinary claim polling, never through wake latency, so the row shapes demand no wake. The recorded cross-engine latency-parity gap stood (the pg tick's in-tx enqueue fires no pg_notify; SQLite's watcher wakes on those commits — tasks/review-wave-4-fixes.md's recorded-for-later note): correctness is covered by the queues row's pinned re-poll safety net on both engines; the row set here pins state outcomes with tolerances bounded well above the boundary resolution, so wake latency is out of scope. [Retired 2026-10-10] The gap itself is closed by tasks/pg-fix-scheduler-fire-wake.md: the pg tick now issues its commit-atomic pg_notify fire wake inside the tick's tx (one coalesced wake per firing schedule per tick, channel = the fired queue's name, best-effort per ADR-006), so both engines wake on a firing tick's commit and the three-way suite story starts symmetric. The row shapes here stay as landed — claim-polling observation, wake latency out of scope — with no re-shape owed.
  • Runner-driving posture details (the extension's riders, beyond what the description pinned): the spawned runners share the row's single store handle via Arc<dyn Store> clones — leadership is storage-level, not handle-level, so two concurrent runners on one handle still contend on the leadership lock (the SQLite engine's own tests open a second handle on the same path; the suite's factory gives one isolated store per row, and the shared-handle shape exercises the same storage-level discipline). A leadership-lock acquire through the writer slot is a short lease on both engines, so the shared handle neither deadlocks nor serializes the tick away. The two-concurrent-runners row takes whichever runner wins (no identity assumption): exactly one Ok(()) and one Err(LeadershipLost) pair, checked by value.
  • Band tolerances (the determinism posture's "bounded well above the boundary resolution"): every fired-count assertion is bounded by an elapsed-boundary band measured from just before registration (unix-second boundary arithmetic slack + one in-progress tick at cancel, +2 s — the engine tests' fire-window tolerance), never a timing value. The catch-up row's lower bound (≥ 3 fires) is the boundary-by-boundary pin — strictly more than a coalescing single-fire would leave; its registration's first-fire grace of one interval is counted in the band.

Summary

Landed: the determinism-posture extension in alkstore-contract-suite/src/properties.rs's module doc (runner-driving rows admitted — spawned run_schedules(stop) tasks with a core StopToken, state outcomes only, elapsed-boundary-band tolerances, never timing-value assertions, never two-task race windows; ADR-017 §2 class 4 posture note, cross-connection wake orchestration still engine-test-side) and the three scheduler rows, version-stamped and re-exported from the suite lib:

  • scheduler_boundary_fires — stamps ADR-009 §3/§4, ADR-020 §3, ADR-019 §4; spawned runner, ≥2 boundary fires observed via ordinary claim polling, the fire row's stamp source pinned (ScheduleOpts priority/max_attempts over the plain-queue derived defaults 300/5/none, payload exact), clean-stop Ok(()) prompt, fired count inside the elapsed-boundary band (no double-fire per boundary under one leader).
  • scheduler_bounded_catchup — stamp ADR-009 §4 (plus ADR-019 §4); runner-less downtime proves schedules never fire without a runner, a 4 s pause yields ≥3 replayed fires (boundary-by-boundary, not coalesced) bounded below the 64-cap and inside the band; the beyond-cap skip-forward disposition recorded in Notes (pinned engine-side on both engines, already implemented there).
  • scheduler_leadership_discipline — stamps ADR-009 §1/§6, ADR-019 §4; two spawned runners on one store, exactly one Ok(())/Err(LeadershipLost) pair by value, fired count inside the elapsed-boundary band with both alive (no duplicated fires).

Wired into alkstore-sqlite/tests/contract_suite.rs (three direct tests) and alkstore-postgres/tests/contract_suite.rs (three harness_row! entries). Suite crate gained tokio (rt + time) as a regular dependency for the runner-driving rows (the unused dev-dependency dropped).

Gates: cargo build; cargo test -p alkstore-sqlite (189 lib + 16 suite rows green); cargo test -p alkstore-postgres serverless green (pg rows skip cleanly); the full pg column green against the harness server (pglo-poc :15432 — 121 lib incl. the previously skipped rows + 16 suite rows + 9 schema); the three new rows re-run twice solo on each engine (determinism check, green both columns both runs); cargo clippy --all-targets -- -D warnings; cargo fmt --check (clean after fmt).