15 KiB
ADR-009: Scheduler collapses into queues — schedule() + runner, @every-only v1
Status
Accepted (2026-10-05, Phase 1 — OQ-09's resolution; resolves the scheduler row transferred by ADR-008 §7)
Context
ADR-002 took the scheduler in as
"documented-thin": the family-wide "who sweeps / renews / reaps, and
when" problem is real, but only just begins to have consumers lean on
it. OQ-09's decision rule (inherited from
queues.md's framing of it): decide against the
inventory rows, not against the whole honker menu — the scheduler is a
first-class mechanism if consumers need inspectable/pausable schedule
objects (add/pause/resume/update/list/remove), and otherwise it
collapses into queues + a schedule() call with the tick machinery
absorbed.
The evidence for the decision:
- alkblobs (sweep cadence) — documented: consumer-inventory.md — needs "run the sweep every N ticks with leader election across a fleet." An interval, not a schedule object.
- alkfs (orphan reaping) — [recalled: alkfs OQ-FS-07,
/workspace/@alkdev/alkfs/docs/research/phase-0.md] — a periodic orphan-cleanup pass. An interval. - No consumer document (pinned, documented, or operator-authority) names any of: wall-clock cron expressions, per-schedule timezones, pausing/resuming schedules without unregistering, listing schedule objects, or updating a schedule's expression in place.
So the collapse question itself answers cleanly. Two secondary questions ride along and this ADR pins them with the collapse:
- Which schedule specs does v1 accept? Honker's parser accepts
5-field cron, seconds-field cron, and
@every <n><unit>— but its cron boundary arithmetic is hard-wired to host local time (chrono::Local, honker cron.rs), which is brittleness a contract should not inherit: a server TZ change silently re-times every stored schedule, and the two engines would diverge if one inherited local time and the other pinned UTC. - What runs the tick? Every consumer punting "the embedder owns cadence" is exactly the hand-rolling the substrate exists to kill — but an ambient runner violates the family's no-ambient-timers posture (alkblobs ADR-005). The line between them is opt-in.
Decision
1. The scheduler collapses into queues
The contract surface is two registration methods on Store plus one
runner; there is no Scheduler handle, no schedule objects, no
inspection surface:
store.schedule(name, spec, queue, payload, opts) -> Result<Schedule>
store.unschedule(name) -> bool
store.run_schedules(stop) -> Result<()> // runs until `stop`
scheduleupserts byname(re-register replaces the whole row) — honker's register semantics, which makes update = re-register and pause = unregister + re-register. Both are honest v1 substitutes; nothing in the inventory needs finer control.Scheduleis a value carrying the registered row (name, spec, queue, opts) — read-back for confirmation, not a mechanism handle.- Name validation: schedule names follow the queue-name rules —
non-empty (
InvalidName) and reserved-prefix rejection (ReservedName) at the entry point;schedule()/unschedule()are name-bearing entry points in the ADR-008 §4 sense (their consumers-side obligation extends the entry-point list: a reserved-prefix schedule name could collide with the machinery's derived names). No charset rule beyond that in v1. (The queue argument ofschedule()gets the same directly-supplied-queue-name validation, added 2026-10-07 by ADR-021 §3 — without it, a schedule row could store a reserved name as its fire target and every boundary fire would enqueue into it, a third write path into the outbox's backing queue contradicting ADR-014's stated guarantee.) run_schedulesis the tick: acquire the leadership lock, loop {renew lock — on loss, return before ticking; fire every due boundary; sleep until the next due boundary orstop}. Losing leadership ends the runner — the consumer's recipe is to respawn it (documented); the lock loss always precedes any stolen fire, so respawn never double-fires a boundary.- Stop and return semantics:
stopis a cancellation token — a core-crate cloneableStopToken(pinned by ADR-019 §4, 2026-10-06; the original "exact type shapes at implementation" deferral is superseded). A leadership loss returnsErr(LeadershipLost)— a distinct, matchable outcome from the cleanOk(())a stop produces; the respawn recipe matches on it. (One taxonomy note:LeadershipLostis a return-shape variant of the runner, pinned by the same act-differently rule as ADR-008 §5 — the caller acts on it by respawning.) (Annotated 2026-10-11 — review 004 item 9; amend-in-place per ADR-017 §2 class 1, no publish tag exists.) The loss arm is the refused renewal (renew→Ok(false)): the lease expired and was re-acquired elsewhere. A failed renewal is not a loss and never returnsLeadershipLost— the lock semantics were never reached: a closed store (respawning a leader against a dead store is misdirected; the engine-wide closed-store posture applies) and storage failure both surface the opaqueDatabasefallback with the engine detail preserved via the source chain. Engines whose renew is infallible (Ok(bool)only) have no failure arm to classify; engines whose renew can fail must propagate theErrunchanged (the SQLite engine's pre-annotation flattening —unwrap_or(false)— was the finding). - Schedules do not fire until a consumer runs
run_schedules— the opt-in is explicit, the no-ambient-timers posture holds, ADR-002's scheduler evidence is served, and consumers who never schedule never pay for the runner.
2. v1 spec grammar: @every <n><unit> only
spec accepts exactly @every <n><unit> with n a positive integer
and <unit> in s | m | h | d (honker's @every grammar,
honker cron.rs). Everything else — 5-field cron, seconds-field cron —
parses as InvalidSpec { spec } and is rejected.
- Why not cron strings: no consumer row names one (checked against
the inventory, not against honker's menu), and accepting them forces
a timezone answer now.
@everydurations are TZ-free — the TZ question dissolves instead of being answered. - Extension path (when a consumer names wall-clock cron via a consumer-inventory row): cron-string support with an explicit TZ decision attached (UTC-pinned contract or a per-schedule TZ parameter — decided by that extension, not assumed here).
- This keeps the SQLite engine off honker's cron boundary machinery
entirely (
@everyadvance isnext = last + interval) — the local-TZ brittleness is not inherited, and a dependency requirement is not created. (Consumer-inventory rows change this only via the row-first discipline.)
3. Enqueue semantics per boundary
At each due boundary the tick enqueues the schedule's payload
(serde_json::Value, uniform with all payloads) into the named queue
with ScheduleOpts { priority, max_attempts, expires } — the
EnqueueOpts fields minus
delay/run-at (a boundary fire is never delayed by its own config; the
schedule is the delay). The enqueued work is ordinary queue work: it
inherits the queue mechanism's delivery guarantees
(ADR-006 queues row). The
scheduler adds no delivery machinery of its own.
4. Boundary guarantee — the scheduler's row in the delivery table
Extends ADR-006's table, resolving the transfer from ADR-008 §7:
| Mechanism | Durability | Replay | Atomicity | Guarantee |
|---|---|---|---|---|
scheduler (@every) |
schedule rows + boundary fires are durable rows | n/a (fire = enqueue; the work's durability is the queues row) | fire and boundary-advance commit atomically under a row lock on the schedule row (crash mid-tick rolls back both — the boundary refires; a rogue ticker cannot double-fire) | at-least-once per elapsed boundary while a leader runs, with bounded catch-up: downtime fires replay boundary-by-boundary up to a fixed contract-constant per-tick cap (64, honker parity), beyond which missed boundaries skip forward to the next future boundary |
- The bounded-catch-up + skip-forward semantics are pinned as the
honest contract: a consumer who pauses runners for a week does not
come back to 10,000 stacked sweep jobs. (For an
@every 60ssweep under an outage, replays beyond the cap are useless anyway.) The cap is a fixed contract constant (64 per task per tick, honker parity) — not a knob; like the backoff cap, no consumer names one, and the skip-forward is documented behavior, not configuration. - One firer per boundary is pinned engine-generically by two
layers: (a) the leadership lock — one leader at a time, a leader
that lost the lock returns before ticking; (b) the boundary
advance is a transactional claim on the schedule row: the fire
transaction takes a row lock on the schedule row, re-checks
next_fire_at <= nowunder the lock, enqueues, and advances — so even a rogue second ticker (an operator bypassing the leadership lock) observes a boundary as already-fired and skips it, on either engine (Postgres:SELECT … FOR UPDATEon the row inside the fire tx; SQLite: the same guarantee via writer serialization —BEGIN IMMEDIATE). Layer (b) is the correctness floor, layer (a) the efficiency (one waiter instead of row-lock contention); crash mid-tick rolls back enqueue + advance together (the boundary refires). - Leader election is a named lock under the reserved prefix —
__alkstore_scheduler(engine-derived name in the lock namespace per ADR-008 §4's rule; consumers can never collide with it because every entry point rejects the reserved prefix). Both engines: each engine's own lock machinery, engine- internal detail. One firer per boundary across all processes sharing the store.
5. Schedule storage is engine-internal
Schedule rows live in engine-owned storage (SQLite: the forked
substrate's __alkstore_scheduler_tasks table per
ADR-011 — honker's scheduler-task
table re-owned and renamed, carrying @every specs in its spec
column, the cron machinery not ported; Postgres: the engine-owned
schema per ADR-010 §8). Not consumer
namespace;
schedule names take the same validation as queue names
(§1 — non-empty, reserved-prefix rejected). No registry
validation against queue names — enqueuing into a queue nobody has
claimed yet is fine (queues are names, not registered objects).
6. Error taxonomy delta
Two variants added by this surface, per ADR-008 §5's act-differently
rule: InvalidSpec { spec } — a schedule spec that doesn't parse
under §2's grammar (the caller fixes the spec string; rejected before
storage, no round trip) — and LeadershipLost — the run_schedules
return when leadership is lost mid-run (the caller acts by
respawning; distinct from clean stop, §1).
Consequences
Positive
- The contract gains the scheduler for two registrations + one runner;
the "who sweeps" problem has a substrate answer that is one
@everyrow and one spawned task. - The tick machinery (leader election, boundary advance, catch-up cap)
is engine-internal and mostly inherited (SQLite: honker's scheduler
machinery with
@everyspecs; Postgres: re-derived, small at this surface size — the row-locked fire tx is the one mechanism the contract pins engine-generically, §4). - No TZ semantics, no cron parser in the contract, no schedule-object CRUD — the surface stays proportionate to the two inventory rows that exist.
Negative
- Honker's richer scheduler (
add/update/pause/resume/list, cron with TZ) is deliberately not consumed at the contract level; consumers needing inspectable schedules need an inventory row + contract extension first. - The opt-in runner is one more thing a deployer must remember: forgot
run_schedules⇒ schedules silently never fire. The mechanism's doc text must carry this (a silent-nothing deployment posture, the honest twin of "no ambient timers"). - The catch-up cap is a consumer-visible loss mode for long outages — documented (skip-forward), not silent: the schedule row's advance makes the gap inspectable.
- The scheduler surface (and this track's
QueueOptsdepth) are post-v1 contract extensions — the first exercises of the extension path ADR-008 leaves open; their versioning discipline (how engine crates track the additions) is OQ-10's, still open. (Resolved 2026-10-06 by ADR-017: nothing was released when these extensions were made, so they factually fall in ADR-017 class 1 — pre-implementation amend-in-place — and ship inside the initial contract v1 text; the post-release extension shape this ADR exercised is ADR-017's class 2.)
References
- OQ-09 (
docs/architecture/open-questions.md) — this ADR's resolution; owns ADR-008 §7's transferred scheduler guarantee row. docs/research/consumer-inventory.md§scheduler — the two rows the decision is made against.- ADR-002 — scheduler in-scope posture this ADR completes (documented-thin → pinned thin).
- ADR-006 — the guarantee table §4 extends; the enqueued work inherits the queues row.
- ADR-008 §1 (scheduler out of v1
pending this decision), §4 (reserved-prefix rule the leadership lock
name follows), §5 (act-differently rule the
InvalidSpecvariant follows), §7 (the guarantee-row transfer). - ADR-010 — the queue semantics depth this ADR's fire semantics compose with; the maintenance recipe (who sweeps) is recorded there with the collapse machinery.
- Honker reference checkout
/workspace/honker@ f4e53c6 — scheduler surface (packages/honker-rs/src/lib.rs,honker-core/src/honker_ops.rs),@everygrammar and local-TZ brittleness (honker-core/src/cron.rs), tick/leader mechanics. - ADR-011 — the SQLite-side schedule storage this §5 describes is the forked substrate's; cron machinery not ported.
- core-contract.md — the spec carrying this surface.