Files

15 KiB

ADR-009: Scheduler collapses into queues — schedule() + runner, @every-only v1

Status

Accepted (2026-10-05, Phase 1 — OQ-09's resolution; resolves the scheduler row transferred by ADR-008 §7)

Context

ADR-002 took the scheduler in as "documented-thin": the family-wide "who sweeps / renews / reaps, and when" problem is real, but only just begins to have consumers lean on it. OQ-09's decision rule (inherited from queues.md's framing of it): decide against the inventory rows, not against the whole honker menu — the scheduler is a first-class mechanism if consumers need inspectable/pausable schedule objects (add/pause/resume/update/list/remove), and otherwise it collapses into queues + a schedule() call with the tick machinery absorbed.

The evidence for the decision:

  • alkblobs (sweep cadence) — documented: consumer-inventory.md — needs "run the sweep every N ticks with leader election across a fleet." An interval, not a schedule object.
  • alkfs (orphan reaping) — [recalled: alkfs OQ-FS-07, /workspace/@alkdev/alkfs/docs/research/phase-0.md] — a periodic orphan-cleanup pass. An interval.
  • No consumer document (pinned, documented, or operator-authority) names any of: wall-clock cron expressions, per-schedule timezones, pausing/resuming schedules without unregistering, listing schedule objects, or updating a schedule's expression in place.

So the collapse question itself answers cleanly. Two secondary questions ride along and this ADR pins them with the collapse:

  • Which schedule specs does v1 accept? Honker's parser accepts 5-field cron, seconds-field cron, and @every <n><unit> — but its cron boundary arithmetic is hard-wired to host local time (chrono::Local, honker cron.rs), which is brittleness a contract should not inherit: a server TZ change silently re-times every stored schedule, and the two engines would diverge if one inherited local time and the other pinned UTC.
  • What runs the tick? Every consumer punting "the embedder owns cadence" is exactly the hand-rolling the substrate exists to kill — but an ambient runner violates the family's no-ambient-timers posture (alkblobs ADR-005). The line between them is opt-in.

Decision

1. The scheduler collapses into queues

The contract surface is two registration methods on Store plus one runner; there is no Scheduler handle, no schedule objects, no inspection surface:

store.schedule(name, spec, queue, payload, opts) -> Result<Schedule>
store.unschedule(name) -> bool
store.run_schedules(stop) -> Result<()>        // runs until `stop`
  • schedule upserts by name (re-register replaces the whole row) — honker's register semantics, which makes update = re-register and pause = unregister + re-register. Both are honest v1 substitutes; nothing in the inventory needs finer control.
  • Schedule is a value carrying the registered row (name, spec, queue, opts) — read-back for confirmation, not a mechanism handle.
  • Name validation: schedule names follow the queue-name rules — non-empty (InvalidName) and reserved-prefix rejection (ReservedName) at the entry point; schedule()/unschedule() are name-bearing entry points in the ADR-008 §4 sense (their consumers-side obligation extends the entry-point list: a reserved-prefix schedule name could collide with the machinery's derived names). No charset rule beyond that in v1. (The queue argument of schedule() gets the same directly-supplied-queue-name validation, added 2026-10-07 by ADR-021 §3 — without it, a schedule row could store a reserved name as its fire target and every boundary fire would enqueue into it, a third write path into the outbox's backing queue contradicting ADR-014's stated guarantee.)
  • run_schedules is the tick: acquire the leadership lock, loop {renew lock — on loss, return before ticking; fire every due boundary; sleep until the next due boundary or stop}. Losing leadership ends the runner — the consumer's recipe is to respawn it (documented); the lock loss always precedes any stolen fire, so respawn never double-fires a boundary.
  • Stop and return semantics: stop is a cancellation token — a core-crate cloneable StopToken (pinned by ADR-019 §4, 2026-10-06; the original "exact type shapes at implementation" deferral is superseded). A leadership loss returns Err(LeadershipLost) — a distinct, matchable outcome from the clean Ok(()) a stop produces; the respawn recipe matches on it. (One taxonomy note: LeadershipLost is a return-shape variant of the runner, pinned by the same act-differently rule as ADR-008 §5 — the caller acts on it by respawning.) (Annotated 2026-10-11 — review 004 item 9; amend-in-place per ADR-017 §2 class 1, no publish tag exists.) The loss arm is the refused renewal (renew → Ok(false)): the lease expired and was re-acquired elsewhere. A failed renewal is not a loss and never returns LeadershipLost — the lock semantics were never reached: a closed store (respawning a leader against a dead store is misdirected; the engine-wide closed-store posture applies) and storage failure both surface the opaque Database fallback with the engine detail preserved via the source chain. Engines whose renew is infallible (Ok(bool) only) have no failure arm to classify; engines whose renew can fail must propagate the Err unchanged (the SQLite engine's pre-annotation flattening — unwrap_or(false) — was the finding).
  • Schedules do not fire until a consumer runs run_schedules — the opt-in is explicit, the no-ambient-timers posture holds, ADR-002's scheduler evidence is served, and consumers who never schedule never pay for the runner.

2. v1 spec grammar: @every <n><unit> only

spec accepts exactly @every <n><unit> with n a positive integer and <unit> in s | m | h | d (honker's @every grammar, honker cron.rs). Everything else — 5-field cron, seconds-field cron — parses as InvalidSpec { spec } and is rejected.

  • Why not cron strings: no consumer row names one (checked against the inventory, not against honker's menu), and accepting them forces a timezone answer now. @every durations are TZ-free — the TZ question dissolves instead of being answered.
  • Extension path (when a consumer names wall-clock cron via a consumer-inventory row): cron-string support with an explicit TZ decision attached (UTC-pinned contract or a per-schedule TZ parameter — decided by that extension, not assumed here).
  • This keeps the SQLite engine off honker's cron boundary machinery entirely (@every advance is next = last + interval) — the local-TZ brittleness is not inherited, and a dependency requirement is not created. (Consumer-inventory rows change this only via the row-first discipline.)

3. Enqueue semantics per boundary

At each due boundary the tick enqueues the schedule's payload (serde_json::Value, uniform with all payloads) into the named queue with ScheduleOpts { priority, max_attempts, expires } — the EnqueueOpts fields minus delay/run-at (a boundary fire is never delayed by its own config; the schedule is the delay). The enqueued work is ordinary queue work: it inherits the queue mechanism's delivery guarantees (ADR-006 queues row). The scheduler adds no delivery machinery of its own.

4. Boundary guarantee — the scheduler's row in the delivery table

Extends ADR-006's table, resolving the transfer from ADR-008 §7:

Mechanism Durability Replay Atomicity Guarantee
scheduler (@every) schedule rows + boundary fires are durable rows n/a (fire = enqueue; the work's durability is the queues row) fire and boundary-advance commit atomically under a row lock on the schedule row (crash mid-tick rolls back both — the boundary refires; a rogue ticker cannot double-fire) at-least-once per elapsed boundary while a leader runs, with bounded catch-up: downtime fires replay boundary-by-boundary up to a fixed contract-constant per-tick cap (64, honker parity), beyond which missed boundaries skip forward to the next future boundary
  • The bounded-catch-up + skip-forward semantics are pinned as the honest contract: a consumer who pauses runners for a week does not come back to 10,000 stacked sweep jobs. (For an @every 60s sweep under an outage, replays beyond the cap are useless anyway.) The cap is a fixed contract constant (64 per task per tick, honker parity) — not a knob; like the backoff cap, no consumer names one, and the skip-forward is documented behavior, not configuration.
  • One firer per boundary is pinned engine-generically by two layers: (a) the leadership lock — one leader at a time, a leader that lost the lock returns before ticking; (b) the boundary advance is a transactional claim on the schedule row: the fire transaction takes a row lock on the schedule row, re-checks next_fire_at <= now under the lock, enqueues, and advances — so even a rogue second ticker (an operator bypassing the leadership lock) observes a boundary as already-fired and skips it, on either engine (Postgres: SELECT … FOR UPDATE on the row inside the fire tx; SQLite: the same guarantee via writer serialization — BEGIN IMMEDIATE). Layer (b) is the correctness floor, layer (a) the efficiency (one waiter instead of row-lock contention); crash mid-tick rolls back enqueue + advance together (the boundary refires).
  • Leader election is a named lock under the reserved prefix — __alkstore_scheduler (engine-derived name in the lock namespace per ADR-008 §4's rule; consumers can never collide with it because every entry point rejects the reserved prefix). Both engines: each engine's own lock machinery, engine- internal detail. One firer per boundary across all processes sharing the store.

5. Schedule storage is engine-internal

Schedule rows live in engine-owned storage (SQLite: the forked substrate's __alkstore_scheduler_tasks table per ADR-011 — honker's scheduler-task table re-owned and renamed, carrying @every specs in its spec column, the cron machinery not ported; Postgres: the engine-owned schema per ADR-010 §8). Not consumer namespace; schedule names take the same validation as queue names (§1 — non-empty, reserved-prefix rejected). No registry validation against queue names — enqueuing into a queue nobody has claimed yet is fine (queues are names, not registered objects).

6. Error taxonomy delta

Two variants added by this surface, per ADR-008 §5's act-differently rule: InvalidSpec { spec } — a schedule spec that doesn't parse under §2's grammar (the caller fixes the spec string; rejected before storage, no round trip) — and LeadershipLost — the run_schedules return when leadership is lost mid-run (the caller acts by respawning; distinct from clean stop, §1).

Consequences

Positive

  • The contract gains the scheduler for two registrations + one runner; the "who sweeps" problem has a substrate answer that is one @every row and one spawned task.
  • The tick machinery (leader election, boundary advance, catch-up cap) is engine-internal and mostly inherited (SQLite: honker's scheduler machinery with @every specs; Postgres: re-derived, small at this surface size — the row-locked fire tx is the one mechanism the contract pins engine-generically, §4).
  • No TZ semantics, no cron parser in the contract, no schedule-object CRUD — the surface stays proportionate to the two inventory rows that exist.

Negative

  • Honker's richer scheduler (add/update/pause/resume/list, cron with TZ) is deliberately not consumed at the contract level; consumers needing inspectable schedules need an inventory row + contract extension first.
  • The opt-in runner is one more thing a deployer must remember: forgot run_schedules ⇒ schedules silently never fire. The mechanism's doc text must carry this (a silent-nothing deployment posture, the honest twin of "no ambient timers").
  • The catch-up cap is a consumer-visible loss mode for long outages — documented (skip-forward), not silent: the schedule row's advance makes the gap inspectable.
  • The scheduler surface (and this track's QueueOpts depth) are post-v1 contract extensions — the first exercises of the extension path ADR-008 leaves open; their versioning discipline (how engine crates track the additions) is OQ-10's, still open. (Resolved 2026-10-06 by ADR-017: nothing was released when these extensions were made, so they factually fall in ADR-017 class 1 — pre-implementation amend-in-place — and ship inside the initial contract v1 text; the post-release extension shape this ADR exercised is ADR-017's class 2.)

References

  • OQ-09 (docs/architecture/open-questions.md) — this ADR's resolution; owns ADR-008 §7's transferred scheduler guarantee row.
  • docs/research/consumer-inventory.md §scheduler — the two rows the decision is made against.
  • ADR-002 — scheduler in-scope posture this ADR completes (documented-thin → pinned thin).
  • ADR-006 — the guarantee table §4 extends; the enqueued work inherits the queues row.
  • ADR-008 §1 (scheduler out of v1 pending this decision), §4 (reserved-prefix rule the leadership lock name follows), §5 (act-differently rule the InvalidSpec variant follows), §7 (the guarantee-row transfer).
  • ADR-010 — the queue semantics depth this ADR's fire semantics compose with; the maintenance recipe (who sweeps) is recorded there with the collapse machinery.
  • Honker reference checkout /workspace/honker @ f4e53c6 — scheduler surface (packages/honker-rs/src/lib.rs, honker-core/src/honker_ops.rs), @every grammar and local-TZ brittleness (honker-core/src/cron.rs), tick/leader mechanics.
  • ADR-011 — the SQLite-side schedule storage this §5 describes is the forked substrate's; cron machinery not ported.
  • core-contract.md — the spec carrying this surface.