Files
alkstore/docs/architecture/decisions/009-scheduler-collapse.md
T
glm-5.3-flash 2949612e2c ADR-012: forked-substrate design — contract-blind boundary, fidelity posture, port deltas
Follow-through on OQ-06/ADR-011: pin the fork's structural decisions
(alkstore-substrate as a vendored path-dep crate, contract-blind API
boundary with contract formulas computed engine-side and pinned
equivalent by the contract suite, keep-the-kept-half API fidelity for
cheap cherry-picks, the W-1/W-2/dead-man's-switch/W-4 port deltas
decided per item, bootstrap re-keying off error-string matching, no
rename migration, deliberate upstream tracking).

Consistency sweep across the doc set for the fork: annotate ADR-003/
005/009/010 and core-contract for superseded ownership facts, fix
schedule-storage table naming (ADR-009 §5, queues.md), re-key ADR-010
§6's notifications hygiene to the at-attach cap the fork scope
realizes, add OQ-11 (scaffold-time residue), and complete both ADR
indexes. Independent review: 0 critical, warnings addressed.
2026-10-05 05:00:55 +00:00

13 KiB

ADR-009: Scheduler collapses into queues — schedule() + runner, @every-only v1

Status

Accepted (2026-10-05, Phase 1 — OQ-09's resolution; resolves the scheduler row transferred by ADR-008 §7)

Context

ADR-002 took the scheduler in as "documented-thin": the family-wide "who sweeps / renews / reaps, and when" problem is real, but only just begins to have consumers lean on it. OQ-09's decision rule (inherited from queues.md's framing of it): decide against the inventory rows, not against the whole honker menu — the scheduler is a first-class mechanism if consumers need inspectable/pausable schedule objects (add/pause/resume/update/list/remove), and otherwise it collapses into queues + a schedule() call with the tick machinery absorbed.

The evidence for the decision:

  • alkblobs (sweep cadence) — documented: consumer-inventory.md — needs "run the sweep every N ticks with leader election across a fleet." An interval, not a schedule object.
  • alkfs (orphan reaping) — [recalled: alkfs OQ-FS-07, /workspace/@alkdev/alkfs/docs/research/phase-0.md] — a periodic orphan-cleanup pass. An interval.
  • No consumer document (pinned, documented, or operator-authority) names any of: wall-clock cron expressions, per-schedule timezones, pausing/resuming schedules without unregistering, listing schedule objects, or updating a schedule's expression in place.

So the collapse question itself answers cleanly. Two secondary questions ride along and this ADR pins them with the collapse:

  • Which schedule specs does v1 accept? Honker's parser accepts 5-field cron, seconds-field cron, and @every <n><unit> — but its cron boundary arithmetic is hard-wired to host local time (chrono::Local, honker cron.rs), which is brittleness a contract should not inherit: a server TZ change silently re-times every stored schedule, and the two engines would diverge if one inherited local time and the other pinned UTC.
  • What runs the tick? Every consumer punting "the embedder owns cadence" is exactly the hand-rolling the substrate exists to kill — but an ambient runner violates the family's no-ambient-timers posture (alkblobs ADR-005). The line between them is opt-in.

Decision

1. The scheduler collapses into queues

The contract surface is two registration methods on Store plus one runner; there is no Scheduler handle, no schedule objects, no inspection surface:

store.schedule(name, spec, queue, payload, opts) -> Result<Schedule>
store.unschedule(name) -> bool
store.run_schedules(stop) -> Result<()>        // runs until `stop`
  • schedule upserts by name (re-register replaces the whole row) — honker's register semantics, which makes update = re-register and pause = unregister + re-register. Both are honest v1 substitutes; nothing in the inventory needs finer control.
  • Schedule is a value carrying the registered row (name, spec, queue, opts) — read-back for confirmation, not a mechanism handle.
  • Name validation: schedule names follow the queue-name rules — non-empty (InvalidName) and reserved-prefix rejection (ReservedName) at the entry point; schedule()/unschedule() are name-bearing entry points in the ADR-008 §4 sense (their consumers-side obligation extends the entry-point list: a reserved-prefix schedule name could collide with the machinery's derived names). No charset rule beyond that in v1.
  • run_schedules is the tick: acquire the leadership lock, loop {renew lock — on loss, return before ticking; fire every due boundary; sleep until the next due boundary or stop}. Losing leadership ends the runner — the consumer's recipe is to respawn it (documented); the lock loss always precedes any stolen fire, so respawn never double-fires a boundary.
  • Stop and return semantics: stop is a cancellation token (a cloneable handle whose flip ends the sleep early — the exact type shapes at implementation, engine-crate docs). A leadership loss returns Err(LeadershipLost) — a distinct, matchable outcome from the clean Ok(()) a stop produces; the respawn recipe matches on it. (One taxonomy note: LeadershipLost is a return-shape variant of the runner, pinned by the same act-differently rule as ADR-008 §5 — the caller acts on it by respawning.)
  • Schedules do not fire until a consumer runs run_schedules — the opt-in is explicit, the no-ambient-timers posture holds, ADR-002's scheduler evidence is served, and consumers who never schedule never pay for the runner.

2. v1 spec grammar: @every <n><unit> only

spec accepts exactly @every <n><unit> with n a positive integer and <unit> in s | m | h | d (honker's @every grammar, honker cron.rs). Everything else — 5-field cron, seconds-field cron — parses as InvalidSpec { spec } and is rejected.

  • Why not cron strings: no consumer row names one (checked against the inventory, not against honker's menu), and accepting them forces a timezone answer now. @every durations are TZ-free — the TZ question dissolves instead of being answered.
  • Extension path (when a consumer names wall-clock cron via a consumer-inventory row): cron-string support with an explicit TZ decision attached (UTC-pinned contract or a per-schedule TZ parameter — decided by that extension, not assumed here).
  • This keeps the SQLite engine off honker's cron boundary machinery entirely (@every advance is next = last + interval) — the local-TZ brittleness is not inherited, and a dependency requirement is not created. (Consumer-inventory rows change this only via the row-first discipline.)

3. Enqueue semantics per boundary

At each due boundary the tick enqueues the schedule's payload (serde_json::Value, uniform with all payloads) into the named queue with ScheduleOpts { priority, max_attempts, expires } — the EnqueueOpts fields minus delay/run-at (a boundary fire is never delayed by its own config; the schedule is the delay). The enqueued work is ordinary queue work: it inherits the queue mechanism's delivery guarantees (ADR-006 queues row). The scheduler adds no delivery machinery of its own.

4. Boundary guarantee — the scheduler's row in the delivery table

Extends ADR-006's table, resolving the transfer from ADR-008 §7:

Mechanism Durability Replay Atomicity Guarantee
scheduler (@every) schedule rows + boundary fires are durable rows n/a (fire = enqueue; the work's durability is the queues row) fire and boundary-advance commit atomically under a row lock on the schedule row (crash mid-tick rolls back both — the boundary refires; a rogue ticker cannot double-fire) at-least-once per elapsed boundary while a leader runs, with bounded catch-up: downtime fires replay boundary-by-boundary up to a fixed contract-constant per-tick cap (64, honker parity), beyond which missed boundaries skip forward to the next future boundary
  • The bounded-catch-up + skip-forward semantics are pinned as the honest contract: a consumer who pauses runners for a week does not come back to 10,000 stacked sweep jobs. (For an @every 60s sweep under an outage, replays beyond the cap are useless anyway.) The cap is a fixed contract constant (64 per task per tick, honker parity) — not a knob; like the backoff cap, no consumer names one, and the skip-forward is documented behavior, not configuration.
  • One firer per boundary is pinned engine-generically by two layers: (a) the leadership lock — one leader at a time, a leader that lost the lock returns before ticking; (b) the boundary advance is a transactional claim on the schedule row: the fire transaction takes a row lock on the schedule row, re-checks next_fire_at <= now under the lock, enqueues, and advances — so even a rogue second ticker (an operator bypassing the leadership lock) observes a boundary as already-fired and skips it, on either engine (Postgres: SELECT … FOR UPDATE on the row inside the fire tx; SQLite: the same guarantee via writer serialization — BEGIN IMMEDIATE). Layer (b) is the correctness floor, layer (a) the efficiency (one waiter instead of row-lock contention); crash mid-tick rolls back enqueue + advance together (the boundary refires).
  • Leader election is a named lock under the reserved prefix — __alkstore_scheduler (engine-derived name in the lock namespace per ADR-008 §4's rule; consumers can never collide with it because every entry point rejects the reserved prefix). Both engines: each engine's own lock machinery, engine- internal detail. One firer per boundary across all processes sharing the store.

5. Schedule storage is engine-internal

Schedule rows live in engine-owned storage (SQLite: the forked substrate's __alkstore_scheduler_tasks table per ADR-011 — honker's scheduler-task table re-owned and renamed, carrying @every specs in its spec column, the cron machinery not ported; Postgres: the engine-owned schema per ADR-010 §8). Not consumer namespace; schedule names take the same validation as queue names (§1 — non-empty, reserved-prefix rejected). No registry validation against queue names — enqueuing into a queue nobody has claimed yet is fine (queues are names, not registered objects).

6. Error taxonomy delta

Two variants added by this surface, per ADR-008 §5's act-differently rule: InvalidSpec { spec } — a schedule spec that doesn't parse under §2's grammar (the caller fixes the spec string; rejected before storage, no round trip) — and LeadershipLost — the run_schedules return when leadership is lost mid-run (the caller acts by respawning; distinct from clean stop, §1).

Consequences

Positive

  • The contract gains the scheduler for two registrations + one runner; the "who sweeps" problem has a substrate answer that is one @every row and one spawned task.
  • The tick machinery (leader election, boundary advance, catch-up cap) is engine-internal and mostly inherited (SQLite: honker's scheduler machinery with @every specs; Postgres: re-derived, small at this surface size — the row-locked fire tx is the one mechanism the contract pins engine-generically, §4).
  • No TZ semantics, no cron parser in the contract, no schedule-object CRUD — the surface stays proportionate to the two inventory rows that exist.

Negative

  • Honker's richer scheduler (add/update/pause/resume/list, cron with TZ) is deliberately not consumed at the contract level; consumers needing inspectable schedules need an inventory row + contract extension first.
  • The opt-in runner is one more thing a deployer must remember: forgot run_schedules ⇒ schedules silently never fire. The mechanism's doc text must carry this (a silent-nothing deployment posture, the honest twin of "no ambient timers").
  • The catch-up cap is a consumer-visible loss mode for long outages — documented (skip-forward), not silent: the schedule row's advance makes the gap inspectable.
  • The scheduler surface (and this track's QueueOpts depth) are post-v1 contract extensions — the first exercises of the extension path ADR-008 leaves open; their versioning discipline (how engine crates track the additions) is OQ-10's, still open.

References

  • OQ-09 (docs/architecture/open-questions.md) — this ADR's resolution; owns ADR-008 §7's transferred scheduler guarantee row.
  • docs/research/consumer-inventory.md §scheduler — the two rows the decision is made against.
  • ADR-002 — scheduler in-scope posture this ADR completes (documented-thin → pinned thin).
  • ADR-006 — the guarantee table §4 extends; the enqueued work inherits the queues row.
  • ADR-008 §1 (scheduler out of v1 pending this decision), §4 (reserved-prefix rule the leadership lock name follows), §5 (act-differently rule the InvalidSpec variant follows), §7 (the guarantee-row transfer).
  • ADR-010 — the queue semantics depth this ADR's fire semantics compose with; the maintenance recipe (who sweeps) is recorded there with the collapse machinery.
  • Honker reference checkout /workspace/honker @ f4e53c6 — scheduler surface (packages/honker-rs/src/lib.rs, honker-core/src/honker_ops.rs), @every grammar and local-TZ brittleness (honker-core/src/cron.rs), tick/leader mechanics.
  • ADR-011 — the SQLite-side schedule storage this §5 describes is the forked substrate's; cron machinery not ported.
  • core-contract.md — the spec carrying this surface.