Files
alkstore/docs/architecture/queues.md
T

16 KiB
Raw Blame History

status, last_updated
status last_updated
draft 2026-10-05

Queues, scheduler, outbox — semantics depth

The queue family's posture is decided (re-derived on both engines; ADR-004 and ADR-005), the hard driver-coupled properties are POC-pinned (transactional enqueue, exactly-once claim under concurrency), and the semantics depth is designed — OQ-05 and OQ-09 resolved by ADR-010 and ADR-009 (2026-10-05). This document states the resulting semantics and the constraints they were made under; the ADRs carry the WHY.

Decided constraints (inherited, not re-opened)

  • Mechanisms: durable at-least-once queues; the outbox is a helper over queues; the scheduler collapses into queues + schedule() + run_schedules (ADR-002, ADR-009).
  • Transactional enqueue: commit-atomic with the caller's business write on both engines (ADR-007); rollback drops the job row with no ghosts. For the outbox's backing queue the tx path is outbox_enqueue_tx (the derived reserved names are unreachable by enqueue_tx — ADR-014).
  • Exactly-once claim: FOR UPDATE SKIP LOCKED (pg, POC-pinned under 4×4 concurrency) / the forked substrate's single-statement claim (SQLite, re-derived on ADR-010 §3a — ADR-011). At-least-once work; exactly-once processing needs the ack + visibility model.
  • Wake-driven consumption: the pg engine's default is LISTEN-driven claim with a re-poll safety net; SQLite's is watcher wake with per-subscription fanout (ADR-006).
  • Job options (the contract v1 skeleton, ADR-008 §1, from the honker-rs surface): delay, priority, max_attempts, expires, run_at at enqueue; ack / retry / fail / heartbeat on the job handle.
  • Cut-flag context: result storage is ADR-002's cut-flag row; ADR-010 §7 reconsidered it as part of this design and left it cut (delete-on-ack keeps completed jobs out of storage; the pg-boss completed-row/output model is that feature under another name). Re-entry stays gated on a consumer-inventory row.

Job lifecycle (ADR-010 §1–§3a)

        enqueue
           │
           ▼
      ┌─────────┐    claim (attempts += 1)     ┌───────────────────┐
      │ pending │ ───────────────────────────► │    processing     │
      └─────────┘ ◄─────────────────────────── └───────────────────┘
           │           retry(err, d)          │ deadline = job-stamped
           │                                  │ visibility timeout;
      cancel deletes the row in EITHER       │ heartbeat(extend) resets it
      state (pending or processing) —        ├─────────────────────────
      not an interrupt; the holder's         │ visibility lapse:
      next ack/heartbeat returns false       │ no transition fires — the
           │                                 │ row is just reclaimable
           ▼                                 │ (a reclaim eats an attempt)
      ┌────────┐   fail(err) / budget         │
      │  dead  │ ◄────── exhausted ───────────┘
      └────────┘
            │        ack deletes the row
            │        (from processing, deadline unexpired)

      Handle-op validity (ADR-010 §2, uniform): ack/heartbeat/retry/
      fail succeed only while the row is processing and the claim's
      deadline is unexpired — lapse refuses all handle ops; a worker
      whose deadline lapsed before reclaim may still complete its
      work, but its ack does not land (the job reprocesses at
      reclaim — at-least-once).
           │
           ▼
      (gone)

      expires lapse (any state) — sweep_expired(q) moves the row to
      dead with last_error='expired'; also enforces dead-letter
      retention. ack/cancel leave no row; get_job sees dead rows.
  • States: pending → processing → dead (+ absence: ack/cancel delete the row). max_attempts frozen at enqueue; attempts counts every claim — fresh or reclaim.
  • Visibility: each claim sets the deadline from the job's stamped visibility_timeout_s (default 300 s, stamped per ADR-010 §3a). heartbeat(extend) is renewal — an absolute reset from now (the new full deadline, not additive); a late heartbeat is refused (never steals the job back from a reclaimer). Deadline lapse refuses all handle ops uniformly (the validity predicate, ADR-010 §2). Between deadline lapse and another worker's reclaim, the original worker may still complete its work — but its ack refuses and the row reprocesses at reclaim — the dual-execution window is the contract's honest at-least-once posture; downstream idempotence is the consumer's job.
  • Reclaim consumes an attempt: a handler that forgets to heartbeat looks like a repeatedly-failing job and dead-letters on budget exhaustion. Stated in the contract text (ADR-010 §2), not discovered in production.
  • Retry: retry(err, None) — the engine computes the queue's curve; retry(err, Some(d)) overrides. fail(err) = immediate dead-letter. Ordering within a queue: priority DESC, ready-time ASC, enqueue order (FIFO under equal priority).
  • get_job sees dead jobs — state, payload, attempts, timestamps, and (dead only) last_error, died_at; post-mortem diagnosis by API, not by SQL.

Retry / backoff (ADR-010 §3, §3a)

  • The curve: equal-jitter exponential from the job's stamped backoff_base_s (default 5 s): delay ∈ [base·2^(a−1)/2, base·2^(a−1)] for attempt a — the range is the definition (not AWS-canonical full jitter); capped at 1 hour (pinned contract constant, no knob). The formula is what both engines compute identically; the range's herding rationale lives in the ADR.
  • Where configured: queue-level QueueOpts { visibility_timeout_s, max_attempts, backoff_base_s, dead_letter_retention_s } (honker-parity defaults: 300 s / 3 / 5 s / none) — stamped onto each job row at enqueue (§3a), never read from a live registry.

Dead-letter (ADR-010 §4)

  • Move, not flag: dead rows physically move to engine-owned dead storage (last_error, died_at) — never scanned by the claim path. Triggers: explicit fail, retry at budget, exhaustion-by-reclaim, expires lapses (swept).
  • Retention: forever by default; dead_letter_retention_s (queue opt) makes sweep_expired enforce a per-queue dead TTL. Honker's no-retention posture kept as default; the mechanism is ours.
  • No redrive API in v1: the recipe is get_job + fresh enqueue. No consumer names an API shape; re-entry needs an OQ.

Sweep / maintenance (ADR-010 §5–§6)

  • sweep_expired(queue) moves every past-expires_at row (pending and processing) to dead, and enforces dead-letter retention — the no-stranded-rows property (ADR-010 §5): a job with an expires deadline is eventually in exactly one of pending/processing/dead, never stuck unreachable. (SQLite-side realization resolved with OQ-06 — the fork fires (ADR-011), the zombie fix lands in owned code.)
  • No ambient sweeper: the engine ships no background maintenance, no default cadence (the no-ambient-timers posture). Correctness of transitions is the engine's; timeliness is the consumer's.
  • The "who sweeps" recipe is the scheduler collapse's payoff (ADR-009): store.schedule("maintenance", "@every 300s", mq, …) + run_schedules(stop) (leader-elected) + a worker handler calling sweep_expired per queue. One schedule row, one worker — the pattern every consumer would have hand-rolled.
  • Multi-process sweeps need no leader lock on a bare sweep_expired call (single-statement atomic per queue); the cadence coordination comes from the scheduler machinery above. The SQLite notifications table's hygiene is engine-internal (not a consumer chore; prune_notifications* stay out of the contract, resolving ADR-008 §8's disposition). The stream log's bounded-growth answer follows the same posture but is consumer-side, because consumers interact with stream rows directly (ADR-015 §5): stream(name).trim_to(offset) — invoked deliberately, cadenced by the collapse recipe like every sweep; no engine-default retention, the replay-forever default documented with the in-contract tool to bound it.

Scheduler (ADR-009)

Collapsed into queues: no Scheduler handle, no schedule objects.

  • Surface: schedule(name, spec, queue, payload, opts) (upsert by name), unschedule(name) -> bool, run_schedules(stop) -> Result<()>. Update = re-register; pause = unregister + re-register.
  • Spec grammar v1: @every <n><unit> only (s|m|h|d). Cron strings rejected with InvalidSpec { spec } (grammar rationale in the ADR; extension path: a consumer-inventory row first).
  • run_schedules is opt-in: schedules never fire without a runner (the no-ambient-timers posture). The runner: leadership lock → {renew (on loss, return before ticking — Err(LeadershipLost), distinct from clean Ok(()) stop) → fire due boundaries → sleep}. Leadership lock: reserved name __alkstore_scheduler (engine-derived name under the reserved prefix, ADR-008 §4).
  • Boundary guarantee (ADR-009 §4): at-least-once per elapsed boundary while a leader runs; fire + boundary-advance commit atomically under a row lock on the schedule row (crash mid-tick refires the boundary; a rogue ticker is defeated by the row-locked advance, not just by the leadership lock); bounded catch-up — missed boundaries replay one-by-one up to a fixed contract constant cap of 64 per task per tick, beyond which remaining boundaries skip forward. Honest loss mode, documented.
  • Each boundary fire enqueues ordinary work: payload into the named queue with ScheduleOpts { priority, max_attempts, expires } — stamped onto the enqueued job per ADR-010 §3a; the queue mechanism's guarantees apply, no separate delivery machinery.
  • Schedule storage is engine-internal (SQLite: the forked substrate's __alkstore_scheduler_tasks table; Postgres: the engine-owned schema) — not a consumer namespace; queues are names, not registered objects, so a schedule may target a queue nothing has claimed yet. Schedule names validate like queue names (non-empty, reserved-prefix rejected).

Namespaces / schema layout (ADR-010 §8)

  • Postgres: all engine-owned tables (job, dead, stream, offsets, schedules, internal) in one engine-owned schema (default alkstore, per-engine option) — co-tenancy-safe with consumer tables (alkblobs ADR-008 precedent; the reserved name namespace governs name-column values inside). Queues are rows in one job table — no per-queue tables/schemas; pg-boss's partition-per-queue opt-in is not inherited (one table + partial indexes matches honker's proven shape and keeps queue(name) a name, not a DDL operation).
  • SQLite: the forked substrate's __alkstore_* table family — storage-internal (ADR-008 §4); the fork (ADR-011) re-owns the names from upstream (_honker_*); nothing consumer-visible changes.
  • Partial indexes pinned identically on both engines: partial indexes matching the claim hot path; dead rows outside it; single clock source (second-precision timestamps); savepoint-guarded multi-statement moves.

Reference material

  • honker's queue design — the SQLite-side lineage design reference (/workspace/honker; ownership resolved by ADR-011 — forked per ADR-012, queue ops re-derived on this design). Full read: docs/research/reference-honker-machinery.md (schema, claim/visibility/retry/dead-letter mechanics with file/line cites; the defect list §8 fed OQ-06 and is the fork's fix/inheritance register), plus docs/research/quality-read-honker-core.md (the dependency-gate read).
  • pg-boss family — /workspace/pgboss-rs @ 98f7d9e (design reference only, ADR-005). Full read: docs/research/reference-pgboss-rs-semantics.md (states, backoff-with-jitter formula, dead-letter-as-queue, maintenance postures; the node-upstream deltas §8).
  • POC ground — poc-pg-posture-findings.md (the minimal queue table + SKIP LOCKED claim measured; the property suite green); poc-sqlite-posture-findings.md (the SQLite twin).
  • Consumers — alkfs OQ-FS-14 (sync outbox), alkblobs ops-surface/gc docs (maintenance cadences), the family-wide "who sweeps" punt (inventory §scheduler — answered by the collapse recipe above).

Design Decisions

ADR Decision Summary
002 Feature scope queues + outbox in; result storage cut-flag (stands, re-checked)
004 pg driver queue machinery re-derived, pg-boss as reference
005 Ownership design-reference postures for the queue family
006 Wake contract wake-driven consumption, guarantee table
007 Tx seam enqueue_tx commit-atomicity
008 Contract v1 queue skeleton + EnqueueOpts pinned; depth was this doc's design space
009 Scheduler collapse queues + schedule() + runner; @every-only; boundary guarantee row
010 Queue depth visibility/heartbeat rules, backoff curve, dead-letter, no-stranded-rows sweep, schema layout
011 Substrate fork SQLite-side queue ops re-derived in owned code; __alkstore_* naming
012 Fork design contract-blind substrate boundary; fidelity posture; port deltas
014 Outbox tx enqueue outbox_enqueue_tx on TxHandle; derived backing queue reached only through the outbox surface

Open Questions

Open questions are tracked in open-questions.md. Key questions affecting this document:

  • OQ-06: honker-core quality read — resolved (2026-10-05, ADR-011): the zombie fix and dead-row visibility — this document's two named SQLite-side gaps — land in owned code via the fork.

Resolved on this document's surface: OQ-05 and OQ-09 (2026-10-05, ADR-010 / ADR-009).