Files
alkstore/docs/architecture/decisions/010-queue-semantics-depth.md

21 KiB
Raw Permalink Blame History

ADR-010: Queue semantics depth — visibility, retry/backoff, dead-letter, sweep, layout

Status

Accepted (2026-10-05, Phase 1 — OQ-05's resolution; composes with ADR-009)

Context

Contract v1 pinned the queue skeleton (ADR-008 §1) and explicitly left the semantics depth to OQ-05: retry policy shape, visibility/renewal mechanics, dead-letter move-vs-flag and retention, sweep/maintenance design, the result-storage cut-flag's disposition, and queue/stream/lock table layout (the co-tenancy collision surface, queues.md).

The evidence base: honker's queue machinery (the SQLite-side incumbent — /workspace/honker @ f4e53c6, whose functions the engine rides directly) and the pg-boss family (the pg-side design reference — /workspace/pgboss-rs @ 98f7d9e standing in for node pg-boss v10; design-reference only per ADR-005). Both models were read in full for this resolution (reference notes: docs/research/reference-honker-machinery.md, docs/research/reference-pgboss-rs-semantics.md). The hard driver-coupled properties (transactional enqueue, exactly-once claim) are already POC-pinned on both engines.

The two references disagree on several semantics points, and honker has real gaps (a dead-letter table get_job can't see; expired processing rows no path can reach — the "zombie" hole; no dead-row retention mechanics; no backoff in the core at all). OQ-05 is the place to fix what re-derivation should fix and inheriting should inherit.

Decision

1. The job state machine: three states, delete-on-ack

Contract job states: pending → processing → dead (+ absence: ack and cancel delete the row).

  • Honker's model, adopted whole — one live table, move-to-dead (physical row move, savepoint-guarded), not the pg-boss seven-state enum. max_attempts is frozen at enqueue; attempts counts every claim.
  • ack deletes (honker parity): completed work leaves no row. The completed-job-with-output model (pg-boss's completed state + output column) is not adopted — it is result storage under another name, and result storage is a cut-flag row (ADR-002); see §7.
  • get_job sees dead jobs. The returned job carries state, payload, attempts, priority, and the timestamps created_at, run_at, claimed_at, expires_at — plus, dead-only, last_error and died_at (no job = None). (Struct disposition 2026-10-07 per ADR-021 §2: claimed_at: Option<i64> — None pre-claim, set on every claim — now stands in ADR-019 §3's pinned struct, which is the field list of record; ADR-019's original omission is corrected there.) This deliberately narrows honker's surface (honker's get_job reads only live rows — post-mortem diagnosis was SQL-only); diagnosis-by-API is the honest fix, and it is what makes move-to-dead (vs flag-in-place) inspectable without raw SQL.
  • cancel is unconditional (honker parity): it deletes the row in either pending or processing, regardless of which worker holds it. Not an interrupt — the holder's next ack/heartbeat returns false, the same shape as expiry. ack_batch is the batch form of ack (count returned; non-claimed ids silently not counted). (Acknowledged 2026-10-07, third review round follow-through: the batch form applies §2's full validity predicate — processing state and unexpired claim deadline — per id, not claim-state alone; an id whose deadline lapsed is silently not counted, identically to an acked-elsewhere or cancelled id. Per-id outcomes are independent; partial success is ordinary.)
  • Claim ordering (same queue): priority DESC, then ready-time (run_at) ascending, then enqueue order — FIFO under equal priority, both references agree; pinned.

2. Visibility timeout and heartbeat: explicit renewal, late-heartbeat refusal

One answer, contract-pinned (honker's design; pg-boss's no-renewal/expire-sweep model is not inherited):

  • visibility_timeout_s is a per-job stamped value (see §3a's resolution rule — default 300 s, honker parity); each claim sets the row's deadline from the job's stamp.
  • heartbeat(extend) is renewal — an absolute reset of the claim deadline from now. There is no progress-signal meaning; a consumer signaling progress uses its own channels. Renewal cadence is the consumer's obligation: handlers that might outlive the visibility timeout must heartbeat inside it.
  • A late heartbeat is refused (deadline already passed ⇒ returns false): it can never steal the job back from a reclaimer. The consequence is the honest at-least-once window — between deadline lapse and another worker's reclaim, the original worker may still complete: dual-execution is possible and idempotence is the consumer's job (matches the locks row's silent-expiry honesty, ADR-008 §7).
  • A reclaim consumes an attempt (honker's counting): a claim is an attempt, whether fresh or a visibility reclaim. Documented explicitly as the contract's counting rule — the footgun is stated, not discovered: a handler that forgets to heartbeat looks like a repeatedly-failing job and dead-letters on budget exhaustion.
  • Deadline lapse itself is lazy: an expired-claim row becomes claimable by ordinary claim (no transition fires on lapse); there is no separate reaper for expired claims, only the budget rules below and the no-stranded-rows sweep (§5).
  • Job-handle op validity predicate (the D-12 disposition, stated uniformly): every handle op (ack/heartbeat/retry/fail) succeeds only when the row is in processing and the caller's claim deadline has not lapsed — ADR-008 §5's false-case list ("expired, acked elsewhere, cancelled") is the same rule one predicate-wide: deadline lapse refuses all ops, not just heartbeat. The dual-execution window stays honest under this: the original worker between lapse and reclaim may still complete its work (the side effects happen), but its ack will not land and the row is reprocessed at reclaim — at-least-once, as documented. The reclaim wins atomically when it races the ack (one statement, same predicate). This state check is engine-pinned in both engines' implementations (upstream's missing-check defect class D-12 is not inherited).

3. Retry and backoff: explicit delay or the queue's curve

  • Job handle: ack(), retry(err, delay), fail(err), and heartbeat(extend) — the v1 skeleton (ADR-008 §1), with heartbeat's meaning pinned by §2 and renew remaining the lock-side term (ADR-008 §8's split).
  • retry(err, None) — the deferred delay — is computed by the engine from the queue's curve; retry(err, Some(d)) overrides it. Honker's core takes a caller-supplied delay with no curve; pg-boss computes in-engine from a base. The contract pins the curve so both engines compute identically.
  • The curve: equal-jitter exponential, base * 2^(attempt-1) with uniform jitter over the lower half of the doubling range — delay ∈ [base·2^(a−1)/2, base·2^(a−1)] — capped at 1 hour (a pinned contract constant; no backoff_max knob — no consumer names one; the explicit-delay override is the escape hatch for bespoke policies). The range, not a jitter-label, is the definition: this is not AWS-canonical full jitter (uniform over [0, cap]), and the formula above is what both engines compute identically. The attempt index a is the value of attempts on the row at the moment the retry is scheduled (post-claim count, including reclaims). Jitter prevents sync-failure herding across a fleet; honker's wrappers have no jitter (a spread-out fleet of failing workers retries in lockstep) and the pg-boss family's jittered exponential is the design adopted.
  • QueueOpts { visibility_timeout_s, max_attempts, backoff_base_s, dead_letter_retention_s } — queue-level defaults (respectively 300 s / 3 / 5 s / none, honker parities except retention which is ours); resolved and stamped per §3a. backoff_base_s feeds the curve. This is the QueueOpts depth contract v1 deliberately withheld (ADR-008 §1) — pinned here as the extension path. Outbox: outbox(name) carries no opts of its own; its backing queue uses a derived QueueOpts set as engine defaults (honker's outbox parities: visibility 60 s, max_attempts 5, backoff base 5 s) — engine-documented, not separately configurable in v1.
  • Explicit fail(err) = immediate dead-letter (honker parity; it is the "stop retrying this" operator). retry at exhausted budget = dead-letter. Exhaustion via reclaim = dead-letter. (The pre-claim sweep of already-exhausted reclaimable rows — a laziness optimization in both upstream lineages — is engine-side and optional; SQLite rides the fork's re-derivation (ADR-011), pg re-derives it (engine-postgres.md).)

3a. QueueOpts resolution: stamped at enqueue, per job

Queues are names, not registered objects — so queue-level QueueOpts need a defined attachment point, and the contract pins one:

  • Every job row carries the resolved opts stamped at enqueue: enqueue resolves EnqueueOpts.max_attempts over the queue handle's max_attempts (the one per-job override the v1 skeleton names), and stamps the result — visibility_timeout_s, backoff_base_s, dead_letter_retention_s, max_attempts — onto the job row. Claims, heartbeats, retries, and sweeps read the job's own stamps, never a live registry.
  • Why stamped: uniform across engines (no per-queue config state to invent on either side — SQLite needs no new table beyond honker's columns, Postgres is job-table columns), and it makes a job's behavior immutable and inspectable (get_job returns the stamps with the rest of the row) no matter which process's handle enqueued it. Two processes opening handles with different QueueOpts for the same queue name do not fight — each enqueue stamps its own opts; the queue name is a routing key, not a configuration owner.
  • The consequence, stated: opts changes apply to future enqueues only; in-flight jobs keep their stamps. This is the honest, durable-row semantics — the same reason max_attempts was already frozen at enqueue in the references.
  • SQLite realization note: honker's claim_batch takes one uniform timeout_s per call, while per-job stamps make deadlines row-local — bridging that gap (post-claim re-stamp, per-row claiming, or another shape) was implementation work over honker's function surface and rode the same OQ-06 assessment as §5's zombie fix. *(Resolved 2026-10-05: the fork fired (ADR-011); the claim statement is re-derived with per-row visibility from the stamps, in owned contract-blind substrate code (ADR-012 §2).) The contract (deadline from the job's stamp) is engine-independent either way.

4. Dead-letter: move, retention = never-expire by default, no redrive API

Move-to-table (both references converge; flag-in-place is not adopted — a dead row must not occupy the claim path's indexes): dead rows physically move to engine-owned dead storage with last_error and died_at. Triggers: explicit fail, retry at budget, exhaustion-by-reclaim, and expires lapses (§5).

  • Retention: dead rows live until explicitly removed by default (honker's posture). QueueOpts.dead_letter_retention_s (default None = forever) makes the sweep enforce per-queue dead TTL (§5) — the retention mechanism is ours (honker has none); the default stays non-magical. last_error source strings are contract-pinned for the engine-triggered cases: "max attempts exceeded" (budget exhaustion, any path) and "expired" (sweep-by-expires); explicit fail(err)/retry(err, …) carry the caller's error string.
  • No redrive/requeue API in v1. The recipe is pinned instead: get_job (now sees dead rows, §1) + fresh enqueue — both primitives already on the surface. No consumer names an API shape for redrive; when one does, it gets a contract extension with its own OQ.
  • The engine-owned dead table/rows are storage-internal — not a consumer namespace (no reserved-prefix question; they're not name-addressable).

5. sweep_expired: the no-stranded-rows property

Contract semantics for sweep_expired (the v1 skeleton's maintenance entry point; the queue-scoped handle form — name on the handle — is pinned by ADR-019 §1, 2026-10-06):

  • It moves to dead (last_error = 'expired') every row of the queue past its expires_at in any state — pending rows (honker parity) and processing rows whose job-level expiry passed. This is the no-stranded-rows property: an enqueued job with an expires deadline is eventually in exactly one of pending, processing, or dead — never stuck unreachable. Honker fails this property: an expired processing row whose worker died is unreachable by claim (predicate requires future expiry), by pre-claim dead-lettering (same), and by sweep_expired (pending only) — the zombie hole, a real defect found in the reference read (D-6 in the quality read's register). This ADR fixes it at the contract level; the Postgres engine enforces it directly; the SQLite engine's realization was implementation work contingent on OQ-06 — (resolved 2026-10-05: the fork fired (ADR-011) and the both-states sweep lands in owned code.)
  • When dead_letter_retention_s is set, the same sweep also deletes dead rows past their retention (the only sweeper dead rows ever have — nothing runs without a caller).
  • Sweep is single-statement atomic per queue (SQLite: writer serialization; Postgres: row-locked UPDATE … RETURNING); no leader lock is required for a bare sweep_expired call — the multi-process recipe gets its coordination from §6's machinery, not from the sweep itself.

6. Sweep/maintenance cadence: no ambient sweeper; the collapse recipe

  • The engine ships no background sweeper, no maintenance thread, no default cadence — the family's no-ambient-timers posture (alkblobs ADR-005) and honker's own verified posture (nothing in honker-core auto-runs; the only implicit maintenance is laziness in claim paths). The engine's responsibility is confined to correctness-of-transition (§1–§5); timeliness is the consumer's.
  • The pinned recipe for the family-wide "who sweeps" problem is the collapse machinery (ADR-009): store.schedule("maintenance", "@every 300s", maintenance_queue, …)
    • run_schedules (leader-elected) + a worker whose handler calls sweep_expired per queue. One schedule row, one worker — the pattern every consumer was going to hand-roll, answered twice in the contract instead of N times above it.
  • Dead-letter and notifications-table hygiene (SQLite engine): prune_notifications/prune_notifications_keep_latest remain out of the contract (ADR-008 §8's disposition, resolved here by disposition): the notifications table is the SQLite wake mechanism's transport detail — consumers interact with wakes, not rows. Its hygiene is engine-internal: an at-attach pruning cap (engine opts) — the mechanism the fork scope realizes (ADR-011); no cadence, ambient or engine-side, contradicts this section's no-ambient-sweeper bullet. The fix for upstream's unbounded notifications growth must not become a consumer's chore. (Annotated 2026-10-05: the earlier "engine-side maintenance cadence" wording was wrong — nothing in the fork scope realizes a cadence; the at-attach cap is the pin.)

7. Result storage: cut-flag stands

Reconsidered per OQ-05's mandate and left cut: no consumer row names outcome-query-by-id; the §1 delete-on-ack decision keeps completed jobs out of storage entirely (the pg-boss output/completed -row model is that feature under another name — adopting it would silent-include the cut flag); the documented workaround (a job writes its own result record atomically with ack) rides the tx seam (ADR-007) cleanly: business row + result write + ack-equivalent in one transaction. Re-entry stays gated on a consumer-inventory row.

8. Table layout: engine-owned schemas/tables; queues are rows, not tables

  • Postgres: all engine-owned tables (job, dead, stream log, offsets, schedule rows, internal state) live in one engine-owned PostgreSQL schema — default alkstore, overridable per-engine option — co-tenanted safely with consumer tables (alkblobs ADR-008's co-tenancy precedent; schema scoping is the collision answer, and the reserved name namespace (ADR-008 §4) governs the name column values inside it). Queue names become row values in one job table — no per-queue tables, no per-queue schemas (pg-boss's partition-per-queue opt-in is not inherited; at this scale one table + partial indexes matches honker's proven shape and keeps queue(name) a name, not a DDL operation — queue creation is not registry-gated, per ADR-009 §5). Per-queue config storage is likewise stamped per §3a.
  • SQLite: honker's _honker_* table family in the caller's database file — storage-internal per ADR-008 §4. (Annotated 2026-10-05: the expectation that job-stamped opts would ride honker's existing columns with no schema change was falsified by the quality read — no stamp columns exist in the upstream schema (D-7). The family's fate is resolved: the fork (ADR-011) re-owns the whole family as __alkstore_*, and the stamp columns + claimed_at are added by the re-derivation on contract v1 (ADR-012 §5). Nothing consumer-visible changes.)

9. Error taxonomy: no delta

The depth adds no new error variants — checked against ADR-008 §5's act-differently rule: no work claimed is a value; deadline operations are booleans; dead-letter is observable state (get_job), not an error thrown; queues are unregistered names, so a payload enqueued to a queue nobody claims yet is not an error (it's durable work awaiting a worker). InvalidSpec { spec } was the scheduler's one variant (ADR-009 §6) and is the only taxonomy delta this track produces.

Consequences

Positive

  • The queue mechanism is fully specified end-to-end: both references' semantics decisions are made once, pinned identically on both engines, with the references' gaps (zombie rows, dead-job invisibility, no retention mechanics, no jitter) fixed at the contract level rather than inherited.
  • The counting/heartbeat rules and the opts-stamping resolution are stated as contract text — the honest at-least-once story (dual-execution window, reclaim-eats-budget) is documented consumer obligation, not per-engine surprise, and per-queue configuration never fights across processes.
  • One engine-owned pg schema keeps co-tenancy mechanical.

Negative

  • Three-state + delete-on-ack means no completion history and no built-in outcome storage — by design (the cut-flag stands, §7); consumers needing audit trails build them on the tx seam.
  • Honker's zombie fix, dead-row get_job visibility, and per-job visibility stamps over the uniform-claim-timeout function surface required either an over-machinery complement or a fork on the SQLite side — resolved by the fork (ADR-011, designed in ADR-012).
  • The backoff curve's 1-hour cap and jitter formula are pinned constants — a consumer needing a different policy uses retry(err, Some(d)) per attempt (correct, but manual).

References

  • OQ-05 (docs/architecture/open-questions.md) — this ADR's resolution (with OQ-09, resolved by ADR-009).
  • docs/research/reference-honker-machinery.md and docs/research/reference-pgboss-rs-semantics.md — the full reference reads (schemas, state machines, defect list with file/line cites) this resolution was made over.
  • ADR-002 — scope (cut-flags §7 stands), ADR-005 — the design-reference postures the two reads serve, ADR-006 — the queues guarantee row this ADR's §2/§3 make precise, ADR-007 — the tx seam §7's workaround rides, ADR-008 — the v1 skeleton this ADR extends, §5's rule §9 applies, §8's dispositions §6 resolves.
  • ADR-009 — the collapse machinery §6's recipe composes with.
  • OQ-06 — the SQLite-side fork candidates (§5, §3a realization note) — resolved by the fork (ADR-011, designed in ADR-012).
  • queues.md, core-contract.md, engine specs — the specs carrying this depth.