docs: resolve OQ-05 + OQ-09 — queue semantics depth (ADR-010) and scheduler collapse (ADR-009)
ADR-009: scheduler collapses into queues — schedule()/unschedule()/ run_schedules (opt-in, no ambient timers), @every-only v1 grammar (dissolves honker's local-TZ cron brittleness), boundary guarantee row (at-least-once per boundary, fixed 64-cap catch-up with skip-forward, row-locked fire tx as engine-generic no-double-fire floor), __alkstore_scheduler leadership lock, InvalidSpec + LeadershipLost taxonomy additions. ADR-010: queue depth pinned engine-uniformly — three-state machine (pending/processing/dead) with delete-on-ack, get_job sees dead rows, heartbeat = renewal with late-heartbeat refusal, reclaim-consumes- an-attempt stated as contract text, equal-jitter exponential backoff (range definitionally pinned, 1 h cap), QueueOpts stamped onto job rows at enqueue (no per-queue registry), move-to-dead dead-letter with retention-sweep support and no redrive API, sweep_expired carries the no-stranded-rows property (fixes honker's expired- processing zombie hole — SQLite-side realization rides OQ-06 as a concrete fork candidate), one engine-owned pg schema, queues are rows not tables, result-storage cut-flag stands. Also: full honker-machinery and pgboss-rs reference reads persisted (docs/research/reference-*.md — the honker defect list pre-stages the OQ-06 quality read), queues.md rewritten from design-space frame to resolved-depth spec, core-contract/engines/README/overview/deployment/ ADR-002 propagated.
This commit is contained in:
1 parent
7ad8ac56bc
commit
79a135c934
13 files changed
+1863
-198
No files matched your search
@@ -43,7 +43,12 @@ has recorded the need):
|
||||
- **scheduler** — cron/`@every` enqueueing into named queues,
|
||||
leader-elected. Documented-thin; the family-wide
|
||||
"who sweeps/renews/reaps" need. May collapse into queues if design
|
||||
shows scheduler = queues + tick (tracked in OQ-09).
|
||||
shows scheduler = queues + tick (tracked in OQ-09). *(Narrowed
|
||||
2026-10-05 by [ADR-009](009-scheduler-collapse.md): collapsed to
|
||||
queues + `schedule()`/`run_schedules`, spec grammar `@every`-only —
|
||||
the "cron" label reflected honker's menu, not any inventory row's
|
||||
body; no consumer names wall-clock cron — grammar pinned by
|
||||
[ADR-009](009-scheduler-collapse.md) §2.)*
|
||||
|
||||
**Cut-flag** (no named consumer; not silently included):
|
||||
|
||||
|
||||
@@ -0,0 +1,244 @@
|
||||
# ADR-009: Scheduler collapses into queues — `schedule()` + runner, `@every`-only v1
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-10-05, Phase 1 — OQ-09's resolution; resolves the
|
||||
scheduler row transferred by [ADR-008](008-contract-v1-pinning.md) §7)
|
||||
|
||||
## Context
|
||||
|
||||
[ADR-002](002-feature-scope.md) took the scheduler in as
|
||||
"documented-thin": the family-wide "who sweeps / renews / reaps, and
|
||||
when" problem is real, but only just begins to have consumers lean on
|
||||
it. OQ-09's decision rule (inherited from
|
||||
[queues.md](../queues.md)'s framing of it): decide against the
|
||||
inventory rows, not against the whole honker menu — the scheduler is a
|
||||
first-class mechanism if consumers need inspectable/pausable schedule
|
||||
*objects* (`add/pause/resume/update/list/remove`), and otherwise it
|
||||
collapses into queues + a `schedule()` call with the tick machinery
|
||||
absorbed.
|
||||
|
||||
The evidence for the decision:
|
||||
|
||||
- **alkblobs (sweep cadence)** — [documented:
|
||||
consumer-inventory.md](../../research/consumer-inventory.md) — needs
|
||||
"run the sweep every N ticks with leader election across a fleet."
|
||||
An interval, not a schedule object.
|
||||
- **alkfs (orphan reaping)** — [recalled: OQ-FS-07] — a periodic
|
||||
orphan-cleanup pass. An interval.
|
||||
- No consumer document (pinned, documented, or operator-authority)
|
||||
names any of: wall-clock cron expressions, per-schedule timezones,
|
||||
pausing/resuming schedules without unregistering, listing schedule
|
||||
objects, or updating a schedule's expression in place.
|
||||
|
||||
So the collapse question itself answers cleanly. Two secondary
|
||||
questions ride along and this ADR pins them with the collapse:
|
||||
|
||||
- **Which schedule specs does v1 accept?** Honker's parser accepts
|
||||
5-field cron, seconds-field cron, and `@every <n><unit>` — but its
|
||||
cron boundary arithmetic is hard-wired to *host local time*
|
||||
(`chrono::Local`, honker cron.rs), which is brittleness a contract
|
||||
should not inherit: a server TZ change silently re-times every
|
||||
stored schedule, and the two engines would diverge if one inherited
|
||||
local time and the other pinned UTC.
|
||||
- **What runs the tick?** Every consumer punting "the embedder owns
|
||||
cadence" is exactly the hand-rolling the substrate exists to kill —
|
||||
but an *ambient* runner violates the family's no-ambient-timers
|
||||
posture (alkblobs ADR-005). The line between them is opt-in.
|
||||
|
||||
## Decision
|
||||
|
||||
### 1. The scheduler collapses into queues
|
||||
|
||||
The contract surface is two registration methods on `Store` plus one
|
||||
runner; there is no `Scheduler` handle, no schedule objects, no
|
||||
inspection surface:
|
||||
|
||||
```text
|
||||
store.schedule(name, spec, queue, payload, opts) -> Result<Schedule>
|
||||
store.unschedule(name) -> bool
|
||||
store.run_schedules(stop) -> Result<()> // runs until `stop`
|
||||
```
|
||||
|
||||
- `schedule` upserts by `name` (re-register replaces the whole row) —
|
||||
honker's register semantics, which makes update = re-register and
|
||||
pause = unregister + re-register. Both are honest v1 substitutes;
|
||||
nothing in the inventory needs finer control.
|
||||
- `Schedule` is a value carrying the registered row (name, spec, queue,
|
||||
opts) — read-back for confirmation, not a mechanism handle.
|
||||
- **Name validation**: schedule names follow the queue-name rules —
|
||||
non-empty (`InvalidName`) and reserved-prefix rejection
|
||||
(`ReservedName`) at the entry point; `schedule()`/`unschedule()` are
|
||||
name-bearing entry points in the
|
||||
[ADR-008](008-contract-v1-pinning.md) §4 sense (their consumers-side
|
||||
obligation extends the entry-point list: a reserved-prefix schedule
|
||||
name could collide with the machinery's derived names). No charset
|
||||
rule beyond that in v1.
|
||||
- `run_schedules` is the tick: acquire the leadership lock, loop
|
||||
{renew lock — on loss, *return before ticking*; fire every due
|
||||
boundary; sleep until the next due boundary or `stop`}. Losing
|
||||
leadership ends the runner — the consumer's recipe is to respawn it
|
||||
(documented); the lock loss always precedes any stolen fire, so
|
||||
respawn never double-fires a boundary.
|
||||
- **Stop and return semantics**: `stop` is a cancellation token
|
||||
(a cloneable handle whose flip ends the sleep early — the exact
|
||||
type shapes at implementation, engine-crate docs). A leadership
|
||||
loss returns `Err(LeadershipLost)` — a distinct, matchable outcome
|
||||
from the clean `Ok(())` a stop produces; the respawn recipe matches
|
||||
on it. (One taxonomy note: `LeadershipLost` is a *return-shape*
|
||||
variant of the runner, pinned by the same act-differently rule as
|
||||
ADR-008 §5 — the caller acts on it by respawning.)
|
||||
- Schedules do not fire until a consumer runs `run_schedules` — the
|
||||
opt-in is explicit, the no-ambient-timers posture holds,
|
||||
[ADR-002](002-feature-scope.md)'s scheduler evidence is served, and
|
||||
consumers who never schedule never pay for the runner.
|
||||
|
||||
### 2. v1 spec grammar: `@every <n><unit>` only
|
||||
|
||||
`spec` accepts exactly `@every <n><unit>` with `n` a positive integer
|
||||
and `<unit>` in `s | m | h | d` (honker's `@every` grammar,
|
||||
honker cron.rs). Everything else — 5-field cron, seconds-field cron —
|
||||
parses as `InvalidSpec { spec }` and is rejected.
|
||||
|
||||
- **Why not cron strings**: no consumer row names one (checked against
|
||||
the inventory, not against honker's menu), and accepting them forces
|
||||
a timezone answer now. `@every` durations are TZ-free — the TZ
|
||||
question dissolves instead of being answered.
|
||||
- **Extension path** (when a consumer names wall-clock cron via a
|
||||
consumer-inventory row): cron-string support with an explicit TZ
|
||||
decision attached (UTC-pinned contract or a per-schedule TZ
|
||||
parameter — decided by that extension, not assumed here).
|
||||
- This keeps the SQLite engine off honker's cron boundary machinery
|
||||
entirely (`@every` advance is `next = last + interval`) — the
|
||||
local-TZ brittleness is not inherited, and a dependency requirement
|
||||
is not created. *(Consumer-inventory rows change this only via the
|
||||
row-first discipline.)*
|
||||
|
||||
### 3. Enqueue semantics per boundary
|
||||
|
||||
At each due boundary the tick enqueues the schedule's `payload`
|
||||
(`serde_json::Value`, uniform with all payloads) into the named queue
|
||||
with `ScheduleOpts { priority, max_attempts, expires }` — the
|
||||
[EnqueueOpts](008-contract-v1-pinning.md) fields minus
|
||||
delay/run-at (a boundary fire is never delayed by its own config; the
|
||||
schedule *is* the delay). The enqueued work is ordinary queue work: it
|
||||
inherits the queue mechanism's delivery guarantees
|
||||
([ADR-006](006-wake-and-delivery-contract.md) queues row). The
|
||||
scheduler adds no delivery machinery of its own.
|
||||
|
||||
### 4. Boundary guarantee — the scheduler's row in the delivery table
|
||||
|
||||
Extends [ADR-006](006-wake-and-delivery-contract.md)'s table,
|
||||
resolving the transfer from [ADR-008](008-contract-v1-pinning.md) §7:
|
||||
|
||||
| Mechanism | Durability | Replay | Atomicity | Guarantee |
|
||||
|---|---|---|---|---|
|
||||
| scheduler (`@every`) | schedule rows + boundary fires are durable rows | n/a (fire = enqueue; the work's durability is the queues row) | fire and boundary-advance commit atomically under a row lock on the schedule row (crash mid-tick rolls back both — the boundary refires; a rogue ticker cannot double-fire) | **at-least-once per elapsed boundary while a leader runs**, with bounded catch-up: downtime fires replay boundary-by-boundary up to a fixed contract-constant per-tick cap (64, honker parity), beyond which missed boundaries skip forward to the next future boundary |
|
||||
|
||||
- The **bounded-catch-up + skip-forward** semantics are pinned as the
|
||||
honest contract: a consumer who pauses runners for a week does not
|
||||
come back to 10,000 stacked sweep jobs. (For an `@every 60s` sweep
|
||||
under an outage, replays beyond the cap are useless anyway.) **The
|
||||
cap is a fixed contract constant (64 per task per tick, honker
|
||||
parity)** — not a knob; like the backoff cap, no consumer names one,
|
||||
and the skip-forward is documented behavior, not configuration.
|
||||
- **One firer per boundary** is pinned engine-generically by two
|
||||
layers: (a) the leadership lock — one leader at a time, a leader
|
||||
that lost the lock returns before ticking; (b) **the boundary
|
||||
advance is a transactional claim on the schedule row**: the fire
|
||||
transaction takes a row lock on the schedule row, re-checks
|
||||
`next_fire_at <= now` under the lock, enqueues, and advances — so
|
||||
even a rogue second ticker (an operator bypassing the leadership
|
||||
lock) observes a boundary as already-fired and skips it, on either
|
||||
engine (Postgres: `SELECT … FOR UPDATE` on the row inside the fire
|
||||
tx; SQLite: the same guarantee via writer serialization —
|
||||
`BEGIN IMMEDIATE`). Layer (b) is the correctness floor, layer (a)
|
||||
the efficiency (one waiter instead of row-lock contention);
|
||||
crash mid-tick rolls back enqueue + advance together (the boundary
|
||||
refires).
|
||||
- **Leader election** is a named lock under the reserved prefix —
|
||||
`__alkstore_scheduler` (engine-derived name in the lock namespace per
|
||||
[ADR-008](008-contract-v1-pinning.md) §4's rule; consumers can never
|
||||
collide with it because every entry point rejects the reserved
|
||||
prefix). Both engines: each engine's own lock machinery, engine-
|
||||
internal detail. One firer per boundary across all processes sharing
|
||||
the store.
|
||||
|
||||
### 5. Schedule storage is engine-internal
|
||||
|
||||
Schedule rows live in engine-owned storage (SQLite: honker's
|
||||
`_honker_scheduler_tasks` — reused as-is, its `cron_expr` column
|
||||
carrying `@every` specs; Postgres: the engine-owned schema per
|
||||
[ADR-010](010-queue-semantics-depth.md) §8). Not consumer namespace;
|
||||
schedule *names* take the same validation as queue names
|
||||
(§1 — non-empty, reserved-prefix rejected). No registry
|
||||
validation against queue names — enqueuing into a queue nobody has
|
||||
claimed yet is fine (queues are names, not registered objects).
|
||||
|
||||
### 6. Error taxonomy delta
|
||||
|
||||
Two variants added by this surface, per ADR-008 §5's act-differently
|
||||
rule: `InvalidSpec { spec }` — a schedule spec that doesn't parse
|
||||
under §2's grammar (the caller fixes the spec string; rejected before
|
||||
storage, no round trip) — and `LeadershipLost` — the `run_schedules`
|
||||
return when leadership is lost mid-run (the caller acts by
|
||||
respawning; distinct from clean stop, §1).
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive**
|
||||
|
||||
- The contract gains the scheduler for two registrations + one runner;
|
||||
the "who sweeps" problem has a substrate answer that is one `@every`
|
||||
row and one spawned task.
|
||||
- The tick machinery (leader election, boundary advance, catch-up cap)
|
||||
is engine-internal and mostly inherited (SQLite: honker's scheduler
|
||||
machinery with `@every` specs; Postgres: re-derived, small at this
|
||||
surface size — the row-locked fire tx is the one mechanism the
|
||||
contract pins engine-generically, §4).
|
||||
- No TZ semantics, no cron parser in the contract, no schedule-object
|
||||
CRUD — the surface stays proportionate to the two inventory rows that
|
||||
exist.
|
||||
|
||||
**Negative**
|
||||
|
||||
- Honker's richer scheduler (`add/update/pause/resume/list`, cron
|
||||
with TZ) is deliberately not consumed at the contract level; consumers
|
||||
needing inspectable schedules need an inventory row + contract
|
||||
extension first.
|
||||
- The opt-in runner is one more thing a deployer must remember: forgot
|
||||
`run_schedules` ⇒ schedules silently never fire. The mechanism's doc
|
||||
text must carry this (a silent-nothing deployment posture, the honest
|
||||
twin of "no ambient timers").
|
||||
- The catch-up cap is a consumer-visible loss mode for long outages —
|
||||
documented (skip-forward), not silent: the schedule row's advance
|
||||
makes the gap inspectable.
|
||||
- The scheduler surface (and this track's `QueueOpts` depth) are
|
||||
**post-v1 contract extensions** — the first exercises of the
|
||||
extension path [ADR-008](008-contract-v1-pinning.md) leaves open;
|
||||
their versioning discipline (how engine crates track the additions)
|
||||
is OQ-10's, still open.
|
||||
|
||||
## References
|
||||
|
||||
- OQ-09 (`docs/architecture/open-questions.md`) — this ADR's
|
||||
resolution; owns ADR-008 §7's transferred scheduler guarantee row.
|
||||
- `docs/research/consumer-inventory.md` §scheduler — the two rows the
|
||||
decision is made against.
|
||||
- [ADR-002](002-feature-scope.md) — scheduler in-scope posture this
|
||||
ADR completes (documented-thin → pinned thin).
|
||||
- [ADR-006](006-wake-and-delivery-contract.md) — the guarantee table
|
||||
§4 extends; the enqueued work inherits the queues row.
|
||||
- [ADR-008](008-contract-v1-pinning.md) §1 (scheduler out of v1
|
||||
pending this decision), §4 (reserved-prefix rule the leadership lock
|
||||
name follows), §5 (act-differently rule the `InvalidSpec` variant
|
||||
follows), §7 (the guarantee-row transfer).
|
||||
- [ADR-010](010-queue-semantics-depth.md) — the queue semantics depth
|
||||
this ADR's fire semantics compose with; the maintenance recipe (who
|
||||
sweeps) is recorded there with the collapse machinery.
|
||||
- Honker reference checkout `/workspace/honker` @ f4e53c6 — scheduler
|
||||
surface (`packages/honker-rs/src/lib.rs`, `honker-core/src/honker_ops.rs`),
|
||||
`@every` grammar and local-TZ brittleness (`honker-core/src/cron.rs`),
|
||||
tick/leader mechanics.
|
||||
- [core-contract.md](../core-contract.md) — the spec carrying this
|
||||
surface.
|
||||
@@ -0,0 +1,353 @@
|
||||
# ADR-010: Queue semantics depth — visibility, retry/backoff, dead-letter, sweep, layout
|
||||
|
||||
## Status
|
||||
|
||||
Accepted (2026-10-05, Phase 1 — OQ-05's resolution; composes with
|
||||
[ADR-009](009-scheduler-collapse.md))
|
||||
|
||||
## Context
|
||||
|
||||
Contract v1 pinned the queue *skeleton* ([ADR-008](008-contract-v1-pinning.md)
|
||||
§1) and explicitly left the semantics depth to OQ-05: retry policy
|
||||
shape, visibility/renewal mechanics, dead-letter move-vs-flag and
|
||||
retention, sweep/maintenance design, the result-storage cut-flag's
|
||||
disposition, and queue/stream/lock **table** layout (the co-tenancy
|
||||
collision surface, [queues.md](../queues.md)).
|
||||
|
||||
The evidence base: honker's queue machinery (the SQLite-side incumbent
|
||||
— `/workspace/honker` @ f4e53c6, whose functions the engine rides
|
||||
directly) and the pg-boss family (the pg-side design reference —
|
||||
`/workspace/pgboss-rs` @ 98f7d9e standing in for node pg-boss v10;
|
||||
design-reference only per [ADR-005](005-dependency-ownership.md)).
|
||||
Both models were read in full for this resolution (reference notes:
|
||||
`docs/research/reference-honker-machinery.md`,
|
||||
`docs/research/reference-pgboss-rs-semantics.md`). The hard
|
||||
driver-coupled properties (transactional enqueue, exactly-once claim)
|
||||
are already POC-pinned on both engines.
|
||||
|
||||
The two references disagree on several semantics points, and honker
|
||||
has real gaps (a dead-letter table `get_job` can't see; expired
|
||||
processing rows no path can reach — the "zombie" hole; no dead-row
|
||||
retention mechanics; no backoff in the core at all). OQ-05 is the
|
||||
place to fix what re-derivation should fix and inheriting should
|
||||
inherit.
|
||||
|
||||
## Decision
|
||||
|
||||
### 1. The job state machine: three states, delete-on-ack
|
||||
|
||||
Contract job states: **`pending` → `processing` → `dead`** (+ absence:
|
||||
`ack` and `cancel` delete the row).
|
||||
|
||||
- Honker's model, adopted whole — one live table, move-to-dead
|
||||
(physical row move, savepoint-guarded), not the pg-boss seven-state
|
||||
enum. `max_attempts` is frozen at enqueue; `attempts` counts
|
||||
every claim.
|
||||
- **`ack` deletes** (honker parity): completed work leaves no row.
|
||||
The completed-job-with-output model (pg-boss's `completed` state +
|
||||
`output` column) is **not** adopted — it is result storage under
|
||||
another name, and result storage is a cut-flag row
|
||||
([ADR-002](002-feature-scope.md)); see §7.
|
||||
- **`get_job` sees dead jobs.** The returned job carries state,
|
||||
payload, attempts, priority, and the timestamps `created_at`,
|
||||
`run_at`, `claimed_at`, `expires_at` — plus, dead-only, `last_error`
|
||||
and `died_at` (no job = `None`). This deliberately *narrows* honker's
|
||||
surface (honker's `get_job` reads only live rows — post-mortem
|
||||
diagnosis was SQL-only); diagnosis-by-API is the honest fix, and it
|
||||
is what makes move-to-dead (vs flag-in-place) inspectable without
|
||||
raw SQL.
|
||||
- **`cancel` is unconditional** (honker parity): it deletes the row in
|
||||
either `pending` or `processing`, regardless of which worker holds
|
||||
it. Not an interrupt — the holder's next ack/heartbeat returns
|
||||
false, the same shape as expiry. `ack_batch` is the batch form of
|
||||
ack (count returned; non-claimed ids silently not counted).
|
||||
- Claim ordering (same queue): `priority DESC`, then ready-time
|
||||
(`run_at`) ascending, then enqueue order — FIFO under equal
|
||||
priority, both references agree; pinned.
|
||||
|
||||
### 2. Visibility timeout and heartbeat: explicit renewal, late-heartbeat refusal
|
||||
|
||||
One answer, contract-pinned (honker's design; pg-boss's
|
||||
no-renewal/expire-sweep model is not inherited):
|
||||
|
||||
- `visibility_timeout_s` is a per-job stamped value (see §3a's
|
||||
resolution rule — default **300 s**, honker parity); each claim sets
|
||||
the row's deadline from the job's stamp.
|
||||
- **`heartbeat(extend)` is renewal** — an absolute reset of the
|
||||
claim deadline from now. There is no progress-signal meaning; a
|
||||
consumer signaling progress uses its own channels. Renewal cadence
|
||||
is the consumer's obligation: handlers that might outlive the
|
||||
visibility timeout must heartbeat inside it.
|
||||
- **A late heartbeat is refused** (deadline already passed ⇒ returns
|
||||
false): it can never steal the job back from a reclaimer. The
|
||||
consequence is the honest at-least-once window — between deadline
|
||||
lapse and another worker's reclaim, the original worker may still
|
||||
complete: dual-execution is possible and idempotence is the
|
||||
consumer's job (matches the locks row's silent-expiry honesty,
|
||||
[ADR-008](008-contract-v1-pinning.md) §7).
|
||||
- **A reclaim consumes an attempt** (honker's counting): a claim is
|
||||
an attempt, whether fresh or a visibility reclaim. Documented
|
||||
explicitly as the contract's counting rule — the footgun is stated,
|
||||
not discovered: a handler that forgets to heartbeat looks like a
|
||||
repeatedly-failing job and dead-letters on budget exhaustion.
|
||||
- Deadline lapse itself is lazy: an expired-claim row becomes
|
||||
claimable by ordinary claim (no transition fires on lapse); there
|
||||
is no separate reaper for expired claims, only the budget rules
|
||||
below and the no-stranded-rows sweep (§5).
|
||||
|
||||
### 3. Retry and backoff: explicit delay or the queue's curve
|
||||
|
||||
- Job handle: `ack()`, `retry(err, delay)`, `fail(err)`, and
|
||||
`heartbeat(extend)` — the v1 skeleton ([ADR-008](008-contract-v1-pinning.md)
|
||||
§1), with `heartbeat`'s meaning pinned by §2 and `renew` remaining
|
||||
the *lock*-side term ([ADR-008](008-contract-v1-pinning.md) §8's
|
||||
split).
|
||||
- **`retry(err, None)`** — the deferred delay — is computed by the
|
||||
engine from the queue's curve; **`retry(err, Some(d))`** overrides
|
||||
it. Honker's core takes a caller-supplied delay with no curve;
|
||||
pg-boss computes in-engine from a base. The contract pins the
|
||||
curve so both engines compute identically.
|
||||
- **The curve: equal-jitter exponential**, `base * 2^(attempt-1)`
|
||||
with uniform jitter over the *lower half* of the doubling range —
|
||||
`delay ∈ [base·2^(a−1)/2, base·2^(a−1)]` — capped at **1 hour** (a
|
||||
pinned contract constant; no `backoff_max` knob — no consumer names
|
||||
one; the explicit-delay override is the escape hatch for bespoke
|
||||
policies). The range, not a jitter-label, is the definition: this
|
||||
is *not* AWS-canonical full jitter (uniform over `[0, cap]`), and
|
||||
the formula above is what both engines compute identically. The
|
||||
attempt index `a` is the value of `attempts` on the row at the
|
||||
moment the retry is scheduled (post-claim count, including
|
||||
reclaims). Jitter prevents sync-failure herding across a fleet;
|
||||
honker's wrappers have no jitter (a spread-out fleet of failing
|
||||
workers retries in lockstep) and the pg-boss family's jittered
|
||||
exponential is the design adopted.
|
||||
- **`QueueOpts { visibility_timeout_s, max_attempts, backoff_base_s,
|
||||
dead_letter_retention_s }`**
|
||||
— queue-level defaults (respectively 300 s / 3 / 5 s / none, honker
|
||||
parities except retention which is ours); resolved and stamped
|
||||
per §3a. `backoff_base_s` feeds the curve. This is the
|
||||
`QueueOpts` depth contract v1 deliberately withheld
|
||||
([ADR-008](008-contract-v1-pinning.md) §1) — pinned here as the
|
||||
extension path. **Outbox**: `outbox(name)` carries no opts of its
|
||||
own; its backing queue uses a *derived* QueueOpts set as engine
|
||||
defaults (honker's outbox parities: visibility 60 s, max_attempts 5,
|
||||
backoff base 5 s) — engine-documented, not separately configurable
|
||||
in v1.
|
||||
- Explicit `fail(err)` = immediate dead-letter (honker parity; it is
|
||||
the "stop retrying this" operator). `retry` at exhausted budget =
|
||||
dead-letter. Exhaustion via reclaim = dead-letter (the pre-claim
|
||||
sweep of already-exhausted reclaimable rows is an engine-side
|
||||
laziness optimization — SQLite rides honker's; pg re-derives).
|
||||
|
||||
### 3a. QueueOpts resolution: stamped at enqueue, per job
|
||||
|
||||
Queues are names, not registered objects — so queue-level `QueueOpts`
|
||||
need a defined attachment point, and the contract pins one:
|
||||
|
||||
- **Every job row carries the resolved opts stamped at enqueue**:
|
||||
enqueue resolves `EnqueueOpts.max_attempts` over the queue handle's
|
||||
`max_attempts` (the one per-job override the v1 skeleton names), and
|
||||
stamps the result — `visibility_timeout_s`, `backoff_base_s`,
|
||||
`dead_letter_retention_s`, `max_attempts` — onto the job row. Claims,
|
||||
heartbeats, retries, and sweeps read the **job's own stamps**, never
|
||||
a live registry.
|
||||
- **Why stamped**: uniform across engines (no per-queue config state
|
||||
to invent on either side — SQLite needs no new table beyond honker's
|
||||
columns, Postgres is job-table columns), and it makes a job's
|
||||
behavior immutable and inspectable (`get_job` returns the stamps
|
||||
with the rest of the row) no matter which process's handle enqueued
|
||||
it. Two processes opening handles with different `QueueOpts` for the
|
||||
same queue name do not fight — each enqueue stamps its own opts;
|
||||
the queue name is a routing key, not a configuration owner.
|
||||
- The consequence, stated: opts changes apply to *future enqueues
|
||||
only*; in-flight jobs keep their stamps. This is the honest,
|
||||
durable-row semantics — the same reason `max_attempts` was already
|
||||
frozen at enqueue in the references.
|
||||
- SQLite realization note: honker's `claim_batch` takes one uniform
|
||||
`timeout_s` per call, while per-job stamps make deadlines row-local —
|
||||
bridging that gap (post-claim re-stamp, per-row claiming, or another
|
||||
shape) is implementation work over honker's function surface and
|
||||
rides the same OQ-06 assessment as §5's zombie fix. The *contract*
|
||||
(deadline from the job's stamp) is engine-independent either way.
|
||||
|
||||
### 4. Dead-letter: move, retention = never-expire by default, no redrive API
|
||||
|
||||
Move-to-table (both references converge; flag-in-place is not
|
||||
adopted — a dead row must not occupy the claim path's indexes):
|
||||
dead rows physically move to engine-owned dead storage with
|
||||
`last_error` and `died_at`. Triggers: explicit `fail`, `retry` at
|
||||
budget, exhaustion-by-reclaim, and `expires` lapses (§5).
|
||||
|
||||
- **Retention**: dead rows live until explicitly removed by default
|
||||
(honker's posture). `QueueOpts.dead_letter_retention_s` (default
|
||||
`None` = forever) makes the sweep enforce per-queue dead TTL (§5) —
|
||||
the retention *mechanism* is ours (honker has none); the default
|
||||
stays non-magical. `last_error` source strings are contract-pinned
|
||||
for the engine-triggered cases: `"max attempts exceeded"` (budget
|
||||
exhaustion, any path) and `"expired"` (sweep-by-`expires`); explicit
|
||||
`fail(err)`/`retry(err, …)` carry the *caller's* error string.
|
||||
- **No redrive/requeue API in v1.** The recipe is pinned instead:
|
||||
`get_job` (now sees dead rows, §1) + fresh `enqueue` — both
|
||||
primitives already on the surface. No consumer names an API shape
|
||||
for redrive; when one does, it gets a contract extension with its
|
||||
own OQ.
|
||||
- The engine-owned dead table/rows are storage-internal — not a
|
||||
consumer namespace (no reserved-prefix question; they're not
|
||||
name-addressable).
|
||||
|
||||
### 5. `sweep_expired`: the no-stranded-rows property
|
||||
|
||||
Contract semantics for `sweep_expired(queue)` (the v1 skeleton's
|
||||
maintenance entry point):
|
||||
|
||||
- It moves to dead (`last_error = 'expired'`) **every row of the
|
||||
queue past its `expires_at` in any state** — pending rows (honker
|
||||
parity) *and* processing rows whose job-level expiry passed. This
|
||||
is the **no-stranded-rows property**: an enqueued job with an
|
||||
`expires` deadline is eventually in exactly one of pending,
|
||||
processing, or dead — never stuck unreachable. Honker fails this
|
||||
property: an expired *processing* row whose worker died is
|
||||
unreachable by claim (predicate requires future expiry), by
|
||||
pre-claim dead-lettering (same), and by `sweep_expired` (pending
|
||||
only) — the zombie hole, a real defect found in the reference read.
|
||||
This ADR fixes it at the contract level; the Postgres engine
|
||||
enforces it directly; the SQLite engine's realization over honker's
|
||||
machinery is implementation work **contingent on OQ-06** (a
|
||||
complement over honker's function surface, or the fork fires and
|
||||
the fix lands in owned code — the concrete fork candidate the
|
||||
quality read should weigh).
|
||||
- When `dead_letter_retention_s` is set, the same sweep also deletes
|
||||
dead rows past their retention (the only sweeper dead rows ever
|
||||
have — nothing runs without a caller).
|
||||
- Sweep is single-statement atomic per queue (SQLite: writer
|
||||
serialization; Postgres: row-locked `UPDATE … RETURNING`); **no
|
||||
leader lock is required** for a bare `sweep_expired` call — the
|
||||
multi-process recipe gets its coordination from §6's machinery, not
|
||||
from the sweep itself.
|
||||
|
||||
### 6. Sweep/maintenance cadence: no ambient sweeper; the collapse recipe
|
||||
|
||||
- **The engine ships no background sweeper, no maintenance thread, no
|
||||
default cadence** — the family's no-ambient-timers posture
|
||||
(alkblobs ADR-005) and honker's own verified posture (nothing in
|
||||
honker-core auto-runs; the only implicit maintenance is laziness in
|
||||
claim paths). The engine's *responsibility* is confined to
|
||||
correctness-of-transition (§1–§5); **timeliness is the
|
||||
consumer's**.
|
||||
- The pinned recipe for the family-wide "who sweeps" problem is the
|
||||
collapse machinery ([ADR-009](009-scheduler-collapse.md)):
|
||||
`store.schedule("maintenance", "@every 300s", maintenance_queue, …)`
|
||||
+ `run_schedules` (leader-elected) + a worker whose handler calls
|
||||
`sweep_expired` per queue. One schedule row, one worker — the
|
||||
pattern every consumer was going to hand-roll, answered twice in
|
||||
the contract instead of N times above it.
|
||||
- Dead-letter and notifications-table hygiene (SQLite engine):
|
||||
`prune_notifications`/`prune_notifications_keep_latest` remain
|
||||
**out of the contract** (ADR-008 §8's disposition, resolved here by
|
||||
disposition): the notifications table is the SQLite wake
|
||||
mechanism's *transport* detail — consumers interact with wakes, not
|
||||
rows. Its hygiene is engine-internal (capped at attach, engine-side
|
||||
maintenance cadence, engine opts) — the fix for honker's unbounded
|
||||
`_honker_notifications` growth must not become a consumer's chore.
|
||||
|
||||
### 7. Result storage: cut-flag stands
|
||||
|
||||
Reconsidered per OQ-05's mandate and **left cut**: no consumer row
|
||||
names outcome-query-by-id; the §1 delete-on-ack decision keeps
|
||||
completed jobs out of storage entirely (the pg-boss `output`/completed
|
||||
-row model is that feature under another name — adopting it would
|
||||
silent-include the cut flag); the documented workaround (a job writes
|
||||
its own result record atomically with ack) rides the tx seam
|
||||
([ADR-007](007-transactional-seam.md)) cleanly: business row + result
|
||||
write + ack-equivalent in one transaction. Re-entry stays gated on a
|
||||
consumer-inventory row.
|
||||
|
||||
### 8. Table layout: engine-owned schemas/tables; queues are rows, not tables
|
||||
|
||||
- **Postgres**: all engine-owned tables (job, dead, stream log,
|
||||
offsets, schedule rows, internal state) live in one **engine-owned
|
||||
PostgreSQL schema** — default `alkstore`, overridable per-engine
|
||||
option — co-tenanted safely with consumer tables (alkblobs ADR-008's
|
||||
co-tenancy precedent; schema scoping is the collision answer, and
|
||||
the reserved *name* namespace ([ADR-008](008-contract-v1-pinning.md)
|
||||
§4) governs the name column values inside it). Queue names become
|
||||
row values in one job table — **no per-queue tables, no per-queue
|
||||
schemas** (pg-boss's partition-per-queue opt-in is not inherited;
|
||||
at this scale one table + partial indexes matches honker's proven
|
||||
shape and keeps `queue(name)` a name, not a DDL operation — queue
|
||||
creation is not registry-gated, per
|
||||
[ADR-009](009-scheduler-collapse.md) §5). Per-queue config storage
|
||||
is likewise unnecessary (opts stamp onto job rows, §3a).
|
||||
- **SQLite**: honker's `_honker_*` table family in the caller's
|
||||
database file — storage-internal per
|
||||
[ADR-008](008-contract-v1-pinning.md) §4; no new tables minted by
|
||||
this ADR (job-stamped opts ride honker's existing columns and the
|
||||
contract's stated resolution rule, no schema change needed — this
|
||||
is exactly the seam the quality read (OQ-06) assesses). The
|
||||
family's exact fate rides that read: a fork re-owns the names,
|
||||
nothing consumer-visible changes.
|
||||
|
||||
### 9. Error taxonomy: no delta
|
||||
|
||||
The depth adds **no new error variants** — checked against
|
||||
[ADR-008](008-contract-v1-pinning.md) §5's act-differently rule: no
|
||||
work claimed is a value; deadline operations are booleans; dead-letter
|
||||
is observable state (`get_job`), not an error thrown; queues are
|
||||
unregistered names, so a payload enqueued to a queue nobody claims
|
||||
yet is not an error (it's durable work awaiting a worker).
|
||||
`InvalidSpec { spec }` was the scheduler's one variant
|
||||
([ADR-009](009-scheduler-collapse.md) §6) and is the only taxonomy
|
||||
delta this track produces.
|
||||
|
||||
## Consequences
|
||||
|
||||
**Positive**
|
||||
|
||||
- The queue mechanism is fully specified end-to-end: both references'
|
||||
semantics decisions are made once, pinned identically on both
|
||||
engines, with the references' gaps (zombie rows, dead-job
|
||||
invisibility, no retention mechanics, no jitter) fixed at the
|
||||
contract level rather than inherited.
|
||||
- The counting/heartbeat rules and the opts-stamping resolution are
|
||||
stated as contract text — the honest at-least-once story
|
||||
(dual-execution window, reclaim-eats-budget) is documented consumer
|
||||
obligation, not per-engine surprise, and per-queue configuration
|
||||
never fights across processes.
|
||||
- One engine-owned pg schema keeps co-tenancy mechanical.
|
||||
|
||||
**Negative**
|
||||
|
||||
- Three-state + delete-on-ack means no completion history and no
|
||||
built-in outcome storage — by design (the cut-flag stands, §7);
|
||||
consumers needing audit trails build them on the tx seam.
|
||||
- Honker's zombie fix, dead-row `get_job` visibility, and per-job
|
||||
visibility stamps over the uniform-claim-timeout function surface
|
||||
require either an over-machinery complement or a fork on the SQLite
|
||||
side — the fork calculus gains concrete candidates to weigh
|
||||
(OQ-06).
|
||||
- The backoff curve's 1-hour cap and jitter formula are pinned
|
||||
constants — a consumer needing a different policy uses
|
||||
`retry(err, Some(d))` per attempt (correct, but manual).
|
||||
|
||||
## References
|
||||
|
||||
- OQ-05 (`docs/architecture/open-questions.md`) — this ADR's
|
||||
resolution (with OQ-09, resolved by [ADR-009](009-scheduler-collapse.md)).
|
||||
- `docs/research/reference-honker-machinery.md` and
|
||||
`docs/research/reference-pgboss-rs-semantics.md` — the full
|
||||
reference reads (schemas, state machines, defect list with
|
||||
file/line cites) this resolution was made over.
|
||||
- [ADR-002](002-feature-scope.md) — scope (cut-flags §7 stands),
|
||||
[ADR-005](005-dependency-ownership.md) — the design-reference
|
||||
postures the two reads serve,
|
||||
[ADR-006](006-wake-and-delivery-contract.md) — the queues guarantee
|
||||
row this ADR's §2/§3 make precise,
|
||||
[ADR-007](007-transactional-seam.md) — the tx seam §7's workaround
|
||||
rides,
|
||||
[ADR-008](008-contract-v1-pinning.md) — the v1 skeleton this ADR
|
||||
extends, §5's rule §9 applies, §8's dispositions §6 resolves.
|
||||
- [ADR-009](009-scheduler-collapse.md) — the collapse machinery §6's
|
||||
recipe composes with.
|
||||
- OQ-06 — the SQLite-side fork candidates (§5, §3a realization note).
|
||||
- [queues.md](../queues.md), [core-contract.md](../core-contract.md),
|
||||
engine specs — the specs carrying this depth.
|
||||
Reference in new issue
Block a user