--- id: suite-queue-depth-rows name: Contract-suite rows — job-handle validity predicate + ADR-010 queue depth status: completed depends_on: [] scope: moderate risk: medium impact: phase level: implementation tags: [wave-5, contract-suite, queues] --- ## Description Add the queue-depth contract rows to `alkstore-contract-suite/src/properties.rs` and wire them into both engines' suite targets (`alkstore-sqlite/tests/contract_suite.rs` directly; `alkstore-postgres/tests/contract_suite.rs` via the `harness_row!` macro), discharging the backlog rows "Job-handle validity predicate on both engines" and the ADR-010 depth properties (core-contract.md §Verification backlog). One owner per property (ADR-012 §2 mirror): the rows are engine-independent, run against both factories, version-stamped per the crate's convention. Rows to add: - **`job_handle_validity_predicate`** — the uniform predicate (ADR-010 §2): an op succeeds only while the row is `processing` and the claim deadline is unexpired. Pin: the late-heartbeat boundary (heartbeat refused exactly when the deadline has lapsed — a short `visibility_timeout_s` stamp plus a tolerance sleep past it; the sleep-then-assert-state shape is sequential single-store and fits the suite's determinism posture), the post-lapse ack refusal (at-least-once: the worker's ack does not land after lapse), `ack_batch` applying the predicate per id (lapsed or non-claimed ids silently not counted — the returned count reflects only live acks), `retry`'s refusal being `false` with the row untouched, and the engine-default strings landing identically (`fail(None)` → `"failed"`, exhaustion → `"max attempts exceeded"`). - **`queue_depth_reclaim_and_dead_letter`** — the ADR-010 depth properties: a visibility reclaim consumes an attempt (expired claim → reclaim → `attempts` incremented, `claimed_at` refreshed), exhaustion moves the row to dead storage (`get_job` sees the dead row with `last_error`/`died_at`), and `cancel` is an unconditional delete (not an interrupt). - **`sweep_no_stranded_rows`** — `sweep_expired()` moves *every* past-expiry row (any state — pending-expired, processing-lapsed, dead-past-retention) and enforces dead-letter retention; the no-stranded-rows property (ADR-010 §5). Include the M-1 follow-up pin here (the waves-1–2 general review's suggested wave-5 add): the sweep's retention half rolling back with the sweep on failure — **with an explicit disposition**: if a retention-DELETE failure is not injectable through the contract surface (it likely is not — the failure is storage-level), pin that leg engine-side (SQLite substrate test seam) and record the disposition in this task's Notes for the review gate to audit. The plan ordered "M-1's retention-failure test → wave 5's contract-suite rows"; where the pin actually lands is the implementer's honest call, documented. Timing posture: short stamps (1 s visibility) + tolerance sleeps (bounded well above the transition's resolution), asserting state outcomes never timing values. No two-task races. ## Acceptance Criteria - [x] Three rows exist in `properties.rs`, `pub async fn(&dyn StoreFactory)`, version-stamped (ADR-010 §2/§5, ADR-019 §3, ADR-008 §5 as applicable; stamps cite ADR § per the `version_stamp.rs` convention) - [x] Rows wired into both engines' `contract_suite.rs` targets - [x] SQLite column green server-less; pg column green against the harness server - [x] The M-1 retention-failure pin's location decided and recorded in Notes (suite row or engine-side, with the reason) - [x] `cargo test -p alkstore-sqlite -p alkstore-postgres` green (pg rows skip cleanly server-less); clippy `-D warnings`; fmt clean ## References - docs/architecture/core-contract.md §Verification backlog (job-handle validity predicate; ADR-010 depth properties) - docs/architecture/decisions/010-queue-semantics-depth.md §2/§5 - docs/architecture/decisions/019-mechanism-handle-surfaces.md §3 - docs/reviews/001-waves-1-2-general-review.md (M-1, §4's suggested adds) - tasks/sqlite-engine-integration.md (the wave-5 lesson: rows are mock-blind until a real single-writer engine runs them) ## Notes > Decisions of record the description didn't pin: - **M-1 retention-failure disposition: pinned engine-side (SQLite substrate), not as a suite row.** The retention-DELETE failure is a storage-level fault with no injection seam through the contract surface (`StoreFactory` + the contract traits offer no way to fail one statement of one engine's sweep while the store stays live) — so the pin landed as `sweep_rolls_back_the_retention_half_with_the_move` in the SQLite substrate's queue test seam (a trigger-raised ABORT on `__alkstore_dead` DELETE — the exact mechanism review 001 §4 sketched): a sweep with both an expired live row and an aged-out dead row fails entire (the move half's effects roll back with the retention half; unblocked, the re-sweep accounts 1 moved + 1 retention-deleted — the sum). The pg twin was **not** given a live-injected leg: the engine has no client-side fault seam either, and its sweep frame is structurally the same guarantee verified by code-read — both halves (move + retention) run inside one `in_tx` `BEGIN…COMMIT/ROLLBACK` frame (queue.rs `sweep_expired`), so a retention-DELETE failure rolls the move back via the frame's error arm. Review gate should audit this disposition: pg's leg rests on frame structure, not a live fault test. - **The exhaustion-string pin is split across the two rows by trigger**: `fail(None)` → `"failed"` and retry-at-exhausted-budget → `"max attempts exceeded"` pin in `job_handle_validity_predicate` (both reachable without a lapse); *reclaim*-exhaustion → `"max attempts exceeded"` pins in `queue_depth_reclaim_and_dead_letter` where that trigger actually occurs (the pre-claim sweep at the budget-spent claim). The row-1 text's single-line mention of exhaustion thus rides the row that owns the exercising path. - **The `ack_batch` leg carries its own queue and a renewal, not a post-lapse fresh claim** — any later claim in a queue holding lapsed rows reclaims them FIFO-first (the predicate makes a lapsed claim reclaimable, so a "fresh" claim in the same queue picks the lapsed row ahead of the fresh one; found live on the first run against the real engine — the wave-3 lesson repeating). The leg therefore keeps one id's claim live through a heartbeat renewal while its sibling lapses on the 1 s stamp, in a dedicated short-visibility queue: three ids (live / lapsed / nonexistent), count 1, the lapsed row still `processing` afterward. - **cancel legs extend the task's list slightly in the same row**: cancel of a *dead* id returns `false` and deletes nothing (outside cancel's live-row reach), alongside the pinned processing-cancel (holder's next op refuses the lapse shape), pending-cancel (never claimable), and missing-id `false` legs. - **Timing posture realization**: sleeps are 2.5 s for 1 s visibility/claim stamps and 3.5 s for `expires: Some(2)` stamps (the extra second keeps the claim window — `expires_at > now` at claim — deterministic against second-resolution clocks); every assertion is a state outcome, none a duration. The deadline-boundary legs pin refusal-state both sides of the lapse (in-window ops land; post-lapse ops refuse) without asserting the boundary instant. - **Row 2's reclaim leg also pins the original holder refusing after the reclaim** (heartbeat → `false` with the row now `w2`'s) — the reclaim side of the same predicate, in the row that owns reclaims. ## Summary > Landed: the three queue-depth contract rows in > `alkstore-contract-suite/src/properties.rs` — > `job_handle_validity_predicate` (in-window heartbeat/ack landing; > late-heartbeat, post-lapse ack, retry, and fail refusals past a 1 s > stamp with the row untouched; `ack_batch` per-id predicate — live > 1, lapsed 0, nonexistent 0; `fail(None)` → `"failed"` and > retry-at-budget → `"max attempts exceeded"`), > `queue_depth_reclaim_and_dead_letter` (reclaim consumes an attempt > with `claimed_at`/deadline refreshed; original holder refuses; > reclaim-exhaustion dead-letters with the pre-claim sweep — dead row > `get_job`-visible with `"max attempts exceeded"` + `died_at`; > cancel unconditional delete with the not-an-interrupt refusal shape, > pending/dead/missing arms), and `sweep_no_stranded_rows` > (pending-expired + processing-expired both move with `"expired"`, > unexpired/never-expiring untouched; retention TTL deletes with the > moved+deleted sum; retention `None` = forever). All version-stamped > (ADR-010 §1–§5, ADR-019 §3, ADR-008 §5), re-exported from the suite > lib, and wired into both engines' `contract_suite.rs` targets > (SQLite three new `#[tokio::test]`s; pg three `harness_row!`s). > > M-1 follow-up: the sweep's retention-half rollback pinned in the > SQLite substrate (`block_dead_deletes` trigger seam + > `sweep_rolls_back_the_retention_half_with_the_move`) — see Notes for > the disposition. > > Verified: workspace `cargo build` and `cargo test` green (25 core + > 189 + 121 SQLite + 121 pg lib + 9 schema + harness crates); > SQLite contract suite 13/13 (three new rows, 8.5 s total); pg > contract suite 13/13 against the harness server (`pglo-poc` :15432, > `postgres/poc`/`blobs`) twice consecutively (15.3 s each — the sleep > legs really ran, not skipped); the substrate rollback tests 4/4; > clippy `--workspace --all-targets -D warnings` clean; > `cargo fmt --check` clean.