Files
alkstore/tasks/suite-queue-depth-rows.md
glm-5.3-flash 5c9ae6a7fc Contract-suite queue-depth rows (task suite-queue-depth-rows): the job-handle validity predicate row (in-window heartbeat/ack landing, late-heartbeat/post-lapse-ack/retry/fail refusals past a 1 s stamp with the row untouched, ack_batch per-id predicate live-1/lapsed-0/nonexistent-0, fail(None)→"failed" and retry-at-budget→"max attempts exceeded"), the ADR-010 depth row (reclaim consumes an attempt with claimed_at/deadline refreshed and the original holder refusing, reclaim-exhaustion dead-lettering with the pre-claim sweep — get_job-visible "max attempts exceeded"+died_at, cancel unconditional delete with the not-an-interrupt refusal shape plus pending/dead/missing arms), and the no-stranded-rows sweep row (both states move with "expired", unexpired/never-expiring untouched, retention TTL enforcing with the moved+deleted sum, None=forever) — each version-stamped per the suite convention (ADR-010 §1–§5, ADR-019 §3, ADR-008 §5), wired into both engines' suite targets (SQLite tokio tests, pg harness_row!s). M-1's retention-failure pin dispositioned engine-side per the task's honest call: the storage-level DELETE failure is not injectable through the contract surface, so it lands in the SQLite substrate's trigger seam (sweep_rolls_back_the_retention_half_with_the_move — the move half rolls back with the retention half), with the pg twin verified structurally (both halves inside in_tx's frame) and the disposition recorded in the task Notes. Verified: sqlite suite 13/13, pg suite 13/13 vs harness twice (postgres/poc@:15432), substrate sweep-rollback tests 4/4, workspace build/test green, clippy -D warnings, fmt clean
2026-10-10 05:50:26 +00:00

9.5 KiB
Raw Permalink Blame History

id, name, status, depends_on, scope, risk, impact, level, tags
id name status depends_on scope risk impact level tags
suite-queue-depth-rows Contract-suite rows — job-handle validity predicate + ADR-010 queue depth completed
moderate medium phase implementation
wave-5
contract-suite
queues

Description

Add the queue-depth contract rows to alkstore-contract-suite/src/properties.rs and wire them into both engines' suite targets (alkstore-sqlite/tests/contract_suite.rs directly; alkstore-postgres/tests/contract_suite.rs via the harness_row! macro), discharging the backlog rows "Job-handle validity predicate on both engines" and the ADR-010 depth properties (core-contract.md §Verification backlog). One owner per property (ADR-012 §2 mirror): the rows are engine-independent, run against both factories, version-stamped per the crate's convention.

Rows to add:

  • job_handle_validity_predicate — the uniform predicate (ADR-010 §2): an op succeeds only while the row is processing and the claim deadline is unexpired. Pin: the late-heartbeat boundary (heartbeat refused exactly when the deadline has lapsed — a short visibility_timeout_s stamp plus a tolerance sleep past it; the sleep-then-assert-state shape is sequential single-store and fits the suite's determinism posture), the post-lapse ack refusal (at-least-once: the worker's ack does not land after lapse), ack_batch applying the predicate per id (lapsed or non-claimed ids silently not counted — the returned count reflects only live acks), retry's refusal being false with the row untouched, and the engine-default strings landing identically (fail(None) → "failed", exhaustion → "max attempts exceeded").
  • queue_depth_reclaim_and_dead_letter — the ADR-010 depth properties: a visibility reclaim consumes an attempt (expired claim → reclaim → attempts incremented, claimed_at refreshed), exhaustion moves the row to dead storage (get_job sees the dead row with last_error/died_at), and cancel is an unconditional delete (not an interrupt).
  • sweep_no_stranded_rows — sweep_expired() moves every past-expiry row (any state — pending-expired, processing-lapsed, dead-past-retention) and enforces dead-letter retention; the no-stranded-rows property (ADR-010 §5). Include the M-1 follow-up pin here (the waves-1–2 general review's suggested wave-5 add): the sweep's retention half rolling back with the sweep on failure — with an explicit disposition: if a retention-DELETE failure is not injectable through the contract surface (it likely is not — the failure is storage-level), pin that leg engine-side (SQLite substrate test seam) and record the disposition in this task's Notes for the review gate to audit. The plan ordered "M-1's retention-failure test → wave 5's contract-suite rows"; where the pin actually lands is the implementer's honest call, documented.

Timing posture: short stamps (1 s visibility) + tolerance sleeps (bounded well above the transition's resolution), asserting state outcomes never timing values. No two-task races.

Acceptance Criteria

  • Three rows exist in properties.rs, pub async fn(&dyn StoreFactory), version-stamped (ADR-010 §2/§5, ADR-019 §3, ADR-008 §5 as applicable; stamps cite ADR § per the version_stamp.rs convention)
  • Rows wired into both engines' contract_suite.rs targets
  • SQLite column green server-less; pg column green against the harness server
  • The M-1 retention-failure pin's location decided and recorded in Notes (suite row or engine-side, with the reason)
  • cargo test -p alkstore-sqlite -p alkstore-postgres green (pg rows skip cleanly server-less); clippy -D warnings; fmt clean

References

  • docs/architecture/core-contract.md §Verification backlog (job-handle validity predicate; ADR-010 depth properties)
  • docs/architecture/decisions/010-queue-semantics-depth.md §2/§5
  • docs/architecture/decisions/019-mechanism-handle-surfaces.md §3
  • docs/reviews/001-waves-1-2-general-review.md (M-1, §4's suggested adds)
  • tasks/sqlite-engine-integration.md (the wave-5 lesson: rows are mock-blind until a real single-writer engine runs them)

Notes

Decisions of record the description didn't pin:

  • M-1 retention-failure disposition: pinned engine-side (SQLite substrate), not as a suite row. The retention-DELETE failure is a storage-level fault with no injection seam through the contract surface (StoreFactory + the contract traits offer no way to fail one statement of one engine's sweep while the store stays live) — so the pin landed as sweep_rolls_back_the_retention_half_with_the_move in the SQLite substrate's queue test seam (a trigger-raised ABORT on __alkstore_dead DELETE — the exact mechanism review 001 §4 sketched): a sweep with both an expired live row and an aged-out dead row fails entire (the move half's effects roll back with the retention half; unblocked, the re-sweep accounts 1 moved + 1 retention-deleted — the sum). The pg twin was not given a live-injected leg: the engine has no client-side fault seam either, and its sweep frame is structurally the same guarantee verified by code-read — both halves (move + retention) run inside one in_tx BEGIN…COMMIT/ROLLBACK frame (queue.rs sweep_expired), so a retention-DELETE failure rolls the move back via the frame's error arm. Review gate should audit this disposition: pg's leg rests on frame structure, not a live fault test.
  • The exhaustion-string pin is split across the two rows by trigger: fail(None) → "failed" and retry-at-exhausted-budget → "max attempts exceeded" pin in job_handle_validity_predicate (both reachable without a lapse); reclaim-exhaustion → "max attempts exceeded" pins in queue_depth_reclaim_and_dead_letter where that trigger actually occurs (the pre-claim sweep at the budget-spent claim). The row-1 text's single-line mention of exhaustion thus rides the row that owns the exercising path.
  • The ack_batch leg carries its own queue and a renewal, not a post-lapse fresh claim — any later claim in a queue holding lapsed rows reclaims them FIFO-first (the predicate makes a lapsed claim reclaimable, so a "fresh" claim in the same queue picks the lapsed row ahead of the fresh one; found live on the first run against the real engine — the wave-3 lesson repeating). The leg therefore keeps one id's claim live through a heartbeat renewal while its sibling lapses on the 1 s stamp, in a dedicated short-visibility queue: three ids (live / lapsed / nonexistent), count 1, the lapsed row still processing afterward.
  • cancel legs extend the task's list slightly in the same row: cancel of a dead id returns false and deletes nothing (outside cancel's live-row reach), alongside the pinned processing-cancel (holder's next op refuses the lapse shape), pending-cancel (never claimable), and missing-id false legs.
  • Timing posture realization: sleeps are 2.5 s for 1 s visibility/claim stamps and 3.5 s for expires: Some(2) stamps (the extra second keeps the claim window — expires_at > now at claim — deterministic against second-resolution clocks); every assertion is a state outcome, none a duration. The deadline-boundary legs pin refusal-state both sides of the lapse (in-window ops land; post-lapse ops refuse) without asserting the boundary instant.
  • Row 2's reclaim leg also pins the original holder refusing after the reclaim (heartbeat → false with the row now w2's) — the reclaim side of the same predicate, in the row that owns reclaims.

Summary

Landed: the three queue-depth contract rows in alkstore-contract-suite/src/properties.rs — job_handle_validity_predicate (in-window heartbeat/ack landing; late-heartbeat, post-lapse ack, retry, and fail refusals past a 1 s stamp with the row untouched; ack_batch per-id predicate — live 1, lapsed 0, nonexistent 0; fail(None) → "failed" and retry-at-budget → "max attempts exceeded"), queue_depth_reclaim_and_dead_letter (reclaim consumes an attempt with claimed_at/deadline refreshed; original holder refuses; reclaim-exhaustion dead-letters with the pre-claim sweep — dead row get_job-visible with "max attempts exceeded" + died_at; cancel unconditional delete with the not-an-interrupt refusal shape, pending/dead/missing arms), and sweep_no_stranded_rows (pending-expired + processing-expired both move with "expired", unexpired/never-expiring untouched; retention TTL deletes with the moved+deleted sum; retention None = forever). All version-stamped (ADR-010 §1–§5, ADR-019 §3, ADR-008 §5), re-exported from the suite lib, and wired into both engines' contract_suite.rs targets (SQLite three new #[tokio::test]s; pg three harness_row!s).

M-1 follow-up: the sweep's retention-half rollback pinned in the SQLite substrate (block_dead_deletes trigger seam + sweep_rolls_back_the_retention_half_with_the_move) — see Notes for the disposition.

Verified: workspace cargo build and cargo test green (25 core + 189 + 121 SQLite + 121 pg lib + 9 schema + harness crates); SQLite contract suite 13/13 (three new rows, 8.5 s total); pg contract suite 13/13 against the harness server (pglo-poc :15432, postgres/poc/blobs) twice consecutively (15.3 s each — the sleep legs really ran, not skipped); the substrate rollback tests 4/4; clippy --workspace --all-targets -D warnings clean; cargo fmt --check clean.