Follow-through on OQ-06/ADR-011: pin the fork's structural decisions (alkstore-substrate as a vendored path-dep crate, contract-blind API boundary with contract formulas computed engine-side and pinned equivalent by the contract suite, keep-the-kept-half API fidelity for cheap cherry-picks, the W-1/W-2/dead-man's-switch/W-4 port deltas decided per item, bootstrap re-keying off error-string matching, no rename migration, deliberate upstream tracking). Consistency sweep across the doc set for the fork: annotate ADR-003/ 005/009/010 and core-contract for superseded ownership facts, fix schedule-storage table naming (ADR-009 §5, queues.md), re-key ADR-010 §6's notifications hygiene to the at-attach cap the fork scope realizes, add OQ-11 (scaffold-time residue), and complete both ADR indexes. Independent review: 0 critical, warnings addressed.
4.7 KiB
ADR-004: Postgres engine — tokio-postgres + deadpool-postgres, hand-rolled LISTEN forwarder
Status
Accepted
Context
The Postgres engine's driver question (OQ-ST-03, Postgres half) had
three candidate families: sqlx (single API across engines, and the
driver pgboss-rs carries), tokio-postgres + deadpool-postgres (the
alkblobs POC evidence base, natively async), and adopt/fork pgboss-rs
(bringing sqlx where the queue lives). POC #2
(docs/research/poc-pg-posture-findings.md, ran 2026-10-04) validated
the tokio-postgres posture end-to-end and measured the pieces the
contract depends on. The sqlite-arm comparison facts from
[ADR-003] apply on this side too: pgboss-rs brings sqlx (a second
driver per binary), has no LISTEN/NOTIFY at all (verified against the
checkout — consumption is fetch_job polling), so the push-reactivity
half is this crate's work regardless of fork-or-adopt.
Decision
The Postgres engine uses:
-
tokio-postgres 0.7.x + deadpool-postgres 0.14.x — pooled connections for queries/claims (exactly-once via
FOR UPDATE SKIP LOCKED), natively async (client isSend + Sync— no bridge, no spawn_blocking, the tx handle holds the pooled object directly). -
A hand-rolled LISTEN forwarder (the POC's ~90-line shape) — a dedicated non-pooled listener connection per process, a poll-message loop fanning out to a bounded broadcast channel, immediate reconnect with exponential backoff (50 ms → 2 s cap), re-LISTEN from the channel list after every reconnect, and a synthetic reconnect-wake on a reserved channel closing the no-replay hole ([ADR-006]).
postgres-notify0.3.8 was evaluated in-probe and passed over (derive-not-adopt: lazy reconnect, connect_script skipped at initial connect, unquoted-identifier LISTENs, single-maintainer posture; it stays a recorded fallback if upstream improves). -
Queue machinery re-derived on this driver with the pg-boss schema family as design reference — not adopted/forked as a dependency ([ADR-005]; semantics depth is [queues.md]'s work, OQ-ST-05). The minimal-queue ground is measured: enqueue/claim/ack with
SKIP LOCKEDis ~40 lines of SQL and passed the full property suite. -
Default consumption posture: LISTEN-driven claim with a re-poll safety net; poll-only remains the fallback (interval-tunable) when a LISTEN connection is unavailable or unwanted. Measured: claim latency p50 3–6 ms LISTEN-driven vs 32–50 ms at a 50 ms poll.
Structural constraints the engine owns (all test-pinned in POC #2):
- Pooled connections cannot carry LISTEN (deadpool#360 —
registers server-side, never delivers). The listener connection is a
per-process budget line outside the pool (
max_size + 1). pg_notifypayloads are ≤ 8000 bytes; the engine checks the limit client-side and returns a typed error before the round-trip. Large payloads ride a table row with the id in the notification (the outbox shape).- A query-vs-poll starvation deadlock exists if the listener's poll
loop isn't running before the first client query on the connection;
and dropping the
Clientcloses the server session even if theConnectiontask survives. Both are hand-rolled-forwarder failure modes, owned and test-pinned by the engine.
Consequences
Positive
- Single-driver engine, natively async, zero version conflicts, no vendoring.
- The push channel pgboss-rs lacks, this engine gets — measured at 5–16× claim-latency improvement over even a 50 ms poll.
- Multi-host is native (no single-host assumption anywhere; property tests ran all-through-network).
- The forwarder's two deadlock pitfalls are learned territory with pinned tests, not unknowns.
Negative
- The listener forwarder is ours to maintain (~90 lines + the two
pinned failure modes) — the cost of not adopting
postgres-notify. - Per-process connection budget grows by one listener connection per process that listens (a deployment-matrix row; consumers sizing pools must account for it).
- The pg queue schema is re-derived work; the pg-boss family is reference, not free.
References
docs/research/poc-pg-posture-findings.md— measurements and the pinned property suite.- OQ-ST-03, OQ-ST-05's posture half (
docs/research/phase-0.md). - ADR-001 — single-driver binaries.
- ADR-005 — the ownership calculus this decision applies per-subsystem.
- ADR-006 — the wake contract the forwarder implements on this engine.
- engine-postgres.md — the engine spec. [ADR-003]: 003-sqlite-driver.md [ADR-005]: 005-dependency-ownership.md [ADR-006]: 006-wake-and-delivery-contract.md [queues.md]: ../queues.md