docs/architecture/ now exists: README index, overview, five component specs (core-contract, engine-sqlite, engine-postgres, queues, deployment), ADR-001..007 carrying the Phase 0 resolved decisions (crate split, feature scope, per-engine drivers, dependency ownership, wake contract, tx seam), and the centralized open-questions tracker promotion: OQ-ST-01..08 mirror to OQ-01..08 one-to-one with statuses/resolutions carried; new Phase 1 questions append (OQ-09 scheduler collapse, OQ-10 contract versioning). Open Phase 1 work: OQ-04 contract pinning (high), OQ-05 queue semantics depth (high), OQ-06 honker-core quality read (high; fork-trigger gate), OQ-08 capability surface, OQ-09, OQ-10. Erratum fixed in phase-0 OQ-ST-04 (thread-affinity friction is SQLite-side, previously garbled as pg-side) and a stale scheduler- boundary pointer corrected in consumer-inventory.md. Two review passes run (findings: OQ-promotion numbering faithfulness, ADR back-reference sync) — all critical/warning findings resolved.
4.6 KiB
ADR-004: Postgres engine — tokio-postgres + deadpool-postgres, hand-rolled LISTEN forwarder
Status
Accepted
Context
The Postgres engine's driver question (OQ-ST-03, Postgres half) had
three candidate families: sqlx (single API across engines, and the
driver pgboss-rs carries), tokio-postgres + deadpool-postgres (the
alkblobs POC evidence base, natively async), and adopt/fork pgboss-rs
(bringing sqlx where the queue lives). POC #2
(docs/research/poc-pg-posture-findings.md, ran 2026-10-04) validated
the tokio-postgres posture end-to-end and measured the pieces the
contract depends on. The sqlite-arm comparison facts from
[ADR-003] apply on this side too: pgboss-rs brings sqlx (a second
driver per binary), has no LISTEN/NOTIFY at all (verified against the
checkout — consumption is fetch_job polling), so the push-reactivity
half is this crate's work regardless of fork-or-adopt.
Decision
The Postgres engine uses:
-
tokio-postgres 0.7.x + deadpool-postgres 0.14.x — pooled connections for queries/claims (exactly-once via
FOR UPDATE SKIP LOCKED), natively async (client isSend + Sync— no bridge, no spawn_blocking, the tx handle holds the pooled object directly). -
A hand-rolled LISTEN forwarder (the POC's ~90-line shape) — a dedicated non-pooled listener connection per process, a poll-message loop fanning out to a bounded broadcast channel, immediate reconnect with exponential backoff (50 ms → 2 s cap), re-LISTEN from the channel list after every reconnect, and a synthetic reconnect-wake on a reserved channel closing the no-replay hole ([ADR-006]).
postgres-notify0.3.8 was evaluated in-probe and passed over (derive-not-adopt: lazy reconnect, connect_script skipped at initial connect, unquoted-identifier LISTENs, single-maintainer posture; it stays a recorded fallback if upstream improves). -
Queue machinery re-derived on this driver with the pg-boss schema family as design reference — not adopted/forked as a dependency ([ADR-005]; semantics depth is [queues.md]'s work, OQ-ST-05). The minimal-queue ground is measured: enqueue/claim/ack with
SKIP LOCKEDis ~40 lines of SQL and passed the full property suite. -
Default consumption posture: LISTEN-driven claim with a re-poll safety net; poll-only remains the fallback (interval-tunable) when a LISTEN connection is unavailable or unwanted. Measured: claim latency p50 3–6 ms LISTEN-driven vs 32–50 ms at a 50 ms poll.
Structural constraints the engine owns (all test-pinned in POC #2):
- Pooled connections cannot carry LISTEN (deadpool#360 —
registers server-side, never delivers). The listener connection is a
per-process budget line outside the pool (
max_size + 1). pg_notifypayloads are ≤ 8000 bytes; the engine checks the limit client-side and returns a typed error before the round-trip. Large payloads ride a table row with the id in the notification (the outbox shape).- A query-vs-poll starvation deadlock exists if the listener's poll
loop isn't running before the first client query on the connection;
and dropping the
Clientcloses the server session even if theConnectiontask survives. Both are hand-rolled-forwarder failure modes, owned and test-pinned by the engine.
Consequences
Positive
- Single-driver engine, natively async, zero version conflicts, no vendoring.
- The push channel pgboss-rs lacks, this engine gets — measured at 5–16× claim-latency improvement over even a 50 ms poll.
- Multi-host is native (no single-host assumption anywhere; property tests ran all-through-network).
- The forwarder's two deadlock pitfalls are learned territory with pinned tests, not unknowns.
Negative
- The listener forwarder is ours to maintain (~90 lines + the two
pinned failure modes) — the cost of not adopting
postgres-notify. - Per-process connection budget grows by one listener connection per process that listens (a deployment-matrix row; consumers sizing pools must account for it).
- The pg queue schema is re-derived work; the pg-boss family is reference, not free.
References
docs/research/poc-pg-posture-findings.md— measurements and the pinned property suite.- OQ-ST-03, OQ-ST-05's posture half (
docs/research/phase-0.md). - ADR-001 — single-driver binaries.
- ADR-005 — the ownership calculus this decision applies per-subsystem.
- ADR-006 — the wake contract the forwarder implements on this engine.
- engine-postgres.md — the engine spec.