Files
alkstore/docs/architecture/decisions/004-postgres-driver.md
T
glm-5.3-flash 4391f6e879 docs: open Phase 1 — architecture spec set over the Phase 0 evidence
docs/architecture/ now exists: README index, overview, five component
specs (core-contract, engine-sqlite, engine-postgres, queues,
deployment), ADR-001..007 carrying the Phase 0 resolved decisions
(crate split, feature scope, per-engine drivers, dependency
ownership, wake contract, tx seam), and the centralized
open-questions tracker promotion: OQ-ST-01..08 mirror to OQ-01..08
one-to-one with statuses/resolutions carried; new Phase 1 questions
append (OQ-09 scheduler collapse, OQ-10 contract versioning).
Open Phase 1 work: OQ-04 contract pinning (high), OQ-05 queue
semantics depth (high), OQ-06 honker-core quality read (high;
fork-trigger gate), OQ-08 capability surface, OQ-09, OQ-10.

Erratum fixed in phase-0 OQ-ST-04 (thread-affinity friction is
SQLite-side, previously garbled as pg-side) and a stale scheduler-
boundary pointer corrected in consumer-inventory.md. Two review
passes run (findings: OQ-promotion numbering faithfulness, ADR
back-reference sync) — all critical/warning findings resolved.
2026-10-04 18:13:10 +00:00

4.6 KiB
Raw Blame History

ADR-004: Postgres engine — tokio-postgres + deadpool-postgres, hand-rolled LISTEN forwarder

Status

Accepted

Context

The Postgres engine's driver question (OQ-ST-03, Postgres half) had three candidate families: sqlx (single API across engines, and the driver pgboss-rs carries), tokio-postgres + deadpool-postgres (the alkblobs POC evidence base, natively async), and adopt/fork pgboss-rs (bringing sqlx where the queue lives). POC #2 (docs/research/poc-pg-posture-findings.md, ran 2026-10-04) validated the tokio-postgres posture end-to-end and measured the pieces the contract depends on. The sqlite-arm comparison facts from [ADR-003] apply on this side too: pgboss-rs brings sqlx (a second driver per binary), has no LISTEN/NOTIFY at all (verified against the checkout — consumption is fetch_job polling), so the push-reactivity half is this crate's work regardless of fork-or-adopt.

Decision

The Postgres engine uses:

  • tokio-postgres 0.7.x + deadpool-postgres 0.14.x — pooled connections for queries/claims (exactly-once via FOR UPDATE SKIP LOCKED), natively async (client is Send + Sync — no bridge, no spawn_blocking, the tx handle holds the pooled object directly).

  • A hand-rolled LISTEN forwarder (the POC's ~90-line shape) — a dedicated non-pooled listener connection per process, a poll-message loop fanning out to a bounded broadcast channel, immediate reconnect with exponential backoff (50 ms → 2 s cap), re-LISTEN from the channel list after every reconnect, and a synthetic reconnect-wake on a reserved channel closing the no-replay hole ([ADR-006]). postgres-notify 0.3.8 was evaluated in-probe and passed over (derive-not-adopt: lazy reconnect, connect_script skipped at initial connect, unquoted-identifier LISTENs, single-maintainer posture; it stays a recorded fallback if upstream improves).

  • Queue machinery re-derived on this driver with the pg-boss schema family as design reference — not adopted/forked as a dependency ([ADR-005]; semantics depth is [queues.md]'s work, OQ-ST-05). The minimal-queue ground is measured: enqueue/claim/ack with SKIP LOCKED is ~40 lines of SQL and passed the full property suite.

  • Default consumption posture: LISTEN-driven claim with a re-poll safety net; poll-only remains the fallback (interval-tunable) when a LISTEN connection is unavailable or unwanted. Measured: claim latency p50 3–6 ms LISTEN-driven vs 32–50 ms at a 50 ms poll.

Structural constraints the engine owns (all test-pinned in POC #2):

  • Pooled connections cannot carry LISTEN (deadpool#360 — registers server-side, never delivers). The listener connection is a per-process budget line outside the pool (max_size + 1).
  • pg_notify payloads are ≤ 8000 bytes; the engine checks the limit client-side and returns a typed error before the round-trip. Large payloads ride a table row with the id in the notification (the outbox shape).
  • A query-vs-poll starvation deadlock exists if the listener's poll loop isn't running before the first client query on the connection; and dropping the Client closes the server session even if the Connection task survives. Both are hand-rolled-forwarder failure modes, owned and test-pinned by the engine.

Consequences

Positive

  • Single-driver engine, natively async, zero version conflicts, no vendoring.
  • The push channel pgboss-rs lacks, this engine gets — measured at 5–16× claim-latency improvement over even a 50 ms poll.
  • Multi-host is native (no single-host assumption anywhere; property tests ran all-through-network).
  • The forwarder's two deadlock pitfalls are learned territory with pinned tests, not unknowns.

Negative

  • The listener forwarder is ours to maintain (~90 lines + the two pinned failure modes) — the cost of not adopting postgres-notify.
  • Per-process connection budget grows by one listener connection per process that listens (a deployment-matrix row; consumers sizing pools must account for it).
  • The pg queue schema is re-derived work; the pg-boss family is reference, not free.

References

  • docs/research/poc-pg-posture-findings.md — measurements and the pinned property suite.
  • OQ-ST-03, OQ-ST-05's posture half (docs/research/phase-0.md).
  • ADR-001 — single-driver binaries.
  • ADR-005 — the ownership calculus this decision applies per-subsystem.
  • ADR-006 — the wake contract the forwarder implements on this engine.
  • engine-postgres.md — the engine spec.