Files
alkstore/docs/architecture/decisions/004-postgres-driver.md
T
glm-5.3-flash 2949612e2c ADR-012: forked-substrate design — contract-blind boundary, fidelity posture, port deltas
Follow-through on OQ-06/ADR-011: pin the fork's structural decisions
(alkstore-substrate as a vendored path-dep crate, contract-blind API
boundary with contract formulas computed engine-side and pinned
equivalent by the contract suite, keep-the-kept-half API fidelity for
cheap cherry-picks, the W-1/W-2/dead-man's-switch/W-4 port deltas
decided per item, bootstrap re-keying off error-string matching, no
rename migration, deliberate upstream tracking).

Consistency sweep across the doc set for the fork: annotate ADR-003/
005/009/010 and core-contract for superseded ownership facts, fix
schedule-storage table naming (ADR-009 §5, queues.md), re-key ADR-010
§6's notifications hygiene to the at-attach cap the fork scope
realizes, add OQ-11 (scaffold-time residue), and complete both ADR
indexes. Independent review: 0 critical, warnings addressed.
2026-10-05 05:00:55 +00:00

4.7 KiB
Raw Blame History

ADR-004: Postgres engine — tokio-postgres + deadpool-postgres, hand-rolled LISTEN forwarder

Status

Accepted

Context

The Postgres engine's driver question (OQ-ST-03, Postgres half) had three candidate families: sqlx (single API across engines, and the driver pgboss-rs carries), tokio-postgres + deadpool-postgres (the alkblobs POC evidence base, natively async), and adopt/fork pgboss-rs (bringing sqlx where the queue lives). POC #2 (docs/research/poc-pg-posture-findings.md, ran 2026-10-04) validated the tokio-postgres posture end-to-end and measured the pieces the contract depends on. The sqlite-arm comparison facts from [ADR-003] apply on this side too: pgboss-rs brings sqlx (a second driver per binary), has no LISTEN/NOTIFY at all (verified against the checkout — consumption is fetch_job polling), so the push-reactivity half is this crate's work regardless of fork-or-adopt.

Decision

The Postgres engine uses:

  • tokio-postgres 0.7.x + deadpool-postgres 0.14.x — pooled connections for queries/claims (exactly-once via FOR UPDATE SKIP LOCKED), natively async (client is Send + Sync — no bridge, no spawn_blocking, the tx handle holds the pooled object directly).

  • A hand-rolled LISTEN forwarder (the POC's ~90-line shape) — a dedicated non-pooled listener connection per process, a poll-message loop fanning out to a bounded broadcast channel, immediate reconnect with exponential backoff (50 ms → 2 s cap), re-LISTEN from the channel list after every reconnect, and a synthetic reconnect-wake on a reserved channel closing the no-replay hole ([ADR-006]). postgres-notify 0.3.8 was evaluated in-probe and passed over (derive-not-adopt: lazy reconnect, connect_script skipped at initial connect, unquoted-identifier LISTENs, single-maintainer posture; it stays a recorded fallback if upstream improves).

  • Queue machinery re-derived on this driver with the pg-boss schema family as design reference — not adopted/forked as a dependency ([ADR-005]; semantics depth is [queues.md]'s work, OQ-ST-05). The minimal-queue ground is measured: enqueue/claim/ack with SKIP LOCKED is ~40 lines of SQL and passed the full property suite.

  • Default consumption posture: LISTEN-driven claim with a re-poll safety net; poll-only remains the fallback (interval-tunable) when a LISTEN connection is unavailable or unwanted. Measured: claim latency p50 3–6 ms LISTEN-driven vs 32–50 ms at a 50 ms poll.

Structural constraints the engine owns (all test-pinned in POC #2):

  • Pooled connections cannot carry LISTEN (deadpool#360 — registers server-side, never delivers). The listener connection is a per-process budget line outside the pool (max_size + 1).
  • pg_notify payloads are ≤ 8000 bytes; the engine checks the limit client-side and returns a typed error before the round-trip. Large payloads ride a table row with the id in the notification (the outbox shape).
  • A query-vs-poll starvation deadlock exists if the listener's poll loop isn't running before the first client query on the connection; and dropping the Client closes the server session even if the Connection task survives. Both are hand-rolled-forwarder failure modes, owned and test-pinned by the engine.

Consequences

Positive

  • Single-driver engine, natively async, zero version conflicts, no vendoring.
  • The push channel pgboss-rs lacks, this engine gets — measured at 5–16× claim-latency improvement over even a 50 ms poll.
  • Multi-host is native (no single-host assumption anywhere; property tests ran all-through-network).
  • The forwarder's two deadlock pitfalls are learned territory with pinned tests, not unknowns.

Negative

  • The listener forwarder is ours to maintain (~90 lines + the two pinned failure modes) — the cost of not adopting postgres-notify.
  • Per-process connection budget grows by one listener connection per process that listens (a deployment-matrix row; consumers sizing pools must account for it).
  • The pg queue schema is re-derived work; the pg-boss family is reference, not free.

References

  • docs/research/poc-pg-posture-findings.md — measurements and the pinned property suite.
  • OQ-ST-03, OQ-ST-05's posture half (docs/research/phase-0.md).
  • ADR-001 — single-driver binaries.
  • ADR-005 — the ownership calculus this decision applies per-subsystem.
  • ADR-006 — the wake contract the forwarder implements on this engine.
  • engine-postgres.md — the engine spec. [ADR-003]: 003-sqlite-driver.md [ADR-005]: 005-dependency-ownership.md [ADR-006]: 006-wake-and-delivery-contract.md [queues.md]: ../queues.md