Files
alkstore/docs/architecture/engine-postgres.md
T

9.2 KiB
Raw Blame History

status, last_updated
status last_updated
draft 2026-10-06

Postgres engine

The alkstore-postgres engine implements core-contract.md on tokio-postgres + deadpool-postgres. This spec records WHAT the engine is internally — connection architecture, the LISTEN forwarder, queue machinery re-derivation — not code-level HOW. Decisions live in ADRs; contract obligations live in the core spec.

Identity and posture

  • Single driver, natively async: tokio-postgres 0.7.x + deadpool-postgres 0.14.x (ADR-004). No bridge, no spawn_blocking; the tx handle holds the pooled object directly (client is Send + Sync — POC #2 compile-probe verified).
  • Multi-host by nature: connections are per-process state; nothing assumes a shared host (ADR-004, verified all-through-network in POC #2). See deployment.md.
  • Queue machinery re-derived on this driver with the pg-boss schema family as design reference (ADR-005); semantics depth pinned by ADR-010 — all engine-owned tables (job, dead, stream, offsets, schedule) in one PostgreSQL schema (default alkstore), queues as rows, no per-queue tables.

Connection architecture

  • Pool (deadpool) — queries, claims, and all non-transactional work. Per-connection statement cache, RecyclingMethod::Fast (no DISCARD ALL recycling; claim SQL re-prepared implicitly with zero errors at POC scale).
  • Listener connection — one dedicated, non-pooled connection per process that listens. Pooled connections cannot carry LISTEN (deadpool#360 — registration succeeds, delivery is impossible; the client-wrapper exposes no notification surface, source-verified and test-pinned as pooled_listen_registers_but_cannot_deliver). The listener is therefore a per-process budget line outside the pool: max_size + 1 per LISTEN-ing process (deployment.md).
  • Forwarder — the listener's loop: poll_message fanning out into a bounded broadcast channel (lag surfaced, not silent), re-LISTEN from the channel list after every reconnect (exponential backoff 50 ms → 2 s cap), and the synthetic reconnect-wake on the reserved channel (__alkstore_listener_reconnected__, ADR-008 §4) — broadcast to every subscriber's receiver, per the wake contract's reconnect-recovery semantics (ADR-006).
  • One listener serves N channels and N subscribers; re-attach is a broadcast re-subscribe (no server round-trips); per-channel connections are never warranted at this scale (POC-verified).

Mapping the contract

Contract piece Engine realization
notify / listen pg_notify(...) inside the caller's tx (delivers at commit — native commit-atomicity, ADR-007); listen() via LISTEN on the forwarder's connection, fanout to receivers
streams durable event table + per-consumer offset cursors; pg_notify as the wake trigger (ADR-006 mechanism split: durable row, LISTEN wake — the pg-boss-family shape). Depth per ADR-015: the event table carries the nullable key column (carried metadata), reads yield offset ASC (global FIFO — bigserial offsets, immutable), publish_with_key_tx validates + inserts inside the caller's tx, trim_to is a DELETE … WHERE offset <= ? on a pool connection (wakes nothing)
queues re-derived queue table + FOR UPDATE SKIP LOCKED claim + LISTEN-driven wake with re-poll safety net (default consumption posture, measured 5–16× vs 50 ms poll; poll-only fallback); states/dead-letter/backoff/visibility per ADR-010, all in the engine-owned schema; the curve/stamps arithmetic is computed engine-side per ADR-012 §2 (equivalence with the SQLite engine pinned by the contract suite)
named locks advisory-lock-semantics TTL locks (pg-boss-family design reference; guarantee row pinned by ADR-008 §7)
scheduler / outbox collapse shape (ADR-009): schedule rows in the engine-owned schema, tick re-derived (boundary advance + 64-boundary catch-up cap, honker parity), leadership via the engine's lock machinery on __alkstore_scheduler; outbox = helper over queues
begin_tx pool checkout + BEGIN, returning the caller-held handle (ADR-007)
handle ops straight .awaits through the held object; commit/rollback returns the object to the pool

Owned failure modes (all test-pinned in POC #2)

Hand-rolling the forwarder means owning its pitfalls — they are learned territory, pinned as passing tests:

  1. Query-vs-poll starvation deadlock — the poll loop must be running before the first client query on the listener connection.
  2. Client-drop closes the server session — a long-lived listener keeps its Client alive for the listener's lifetime.
  3. The no-replay hole — commit during a connection gap is never re-delivered; recovery = reconnect + synthetic wake + consumer re-read (ADR-006). The !saw_replay test pins the honesty.
  4. Payload boundary — pg_notify ≤ 8000 bytes; client-side checked, typed error before the round-trip. Large payloads ride a table row with the id in the notification (outbox shape).
  5. Read-your-writes — read committed default verified in both directions (in-tx and post-commit).

Constraint also carried: mixed rusqlite+sqlx binaries would need a vendored patch today — excluded by construction in this engine (single driver, ADR-001), noted for the record in ADR-003.

Design Decisions

ADR Decision Summary
001 Crate split single-driver engine crate
002 Feature scope which rows this engine serves
004 Driver tokio-postgres + deadpool; hand-rolled forwarder; re-derived queues
005 Ownership published libs as-is; postgres-notify derive-not-adopt
006 Wake contract LISTEN push, no replay, synthetic reconnect-wake
007 Tx seam direct pooled-object handle, no bridging
008 Contract v1 pinned surface; reserved reconnect-wake channel string; PayloadTooLarge taxonomy variant
009 Scheduler collapse schedule rows in the engine schema, re-derived tick, row-locked fire tx, __alkstore_scheduler leadership
010 Queue depth job-stamped opts, equal-jitter backoff, dead-letter move, no-stranded-rows sweep, one engine-owned schema
012 Fork design contract-blind substrate (SQLite side); pg engine owns its own curve/stamps arithmetic, equivalence pinned by the contract suite
014 Outbox tx enqueue outbox_enqueue_tx on TxHandle; derivation engine-side inside the caller's tx
015 Streams depth nullable key column (carried metadata); bigserial offsets, global-FIFO reads; keyed tx publish; trim_to as a pool-connection delete
016 Deployment honesty no runtime capability surface — this engine's multi-host posture is stated by its crate identity and docs; PayloadTooLarge occurrence pinned contract-suite (this engine produces it, client-side pre-round-trip)

Open Questions

Open questions are tracked in open-questions.md. Key questions affecting this document:

  • OQ-08: capability surface — resolved (2026-10-06, ADR-016): no runtime capability surface; this engine's multi-host posture lives in its crate identity + docs and the deployment matrix; the engine's one runtime-visible asymmetry (PayloadTooLarge) was already contract-pinned.
  • OQ-12: streams depth — resolved (2026-10-05, ADR-015): key = carried metadata; global-FIFO-by-offset ordering row (bigserial offsets, immutable); StreamEvent shape pinned (topic → stream in the contract type); publish_with_key_tx inside the caller's tx; trim_to on the stream handle — no engine-default retention.

Resolved: OQ-09 (scheduler collapse — ADR-009), OQ-05 (queue semantics depth — ADR-010), OQ-12 (streams depth — ADR-015), and OQ-08 (capability surface — none; ADR-016), 2026-10-05/06.