Findings: unified surface holds on tokio-postgres with the transactional property intact (in-tx NOTIFY is commit-atomic; rollback drops all); LISTEN wake beats poll 5-16x at p50 with 300/300 isolated delivery; pooled-LISTEN discard (deadpool#360) verified and pinned as our own test; postgres-notify 0.3.8 evaluated and passed over (lazy reconnect, no initial-connect script, unquoted identifier LISTENs) in favor of the ~90-line hand-rolled forwarder with test-pinned pitfalls. sqlx PgListener fallback retired unfired.
15 KiB
status, title, last_updated
| status | title | last_updated |
|---|---|---|
| findings-filed | POC #2 — Postgres engine posture: LISTEN/NOTIFY wiring, the pool tx-seam, and reactive parity on tokio-postgres | 2026-10-04 |
POC: Postgres engine posture — spec
POC register #2 (phase-0.md). Feeds OQ-ST-03's Postgres half (the driver question's remaining open part) and OQ-ST-04's Postgres wake side (the reactive contract). Code: standalone crate in the global workspace (
/workspace/alkstore-pg-posture-poc; harness conventions cloned from/workspace/alkstore-sqlite-posture-poc). Findings land indocs/research/poc-pg-posture-findings.mdhere regardless. Server: dockerized postgres (POC-harness convention from the alkblobs POCs —postgres:16-alpineon :15432, config knobs stated per-probe). Pre-spec research input (2026-10-04): the dedicated-LISTEN-connection requirement (verified against deadpool-rs/deadpool#360, open upstream issue — pooled connections discard async notifications) and thepostgres-notify0.3.8 candidate (verified: reconnection + multi- channel subscribe) — both folded in below; neither's behavioral claims are trusted until probed.Ran 2026-10-04: PASS — findings in poc-pg-posture-findings.md. Both pre-spec research claims verified in-probe and pinned as tests (#360: pooled LISTEN registers but never delivers; postgres-notify: reconnect + re-LISTEN verified with recorded quirks — lazy reconnect, no script at initial connect, unquoted identifiers). The spec's decision gate passed on all three conditions; OQ-ST-03 closed; the sqlx PgListener fallback retired unfired.
What this POC must decide
OQ-ST-03's SQLite half is resolved (POC #1: honker-core on our rusqlite). The Postgres half asks the structural twin of the same question, with three differences that make it not a rerun:
- The queue machinery is greenfield either way. pgboss-rs (the
schema-family port) has no LISTEN/NOTIFY (verified 2026-10-03) and
is sqlx-coupled; under OQ-ST-02's per-engine-crate split it brings
no driver synergy. So the engine crate writes queue tables + claim/
ack + retry/dead-letter itself on some driver — this POC grounds
that build on tokio-postgres (the alkblobs POC #5/#7-validated
stack: deadpool-postgres pool, per-connection prepared-statement
discipline,
synchronous_commitas the durability knob). - The wake mechanism is native server push (LISTEN/NOTIFY), structurally
different from SQLite's
data_versionpolling: connection-bound, delivered on a dedicated connection, no retry/visibility semantics. The reactive contract's honest-shape question (OQ-ST-04) needs measured ground: what LISTEN plumbing costs, how it fails, and what a transaction-scoped interplay (NOTIFY fires only on commit of the sending transaction — the pg analogue of honker's notify-in-tx atomicity) looks like over a pool. - The tx-seam shape differs from SQLite's resolved one. SQLite
(POC #1): an owned writer-slot lease, ops via
spawn_blocking. Postgres: interactive transactions handed out by the pool (deadpoolgivesObject→transaction()), LISTEN connections are separate long-lived connections — the*_txseam and the listener lifecycle have to compose on one driver without one emulating the other's weaknesses.
The question, concretely: does tokio-postgres carry the unified
surface (enqueue_tx / claim / ack / notify_tx / listen / stream
offset ops / lock ops) with the transactional property intact, at
seam/pool/wake costs comparable to what POC #1 pinned for SQLite — and
what are the honest posture deltas (pool sizing vs LISTEN connection
dedication, synchronous_commit posture, notification payload limits)?
If yes, OQ-ST-03 closes with "per-engine drivers: rusqlite+honker-core
(SQLite) / tokio-postgres+deadpool (Postgres)" and OQ-ST-04's
contract-pinning has measured ground on both sides. If no — if LISTEN
over deadpool proves fragile or the tx-seam shape fights the pool — the
findings name the fight (sqlx's PgListener as the alternative is a
fallback posture, evaluated only on that failure).
The one arm, three sub-modules
Unlike POC #1's A/B split, this POC has one driver posture with three sub-modules to validate — the comparison axis is within the engine (listen-vs-poll, tx-seam shape), not between drivers:
Sub-module L (listen plumbing)
- Upstream-verified constraint (2026-10-04, from the deadpool issue
tracker — deadpool-rs/deadpool#360, open): pooled connections
cannot carry LISTEN. deadpool's connect task awaits the underlying
tokio_postgres::Connectionto completion, which discards the async notification messages onlyConnection::poll_messageexposes — and any LISTEN on a pooled connection is lost at recycle. So the shape is confirmed required, not just preferred: a dedicated, non-pooledtokio-postgresconnection per listener process, with a forwarding task pollingpoll_messageand fanning notifications out to per-subscriber tokio mpsc channels.pg_notifyas a sender rides pooled connections freely (it is ordinary SQL). The POC asserts the discarded-notifications failure empirically (LISTEN issued on a pooled connection, verify no delivery) so the constraint is in our evidence base with our own test, not only the issue's word. - Two implementation postures for the listener substrate, compared:
- (a) hand-rolled forwarder — the thin forwarding task above, ours: reconnect policy, channel re-LISTEN after reconnect, and backpressure are all ours (~the size of a small module; the POC #1 Arm-B watcher taught what "ours" costs — every failure mode is maintained by us).
- (b)
postgres-notify0.3.8 (MIT, tokio-postgres-based, published, verified 2026-10-04: auto-reconnect with exponential backoff + jitter, multi-channelsubscribe_notify, aconnect_scripthook executed on (re)connect — which is the LISTEN-restoration mechanism — query timeout + cancellation, and documented callback-panic safety properties). A turnkey shape for exactly this sub-module. Caveat to verify in-probe: docs do not explicitly promise subscription restoration across reconnect — theconnect_scriptis the mechanism but restoration behavior (does it re-issue our LISTENs) must be verified by test, plus the single-maintainer/small-crate posture recorded per the AGENTS.md ownership rules. It also owns its connection for queries too (PGRobustClient wraps the whole client) — the engine crate would use it as the dedicated-listener-side client only, not the query path. Both are implemented in-probe; the comparison axes are exactly the ones sub-module W probes (reconnect honesty, fan-out, latency) plus dependency-posture cost. Preference found here is OQ-ST-04/OQ-ST-06 input (it is a per-subsystem adopt-or-derive vote, the POC-level answer to the same calculus as honker-core).
- Multi-channel: one LISTEN connection serving N channels (
LISTENaccepts multiple registrations per connection) — measure whether the dedicated-connection-per-channel shape is ever warranted (connection budget vs fan-out complexity). - Payload:
pg_notifycarries ≤ 8000 bytes. Probe the boundary behavior (payload size limits, the "payload too large" error path) and record the honest contract: notify payloads are hints; large payloads ride a table row + the notify carries the row id (the honker-outbox shape, pg-side).
Sub-module T (tx-seam over the pool)
deadpool-postgrespool;enqueue_tx/publish_tx/save_offset_txreceive a caller-held transaction. The property the SQLite side pinned (business write + enqueue + notify in one caller tx, rollback drops both) must hold withNOTIFYissued inside the caller's tx (Postgres delivers only on commit — verify, don't assume).- The seam question: transaction ownership vs transaction handle.
SQLite's lease shape cannot port (no slot model under a pool). Two
candidate shapes, both implemented and compared:
- (a) caller-owned tx handle — the trait hands the caller a
TxHandle(deadpoolTransaction<'_>wrapped) and*_txmethods take it by reference; the caller drives begin/commit. - (b) closure-scoped — the trait exposes
with_tx(|tx| async { ... })(pool checkout + begin + commit/ rollback inside the closure), and the business write also happens through the same handle — the caller never holds the raw transaction. Both are viable-looking; the POC measures and records the ergonomic and correctness trade (deadlock risk, spawn_blocking needs, what the caching-subscriber pattern wants). This is direct OQ-ST-04 contract input.
- (a) caller-owned tx handle — the trait hands the caller a
- Prepared-statement discipline per B2/B5 (POC #5): per-connection caching across pool recycling (deadpool's statement cache posture) — verify the "prepared statement s1 does not exist" failure mode does not recur with the queue SQL shape.
Sub-module W (wake-vs-poll parity)
- The queue consumption path two ways: (1) LISTEN-driven — a queue-table notify on commit wakes claimants (the push channel pgboss-rs lacks; our machinery), with claim on wake + re-poll safety net; (2) poll-only — interval polling (the pgboss-rs posture). Measure claim latency (enqueue→claim) both ways, p50/p99.
pg_notifyvisibility timing: the notifier's NOTIFY is delivered to listeners at tx commit, but the notifying connection's subsequent reads may or may not see their own effects timing-wise — probe reader-visibility (the "read-your-writes over the pool" boundary underread committed, the honest default).- LISTEN connection failure modes: drop the listen connection
mid-subscription (network kill), verify reconnect + the replay hole
it implies (LISTEN has no replay — anything committed while the
listener was down is not re-delivered; the honest contract is
opaque-wake + re-read, so recovery = on reconnect, wake all
subscribers once — and, under listener posture (b),
connect_script-driven re-LISTEN must be verified to run). Re-attach storm: N listeners × M channels on one re-connecting connection.
Instruments
- Seam/cost probe (
pgdiag-st-1): the POC #1 seam workload's pg twin —begin + enqueue_tx + commitper iteration through the pool, n=3000,synchronous_commit=on(ship config) and=off(the B5 knob, both reported), p50/p90/p99. Compare shape (not absolute number) against POC #1's SQLite table — the cross-engine relative claim OQ-ST-04's contract must absorb. Plus the raw floor re-measure (prepared round-trip ~150–500 µs per POC #5 B2) as methodology cross-check. - Wake probe (
pgdiag-st-2): N=300 notify-commits at 25 ms spacing, listener attach → wake latency (p50/p99/max); burst-stress (30 rapid commits — LISTEN does not coalesce per-tick the way data_version polling does; measure whether bursts each deliver or the notifications() stream backs up); payload-size boundary; channel fan-out (4 subscribers, same channel, every one wakes). - Property tests (
tests/contract.rs— the POC #1 suite's pg twin, all shared assertions reused): commit-atomicity (enqueue+notify+business write, rollback drops all — ghost-claim check against the queue table), exactly-once claim under 4 concurrent claimants, offset save/read through the caller tx, lock acquire/release/renew with TTL, listener-reconnect recovery (kill listen conn, commit during the gap, verify reconnect wakes + state re-read correct), and the pooled-LISTEN-discard assertion (LISTEN on a pool connection, verify nothing is delivered — pinning deadpool#360's behavior in our own evidence base; if upstream ever fixes it, the test flips and the fix becomes usable — but the engine design must not count on it). - Poll-vs-listen claim-latency probe (
pgdiag-st-3): enqueue→ claim latency under both consumption postures at 1/8/32 claimant workers — the honest comparison that justifies (or retires) the LISTEN-driven claim path as the engine's default. - Pool/posture probe (
pgdiag-st-4): deadpool pool sizing vs the LISTEN connection budget (the dedicated listen connection is outside the pool by the #360 constraint — the probe validates the budget accounting: pool max_size + one listener connection per process); fresh-session cost re-verify (~19–25 ms per B2) against the pooling discipline;synchronous_commitper-session posture (SET on checkout vs system config) mechanics.
Decision gate
Sub-module findings constitute OQ-ST-03's Postgres half resolution iff:
- Contract: every shared property holds (the commit-atomicity property via in-tx NOTIFY is the load-bearing one; exactly-once claim under concurrency the second);
- Seam: the caller-tx seam shape survives with measured costs (either candidate shape; the findings name the preferred one with evidence — that preference is OQ-ST-04 input, not a Phase 0 ADR);
- Wake: LISTEN plumbing is robust enough to be the engine's wake story (reconnect recovery honest, fan-out correct, latency ≤ the sqlite watcher's order — LISTEN should beat poll; if it doesn't, that is a finding that reshapes OQ-ST-04's contract);
Failure on any: the findings name the failure precisely, the sqlx
PgListener fallback posture gets its own (smaller) evaluation round,
and OQ-ST-03's pg half stays open with a narrowed question.
Out of scope
- The queue semantics depth (retry/backoff/dead-letter/sweep tuning — OQ-ST-05's design work; this POC uses a minimal queue table: enqueue, claim, ack — enough for the property tests).
- The full pg-boss schema family (job states beyond the minimal set,
pgboss-rs's DDL compatibility) — schema design reference status is OQ-ST-05; this POC does not adopt or validate pgboss-rs' code. - The sqlite engine (POC #1's ground; its seam/cost numbers are used only as the cross-engine reference points).
- Multi-host stress (failover, pooling across machines) — OQ-ST-08's deployment-matrix work; this POC's dockerized server rides the established harness convention.
- Streams' full contract (offset save/read verified; per-consumer replay windows/filtering is OQ-ST-04 design work).
Register note
This POC is #2 in alkstore's register (phase-0.md), specified 2026-10-04. Sequencing rationale: it completes OQ-ST-03's driver resolution (the SQLite half is POC #1's) and gives OQ-ST-04's contract-pinning measured ground on both engines — after which the reactive-contract and ownership questions (OQ-ST-04/05/06) are paper-decisions over a filled evidence base, which is what Phase 1 should inherit.