docs: fold verified LISTEN research into POC #2 spec
Two claims verified independently before folding: (1) deadpool-postgres discards async notifications — upstream deadpool-rs/deadpool#360 (open, Oct 2024) confirms the pool's connect task awaits the connection to completion, dropping what only poll_message exposes; pooled LISTEN is lost at recycle. The dedicated non-pooled LISTEN connection is now upstream-verified required, not spec-preferred. (2) postgres-notify 0.3.8 exists as described (MIT, tokio-postgres, auto-reconnect with backoff+jitter, multi-channel subscribe_notify, connect_script hook) — admitted as sub-module L's second arm (hand-rolled forwarder vs turnkey listener substrate), with in-probe verification required: docs don't explicitly promise subscription restoration across reconnect; connect_script is the mechanism; single-maintainer posture recorded. New property test pinned: pooled-LISTEN-discard assertion (our own evidence for the #360 behavior, flips if upstream fixes). Probe 5 updated to budget accounting (listener conn outside the pool).
This commit is contained in:
1 parent
e18281735e
commit
4c144f8f7f
1 file changed
+59
-9
@@ -14,6 +14,12 @@ last_updated: 2026-10-04
|
||||
> in `docs/research/poc-pg-posture-findings.md` here regardless.
|
||||
> Server: dockerized postgres (POC-harness convention from the alkblobs
|
||||
> POCs — `postgres:16-alpine` on :15432, config knobs stated per-probe).
|
||||
> Pre-spec research input (2026-10-04): the dedicated-LISTEN-connection
|
||||
> requirement (verified against deadpool-rs/deadpool#360, open upstream
|
||||
> issue — pooled connections discard async notifications) and the
|
||||
> `postgres-notify` 0.3.8 candidate (verified: reconnection + multi-
|
||||
> channel subscribe) — both folded in below; neither's *behavioral*
|
||||
> claims are trusted until probed.
|
||||
|
||||
## What this POC must decide
|
||||
|
||||
@@ -66,11 +72,46 @@ sub-modules to validate — the comparison axis is *within* the engine
|
||||
|
||||
### Sub-module L (listen plumbing)
|
||||
|
||||
- One dedicated `tokio-postgres` connection per process running
|
||||
`LISTEN <channel>`; notifications surface via the connection's
|
||||
`notifications()` stream. Fan-out to per-subscriber tokio mpsc
|
||||
channels (the `listen()` contract: multiple subscribers per channel,
|
||||
each getting every notification on that channel after attach).
|
||||
- **Upstream-verified constraint (2026-10-04, from the deadpool issue
|
||||
tracker — deadpool-rs/deadpool#360, open):** pooled connections
|
||||
cannot carry LISTEN. deadpool's connect task awaits the underlying
|
||||
`tokio_postgres::Connection` to completion, which discards the async
|
||||
notification messages only `Connection::poll_message` exposes — and
|
||||
any LISTEN on a pooled connection is lost at recycle. So the shape is
|
||||
confirmed *required*, not just preferred: **a dedicated, non-pooled
|
||||
`tokio-postgres` connection per listener process**, with a forwarding
|
||||
task polling `poll_message` and fanning notifications out to
|
||||
per-subscriber tokio mpsc channels. `pg_notify` as a *sender* rides
|
||||
pooled connections freely (it is ordinary SQL). The POC asserts the
|
||||
discarded-notifications failure empirically (LISTEN issued on a
|
||||
pooled connection, verify no delivery) so the constraint is in our
|
||||
evidence base with our own test, not only the issue's word.
|
||||
- **Two implementation postures for the listener substrate, compared:**
|
||||
- **(a) hand-rolled forwarder** — the thin forwarding task above,
|
||||
ours: reconnect policy, channel re-LISTEN after reconnect, and
|
||||
backpressure are all ours (~the size of a small module; the POC #1
|
||||
Arm-B watcher taught what "ours" costs — every failure mode is
|
||||
maintained by us).
|
||||
- **(b) `postgres-notify` 0.3.8** (MIT, tokio-postgres-based,
|
||||
published, verified 2026-10-04: auto-reconnect with exponential
|
||||
backoff + jitter, multi-channel `subscribe_notify`, a
|
||||
`connect_script` hook executed on (re)connect — which is the
|
||||
LISTEN-restoration mechanism — query timeout + cancellation, and
|
||||
documented callback-panic safety properties). A turnkey shape for
|
||||
exactly this sub-module. *Caveat to verify in-probe:* docs do not
|
||||
explicitly promise subscription restoration across reconnect —
|
||||
the `connect_script` is the mechanism but restoration behavior
|
||||
(does it re-issue our LISTENs) must be verified by test, plus the
|
||||
single-maintainer/small-crate posture recorded per the AGENTS.md
|
||||
ownership rules. It also *owns* its connection for queries too
|
||||
(PGRobustClient wraps the whole client) — the engine crate would
|
||||
use it as the dedicated-listener-side client only, not the query
|
||||
path.
|
||||
Both are implemented in-probe; the comparison axes are exactly the
|
||||
ones sub-module W probes (reconnect honesty, fan-out, latency) plus
|
||||
dependency-posture cost. Preference found here is OQ-ST-04/OQ-ST-06
|
||||
input (it is a per-subsystem adopt-or-derive vote, the POC-level
|
||||
answer to the same calculus as honker-core).
|
||||
- Multi-channel: one LISTEN connection serving N channels (`LISTEN`
|
||||
accepts multiple registrations per connection) — measure whether the
|
||||
dedicated-connection-per-channel shape is ever warranted (connection
|
||||
@@ -125,7 +166,9 @@ sub-modules to validate — the comparison axis is *within* the engine
|
||||
it implies (LISTEN has no replay — anything committed while the
|
||||
listener was down is *not* re-delivered; the honest contract is
|
||||
opaque-wake + re-read, so recovery = on reconnect, wake all
|
||||
subscribers once). Re-attach storm: N listeners × M channels on one
|
||||
subscribers once — and, under listener posture (b),
|
||||
`connect_script`-driven re-LISTEN must be verified to run). Re-attach
|
||||
storm: N listeners × M channels on one
|
||||
re-connecting connection.
|
||||
|
||||
## Instruments
|
||||
@@ -151,14 +194,21 @@ sub-modules to validate — the comparison axis is *within* the engine
|
||||
concurrent claimants, offset save/read through the caller tx,
|
||||
lock acquire/release/renew with TTL, listener-reconnect
|
||||
recovery (kill listen conn, commit during the gap, verify
|
||||
reconnect wakes + state re-read correct).
|
||||
reconnect wakes + state re-read correct), and the
|
||||
**pooled-LISTEN-discard assertion** (LISTEN on a pool connection,
|
||||
verify nothing is delivered — pinning deadpool#360's behavior in
|
||||
our own evidence base; if upstream ever fixes it, the test flips
|
||||
and the fix becomes usable — but the engine design must not count
|
||||
on it).
|
||||
4. **Poll-vs-listen claim-latency probe** (`pgdiag-st-3`): enqueue→
|
||||
claim latency under both consumption postures at 1/8/32 claimant
|
||||
workers — the honest comparison that justifies (or retires) the
|
||||
LISTEN-driven claim path as the engine's default.
|
||||
5. **Pool/posture probe** (`pgdiag-st-4`): deadpool pool sizing vs
|
||||
the LISTEN connection budget (does the dedicated listen conn count
|
||||
against max_size? separate pool?); fresh-session cost re-verify
|
||||
the LISTEN connection budget (the dedicated listen connection is
|
||||
*outside* the pool by the #360 constraint — the probe validates the
|
||||
budget accounting: pool max_size + one listener connection per
|
||||
process); fresh-session cost re-verify
|
||||
(~19–25 ms per B2) against the pooling discipline; `synchronous_commit`
|
||||
per-session posture (SET on checkout vs system config) mechanics.
|
||||
|
||||
|
||||
Reference in new issue
Block a user