docs: fold verified LISTEN research into POC #2 spec

Two claims verified independently before folding: (1) deadpool-postgres
discards async notifications — upstream deadpool-rs/deadpool#360 (open,
Oct 2024) confirms the pool's connect task awaits the connection to
completion, dropping what only poll_message exposes; pooled LISTEN is
lost at recycle. The dedicated non-pooled LISTEN connection is now
upstream-verified required, not spec-preferred. (2) postgres-notify
0.3.8 exists as described (MIT, tokio-postgres, auto-reconnect with
backoff+jitter, multi-channel subscribe_notify, connect_script hook)
— admitted as sub-module L's second arm (hand-rolled forwarder vs
turnkey listener substrate), with in-probe verification required:
docs don't explicitly promise subscription restoration across
reconnect; connect_script is the mechanism; single-maintainer posture
recorded. New property test pinned: pooled-LISTEN-discard assertion
(our own evidence for the #360 behavior, flips if upstream fixes).
Probe 5 updated to budget accounting (listener conn outside the pool).
This commit is contained in:
glm-5.3-flash committed 2026-10-04 16:08:45 +00:00
1 parent e18281735e
commit 4c144f8f7f
1 file changed
+59 -9
+59 -9
View File
@@ -14,6 +14,12 @@ last_updated: 2026-10-04
> in `docs/research/poc-pg-posture-findings.md` here regardless.
> Server: dockerized postgres (POC-harness convention from the alkblobs
> POCs — `postgres:16-alpine` on :15432, config knobs stated per-probe).
> Pre-spec research input (2026-10-04): the dedicated-LISTEN-connection
> requirement (verified against deadpool-rs/deadpool#360, open upstream
> issue — pooled connections discard async notifications) and the
> `postgres-notify` 0.3.8 candidate (verified: reconnection + multi-
> channel subscribe) — both folded in below; neither's *behavioral*
> claims are trusted until probed.
## What this POC must decide
@@ -66,11 +72,46 @@ sub-modules to validate — the comparison axis is *within* the engine
### Sub-module L (listen plumbing)
- One dedicated `tokio-postgres` connection per process running
`LISTEN <channel>`; notifications surface via the connection's
`notifications()` stream. Fan-out to per-subscriber tokio mpsc
channels (the `listen()` contract: multiple subscribers per channel,
each getting every notification on that channel after attach).
- **Upstream-verified constraint (2026-10-04, from the deadpool issue
tracker — deadpool-rs/deadpool#360, open):** pooled connections
cannot carry LISTEN. deadpool's connect task awaits the underlying
`tokio_postgres::Connection` to completion, which discards the async
notification messages only `Connection::poll_message` exposes — and
any LISTEN on a pooled connection is lost at recycle. So the shape is
confirmed *required*, not just preferred: **a dedicated, non-pooled
`tokio-postgres` connection per listener process**, with a forwarding
task polling `poll_message` and fanning notifications out to
per-subscriber tokio mpsc channels. `pg_notify` as a *sender* rides
pooled connections freely (it is ordinary SQL). The POC asserts the
discarded-notifications failure empirically (LISTEN issued on a
pooled connection, verify no delivery) so the constraint is in our
evidence base with our own test, not only the issue's word.
- **Two implementation postures for the listener substrate, compared:**
- **(a) hand-rolled forwarder** — the thin forwarding task above,
ours: reconnect policy, channel re-LISTEN after reconnect, and
backpressure are all ours (~the size of a small module; the POC #1
Arm-B watcher taught what "ours" costs — every failure mode is
maintained by us).
- **(b) `postgres-notify` 0.3.8** (MIT, tokio-postgres-based,
published, verified 2026-10-04: auto-reconnect with exponential
backoff + jitter, multi-channel `subscribe_notify`, a
`connect_script` hook executed on (re)connect — which is the
LISTEN-restoration mechanism — query timeout + cancellation, and
documented callback-panic safety properties). A turnkey shape for
exactly this sub-module. *Caveat to verify in-probe:* docs do not
explicitly promise subscription restoration across reconnect —
the `connect_script` is the mechanism but restoration behavior
(does it re-issue our LISTENs) must be verified by test, plus the
single-maintainer/small-crate posture recorded per the AGENTS.md
ownership rules. It also *owns* its connection for queries too
(PGRobustClient wraps the whole client) — the engine crate would
use it as the dedicated-listener-side client only, not the query
path.
Both are implemented in-probe; the comparison axes are exactly the
ones sub-module W probes (reconnect honesty, fan-out, latency) plus
dependency-posture cost. Preference found here is OQ-ST-04/OQ-ST-06
input (it is a per-subsystem adopt-or-derive vote, the POC-level
answer to the same calculus as honker-core).
- Multi-channel: one LISTEN connection serving N channels (`LISTEN`
accepts multiple registrations per connection) — measure whether the
dedicated-connection-per-channel shape is ever warranted (connection
@@ -125,7 +166,9 @@ sub-modules to validate — the comparison axis is *within* the engine
it implies (LISTEN has no replay — anything committed while the
listener was down is *not* re-delivered; the honest contract is
opaque-wake + re-read, so recovery = on reconnect, wake all
subscribers once). Re-attach storm: N listeners × M channels on one
subscribers once — and, under listener posture (b),
`connect_script`-driven re-LISTEN must be verified to run). Re-attach
storm: N listeners × M channels on one
re-connecting connection.
## Instruments
@@ -151,14 +194,21 @@ sub-modules to validate — the comparison axis is *within* the engine
concurrent claimants, offset save/read through the caller tx,
lock acquire/release/renew with TTL, listener-reconnect
recovery (kill listen conn, commit during the gap, verify
reconnect wakes + state re-read correct).
reconnect wakes + state re-read correct), and the
**pooled-LISTEN-discard assertion** (LISTEN on a pool connection,
verify nothing is delivered — pinning deadpool#360's behavior in
our own evidence base; if upstream ever fixes it, the test flips
and the fix becomes usable — but the engine design must not count
on it).
4. **Poll-vs-listen claim-latency probe** (`pgdiag-st-3`): enqueue→
claim latency under both consumption postures at 1/8/32 claimant
workers — the honest comparison that justifies (or retires) the
LISTEN-driven claim path as the engine's default.
5. **Pool/posture probe** (`pgdiag-st-4`): deadpool pool sizing vs
the LISTEN connection budget (does the dedicated listen conn count
against max_size? separate pool?); fresh-session cost re-verify
the LISTEN connection budget (the dedicated listen connection is
*outside* the pool by the #360 constraint — the probe validates the
budget accounting: pool max_size + one listener connection per
process); fresh-session cost re-verify
(~19–25 ms per B2) against the pooling discipline; `synchronous_commit`
per-session posture (SET on checkout vs system config) mechanics.