From 80e6af0a0a2dc0843fa98020abeb6edf35b2ee6c Mon Sep 17 00:00:00 2001 From: "glm-5.3-flash" Date: Sun, 4 Oct 2026 17:25:27 +0000 Subject: [PATCH] =?UTF-8?q?docs:=20POC=20#2=20ran=20and=20passed=20?= =?UTF-8?q?=E2=80=94=20OQ-ST-03=20closed=20(per-engine=20drivers:=20tokio-?= =?UTF-8?q?postgres+deadpool=20for=20pg),=20OQ-ST-04=20ground=20complete?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Findings: unified surface holds on tokio-postgres with the transactional property intact (in-tx NOTIFY is commit-atomic; rollback drops all); LISTEN wake beats poll 5-16x at p50 with 300/300 isolated delivery; pooled-LISTEN discard (deadpool#360) verified and pinned as our own test; postgres-notify 0.3.8 evaluated and passed over (lazy reconnect, no initial-connect script, unquoted identifier LISTENs) in favor of the ~90-line hand-rolled forwarder with test-pinned pitfalls. sqlx PgListener fallback retired unfired. --- docs/research/phase-0.md | 126 ++++-- docs/research/poc-pg-posture-findings.md | 484 +++++++++++++++++++++++ docs/research/poc-pg-posture-spec.md | 11 +- 3 files changed, 584 insertions(+), 37 deletions(-) create mode 100644 docs/research/poc-pg-posture-findings.md diff --git a/docs/research/phase-0.md b/docs/research/phase-0.md index 4f69680..d16e490 100644 --- a/docs/research/phase-0.md +++ b/docs/research/phase-0.md @@ -1,14 +1,11 @@ --- status: draft -last_updated: 2026-10-04 (consumer inventory landed — OQ-ST-01 answered -per-feature from the paused consumers' documents; phase-0 plan step 1 -done; streams upgraded to in-scope by operator-authority record; OQ-ST-02 -resolved: reactive-core + engine crates, operator decision; OQ-ST-03's -SQLite half resolved by POC #1 (honker-core on our rusqlite — findings -in poc-sqlite-posture-findings.md); OQ-ST-04/05/06 carry POC #1 input; -POC #2 (pg posture: LISTEN/pool/tx-seam) specified — running it closes -OQ-ST-03. Interface finding and driver tension from 2026-10-03 remain -trusted-but-unverified working input except where POC #1 verified them.) +last_updated: 2026-10-04 (POC #2 ran and passed — OQ-ST-03 resolved: +per-engine drivers, rusqlite+honker-core for SQLite / tokio-postgres+ +deadpool-postgres for Postgres, findings in poc-pg-posture-findings.md; +OQ-ST-04 now has measured ground on both engines, wake contract +numbers + tx-seam shapes + listener-recovery semantics recorded from +POC #2; OQ-ST-05/06 carry both POCs' per-subsystem votes) --- # alkstore — Phase 0 (Exploration) @@ -483,16 +480,31 @@ inherited watcher is tighter than a re-derived one (p50 1.40 vs 2.15 ms, max 29 vs 172 ms, with battle-tested failure handling), and the `.so` runtime dependency is packaging cost A doesn't pay for no compensating advantage. The transactional contract holds identically -on both (it is SQLite's property, not the posture's). The Postgres -half remains the open part of OQ-ST-03 (tokio-postgres per the POC -evidence base is the working lean; the OQ is not closed until the pg -side's LISTEN/pool/tx-seam story is validated the same way). New +on both (it is SQLite's property, not the posture's). New constraints recorded for the engine crate regardless: honker-core=0.5.0 pins rusqlite ^0.40.1 whose rustc requirement (≥1.99) is a deployment note, and mixed rusqlite+sqlx binaries currently need a vendored one-line libsqlite3-sys patch — OQ-ST-02's per-engine-crate split is what keeps the engine binary single-driver. +**Resolved (2026-10-04, both halves): per-engine drivers — +rusqlite + honker-core for the SQLite engine; tokio-postgres + +deadpool-postgres for the Postgres engine.** The Postgres half is +POC #2's (findings: `poc-pg-posture-findings.md`): every gate +condition fired affirmatively on the tokio-postgres posture — +unified surface (enqueue_tx/claim/ack/notify_tx/listen/stream-offset/ +locks) with the transactional property intact (in-tx NOTIFY delivers +only on commit; rollback drops all), LISTEN-driven wake beats poll +5–16× at p50 (3–6 ms vs 32–50 ms end-to-end claim latency; isolated +wakes 1.1 ms p50, 300/300 delivered), the tx-seam is *simpler* on pg +than SQLite (tokio-postgres Client is Send+Sync — no spawn_blocking +rigging), pooled-LISTEN discard (deadpool#360) verified and pinned, +and `postgres-notify` 0.3.8 evaluated and passed over (derive-not- +adopt: lazy reconnect, initial-connect script gap, identifier +quoting; the hand-rolled forwarder is ~90 lines and pitfalls are +pinned by tests). The sqlx `PgListener` fallback never fired and is +retired. + ### OQ-ST-04: The reactive abstraction — what does the unified notify surface look like? The two engines' mechanisms are structurally different: SQLite = @@ -545,13 +557,34 @@ lease whose ops each ride `spawn_blocking` (the connection is not `Sync`; holding it across await points is wrong) — Phase 1's contract work starts from that shape plus the POC's `TxHandle` sketch (with its two recorded frictions: the `as_any_mut` downcast and the -thread-affinity of rusqlite tx ops). What remains open is everything -Postgres (transaction-scoped LISTEN over a pool, and the delivery- -guarantee/unification bullets above). +thread-affinity of rusqlite tx ops). + +**POC #2 input (2026-10-04, closing the pg side):** the same +opaque-wake + re-read contract holds on Postgres and the wake +mechanism is *already* the shape the contract wants — LISTEN delivers +push (~1.1 ms p50, 300/300 isolated, no coalescing needed), has no +replay (gap commits recovered by the listener broadcasting a synthetic +reconnect-wake on a reserved channel, verified: subscribers wake and +re-read state correctly through a killed-connection recovery), and +notify is commit-atomic natively (delivers only at tx commit; rollback +drops it — the exact analogue of honker's notify-in-tx property). The +tx-seam resolves to the same *shape* both engines: caller-held tx +handle (`*_tx` methods on the handle); pg's instance is async-native +(tokio-postgres Client is Send+Sync — the handle holds the pooled +connection directly, no spawn_blocking), SQLite's is a bridged writer- +slot lease. The core-crate `TxHandle` trait from POC #1's sketch +stands unchanged; the per-engine difference is bridging mechanism, not +trait shape. Delivery-guarantee contract is now measurable, native on +both sides: notify = fire-and-forget (commit-atomic, no replay), +streams = durable with explicit offsets. What remains open on OQ-ST-04 +is the contract-pinning work itself (which parts of the honker-rs +surface become contract, per the inventory rows) — paper work over a +now-complete evidence base, for Phase 1. Open; this is the second central research question, coupled to OQ-ST-03 -(the driver determines what LISTEN plumbing exists) — SQLite side now -de-risked, Postgres side is the remaining shape work. +(the driver determines what LISTEN plumbing exists) — **both engine +sides now de-risked (POC #1 SQLite, POC #2 Postgres); what remains is +the contract-pinning paper work.** ### OQ-ST-05: Queue semantics — adopt, fork, or re-derive? @@ -571,11 +604,15 @@ Open; inputs: OQ-ST-01's scope vote on queues (inventory: documented need) + OQ-ST-03 resolution. **POC #1 input (2026-10-04):** on the SQLite side the adopt question dissolved — honker-core is consumed as a published-library dependency (the inventory-confirmed feature rows ride -its machinery; OQ-ST-06 holds the fork-vs-reference question). The -Postgres side keeps this OQ's full option space (pgboss-rs vs fork vs -re-derive on tokio-postgres) — with the POC-confirmed constraint that -whichever choice is made, reactivity is built by this crate (the -verified pgboss-rs LISTEN/NOTIFY gap stands). +its machinery; OQ-ST-06 holds the fork-vs-reference question). **POC #2 +input (2026-10-04):** on the Postgres side the re-derive posture is +strengthened — the minimal queue table + `FOR UPDATE SKIP LOCKED` claim ++ LISTEN wake is the driver-coupled hard part and it was proven in +~40 lines of SQL over the pool (all claim/atomicity properties pass); +reactivity is built by this crate either way (the verified pgboss-rs +LISTEN/NOTIFY gap stands). The remaining OQ-ST-05 question is the +*semantics depth* (retry/backoff/dead-letter/sweep design on that +ground), with pgboss-rs as schema/design reference. ### OQ-ST-06: Honker relationship — reference, fork, or vendor? @@ -609,6 +646,14 @@ reference-usage posture works as-is. The remaining fork trigger would be the Phase-1 quality read (the watcher/transactional core assessment) or a needed change upstream won't take — the calculus is unchanged in kind, but the *default* posture is now evidenced: depend on the published crate. +**POC #2 input (2026-10-04):** the pg-side dependencies are +published-library use as-is (tokio-postgres 0.7.18 + deadpool-postgres +0.14.2: clean, zero conflicts, actively maintained); `postgres-notify` +0.3.8 evaluated in-probe and passed over (derive-not-adopt — lazy +reconnect, no connect_script on initial connect, unquoted identifier +LISTENs, single-maintainer posture; the hand-rolled ~90-line forwarder +with test-pinned pitfalls is the preferred shape; this is the OQ-ST-06 +calculus applied per-subsystem, recorded, not a Phase 0 ADR). ### OQ-ST-07: SQLite-side scope — loadable extension, embedded rusqlite, or both? @@ -646,7 +691,7 @@ land in `docs/research/` here. Named per the OQ each feeds: | # | POC | Spec | Findings | |---|---|---|---| | 1 | SQLite engine posture: honker-core-on-rusqlite vs honker-extension-over-sqlx (async seam, watcher, transactional contract, packaging, interop) | [poc-sqlite-posture-spec.md](poc-sqlite-posture-spec.md) | [poc-sqlite-posture-findings.md](poc-sqlite-posture-findings.md) — **passed** (verdict: Arm A; ran 2026-10-04) | -| 2 | Postgres engine posture: LISTEN/NOTIFY plumbing, tx-seam over the pool (caller-owned tx vs closure-scoped), wake-vs-poll claim latency, reconnect recovery | [poc-pg-posture-spec.md](poc-pg-posture-spec.md) | — (specified 2026-10-04; runs the POC #1 harness's pg twin) | +| 2 | Postgres engine posture: LISTEN/NOTIFY plumbing, tx-seam over the pool (caller-owned tx vs closure-scoped), wake-vs-poll claim latency, reconnect recovery | [poc-pg-posture-spec.md](poc-pg-posture-spec.md) | [poc-pg-posture-findings.md](poc-pg-posture-findings.md) — **passed** (verdict: tokio-postgres+deadpool, hand-rolled listener, caller-tx seam; ran 2026-10-04) | ## Phase 0 plan @@ -665,24 +710,27 @@ Expected sequence (deliberately rough): - ~~SQLite driver posture~~ — **POC #1 passed (2026-10-04)**: posture 1 (honker-core on our rusqlite); findings + constraint notes in the register row and OQ-ST-03. - - Postgres side of OQ-ST-03 (tokio-postgres LISTEN/pool/tx-seam - validation) — **POC #2 specified**: - [poc-pg-posture-spec.md](poc-pg-posture-spec.md) — one driver - posture (tokio-postgres + deadpool, the POC #5/#7-validated stack) - with three sub-modules (listen plumbing / tx-seam shapes / - wake-vs-poll parity); on pass, OQ-ST-03 closes with per-engine - drivers. + - ~~Postgres side of OQ-ST-03 (tokio-postgres LISTEN/pool/tx-seam + validation)~~ — **POC #2 passed (2026-10-04)**: tokio-postgres + + deadpool posture validated end-to-end; OQ-ST-03 **closed** with + per-engine drivers; findings in the register row and OQ-ST-03/04. - OQ-ST-04's contract pinning — the honker-rs surface (§Interface finding) is the concrete starting artifact: pinning its contract costs less and is more honest than inventing a parallel shape — now scoped against the inventory's confirmed features rather than the full honker menu, shaped as the core-crate trait surface per - OQ-ST-02's split, with the POC's `TxHandle`/writer-slot-lease - shape as the SQLite-seam starting point. + OQ-ST-02's split, with the tx-seam shape **resolved by both POCs** + (caller-held tx handle; SQLite bridges via writer-slot lease + + spawn_blocking, Postgres holds the pooled connection directly + — same shape, different bridging) — **paper work remains, over a + complete evidence base.** 3. Ownership decisions (OQ-ST-05/06) — adopt/fork/derive per subsystem, - after the driver and shape questions narrow the option space - (OQ-ST-06's default is now evidenced as published-library use; - OQ-ST-05 waits on the pg side of OQ-ST-03). + ~~after the driver and shape questions narrow the option space~~ — + the option space is narrow now (both POCs ran; per-subsystem votes + recorded at OQ-ST-05/06): the remaining work is the *semantics-depth* + design inputs (OQ-ST-05: retry/dead-letter/sweep on the pg ground + POC #2 proved) and the Phase-1 quality read (OQ-ST-06's fork trigger + assessment). 4. Converge; Phase 1 opens with the ADR backlog this register becomes. ## References @@ -721,6 +769,12 @@ Expected sequence (deliberately rough): - alkstore-sqlite-posture-poc — `/workspace/alkstore-sqlite-posture-poc` (standalone POC crate, published-deps-only): POC #1's code — both arms end-to-end, property tests, seam/watcher probes. +- alkstore-pg-posture-poc — `/workspace/alkstore-pg-posture-poc` + (standalone POC crate, published-deps-only): POC #2's code — + the pg engine posture end-to-end (engine + hand-rolled listener + + postgres-notify wrapper), 11-test contract suite, seam/wake/ + burst/claim/pollvlisten/pnlisten probes; harness server + `pglo-poc` (postgres:16-alpine, :15432). - alktty — `/workspace/@alkdev/alktty` (architecture reviewed): REQ-TTY-01 (`docs/architecture/tty-backend.md`) — the async-facing-trait + sync-bridge posture ("backends are not required to be natively diff --git a/docs/research/poc-pg-posture-findings.md b/docs/research/poc-pg-posture-findings.md new file mode 100644 index 0000000..88da7ca --- /dev/null +++ b/docs/research/poc-pg-posture-findings.md @@ -0,0 +1,484 @@ +--- +status: findings +title: "POC #2 findings — Postgres engine posture: LISTEN/NOTIFY wiring, the pool tx-seam, and reactive parity on tokio-postgres" +last_updated: 2026-10-04 +--- + +# POC #2 findings — Postgres engine posture + +**Verdict: PASS — tokio-postgres + deadpool-postgres carries the unified +surface with the transactional property intact, and LISTEN beats poll +decisively. OQ-ST-03 closes with per-engine drivers (rusqlite+honker-core +for SQLite, tokio-postgres+deadpool for Postgres); OQ-ST-04's +contract-pinning now has measured ground on both engines.** All three +sub-modules returned affirmative findings; the sqlx `PgListener` fallback +was never needed and stays retired. Every shared contract property from +POC #1 holds on Postgres, plus the pg-native facts (in-tx NOTIFY delivers +at commit; pooled connections cannot deliver notifications — deadpool#360 +— pinned as our own test). + +POC code: `/workspace/alkstore-pg-posture-poc` (standalone crate, +published-deps-only, same harness conventions as POC #1). Server: +dockerized `postgres:16-alpine` on :15432 (`pglo-poc` container, db +`blobs`, user/password `postgres/poc` — the alkblobs POC harness +convention). Toolchain: rustc/cargo 1.99.0. + +## Versions / dependency surface (verified 2026-10-04) + +- `tokio-postgres` 0.7.18 (updated 2026-09-03, actively maintained, + MIT/Apache-2.0). `with-serde_json-1` feature for JSONB/Value binding. +- `deadpool-postgres` 0.14.2 (pool + per-connection statement cache; + `RecyclingMethod::Fast` does not clobber LISTEN state — it recycles + with a health query only — which is also why a LISTEN-registered + connection *returning to the pool* keeps its registration but loses + the delivery path; see the #360 section). +- `postgres-notify` 0.3.8 (MIT, single maintainer, 5.9k downloads total, + ~1.5k recent; published 2026-01-07 by the sole owner). Verified + in-probe; quirks below under Sub-module L (b). +- No version conflicts, no vendoring needed. The pg engine is + single-driver by construction (per OQ-ST-02's split); no + libsqlite3-sys-style link collision exists on this side. + +## Sub-module L (listen plumbing) — findings + +### The #360 constraint: verified empirically, pinned as a test + +The spec's requirement (dedicated non-pooled LISTEN connection) was +asserted with our own probe rather than trusted from the issue tracker: + +- `tests/contract.rs::pooled_listen_registers_but_cannot_deliver` — + three-part honest probe: (1) `LISTEN` issued through a pooled + deadpool connection **registers** server-side (the session shows it + in `pg_listening_channels()`), (2) the registered notification is + **never deliverable** (no delivery API exists on the pooled object — + verified by source read: deadpool 0.14.2's connect task + `spawn(async move { connection.await })` plain-awaits the + tokio-postgres `Connection` future, whose `Future` impl discards all + async messages; `ClientWrapper` exposes no notification surface + whatsoever), and (3) the control — a dedicated connection driven by + `poll_message` — delivers the identical notification immediately. +- Source-verified against the installed 0.14.2 (not just the issue): + `deadpool-postgres/src/lib.rs` lines ~221 and ~311 — + `create()` spawns the connection task with a bare `connection.await`, + and `ClientWrapper::drop` aborts it. **deadpool#360 remains open.** + +Consequence recorded for the engine: the listener connection is a +*separate budget line*, outside the pool, per process. Pool +`max_size + 1` per LISTEN-ing process (pgdiag-st-4 verified the +accounting end-to-end: pool size 6 observed + 1 listener + probe conns += exactly the server-side count). + +### Hand-rolled forwarder (a) — built and measured (~90 lines) + +`src/listen.rs`: dedicated connection per listener process, a spawned +task driving `Connection::poll_message` and fanning out into a tokio +`broadcast` channel (bounded 1024; lag is surfaced, not silent — the +bridge logs `Lagged(n)` and keeps going). Reconnect policy: exponential +backoff 50 ms → 2 s cap; re-LISTEN after every reconnect (the LISTEN +SQL is re-issued from the channel list on each reconnect), plus a +**synthetic reconnect wake** on a reserved channel after each +successful re-LISTEN — that synthetic wake is how the no-replay hole +is communicated to subscribers (below). + +**Two deadlocks found and fixed during the build — both are exactly +the failure modes a hand-rolled forwarder owns and both are now pinned +by the passing tests:** + +1. **Query-vs-poll starvation deadlock.** With the raw `Connection` + polled for async messages, any *client* query (`batch_execute`) + awaits its response through the same connection — if the poll loop + task is spawned *after* the first client query, the query never + completes. The poll loop must be running *before* the first query. + (The initial version deadlocked the LISTEN itself; caught in-probe, + restructured, pinned by every listen test passing.) +2. **Client-drop closes the connection.** Dropping + `tokio_postgres::Client` terminates the server session even if the + `Connection` task still runs. A long-lived listener must keep the + `Client` alive for the listener's lifetime (`std::mem::forget` in + the one-shot probe path; ownership in the shared listener path). + Verified in-probe (`examples/raw_listen_selfcheck.rs` etc). + +Neither failure mode was visible from the docs; both are the "every +failure mode is maintained by us" cost the spec predicted for the +hand-rolled posture. They are contained (~90 lines, two pitfalls) and +now test-pinned. + +### `postgres-notify` 0.3.8 (b) — built, verified working, with four documented quirks + +`src/pgnotify.rs` wraps `PGRobustClient` as the listener substrate. +Verified in-probe (`pnlisten` command + `examples/pn_reconnect_flow.rs`): + +1. **Initial connect does NOT execute the connect_script** (source: + `PGClient::connect` never runs it; only `reconnect()` does). The + initial LISTEN must be issued explicitly via `subscribe_notify` — + which conveniently registers channels into the config's + `subscriptions` set that the reconnect script re-issues. The spec's + caveat ("restoration behavior must be verified by test") was + sharper than doc'd: **restoration works on reconnect, but the first + LISTEN is entirely yours.** +2. **Reconnect is lazy.** The polling task reports a drop via a + `Disconnected` callback, but `reconnect()` only runs when the next + *query* fails its connection check. A killed listener connection + sits disconnected until the client issues something. The engine's + listener loop would need a periodic keepalive/trigger query to + make recovery timely — more machinery ours-with-a-dependency, not + fewer. +3. **Auto-reconnect + re-LISTEN verified:** backend kill (by pid, via + `pg_terminate_backend`) → `Disconnected` callback → trigger query → + `Reconnect` (attempt #1, backoff 500 ms base) → `Connected` → LISTEN + re-issued from the script → post-reconnect notification delivered. + `PASS pnlisten: reconnect + connect_script re-LISTEN verified`. +4. **`issue_listen`/`issue_unlisten` do not quote identifiers** — + `LISTEN poc:pn` is a syntax error at the server. Channel names + through this crate must be quote-free (letters/digits/underscore). + Also `application_name` (config field) is interpolated into the + connect *script*, so it does not apply to the initial connection + either — server-side kill-by-name targeting misses the initial + connection. + +Comparison verdict (preference recorded, this is OQ-ST-04/OQ-ST-06 +input, not a Phase 0 ADR): **(a) hand-rolled forwarder wins** — the +hand-rolled version is ~90 lines with immediate reconnect (no lazy +trigger dependency), full identifier control (quoted channels — our +`poc:` prefixes work), no dependency posture (single-maintainer crate +with 1.5k recent downloads and behavior quirks the docs misdescribe), +and the two deadlock pitfalls are already *learned and pinned*. The +(b) crate would still force us to own a keepalive + reconnect trigger +loop on top (its reconnect laziness), eroding most of the "turnkey" +advantage. Posture (b) stays recorded as the fallback if a future +upstream fix lands (channel quoting + eager reconnect) and the +maintenance posture improves. + +### Multi-channel and fan-out + +- One LISTEN connection serves N channels: verified (`LISTEN a; + LISTEN b; LISTEN c;` on one connection, all three deliver — + `shared_listener_multichannel.rs`, and every contract test runs + three channels through one listener). No per-channel connection is + ever warranted at this scale; fan-out to per-subscriber receivers is + a broadcast split, ~zero cost. +- 4 subscribers on one channel, every one wakes: verified + (`listen_wake_fanout_and_burst`). + +### Payload boundary + +`pg_notify` ≤ 8000 bytes confirmed. Probe: 7700-byte payload delivers; +the engine checks the limit client-side and returns a typed +`PayloadTooLarge` error before the round-trip for 8100 bytes +(`notify_payload_boundary`). Honest contract recorded: notify payloads +are hints; large payloads ride a table row with the id in the +notification (the honker-outbox shape, pg-native). + +## Sub-module T (tx-seam over the pool) — findings + +Both candidate shapes implemented: + +- **(a) caller-owned tx handle** (`OwnedTxHandle`): the trait's + `begin_tx` checks a pooled connection out, issues `BEGIN`, and hands + the caller a handle owning the `Object`. `*_tx` methods downcast the + handle (`as_any_mut`, same as POC #1's sketch) and issue their SQL + through it. Commit/rollback return the object to the pool. +- **(b) closure-scoped** (`with_tx`): pool checkout + BEGIN + closure + + COMMIT/ROLLBACK inside one future; the caller never holds the + transaction. + +**The load-bearing structural finding: the whole SQLite side's +thread-affinity rigging does not exist on Postgres — deliberately.** +POC #1's SQLite tx handle needed an `Arc>>` +with every op round-tripping `spawn_blocking`, because rusqlite's +connection is neither `Send`-across-await nor `Sync`, and the handle +had to be callable from async trait methods. tokio-postgres's `Client` +(and deadpool's `Object`) **is `Send + Sync`** (request-channel +design: `Arc` + unbounded mpsc), verified by a +compile-time probe in this POC (`tmp sendsync` check: `Client` and +`Object` both asserted `Send + Sync`). So shape (a)'s handle is +`Option` held **directly**, ops are straight `.await`s, no +mutex, no spawn_blocking, no per-op hop cost. The `as_any_mut` +downcast friction from POC #1 remains (same trait shape), but the +thread-affinity friction does not carry over. This is the honest +per-engine delta the unified trait must absorb: same *seam shape*, but +the PostgreSQL handle is free to be held across await points while the +SQLite handle must own a blocking thread. + +- **Commit-atomicity property holds and was verified directly:** + `NOTIFY` issued inside the caller's tx (via `pg_notify($1,$2)`) is + delivered to listeners **only at commit** of the sending transaction + (`notify_fires_only_on_commit`: notify in an open tx → listener + silent for 400 ms → commit → delivery < 3 s; `tx_enqueue_rollback_… + drops_all_ghost_free`: rollback drops the job row, the business row, + *and* the notification — no ghosts anywhere). This is pg's native + statement-vs-commit semantics, exactly the analogue of honker's + notify-in-tx atomicity — no emulation needed. +- **Read-your-writes under `read committed`** (the honest default): + verified both directions (in-tx reads see own writes; post-commit + reads from other pool connections see the commit). No surprises. +- **Shape (a) vs (b) for the engine posture:** (a) is the viable + contract seam (long transactions compose naturally; the caller drives + begin/commit — which is exactly what POC #1's `TxHandle` sketch + assumed). (b) works for self-contained ops but composes poorly with + the caching-subscriber pattern (state can't outlive the closure). + Preference recorded: **(a)**, with (b) available as a wrapper — + note (b) is *implementable over* (a) trivially but not vice versa. + Deadlock risk on (a): none found (the connection is checked out + exclusively; no re-entrant checkout happens inside the handle's ops). +- **Prepared-statement discipline:** deadpool's per-connection + statement cache is keyed per pooled connection and the cache is + re-usable across checkouts; with `RecyclingMethod::Fast` there is no + `DISCARD ALL`/`DEALLOCATE` recycling, so the "prepared statement + s1 does not exist" failure (alkblobs POC-era) does not recur: the + engine's claim SQL is prepared implicitly through the same pooled + connection repeatedly during the probes with zero errors (n=3000 + seam iterations ×2 sync postures + pollvlisten's 12 × 60 claims). + The queue SQL is ordinary SQL — no manual prepare discipline needed + at POC scale; noting the engine crate should keep using + `prepare_cached` if it wants explicit caching (deadpool's + `ClientWrapper::prepare_cached` exists for exactly this). + +## Sub-module W (wake-vs-poll parity) — findings + +### LISTEN-driven wake latency (pgdiag-st-2) + +300 notify-commits at 25 ms spacing, listener attached before first +commit, delivery = wake-on-channel after commit: + +- **p50 = 1.10–1.15 ms, p99 = 1.42–1.45 ms, max = 1.49–2.21 ms, + 300/300 delivered (0 missed).** + +Cross-engine relative claim (contract input): this is *the same +sub-2.5 ms* order as POC #1's SQLite watcher at its 1 ms cadence (p50 +1.40 ms arm-A / 2.15 ms arm-B) — the wake layer is not the Postgres +engine's weakness; it's *tighter* than poll-based wake and +structurally push (nothing polls: the server delivers). + +- **Burst-stress (30 rapid commits, no spacing): 30/30 delivered** + (both release and debug runs; commit burst executes at ~200–520 + commits/s, all 30 per-payload notifications delivered). LISTEN does + not coalesce the way a poller does — each committed NOTIFY is + delivered. The per-payload assert (not per-count) is what our + contract tests pin: `seen.contains(payload_i)` for every i. +- **Re-attach storm** (N listeners × M channels): the shape is one + listener connection with N broadcast subscribers — re-attach is a + broadcast re-subscribe, no server round-trips; the reconnect path + re-issues M LISTENs on the one connection (verified in the reconnect + test; the storm is O(M) SQL statements once per reconnect). + +### LISTEN connection failure and the replay hole + +`listener_reconnect_recovery_and_replay_hole`: listener connection +killed mid-subscription (`pg_terminate_backend` targeted at the +listener's application_name — the listener sets +`application_name('alkstore-pg-poc-listener')` precisely to be +kill-targetable and diagnosable) → + +- the commit during the gap is **not re-delivered** (LISTEN has no + replay — the honest contract the spec predicted holds); +- the listener reconnects (exponential backoff 50 ms base) and + re-LISTENs; +- **a synthetic wake is broadcast on a reserved channel** (`__listener_ + reconnected__`) when the connection re-establishes — this is the + mechanism that turns the no-replay hole into a *recoverable* event: + subscribers get `Disconnected`-ish signal via the wake and re-read + state. Verified: reconnect wake arrives; post-reconnect deliveries + work; state re-read is complete (all three commits — pre-kill, + in-gap, post-reconnect — visible to a fresh consumer read despite + the in-gap notification never being delivered); +- the assertion `!saw_replay` pins the no-replay honesty. + +This is the "on reconnect, wake all subscribers once" recovery the +spec sketched, now implemented and pinned. The engine's subscriber +contract inherits it: wake = opaque hint + re-read; gaps surface as a +reconnect wake, not as lost silence (the SQLite side's missed-wake +stress tests the overtriggering twin of this contract). + +### Poll-vs-listen claim latency (pgdiag-st-3) — the decisive measurement + +Enqueue→claim latency, one job in flight, 1/8/32 claimants, both +postures, release build (debug numbers were consistent): + +| posture | claimants | p50 ms | p90 | p99 | max | +|---|---|---|---|---|---| +| poll (50 ms interval) | 1 | 50.0 | 50.8 | 65.3 | 65.3 | +| poll | 8 | 32.5 | 36.7 | 46.5 | 46.5 | +| poll | 32 | 40.4 | 49.3 | 51.1 | 51.1 | +| **listen** | 1 | **6.0** | 6.4 | 8.6 | 8.6 | +| **listen** | 8 | **3.1** | 7.0 | 7.6 | 7.6 | +| **listen** | 32 | **5.7** | 6.6 | 10.6 | 10.6 | + +LISTEN-driven claiming beats even a *50 ms* poll by **5–16× at p50**; +against pgboss-style 1 s default poll intervals the gap would be +~100–300×. The claimants-scaling is flat-to-better (multiple claimants +absorb wakeup jitter). **The LISTEN-driven claim path with re-poll +safety net is the engine's default consumption posture; poll-only +(interval-tunable) remains the fallback when a LISTEN connection is +unavailable/not wanted.** This decides the pg half of the +reactivity-vs-pgboss gap: the push channel pgboss-rs lacks, this +engine gets, measured. + +## The seam/cost probe (pgdiag-st-1) + +`begin + enqueue_tx + commit` per iteration through the pool, n=3000, +release, sequential (POC #1's seam workload's pg twin): + +- **`synchronous_commit=on` (ship config): p50 = 2.40 ms, p90 = 2.73, + p99 = 3.00, max = 6.93 ms** (0 fails); +- **`synchronous_commit=off`: p50 = 2.26 ms, p90 = 2.66, p99 = 3.02, + max = 40.9 ms** (0 fails). + +Reading (shape, not absolute — different machines/databases): + +- The pg seam is flat and tight at p50–p99 (~2.3–3 ms) — the pool + checkout + BEGIN + INSERT + COMMIT round-trip cost with the + per-connection statement cache warm. The `=off` knob does not + *materially* change p50–p99 here (the local docker fsync floor is + small; the knob would matter more on replicated/farther storage). + The max-tail differences (~7 ms vs ~41 ms) reflect exactly that: + `=on` pays fsync per commit predictably; `=off` batches, occasionally + longer flush batches. **`synchronous_commit` remains a per-session + durability knob with honest trade shape — measured, both postures + reported; per-session SET mechanics verified (pool.rs sets it via + the connect options; `set_sync_commit_session` via SET).** +- Cross-engine: SQLite's seam (same workload shape) was 0.354/0.183 ms + (arm A) vs sqlx's 0.707 ms in POC #1 — i.e., Postgres' transactional + seam on the network stack costs ~3–10× the SQLite one here, which + is the honest expected order (network round-trips vs in-process + file writes). The unified trait absorbs this as "engine costs + differ by orders for the same op" — the per-engine performance + expectations must not be baked into the shared contract numbers. +- Raw floor re-measure (methodology cross-check): prepared SELECT + round-trip p50 = **284 µs** (p99 344 µs) — consistent with the + POC #5-era ~150–500 µs claim, same box, docker-hop counted (the + docker bridge adds overhead vs the earlier bare box; the *shape* + cross-checks). Fresh-session connect: p50 ≈ 23 ms — matches the + ~19–25 ms B2-era figure and validates the pooling discipline's + value (~2.4 ms amortized vs 23 ms fresh). + +## Property tests (the contract suite) — all green + +`tests/contract.rs` — 11 tests, pg twin of POC #1's suite, all shared +assertions reused: + +| Property | Result | +|---|---| +| enqueue+business+notify in one caller tx; rollback drops all (no ghosts: queue, stream, notification) | ✅ | +| exactly-once claim under 4 concurrent producers + 4 concurrent claimants (SKIP LOCKED claim SQL) | ✅ | +| 40 jobs 4 producers → 40 unique claims + full ack + zero ghosts | ✅ | +| stream read/save offset through caller tx; rollback drops; offset advances | ✅ | +| lock acquire/release/renew with TTL; expiry re-acquirable | ✅ | +| fan-out: 4 subscribers × 1 channel, every one wakes | ✅ | +| burst 30 commits: per-payload 30/30 delivered; state re-read complete | ✅ | +| listen-wake queue claim path (commit → notify → wake → claim) | ✅ | +| **in-tx NOTIFY delivers only on commit** (not at statement time) | ✅ | +| pooled-LISTEN-discard (#360): registers but never delivers; dedicated control delivers | ✅ | +| listener-kill → reconnect-wake → no replay of gap notify → post-commit delivery → state re-read complete | ✅ | +| notify payload boundary (7700 ok; 8100 rejected client-side) | ✅ | +| read-your-writes under read committed (in-tx + post-commit) | ✅ | + +## Decision gate — evaluated + +- **Contract:** every shared property holds. The commit-atomicity + property via in-tx NOTIFY is verified natively (`notify_fires_only_ + on_commit` + rollback-ghost tests); exactly-once claim holds under + concurrency via `FOR UPDATE SKIP LOCKED` claim SQL. ✅ +- **Seam:** the caller-tx shape survives with measured costs (2.4 ms + p50 ship config). Preference for shape (a) recorded with evidence. + The `*_tx` seam is *cheaper to hold* on pg than on SQLite (Send+Sync + client; no spawn_blocking round-trips). ✅ +- **Wake:** LISTEN plumbing is robust (reconnect honest, synthetic + reconnect-wake closes the replay hole operationally, fan-out + correct) and *decisively beats poll* (3.1–6 ms vs 32–50 ms at + comparable claimant counts; 300/300 isolated wakes at 1.1 ms p50). + ✅ + +**Verdict: PASS on all three gate conditions. OQ-ST-03 closes with +"per-engine drivers: rusqlite+honker-core (SQLite) / tokio-postgres + +deadpool-postgres (Postgres)". The sqlx PgListener fallback posture is +retired — no condition requiring it fired.** + +## What feeds where + +- **OQ-ST-03 (Postgres half): resolved** — tokio-postgres + deadpool as + recorded; per-engine-crate split (OQ-ST-02) keeps binary drivers + single, eliminating the link-collision class of problems the SQLite + arm documented. +- **OQ-ST-04 (reactive contract, pg side): measured ground.** + - Wake contract numbers: LISTEN p50 ≈ 1.1 ms (push), claim-latency + p50 ≈ 3–6 ms end-to-end; poll-only fallback ≈ interval-bound. + - Listener semantics: no replay; recovery = synthetic reconnect-wake + + re-read; subscriber contract stays opaque-wake + re-read + (identical shape to SQLite's overtriggering watcher — the two + engines now share the *same* wake contract). + - Delivery-guarantee split confirmed native: notify = fire-and-forget + (commit-atomic, at-most-once per listener session, no replay); + streams = durable with offsets. The trait must NOT promise replay + under `listen()` — both engines are honest only as opaque wake. + - Tx-seam: caller-held tx handle (`*_tx` on a handle) is the seam on + both engines; pg's handle is async-native (Send+Sync client) while + SQLite's is a bridged lease. Same shape, different bridging + mechanism — the core-crate `TxHandle` trait from POC #1's sketch + stands, with the pg handle as the trivial-instance case. +- **OQ-ST-05 (queue semantics): re-derive-on-tokio-postgres posture + strengthened** — the minimal queue table + `SKIP LOCKED` claim + + LISTEN wake is ~40 lines of SQL around the pool, the exactly-once + property rides pg's native serialization semantics, and the + pgboss-rs schema family remains design-reference (its *states* + matter only for retry/dead-letter depth, which is OQ-ST-05's design + work; the POC's minimal table sufficed for every property). +- **OQ-ST-08 (multi-host): the pg engine is natively multi-host** — + the property tests ran all-through-network (docker bridge), the + listener/wake machinery is connection-based (per-process), and + nothing assumes single-host. The engine-capability question remains + for the trait surface, but the pg side has no single-host assumption + to remove. +- **Dependency postures (OQ-ST-06-adjacent, per-subsystem votes):** + honker-core = published-library (POC #1); deadpool-postgres + + tokio-postgres = published-library (this POC — clean, zero + conflicts); postgres-notify = **derive-not-adopt** (the 0.3.8 quirks + + single-maintainer posture + our forwarder being already-learned + territory; fallback if upstream improves). + +## Honest caveats + +- Same single-box posture as POC #1 (8-core shared dev box, dockerized + server over the docker bridge). Absolute numbers carry the docker-hop + (~284 µs raw floor here vs the ~150 µs-era bare-metal figure); + relative/shape claims are the deliverable and were consistent across + re-runs (debug and release). +- The LISTEN re-attach "storm" is tested as one-connection-N-channels + re-listen (our engine shape); a *per-listener-connection* storm (N + processes each × M channels at reconnect) is bounded by N × M + separate LISTEN statements — plausible but not measured (the engine + shape never exercises it; recording as assumption with the reasoning). +- The queue table is the minimal spec-out-of-scope shape (enqueue, + claim, ack, visibility column present but not scheduled/retried). + Retry/backoff/dead-letter depth remains OQ-ST-05's design work; the + *transactional and claim* properties this POC pins are the hard + driver-coupled part and they pass. +- `postgres-notify` verification was against 0.3.8's published source + (crate downloaded into the cargo registry cache); behavior quirks + (lazy reconnect, initial-connect script skip, identifier quoting) + are recorded from that source plus in-probe verification, cited as + of 2026-10-04. +- The synthetic reconnect-wake channel is a reserved name + (`__listener_reconnected__`) — engine consumers choosing their own + channel names must not collide; a prefix convention or a separate + meta-channel namespace is Phase 1 contract surface. + +## Artifacts + +- POC crate: `/workspace/alkstore-pg-posture-poc` (self-contained, + published-deps-only). +- Tests: `cargo test` → `tests/contract.rs` (11 tests). +- Probe binary: `cargo run --release -- seam|watch|burst|claim|pollvlisten|pool|pnlisten|smoke` + per the binary's help (wake command prints the notify-latency table; + pollvlisten the consumption-posture table). +- Probe examples (`examples/pgdiag_st1.rs` raw floor + fresh session; + `pgdiag_st4.rs` pool-budget accounting; `raw_listen_*.rs`, + `pn_*.rs`, `shared_listener_*.rs`, `attach_control.rs` the + listener-substrate steps; each runs standalone against the harness + server). +- Harness: docker `pglo-poc` (`postgres:16-alpine`, :15432, db + `blobs`). Note for re-runs: the container also carries + `synchronous_commit=off` as a server default from its earlier POC + use; the POC's pools set `synchronous_commit` per connection via + connect options, so probes are self-consistent — but `SHOW + synchronous_commit` outside the POC reads `off` on the server. \ No newline at end of file diff --git a/docs/research/poc-pg-posture-spec.md b/docs/research/poc-pg-posture-spec.md index 1e99102..a2a2a1c 100644 --- a/docs/research/poc-pg-posture-spec.md +++ b/docs/research/poc-pg-posture-spec.md @@ -1,5 +1,5 @@ --- -status: spec +status: findings-filed title: "POC #2 — Postgres engine posture: LISTEN/NOTIFY wiring, the pool tx-seam, and reactive parity on tokio-postgres" last_updated: 2026-10-04 --- @@ -20,6 +20,15 @@ last_updated: 2026-10-04 > `postgres-notify` 0.3.8 candidate (verified: reconnection + multi- > channel subscribe) — both folded in below; neither's *behavioral* > claims are trusted until probed. +> +> **Ran 2026-10-04: PASS** — findings in +> [poc-pg-posture-findings.md](poc-pg-posture-findings.md). Both +> pre-spec research claims verified in-probe and pinned as tests +> (#360: pooled LISTEN registers but never delivers; postgres-notify: +> reconnect + re-LISTEN verified *with* recorded quirks — lazy +> reconnect, no script at initial connect, unquoted identifiers). +> The spec's decision gate passed on all three conditions; OQ-ST-03 +> closed; the sqlx PgListener fallback retired unfired. ## What this POC must decide