Files
alkstore/tasks/pg-engine-open-opts.md
T

13 KiB

id, name, status, depends_on, scope, risk, impact, level, tags
id name status depends_on scope risk impact level tags
pg-engine-open-opts Postgres engine — `open` constructor, `PgOpts`, pool + listener wiring completed
pg-engine-schema
moderate medium component implementation
wave-4
postgres-engine

Description

Stand up the pg engine's public surface: the open constructor, PgOpts, and the connection architecture of engine-postgres.md's "Connection architecture" section — the structural twin of wave 3's sqlite-engine-open-opts, with the pg deltas.

  • PgOpts (engine-crate type, ADR-008 §6's config split; NOT #[non_exhaustive] — opts structs consumers construct, ADR-017 §3's exemption; Default + Clone): at minimum
    • connection string / config (the tokio-postgres Config-parseable form — the consumer's server address, credentials; the engine never invents defaults for these, open fails Database if unparseable/unreachable),
    • schema: String (default "alkstore" — ADR-010 §8),
    • max_size (the deadpool pool size; default documented, the deployment matrix's budget line — the listener adds +1 outside the pool),
    • synchronous_commit: bool (default true — ship config, ADR-004/deployment.md's measured trade; wired as a per-session SET at connect, POC-verified mechanics).
  • open(config, opts) -> Result<Box<dyn Store>> (concrete PgStore re-exported, mirroring the SQLite engine's boxed-constructor + pub-concrete shape):
    • Build the deadpool pool (RecyclingMethod::Fast — no DISCARD ALL recycling, the POC's zero-error posture).
    • Run the schema task's bootstrap on a pool connection at open (idempotent; open fails Database if bootstrap fails).
    • Spawn the forwarder (the listener task) — the wake substrate every later mechanism task rides. The forwarder is this task's hard part; its full behavior (reconnect, re-LISTEN, reconnect- wake) is the notify-listen task's, but the skeleton lives here: a dedicated non-pooled connection (deadpool#360 — pooled connections cannot deliver), the poll_message loop fanning out into a bounded broadcast channel (lag surfaced, not silent), the two POC-pinned deadlock pitfalls respected (the poll loop running before the first client query on the listener connection; the Client kept alive for the listener's lifetime), and the dynamic channel set (LISTENs re-issued from the live channel list — the notify-listen task adds channels to it).
    • The store struct holds pool + forwarder handle + schema name behind the Store trait; close()/Drop tear down cleanly (drop the listener connection, close the pool).
  • Trait stubs: all Store methods return Err(Database("… wiring lands with the … task")) — the wave-3 posture; the mechanism tasks replace them. with_tx surfaces the begin_tx stub naturally.
  • Error mapping posture established here (the seam task's foundation): tokio-postgres/deadpool errors → Error::Database with the source chain preserved (ADR-008 §5's opaque fallback) — the two mappings (pg_error, pool_error or one helper) the mechanism tasks reuse.
  • No spawn_blocking seam on this engine — natively async (Send+Sync client, POC #2 compile-probe); the engine crate's tokio dep already carries what's needed. State the posture in the crate docs (the SQLite engine's # Posture heading shape: multi-host native per ADR-016, the listener budget line per deployment.md).

Tests against the harness server: open/close round-trip, bootstrap ran (tables exist), pool bounds, forwarder skeleton up (a LISTEN issued through it delivers — the minimal fanout proof, full forwarder behavior in the notify-listen task), unreachable-server open fails Database, opts flow (schema name lands in the created schema's name; synchronous_commit observable via SHOW).

Acceptance Criteria

  • PgOpts with connection config, schema (default alkstore), max_size, synchronous_commit (default on); documented; not #[non_exhaustive]
  • open boots pool + bootstrap + forwarder skeleton; failure at any step is a typed Database error (source chain preserved)
  • Forwarder skeleton: dedicated non-pooled connection, poll loop before first query, bounded broadcast fanout, dynamic channel set; the two POC deadlock pitfalls structurally excluded
  • close()/Drop teardown clean (listener connection dropped, pool closed); post-close ops fail closed
  • Error-mapping helpers exist and are the documented reuse point
  • Crate docs carry the # Posture statements (multi-host, listener budget, no spawn_blocking)
  • cargo test -p alkstore-postgres (harness server), clippy -D warnings, fmt clean; gates green server-less (tests skip)

References

  • docs/architecture/engine-postgres.md (Connection architecture)
  • docs/architecture/decisions/004-postgres-driver.md (the forwarder shape, the two pinned pitfalls, RecyclingMethod::Fast)
  • docs/architecture/decisions/008-contract-v1-pinning.md §6
  • docs/architecture/deployment.md (budgets, synchronous_commit)
  • docs/research/poc-pg-posture-findings.md (Sub-modules L/T; the invocation note's isolation caveat)
  • alkstore-sqlite/src/store.rs, opts.rs (the structural twin's shape)

Notes

Decisions of record made while implementing (the description didn't pin them):

  • The connection config rides open's first argument, not PgOpts (the description's bullet list carried it under both; the signature line open(config, opts) and ADR-008 §6's alkstore_postgres::open(url, PgOpts) pin decide it): PgOpts carries schema/max_size/synchronous_commit only. open is async (connect is; the SQLite twin's sync open was file-only). Fail posture: unparseable config → pg_error (a tokio_postgres::Error::ConfigParse); refused connection → pg_error — both typed Database, source chain preserved, pinned by test.
  • open connects the listener inline, not inside the task — an unreachable server must fail open synchronously (the unreachable-server acceptance row); the forwarder task takes over the established connection and owns reconnects from there.
  • Forwarder skeleton shape (forwarder.rs, the POC's ~90-line shape restructured): a LoopCtx struct carries the loop's state (clippy's too-many-arguments); per connection generation the loop spawns a poll task owning the Connection (poll_message → bounded broadcast 1024, lag surfaced) before any client query — pitfall 1's structural exclusion held permanently (the command select keeps the poll task co-resident with LISTEN/UNLISTEN client queries every generation); the Client is owned by the loop (pitfall 2). Connection generations carry an Option<ListenerConnection> slot; ListenerConnection is the NoTls Connection<Socket, NoTlsStream> alias. Commands (LISTEN/UNLISTEN) ride an mpsc queue the loop issues between polls (reconnect replays them harmlessly alongside the registry re-issue); the shutdown watch flips at close/Drop.
  • Dynamic channel set: a BTreeSet-backed registry (Forwarder::{register,unregister}), registry-write-first ordering (a registration survives a connection death between write and LISTEN — the reconnect re-issues from the snapshot). register fails Database only mid-reconnect (transient, honest state surfaced). subscribe() on the fanout exposes the raw notification broadcast — the notify-listen task bridges receivers onto it. RawNotification { channel } is the skeleton's broadcast content (the Wake mapping is that task's).
  • The synthetic reconnect-wake is broadcast by the skeleton (the reserved channel __alkstore_listener_reconnected__ after every successful re-LISTEN except the first) — the notify-listen task owns the wake-content/receiver semantics; the skeleton's fanout already carries it (pinned by the backend-kill test).
  • Listener application_name is per-instance: alkstore-pg-listener-{pid}-{seq} — the deployment.md ops kill-targetability with a unique suffix so parallel stores/tests are distinguishable in pg_stat_activity (the outside-the-pool and teardown assertions target the store's own instance). Exposed at PgStore::listener_application_name() for the notify-listen task's backend-kill tests.
  • Liveness is verified, not assumed: connected starts true (the handed-in connection is fresh) and a no-channel generation probes with SELECT 1 (a dead connection would otherwise report live).
  • Empty PgOpts::schema falls back to DEFAULT_SCHEMA (defensive; the documented default is .default()-driven).
  • Error mappings (seam.rs, the mechanism tasks' reuse surface): pg_error(tokio_postgres::Error), pool_error(PoolError), build_error(BuildError) (the pool-build failure arm — PoolBuilder::build returns BuildError, not CreatePoolError), database_error(msg) (the SQLite twin's io-Error string helper).
  • No runtime dep additions beyond the existing set (tokio-postgres, deadpool-postgres, tokio, thiserror already carried by the schema task; serde_json added — core's payload type appears in the trait stubs' signatures). TLS: NoTls on both paths (the consumer's sslmode/TLS story is a deployment concern; noted in the forwarder module docs, the notify-listen task revisits listener TLS if a task or ADR ever demands it).
  • Trait stubs carry the closed-store check ahead of the stub error (close()/Drop make every arm fail closed immediately — the post-close fails-closed row is pinned from open onward). Stub messages name their landing task; their text rides the source chain (Database Display is opaque — asserted via the chain).
  • Machinery accessors are pub(crate) (pool, forwarder, schema, listener_application_name) — the mechanism tasks reach them; tests exercise them and the lib-build dead-code gates are #[cfg_attr(not(test), allow(dead_code))]-carried until the wiring tasks land (the wave-3 intermediate posture).
  • Harness reuse: the schema task's env-carried DSN convention (ALKSTORE_PG_HOST/PORT/USER/PASSWORD/DB), schema-per-test isolation, skip-clean server-less; the shared default schema is never dropped by tests; the POC's fanout shape (broadcast 1024, lag-surfaced) and backoff constants (50 ms → 2 s) are re-owned as named constants.

Summary

Stood up alkstore-postgres's public surface: PgOpts (schema — default alkstore per ADR-010 §8; max_size — default DEFAULT_MAX_SIZE = 8, the deployment budget line with the +1 listener outside the pool; synchronous_commit — default on, wired as connect-options -c per-session SET per the POC's verified mechanics; not #[non_exhaustive], Debug+Clone+Default per ADR-017 §3's opts exemption), the async open(config: &str, opts: PgOpts) -> Result<Box<dyn Store>> constructor (parse → pool build with RecyclingMethod::Fast → pool-checkout schema bootstrap → dedicated non-pooled listener connect + forwarder spawn; every failure a typed Error::Database with the source chain preserved), the concrete PgStore re-exported (pool + forwarder handle + schema name held behind the Store trait; close()/Drop shut the forwarder down and close the pool — post-close ops fail closed), the LISTEN forwarder skeleton (forwarder.rs: dedicated non-pooled connection, poll-first structure excluding both POC-pinned deadlock pitfalls, bounded 1024 broadcast fanout with lag surfaced, dynamic channel set with registry-write-first recovery ordering, reconnect loop with the 50 ms → 2 s exponential backoff, synthetic reconnect-wake broadcast on the reserved channel), the seam error mappings (pg_error, pool_error, build_error, database_error — the mechanism tasks' documented reuse point), and the full Store trait as wave-3-style Database("… wiring lands with the … task") stubs with closed-store fail-closed checks. Crate docs carry the # Posture statements (multi-host, the listener budget line, natively async / no spawn_blocking). Verified: 12 new open/opts tests (src/store/open_tests.rs) + the 9 schema tests green against the harness server (boot round-trip, pool bounds + pool saturation leaving the listener untouched, opts flow: schema name in the created schema + synchronous_commit observable via SHOW on both settings, pitfall pins — immediate post-open LISTEN + long-after-open delivery — and the outside-the-pool accounting via per-instance pg_stat_activity, backend-kill reconnect with reserved-wake + post-reconnect delivery, close/drop teardown with server-side session termination + closed pool + fail-closed ops, unparseable-config and unreachable-server Database failures, the stub surface, the seam mappings' source chains); workspace cargo test green server-less (236 tests, pg tests skipping per convention); cargo clippy --all-targets -- -D warnings and cargo fmt --check clean workspace-wide.