Files
glm-5.3-flash 6a6ffd68de docs: the pg LISTEN identifier-truncation wake-collision note (review 004 item 13)
deployment.md's Postgres co-tenancy paragraph and engine-postgres.md's
connection-architecture wake-substrate section now carry the 63-byte
identifier-truncation collision note as dated annotations: two names
longer than 63 bytes sharing a 63-byte prefix silently register the
same server-side wake channel; pg_notify (the parameter path, untrun-
cated, rejecting >= 64-byte channels) is unaffected, so the collapse is
on the registration side. Mitigation: keep mechanism names at most 63
bytes (byte count). The naming contract's deliberate no-cap verdict is
untouched. Discharges review 004 'findings recorded, not fixed' item
13 (task docs-pg-identifier-truncation-note).
2026-10-11 15:03:55 +00:00

22 KiB
Raw Permalink Blame History

status, last_updated
status last_updated
stable 2026-10-11 (the pg identifier-truncation wake-collision note — review 004 item 13's doc-fix disposition — added to the Postgres engine section). Earlier the same day: the log-facade decision of record landed — [ADR-025](decisions/025-log-facade.md): engine diagnostics route through the `tracing` facade, silent by default, subscriber wiring on the consumer/binary; the diagnostics-posture section added below. Earlier the same day: the deployment-matrix release final pass — every row verified true as implemented across the three engines, the env-variable boundary pinned, the growth-posture acceptance folded into the release text, the fixed `synchronous = NORMAL` posture corrected, internal task/wave pointers removed)

Deployment

What a deployer must know to size, run, and reason about alkstore engines: host semantics, connection budgets, durability knobs, and where engine differences may honestly surface in the contract. The capability-surface decision is resolved (ADR-016, 2026-10-06): no runtime capability surface — this document's matrix is (with the engine crates' own docs) where the honest boundary lives; this document holds the facts.

Host semantics

Engine Host posture Notes
SQLite single-machine, file-backed NFS two-writers unsupported — a .db on network storage is not a supportable deployment posture, inherited from the upstream fork (ADR-016, ADR-003). Cross-process on one host is verified territory (data_version is cross-process by nature).
Postgres multi-host native Nothing assumes a shared host; the engine ran fully across a network bridge (docker) with the same properties, measured (ADR-004).
Mem single-process, ephemeral (ADR-024 §2) In-memory, per-instance; no cross-process story at all — two engine instances share nothing by design. No durability, by design: loss on process exit is the documented posture, and the dependency on alkstore-mem is the ephemerality declaration (honest for workloads that accept it). Never fleet-valid — the shared-pool predicate (ADR-010 §3) fails structurally; there is no pool. Compile-clean on wasm32-unknown-unknown (ADR-024 §5) — the family's member that runs in the sandbox.

The unified trait must not pretend SQLite is multi-host — and it does not: ADR-016 resolves that honesty to compile-time engine identity (the engine crate a binary depends on is the deployment statement) plus this documented matrix. There is no Store::capabilities() — the trait's shape constrains where capability differences can surface (ADR-006), and the resolution is: nowhere at runtime.

Options for OQ-08, with their outcome (ADR-016):

  1. Compile-time only — a consumer chooses an engine crate at dependency time; the engine's docs carry its deployment facts. Smallest contract; nothing runtime to match on. Adopted — together with (3); the two compose (engine docs serve the consumer choosing the dependency, the matrix serves the operator choosing the topology).
  2. Store::capabilities() — a runtime description (payload limits, wake cadence knobs, host semantics). Lets a consumer adapt (e.g., chunk large notify payloads) but adds a contract surface all engines must keep honest. Rejected — field-by-field under ADR-008 §5's act-differently rule, and no consumer-inventory row names a runtime-adapt need (ADR-016 §2).
  3. Deployment matrix only (this document) — no API surface. The honest-middle choice; matches the ecosystem's doc-first posture but provides no programmatic guard. Adopted (with (1)) — the "programmatic guard" gap is closed where it can honestly be: the engine-crate dependency edge cannot drift out of sync with the truth it states; the misconfiguration case (SQLite as shared network storage) follows the family's deployment-asserts-truth posture — documented detection symptom, no fabricated runtime machinery (ADR-016 §3).

Connection budgets

SQLite engine

  • Connections are in-process (writer + reader pool + watcher thread owning a connection). No external budget lines; the file lock is the OS-level resource.

Postgres engine

Connection class Count Notes
Pool max_size per process claims/queries via deadpool
Listener +1 per LISTEN-ing process non-pooled, dedicated; pooled connections cannot carry LISTEN (deadpool#360, test-pinned — ADR-004)
— — Sizing rule: max_size + 1 per process; verify end-to-end accounting (pg_stat_activity shows the exact server-side count)

Shared-server co-tenancy (the family's queue precedent — consumer tables co-tenant the pg instance) is supported and expected — the naming / reserved namespace contract protects reserved names; queue/stream/lock/schedule tables live in one engine-owned PostgreSQL schema (default alkstore, per-engine option) — layout per ADR-010 §8 (resolved from queues.md's namespace bullet).

One pg-specific naming caveat, dated per the family's amendment convention (annotated 2026-10-11 — review 004 item 13, recorded "not fixed" in the review with a fold-into-the-next-doc-pass disposition; this note is that fix): the wake machinery interpolates the channel as a quoted SQL identifier on LISTEN/UNLISTEN, and PostgreSQL truncates identifiers to 63 bytes (NAMEDATALEN − 1) server-side — with only a NOTICE, never an error — so two distinct names longer than 63 bytes sharing a 63-byte prefix register the same server-side wake channel (a silent cross-channel collision). The pg_notify path is unaffected by the truncation — it carries the channel as a text parameter, not an identifier, and the server rejects channels of 64 bytes or more there ("channel name too long") — so the collapse is on the registration side. The naming contract deliberately caps no name (the length-independence verdict stands), so nothing engine-side warns. Mitigation: keep mechanism names under the PostgreSQL identifier limit — at most 63 bytes; the limit counts bytes, not characters, so multibyte names reach it sooner. Distinct names within the limit cannot collide.

Mem engine

  • No connections — there is nothing to budget. The engine instance owns its state in-process; sharing nothing is the isolation design (ADR-024 §4's instance rule; the two-opens-one-engine constructor is rejected, re-entry via a consumer-inventory row).

TLS posture (v1)

The pg engine hardwires NoTls on every connection path in v1 — the pooled connections, the dedicated listener connection, and every reconnect attempt alike; the engine does not ride the consumer's Config sslmode setting, and a DSN carrying sslmode=require (or any TLS demand) fails at connect. The listener's NoTls posture is test-pinned in the type it carries (ADR-004); v1 TLS should therefore be treated as effectively unavailable — network confidentiality between the process and the Postgres server must come from the deployment topology itself (private network, egress rules, a local unix socket or sidecar), not from driver TLS. Encrypting traffic between engine and server is a post-v1 deployment concern (wire a TLS connector through the pool and listener construction); until then the boundary above is the honest statement, per ADR-016's spirit — the matrix states the limitation rather than having code pretend otherwise.

Consumer-obligation notes on engine options

Consumer-constructed numeric option values are trusted as given — the engine does not validate them against domain extents. The engine would be a second normative home for semantics the contract deliberately leaves to the caller's constants; the ADR-023 §2 domain table covers the trait-surface arguments, not these. This applies to the queue stamps (QueueOpts: visibility_timeout_s, dead_letter_retention_s, max_attempts) on all three engines:

  • Negative or zero visibility_timeout_s stamps claims whose deadline is already past — every claim is instantly reclaimable (the dual-execution window is the consumer's documented budget ADR-010 §2).
  • Zero or negative max_attempts stamps rows the claim statement never hands out — each is dead-lettered with reason "max attempts exceeded" at the next claim call on that queue (the pre-claim sweep's attempts >= max_attempts arm), i.e. the queue silently discards its work.
  • Negative dead_letter_retention_s deletes every dead row at the next sweep_expired call (retention is driver-free — nothing runs without a caller).
  • The stamps ride future enqueues only (QueueOpts's documented shape) — repair the opts before enqueueing rather than repairing rows afterwards.

The same trusted-as-given shape applies to the engine opts' other numeric fields — with the exceptions the code actually carries: PgOpts::max_size's positive-integer contract (0 is typed-failed at open, no round trip) and SqliteOpts::poll_interval's non-zero contract (0 rejected at open); SqliteOpts::max_readers clamps 0 up to one reader. The queue stamps above are the unvalidated surface. The consumer-obligation shape here is the visibility-budgeting precedent (ADR-010 §2): the constant is the consumer's, the behavior of a mis-set one is documented, nothing runtime is fabricated.

Configuration boundary

Env reads are harness-only; production configuration is the opts-constructor surface. No production path in any of the three engine crates or the core consults the process environment — grep re-verified across the workspace at release: every env::var read lives in a test target (the pg test modules' DSN plumbing; test temp files use std::env::temp_dir(), not configuration). Connection settings, credentials, and knobs arrive exclusively through open's arguments / option structs (PgOpts, SqliteOpts, the mem constructor), handed in by the consumer's own config layer — so no credential (or any setting) can arrive implicitly. If a future consumer-facing surface ever wants env conveniences, the family pattern is caller-side: the app's config layer reads env and passes the values into the opts.

Diagnostics posture

Engine diagnostics route through the tracing facade only, in the engine crates (ADR-025, 2026-10-11 — the family posture: all four family repos emit through tracing = "0.1" facade-only). What an operator needs:

  • Silent by default. An engine emits nothing until a subscriber is installed — the library crates carry no subscriber and no eprintln!; a consumer running without one sees no engine diagnostics. This is deliberate (published-crate consumers did not ask for stderr noise) — but it also means an integration that wants visibility must wire it at the consuming binary.
  • Operator recipe: a tracing-subscriber (env-filter) at the binary/init point; e.g. RUST_LOG=warn for the degraded-operation warnings (wake failures, retry/backoff, reconnect failures), RUST_LOG=error for the unrecoverable-loss events only (watcher death, a discarded pooled client). Core never emits anything — the contract's failure reporting is the typed error taxonomy alone.
  • Secret-free rule of record: no DSN, password, token, or credential material in any engine event — diagnostics name error classes, mechanism identifiers (channel/queue/schedule names), and backoff/attempt state, never connection material. The review's verified property, promoted to standing rule.
  • Listener kill-targetability (pg) still comes from the application_name setting (ADR-004, the toolchain row below) — diagnosability joins it: with a subscriber wired, failed listener reconnects and poll errors surface as events instead of vanishing.

Durability knobs

Engine Knob Shape
SQLite synchronous not a knob — the engine hardwires synchronous = NORMAL at connection bootstrap (WAL shipped, ADR-003); commit fsyncs land at WAL checkpoints (POC-measured: checkpoint spikes roughly every ~1000 commits, hundreds of ms each); there is no consumer surface to raise the sync level (a FULL option would be a post-v1 engine-surface addition)
SQLite poll_interval the watcher's data_version poll cadence — 1 ms shipping default (ADR-023 §4, the measured-wake-latency posture), carried on SqliteOpts; the idle cost is ~1000 poll reads/sec/instance; raising the interval trades wake latency (interval-bound) for idle CPU — the tuning recipe for latency-tolerant deployments
Postgres synchronous_commit per-session knob; on is ship config (p50 2.40 ms seam); off trades max-tail (40.9 ms) for slightly better p50 — measured, honest trade (ADR-004); session-level SET mechanics POC-verified
Mem — (none) There is no durability knob and no durability: the honest row is the absence itself (ADR-024 §2). Wake delivery is an in-process channel send at the commit boundary — no network or driver round-trip exists to budget; no tuning surface exists

These are engine-configuration concerns, not trait surface. What part of engine config is contract-level shape vs engine-crate docs is decided: constructors and option structs live in the engine crates; the contract is the trait the constructor returns (ADR-008 §6).

Growth postures (consciously accepted)

Three engine-internal table/state families grow without internal hygiene caps. This is the decision of record, made against the implemented engines (not the design docs): all three are accepted for v1, per engine, with the rationale below. None carries a v1 engine-side cap beyond what is described; engine-side hygiene (a background sweeper, an integrated sweep cadence) is a deliberate non-goal — adding it would be a scope change.

Growth path SQLite Postgres Mem
Dead letters (__alkstore_dead / pg dead table / mem's dead table) grows forever unless the consumer opts in and sweeps — accepted same shape — accepted grows in memory; erased at process exit; retention still sweep-on-call — accepted
Notify transport (__alkstore_notifications) the engine's only unconstrained-growth-with-hygiene table — pruned to its newest 10,000 rows at attach only (the at-attach cap, ADR-010 §6); an idle-but-listening process never prunes — accepted with the cadence caveat no table, no path — pg_notify is transient transport; nothing persists (≤ 8000 bytes per notify, ADR-004) no storage — wakes deliver at the commit boundary to open receivers; per-subscriber feeds are unbounded in-memory, ending at engine drop
Stream log durable-log posture by contract (ADR-015); trim_to is the consumer-invoked cap — accepted same durable-log shape — accepted in-memory event log, process-lifetime-bounded; trim_to available

The per-path rationale:

  1. Dead letters — the retention stamp (QueueOpts::dead_letter_retention_s) is opt-in; the default None is documented and suite-tested as "dead rows live forever". Accepted on top of that: even with a retention set, its enforcement has exactly one trigger — a caller-driven sweep_expired() call (the Queue trait method; the engines run no background retention sweeper). A consumer who wants dead-letter hygiene budgets the sweep (a @every schedule is the natural recipe). Growth is therefore fully opted-in twice: retention off by default, and no automatic cadence when on.
  2. Notifications — SQLite-only growth. The at-attach prune cap (oldest-first trim to 10,000 rows, every connection open; the substrate provenance register's D-16 port delta, ADR-010 §6) is the only prune site: there is no background sweeper and no consumer- visible prune operation (the table is engine-internal wake transport, kept off the contract by ADR-008 §4's disposition). An idle-but-listening process attaches never, so its __alkstore_notifications table grows past the cap until the next attach — a reconnect, a second open of the same file, or a restart. Accepted for v1 with that cadence sharpened in the SQLite engine spec (workloads with long-lived notify-heavy processes should budget an attach event or size the table deliberately).
  3. Stream log — growth is the contract's durable-log posture, not an accident: events persist until deleted, replay works across consumers, and the bounded-growth op is trim_to on the stream handle, consumer-invoked by design (ADR-015 §5: the default is unbounded, documented squarely). Accepted as contract posture; nothing engine-side to add.

The mem engine's entries are postures new since the paths were first flagged (its charter, ADR-024, postdates them): everything the durable engines keep in tables mem keeps in guarded in-process state, so its growth — dead letters, the stream log, and unconsumed wake feeds alike — is process-lifetime-bounded by design, ending at engine drop or process exit. Accepted structurally rather than via caps.

Toolchain / platform notes

Note Engine Affects
rusqlite 0.40.x needs rustc ≥ 1.99 SQLite any binary linking alkstore-sqlite (ADR-003)
bundled-sqlite adds a C build (~10 s dev, cacheable) SQLite build/CI time
libsqlite3-sys collision with sqlx today any mixed-driver binary structurally avoided (ADR-001 single-driver rule)
Listener application_name set for diagnosability (kill-targetable) Postgres ops runbooks (ADR-004)
alkstore-mem compiles clean on wasm32-unknown-unknown (ADR-024 §5); tokio features limited to the wasm-supported subset (sync + time); current-thread-runtime-clean internals (the contract suite's mem column runs under that flavor) Mem wasm targets, sandboxes, downstream ffi/napi/python adapters — the only engine that can compile to wasm (rusqlite/tokio-postgres are structurally out)

Consumer-facing latency profile (indicative, POC-measured)

From both POCs (single-box, relative shapes are the deliverable — ADR-003 and ADR-004 carry the full tables):

  • Tx seam: SQLite ~0.35 ms p50; Postgres ~2.4 ms p50 (ship config).
  • Wake: ~1.1–2.2 ms p50 both engines at default cadence.
  • Queue claim: LISTEN-driven 3–6 ms p50 (pg); poll-only interval-bound (32–50 ms at a 50 ms poll).
  • Mem: no POC measured it (the engine postdates the POCs); its wake and tx costs are in-process channel/lock operations, not the seam/network classes above — see engine-mem.md's architecture for what it realizes instead.
  • Absolute numbers will differ per hardware/network; they set expectations of order, not SLAs — contract docs must not bake them in (ADR-007's pg-POC note).

Design Decisions

ADR Decision Summary
001 Crate split single-driver binaries shape the matrix
003 SQLite driver bundling, toolchain floor
004 Postgres driver listener budget line, forwarder posture
006 Wake contract where capability differences may surface
008 Contract v1 constructor/options in engine crates; no capability surface in v1 (OQ-08; resolved by ADR-016)
016 Deployment honesty no runtime capability surface — compile-time identity + this matrix; PayloadTooLarge is the one runtime asymmetry carriage
021 Third review round drop = rollback keeps the connection budgets exact under error paths (SQLite lease releases; pg client re-pools)
024 Mem engine the matrix's third engine — honest-ephemeral: single-process, no durability, never fleet-valid, wasm32 compile-clean; no connection budget, no knobs
025 Log facade the tracing facade in the engine crates, silent by default, subscriber wiring on the consumer/binary; core stays dep-free and silent

Open Questions

Open questions are tracked in open-questions.md. Key questions affecting this document:

  • OQ-08: capability-surface shape — resolved (2026-10-06, ADR-016): no runtime capability surface; compile-time engine identity + this matrix; re-entry via a consumer-inventory row naming a runtime-adapt need.