Files
alkstore/docs/research/phase-0.md
T

44 KiB
Raw Blame History


status: draft last_updated: 2026-10-04 (Phase 0 complete and the register promoted: docs/architecture/ opened with ADR-001..007 carrying the resolved decisions and the OQ tracker mirroring OQ-ST-01..08 1:1. Phase 0 research itself is closed; remaining work is in docs/architecture/ open-questions.md. One erratum fixed in OQ-ST-04's tx-friction line — thread-affinity is SQLite-side, previously garbled as "pg-side".)

alkstore — Phase 0 (Exploration)

This document captures Phase 0 (Exploration) for the alkstore crate: vision, guiding principles, prior art, the open-question register (OQ-ST-01..NN), the POC register, and the convergence the phase objective asks for. Phase 0's objective per docs/sdd_process.md: capture vision and guiding principles; research options; validate approaches; converge on a recommended approach. All of that objective is now met: the scope question (OQ-ST-01) is answered per-feature from the consumers' documents, the crate-split (OQ-ST-02) is decided, the driver question (OQ-ST-03) is closed POC-backed on both engines, the reactive contract (OQ-ST-04) is de-risked down to paper work, and the ownership postures (OQ-ST-05/06) have recorded per-subsystem votes. §Convergence collects the recommendation; the OQ register's open residue is Phase 1 architecture work, not further research.

Context for why this crate starts now: alkblobs (/workspace/@alkdev/alkblobs — spec + POCs only, paused mid-planning) hit repeated circular hedging in its Phase 0, and a root cause was that the storage substrate it deploys onto was itself disjoint and fuzzy — a repo pattern with a default in-memory adapter across the alk* ecosystem, cache-invalidation patches in hot paths, non-invalidated caches where delay was tolerable, no single definition of "how does a change in the database become visible to other processes/connections?" alkblobs paused partly to let this crate answer that first. alkstore is the attempt to make that substrate real once, so downstream stores don't re-derive it.

Vision and guiding principles

One sentence (draft): one reactive store interface over SQLite and Postgres — durable pub/sub notify, queues, streams, and the transactional integration (write + enqueue in one transaction) that honker delivers on SQLite — with the Postgres side building on the natively-available machinery (pg_notify/LISTEN, and the pgboss job-queue schema family) rather than emulating it.

The honker relationship. /workspace/honker (reference checkout, not for direct use as a dependency; alpha-quality per its own README, MIT/Apache-2.0 dual) adds Postgres-style NOTIFY/LISTEN semantics to SQLite without a broker: durable at-least-once queues with retries/delay/priority/visibility timeouts/dead-letter, durable streams with per-consumer offsets, cron/@every scheduling, named locks, rate limits, transactional outbox — all as INSERTs inside the caller's transaction, with the cross-process wake delivered by a shared watcher that polls PRAGMA data_version (default 1 ms → single-digit-ms delivery) and re-reads indexed state after every wake. Its own Prior Art section names the lineage: pg_notify, pg-boss, Oban, Huey.

The goal is not "port honker to Postgres." The goal is the interface: one store API whose consumer code (queues, streams, notify) looks the same whether the backing engine is SQLite or Postgres, while each engine uses its own native wake/delivery story under the hood. The honker docs recommend pgboss + pg_notify for the Postgres equivalent — that recommendation is the design brief for this crate's Postgres engine.

The interface finding

(2026-10-03, from the four honker.dev guides + packages/honker-rs/src/lib.rs, the Rust binding, v0.5.0.) Honker's surface is already the unified-interface candidate. Its Rust binding exposes exactly the surface this crate wants, engine-clean:

  • db.queue(name, QueueOpts) → enqueue / enqueue_tx / claim_one / claim_batch / ack_batch / cancel / get_job / sweep_expired / claim_waker, with job.ack / retry / fail / heartbeat and EnqueueOpts {delay, priority, max_attempts, expires, ...} — semantically the pg-boss model (visibility timeouts, retries, dead-letter via move-to-_honker_dead), not a LISTEN-emulation.
  • db.stream(name) → publish / publish_tx / publish_with_key / read_since / read_from_consumer / save_offset(_tx) / get_offset / subscribe(consumer) — offsets are explicit, transaction-aware (save_offset_tx) for the exactly-once-within-a-business-tx shape, and replay-on-reconnect is the default.
  • db.notify(channel, payload) / notify_tx / db.listen(channel) — the pg_notify-analogue fire-and-forget signal layer ("fire-and-forget, no replay, no guarantees" — streams are the durable cousin), listener starts from MAX(id) at attach, no historical replay.
  • db.scheduler() → add/pause/resume/update/list/remove/tick/run — cron + @every enqueueing into named queues, leader-elected via advisory lock with TTL heartbeat, missed-boundary catch-up.
  • db.outbox(name) — the transactional outbox helper (enqueue + run_once delivery worker).
  • db.try_lock / try_rate_limit / save_result / get_result / sweep_results — the coordination/adjacent-tools surface.

The implication flips the framing of the unified-API work: it is not "invent a shape both engines fit" — it is "this shape both engines can fit" (pg-boss's queue model is already the native Postgres tooling model; streams/offsets and notify have direct Postgres counterparts) and the work is pinning which parts of the shape are the crate's contract — the delivery-guarantee differences honker's own guide documents per-binding (auto-checkpoint cadence vs manual offset save; the processing-guarantees table) are exactly the seams a single-crate version must clean up. Honker-rs is the concrete prior art for that pinning exercise.

Postgres side: pg_notify gives fast triggers with no retry or visibility semantics; pg-boss/Oban are the durable-layer gold standards — "If you already run Postgres, use the Postgres tools."

Consumer shape (from the paused alkblobs planning): any crate that today uses the ecosystem's repo-pattern + in-memory adapter should be able to swap in an alkstore-backed engine and get durability + true cross-process/cross-instance reactivity. That means the reactive surface must compose with client-side caching: a subscriber that also holds a cache can invalidate on notification instead of re-querying or re-polling — the hot-path pattern the ecosystem already uses, given a real invalidation source.

Guiding principles:

  1. One interface, two engines, native underneath. The abstraction layer unifies the consumer-visible features; the engines stay dialects, not two emulations of one dialect. SQLite follows honker's design (queue in the same file, same transaction, watcher-based wake); Postgres follows pg-boss' design (schema-based job tables + pg_notify-driven wake). A "lowest common denominator" unification (both sides polling, both sides emulating LISTEN) is explicitly the failure mode to avoid — it would re-create the fuzziness this crate exists to remove.
  2. Transactional local-adjacency is the load-bearing property. Honker's core claim: enqueue/publish/notify in the same transaction as the business write; rollback drops both. The unified surface must preserve this on any engine, because the ecosystem's repo pattern assumes it (a business write that loses its side-effect notification is the dual-write problem honker names). Both POCs verified this property holds natively on their engines — it is now evidence, not aspiration.
  3. Ownership of the whole stack. honker is third-party; pgboss-rs is third-party. Whether any of them are adopted, forked, or used as schema/design reference only is a deliberate per-question decision — not inherited by adjacency. The workspace precedent is the targeted fork (alksocks' fast-socks5 extraction: adopt the design, own the code, port to our conventions).
  4. Substrate-agnostic consumer API, engine-specific setup. A consumer opens a Store from a connection string / file path and gets the same trait surface. Which engine is behind what can vary (per-deployment config), but consumer code must not branch on engine type.
  5. No panics, tokio, thiserror, lean base crate, feature-gated optional engines — family-standard, pre-committed (AGENTS.md).

The driver conflict (Phase 0's central tension — now resolved)

The immediate design fork, flagged by the user early in Phase 0: a single crate with both engines means a driver decision, and the reactivity story is entangled with it. The facts that made the tension real:

  • The alkblobs POCs used tokio-postgres + deadpool-postgres (validated in poc-postgres-kv-findings.md, poc-pglo-findings.md — including Large Objects work), while pgboss-rs (/workspace/pgboss-rs @ 98f7d9e) uses sqlx, and honker-core uses rusqlite.
  • pgboss-rs currently has no LISTEN/NOTIFY at all (verified 2026-10-03 against the checkout — src/ contains no LISTEN/NOTIFY usage; consumption is fetch_job polling). So even "use pgboss for the queue" does not deliver reactivity — LISTEN/NOTIFY wiring would be new work either way, and the driver choice determines whose LISTEN plumbing.
  • Honker's reactivity on SQLite is a watcher polling PRAGMA data_version — a fundamentally different mechanism from LISTEN/NOTIFY. The unified reactive trait must abstract over both without collapsing to the polling behavior of the weaker side.
  • honker-rs is sync (std threads + blocking iterators; parking_lot + rusqlite, no tokio) while the family standard is tokio-async — so even the SQLite side looked like a port-and-adapt, and the async question was entangled with whether rusqlite-in-a-pool or a native-async driver (sqlx sqlite) is the right shape.

Two corrections sharpened the framing before the POCs landed: the honker-rs interface is largely driver-independent, with the transactional seam (*_tx methods assuming a live transaction handle from the caller's driver) as the one genuinely driver-coupled design point; and the alktty REQ-TTY-01 family precedent ("backends are not required to be natively async" — bridge-at-the-seam is a supported posture, not a workaround) already blesses the sync-machinery + async-facing-trait shape, weakening native-async's main differentiator.

The tension dissolved with the POCs (details and measurements at OQ-ST-03): per-engine drivers under OQ-ST-02's crate split, with the bridged-rusqlite seam measurably faster than sqlx's native async on SQLite, and tokio-postgres natively async-native on Postgres. What pgboss-rs genuinely offers (schema DDL, job states, retry semantics, the node-compatible API) is design reference regardless of driver; the push channel is this crate's own work either way (OQ-ST-05).

Prior art

Notes below are from reading the checkouts on 2026-10-03; both external projects are reference checkouts — read freely, but not for direct use as a dependency. We use the published version of anything that lives in the global workspace unless we vendor or fork it (the alksocks fast-socks5 precedent); if adoption ever requires a fork, forking is normal work we own, not an exception. Provenance/licensing gets recorded per AGENTS.md §3 when code is adopted, not while only reading.

honker — the SQLite-side template

/workspace/honker (checkout @ f4e53c6; SQLite extension + bindings). What matters for this crate:

  • The full feature set to match on Postgres (its §What It Does): notify/listen across processes, durable at-least-once queues (retries, delayed jobs, priority, visibility timeouts, dead-letter rows, result storage), durable streams with per-consumer offsets, cron/@every scheduling, named locks, rate limits, transactional outbox helpers. Deliberately excluded there: workflow DAGs, task chains/chords, multi-writer replication, cross-machine locking — scope line likely inherited, to be confirmed.
  • The wake mechanism — PRAGMA data_version polling watcher (default 1 ms; raise for idle CPU), re-read indexed state after wake, overtriggering on purpose ("one indexed SELECT is cheap; a missed wake is a correctness bug."). Optional kernel-events and WAL shared-memory backends exist in source builds.
  • Single-machine honesty — file-backed, one host; NFS-two-writers explicitly not supported. This posture needs an explicit Postgres counterpart (multi-host is Postgres' normal case, so the interface must not bake SQLite's single-host assumption into the shared surface).
  • The transactional enqueue shape — every feature is an INSERT inside the caller's transaction. This is the pattern the unified API must keep visible and cheap.
  • The honker-rs binding is the concrete interface prior art (v0.5.0, packages/honker-rs, read 2026-10-03): the full surface per §Interface finding. Notable honest limitations documented by its own guides — the per-binding processing-guarantees table (auto-checkpoint cadence vs manual offset save; several bindings "may persist an offset on a cadence... without knowing whether downstream application work committed"), the Node reverse-order consumer-checkpoint bug, per-binding feature gaps (JVM missing cancel/get_job, Go/Bun/C++ missing typed pruning) — are exactly the seams a single-crate version designed-for-the-contract from day one can clean up. Its sync-only shape (std threads, blocking iterators, no tokio) is a port-and-adapt constraint, not an adopt candidate as-is.

pgboss-rs — the Postgres queue family reference

/workspace/pgboss-rs (checkout @ 98f7d9e; v0.1.0-rc6, MIT/Apache-2.0 dual). Ported from node pg-boss: builder-based queue/job API, retry/delay/priority/singleton/dead-letter concepts, sqlx 0.8, schema-scoped DDL.

  • Verified gap (2026-10-03): no LISTEN/NOTIFY anywhere in src/ — consumption is polling fetch_job. Any push-reactivity is new work, not an adoption freebie. This is the substantive difference between pgboss-rs and what alkstore needs: regardless of fork vs re-derive, reactivity is ours to build on the Postgres side either way.
  • Its value as reference: the pg-boss schema family (job states, maintenance/dead-letter behavior) is battle-tested against real Postgres semantics — worth borrowing as design, independent of the driver decision.
  • The node original (pg-boss) is the upstream of record for semantics the port may have dropped; compare against it when adopting queue semantics.

Honker's Postgres-side recommendation

The honker README's own posture: if you run Postgres, use the Postgres tools. pg_notify + pgboss is the recommended assembly. The design brief: the queue machinery from the pg-boss family, the push semantics from LISTEN/NOTIFY, the unified API shape from honker's Rust binding.

The alk* repo pattern — what this crate replaces

The ecosystem's current shape: a repository trait with a default in-memory adapter; cache-invalidation wiring in hot paths; uninvalidated (non-reactive) caches where delay was acceptable; each project composing these slightly differently. No persistence-backed reactive substrate exists in the family — alkblobs was the first project to try to plan against one, found it missing, and paused. This crate's reason to exist is precisely that that substrate should exist once, well, instead of per-project approximations.

alkcall — the substrate (not a dependency of the store layer)

/workspace/@alkdev/alkcall (pure protocol crate, no transport). The alk* crates (alktty, alktunnels, alksocks) are its consumers; a future alkstore ops/protocol surface (if this crate ever exposes store access over alkcall channels) rides the same substrate. Like the alkblobs split (store layer stays substrate-free), the store layer here stays alkcall-free; any networked surface is an ops module/sibling concern and a separate decision.

Convergence

Phase 0's objective — converge on a recommended approach — is met. The recommendation, assembled from the OQ resolutions below (each carries its own evidence):

Shape (OQ-ST-02): a reactive-core crate carrying the trait surface/types, plus per-engine crates implementing it (SQLite, Postgres; a mem-shaped test engine as a third impl if useful). The split makes the engines' real asymmetry of work structural: the SQLite engine rides honker's existing machinery; the Postgres engine is the build-heavy side; any future engine is additive. It also keeps each engine binary single-driver, which the dependency constraints below effectively require.

Engines (OQ-ST-03):

  • SQLite engine: rusqlite + published honker-core 0.5.0 (POC posture 1) — we own connection/schema/watcher wiring per the alknet-filesystem POC's shape; the async seam is the bridge-at-the-seam posture (writer-slot + spawn_blocking, REQ-TTY-01 precedent), measured ~2× faster at p50 than sqlx's native async; honker-core's SharedUpdateWatcher is inherited for wake (p50 ≈ 1.4 ms, battle-tested failure handling). No .so runtime artifact, no vendored patches.
  • Postgres engine: tokio-postgres + deadpool-postgres (POC #2) — pooled connections for queries/claims (FOR UPDATE SKIP LOCKED), a dedicated non-pooled listener connection with a hand-rolled ~90-line LISTEN forwarder (immediate reconnect, synthetic reconnect-wake on a reserved channel closing the no-replay hole, identifier quoting), LISTEN-driven claim beats poll 5–16× at p50. postgres-notify evaluated and passed over (derive-not-adopt).

Contract starting shape (OQ-ST-04): the honker-rs surface (§Interface finding), scoped to the inventory-confirmed features, is the starting artifact for contract pinning. The load-bearing pieces hold identically on both engines, POC-verified: the opaque-wake + re-read listener contract; notify = fire-and-forget commit-atomic (no replay) vs streams = durable with explicit per-consumer offsets; and the caller-held tx handle (*_tx on the handle) whose per-engine difference is bridging mechanism, not trait shape.

Ownership (OQ-ST-05/06): published-library dependencies as the default posture — honker-core (SQLite), tokio-postgres + deadpool-postgres (Postgres); the pg queue machinery is re-derived on our driver with the pg-boss schema family as design reference; the hand-rolled listener forwarder replaces postgres-notify. Named fork triggers remain: a Phase 1 quality read of honker-core's watcher/transactional core, or a needed change upstream won't take.

Scope (OQ-ST-01): notify/listen, named locks, queues, outbox, scheduler, streams are in (with the inventory's per-row evidence grades); rate limits and result storage are cut-flags; the loadable extension surface is out (OQ-ST-07); honker's exclusion lines (DAGs, task chains/chords, multi-writer replication, distributed locking) stay out.

Phase 1 inherits, as architecture work over complete evidence: the contract-pinning itself (OQ-ST-04's remainder — which surface parts become contract, per the inventory rows, and the per-engine capability surface, OQ-ST-08); queue semantics depth (retry/backoff/dead-letter/sweep design, OQ-ST-05); the honker-core quality read (OQ-ST-06's fork trigger); and the deployment-matrix / capability-flags decision (OQ-ST-08). No further Phase 0 research is required. The known deployment constraints to design around: honker-core 0.5.0 pins rusqlite ^0.40.1 whose rustc requirement (≥1.99) is a deployment note; pooled connections cannot carry LISTEN (deadpool#360) so the listener connection is a per-process budget line outside the pool; notify payloads are ≤ 8000 bytes (large payloads ride a table row with the id in the notification — the honker-outbox shape); the reserved reconnect-wake channel name needs a namespace convention in the contract.

Open Questions

Register in docs/research/phase-0.md; IDs OQ-ST-NN (stable, append only). Promotion target: Phase 1 docs/architecture/open-questions.md — promoted 2026-10-04: OQ-ST-01..08 mirror 1:1 to OQ-01..08 there (their statuses and resolutions carried; resolved decisions carried into ADRs 001–007), and new Phase 1 questions append from OQ-09. Status conventions: resolved (evidence recorded here); open — <work-type> where the work-type names the remaining work and the entry is de-risked (the remainder is Phase 1 architecture work, not further research); open (genuinely open).

OQ-ST-01: Scope boundary — which honker features are in-scope?

Status: resolved (2026-10-04). Answered by the consumer inventory — docs/research/consumer-inventory.md; scope votes shrink to named rows, not the whole feature list. The original framing ("blocked on a consumer-driven inventory pass") was circular hedging: the consumers are paused, so the input would never arrive — but their documents are stable evidence, and the inventory walks them per feature.

  • In scope, first-class: notify/listen (pinned: alkfs path-tree invalidation), streams (operator-authority record — type-filtered event watching from several places, e.g. repo-change subscriptions; a reactivity requirement notify cannot serve honestly, being fire-and-forget).
  • In scope: named locks (pinned: alkblobs fleet sweeper lock; documented: alkfs OQ-FS-05 writer coordination), queues + the outbox helper (documented: alkfs sync/fetch-on-miss outbox; alkblobs embedder-owned maintenance cadence), scheduler (documented-thin: the family-wide "who sweeps/renews/reaps" problem, possibly collapsing into queues — watch at OQ-ST-04).
  • Cut-flags (no named consumer; carried per the keep-until-implementation posture, cut later rather than silently included): rate limits (alkgit enforces budgets in its own wire layer — an in-crate alternative exists), result storage.
  • Out: honker's exclusion lines (DAGs, task chains/chords, multi-writer replication, distributed locking) — no consumer names these either; they stay out unless a consumer document grows one.

New consumers (alksftp, the alknet rewrite) add a row to the inventory before being assumed into scope.

OQ-ST-02: Crate scope — one store crate, or reactive-core + engines?

Status: resolved (2026-10-04, operator decision): reactive-core + engine crates — a core crate carrying the trait surface/types, per-engine crates implementing it (sqlite, postgres; mem-shaped test engine as a third impl if useful). Options considered: single crate with feature-gated engines (the alk* feature-gate pattern); a core trait crate + per-engine crates; engine crates consuming a thin core.

The reasoning, recorded because it overrode the inventory's lean: the split isolates the engines' real asymmetry of work — the SQLite engine rides honker's existing machinery as the baseline (port-and-adapt sync→async), the Postgres engine is the build-heavy side (LISTEN/NOTIFY wiring + pg-boss-family schema work, OQ-ST-05) — and it makes any future engine (alkfs's in-tree needs, an ops-surface engine) additive rather than a feature-graph edit to one crate. Base-crate-lean becomes structural rather than a feature-discipline. It also keeps each engine binary single-driver, which the libsqlite3-sys link-collision constraint (OQ-ST-03) effectively requires.

The inventory's uniform-feature-family fact still stands, not contradicted: it reads as "the core contract can stay small — one feature family, both engines," not as an argument for the crates to merge (its single-crate lean was inductive from that fact; the structural reasoning here supersedes it — correction recorded in consumer-inventory.md too).

OQ-ST-03: Driver story — sqlx, tokio-postgres, or per-engine drivers?

Status: resolved (2026-10-04, both halves, POC-backed): per-engine drivers — rusqlite + honker-core 0.5.0 (SQLite engine); tokio-postgres 0.7.18 + deadpool-postgres 0.14.2 (Postgres engine) — under OQ-ST-02's per-engine-crate split.

Original option list (the decision as first framed): one driver across engines (sqlx: both engines native, one API — but the alkblobs POC evidence is tokio-postgres); per-engine drivers under a unified trait (tokio-postgres + deadpool-postgres, POC-validated in alkblobs findings, + rusqlite, honker's choice); adopt/fork pgboss-rs (brings sqlx along where the queue lives). Honest unknowns at framing time: does sqlx support SQLite data_version/extension-style machinery equally well; does a unified trait over (tokio-postgres, rusqlite) pay more trait-fitting cost than sqlx's single-API convenience costs elsewhere; extension loading under sqlx vs rusqlite; and the async question honker-rs's sync shape forces — rusqlite-in-a-pool with a bridge, or a native-async driver?

The SQLite option space, named explicitly (2026-10-04, operator + verified against the checkout @ f4e53c6) — three distinct postures, not one "rusqlite vs sqlx" axis:

  1. honker-core on our own rusqlite connection (the attach_honker_functions shape — the alknet-filesystem POC's actual usage): we own the connection, the schema bootstrap, and the watcher; honker supplies the SQL-function machinery.
  2. honker-rs as the crate's SQLite substrate (Database::open, typed Queue/Stream/Transaction primitives): maximum reuse, least control — honker-rs opens and holds its own connections, its Database wraps a connection mutex (transactions pin the mutex; same-thread *_tx methods only), and the whole engine is sync under OUR async core (bridge at every seam).
  3. raw SQL over sqlx-sqlite with the honker loadable extension (SqliteConnectOptions::extension(ext) + SELECT honker_bootstrap(), then every feature is plain SQL callable through SqliteExecutor<'e> — pool, connection, AND Transaction alike). Verified in-harness: the pattern is CI-proven in the honker checkout itself (scripts/proof/orm/rust — async business-write + honker_enqueue inside a conn.begin() tx, commit-visibility + rollback-drops-job asserted).

Family precedent bearing on the async sub-question (2026-10-04): the async-facing-trait + sync-bridge posture is family-standard, twice over — alktty REQ-TTY-01 ("backends are not required to be natively async": blocking work on dedicated threads or spawn_blocking feeding tokio channels is a documented, supported implementation strategy, not a workaround; the wezterm/portable_pty pattern), and alkblobs store-api.md (blocking file work in spawn_blocking inside engine impls). This reframed the sync→async port: less "rewrite onto a native-async driver," more "keep the sync machinery and bridge at the trait seam" — weakening sqlx's main differentiator for the SQLite side.

Resolution, SQLite half — POC #1 (2026-10-04; findings: poc-sqlite-posture-findings.md): posture 1 — honker-core on our rusqlite. All three gate conditions fired in A's favor, mildly: the bridged path is ~2× sqlx's native-async at p50 (0.354 vs 0.707 ms on the tx-enqueue workload; B's premise measured false), honker-core's inherited watcher is tighter than a re-derived one (p50 1.40 vs 2.15 ms, max 29 vs 172 ms, with battle-tested failure handling), and the .so runtime dependency is packaging cost A doesn't pay for no compensating advantage. The transactional contract holds identically on both (it is SQLite's property, not the posture's). Constraints recorded for the engine crate regardless: honker-core=0.5.0 pins rusqlite ^0.40.1 whose rustc requirement (≥1.99) is a deployment note, and mixed rusqlite+sqlx binaries currently need a vendored one-line libsqlite3-sys patch — OQ-ST-02's per-engine-crate split is what keeps the engine binary single-driver.

Resolution, Postgres half — POC #2 (2026-10-04; findings: poc-pg-posture-findings.md): tokio-postgres + deadpool-postgres. All three gate conditions held: the transactional property is native (in-tx NOTIFY delivers only at commit; rollback drops job row + business row + notification), exactly-once claim via FOR UPDATE SKIP LOCKED, and the LISTEN wake layer is push (p50 1.1 ms, 300/300; claim latency 3–6 ms vs poll-only 32–50 ms — 5–16×, the push channel pgboss-rs lacks, measured). Structural findings the contract must absorb: pooled connections cannot carry LISTEN (deadpool#360 — pinned as our own test; the dedicated listener connection is a per-process budget line outside the pool), and the *_tx seam differs from SQLite's exactly where expected (pg's client is Send+Sync — the tx handle is held directly across awaits; SQLite's is a bridged writer-slot lease — same seam shape, different bridging). Listener substrate vote: hand-rolled ~90-line forwarder (immediate reconnect, quoting control, no dependency posture) over postgres-notify 0.3.8 (lazy reconnect, connect_script skipped at initial connect, unquoted-identifier LISTEN — derive-not-adopt; fallback if upstream improves). The sqlx PgListener fallback retired unfired.

OQ-ST-04: The reactive abstraction — what does the unified notify surface look like?

Status: open — contract work (de-risked; both engine sides POC-verified). The remainder is contract-pinning paper work over a complete evidence base — Phase 1, not further research.

Background: the two engines' wake mechanisms are structurally different — SQLite = watcher polling PRAGMA data_version (deliver on commit; no server-side push exists), Postgres = LISTEN/NOTIFY (server push, connection-bound, no retry/visibility semantics). The reactive trait must have a shape both implement without one emulating the other's weaknesses.

The honker-rs surface (§Interface finding) is the concrete starting point — the work decomposes into contract-pinning rather than shape-invention:

  • Which parts of the honker-rs surface become the crate's contract: the notify/listen pair, the stream/offset/consumer model (in scope per the inventory — subscriptions are the durable reactivity half notify can't serve), the queue claim/ack/visibility model, locks, outbox, scheduler — all have named consumers now; rate-limits have an in-crate alternative mechanism (alkgit's wire layer) — subset, renamed/regrouped, decided against the inventory rows rather than against the whole honker menu.
  • What is the delivery-guarantee contract, per mechanism (honker's own guide table shows how easily per-binding auto-checkpoint vs manual-save ambiguity produces different guarantees under one function name — the single-crate version must pick one answer, not inherit the table)?
  • Listener semantics: honker starts from MAX(id) and replays nothing; Postgres LISTEN has no replay either but delivers via a dedicated connection with its own lifecycle. Does listen() abstract over both honestly (opaque wake + re-read contract) or promise durability it only has on one engine (that's what streams are for)?
  • The transactional seam (enqueue_tx/publish_tx/save_offset_tx) across two transaction models — the driver-coupled point (§The driver conflict).
  • How a caching subscriber receives sufficient invalidation information (keys? table/channel names? opaque wake + re-read contract?) — rides the same contract decision.

Evidence now in hand (both POCs, 2026-10-04):

  • The opaque-wake + re-read contract holds on both engines, unchanged. SQLite: same data_version mechanism under either posture (wake coalescing verified correct by design, p50 1.4–2.2 ms at the default 1 ms cadence, missed-wake stress passes with correct re-reads on both). Postgres: LISTEN delivers push (~1.1 ms p50, 300/300 isolated, no coalescing needed), no replay, and the reconnect gap is made recoverable by the listener broadcasting a synthetic reconnect-wake on a reserved channel (verified through a killed-connection recovery — subscribers wake and re-read state completely despite the in-gap notification never being delivered). The two engines now share the same wake contract.
  • notify is commit-atomic natively on both (in-tx NOTIFY delivers only at commit; rollback drops it — the exact analogue of honker's notify-in-tx property). Delivery-guarantee split is measurable and native: notify = fire-and-forget (commit-atomic, at-most-once per listener session, no replay); streams = durable with explicit offsets. The trait must NOT promise replay under listen().
  • The *_tx seam resolves to the same shape both engines: caller-held tx handle (*_tx methods on the handle); pg's instance is async-native (tokio-postgres Client is Send+Sync — the handle holds the pooled connection directly, no spawn_blocking), SQLite's is a bridged writer-slot lease. The core-crate TxHandle trait from POC #1's sketch stands unchanged; the per-engine difference is bridging mechanism, not trait shape. Phase 1 starts from that shape plus its two recorded frictions (the as_any_mut downcast and the thread-affinity of rusqlite tx ops — the latter SQLite-side only; erratum 2026-10-04, this line previously said "pg-side only," which garbled POC #2's finding that the affinity friction does not carry over to pg).

OQ-ST-05: Queue semantics — adopt, fork, or re-derive?

Status: open — design work (posture resolved; the remainder is semantics-depth design on measured ground — Phase 1, not further research).

If queues land in scope (OQ-ST-01: they do, documented need), the pg-boss schema family is the Postgres-side incumbent and honker's queue design is the SQLite-side one. Options: adopt pgboss-rs as a dependency (new feature-gated option); targeted-fork the relevant subsystem (alksocks precedent, ported to our conventions); schema/design-reference only (re-derive on our driver). Fork-vs-derive depends on how much of pgboss-rs is queue-machinery vs driver-wiring (the sqlx coupling — OQ-ST-03), on our tolerance for the alpha-state rc port, and on the verified gap (§Prior art → pgboss-rs): the push-reactivity half has to be built on top of any choice, so the queue-machinery reuse value is the honest comparison point, not the whole.

Posture evidence (both POCs, 2026-10-04):

  • SQLite side: the adopt question dissolved — honker-core is consumed as a published-library dependency; the inventory-confirmed feature rows ride its machinery (fork-vs-reference for that consumption is OQ-ST-06's calculus).
  • Postgres side: the re-derive posture is strengthened — the minimal queue table + FOR UPDATE SKIP LOCKED claim + LISTEN wake is ~40 lines of SQL over the pool (all claim/atomicity properties pass in POC #2's suite); reactivity is built by this crate either way (the verified pgboss-rs LISTEN/NOTIFY gap stands). The pg engine's default consumption posture is LISTEN-driven claim with a re-poll safety net; poll-only remains the fallback (measured: p50 3–6 ms vs 32–50 ms).

Remaining question: the semantics depth — retry/backoff/dead-letter/ sweep design on that ground, with pgboss-rs (and the node original) as schema/design reference.

OQ-ST-06: Honker relationship — reference, fork, or vendor?

Status: open — quality-read gate (default posture evidenced; the remainder is the Phase 1 fork-trigger assessment).

Options: design-reference only (read, don't copy); targeted fork of honker-core's engine machinery; vendor the extension. Honker is alpha-quality per its own README, MIT/Apache-2.0 dual-licensed, and covers only the SQLite side — but it embodies exactly the watcher/ transactional design this crate wants on SQLite, and honker-rs demonstrates the interface shape is sound.

Refinements from the 2026-10-03 reading:

  • honker-rs is sync-only (std threads, blocking iterators) — the tokio port is required work under any fork posture, which changes the fork-vs-reference calculus (a fork is already a serious port).
  • The crate likely needs only the core engine machinery (honker-core minus the extension C surface — see OQ-ST-07), a smaller extraction than the whole project.
  • Honker's own documented per-binding inconsistencies (the processing-guarantees table, OQ-ST-04) suggest extracting design+ semantics with our contract pinned, rather than preserving its behavior verbatim — closer to the alkblobs "borrow conclusions, not wire surface" principle than to alksocks' verbatim extraction.

Default posture, evidenced by POC #1 (2026-10-04): depend on the published crate. Under the resolved SQLite posture, honker-core is consumed as a published library (Writer/Readers/SharedUpdateWatcher/ attach_* — nearly all its surface minus the experimental backends), not vendored or forked to ship; 0.5.0 is published with clean deps and the reference-usage posture works as-is. The fork trigger is now specifically: the Phase-1 quality read of honker-core's watcher/transactional core, or a needed change upstream won't take.

Per-subsystem dependency votes (POC #2, 2026-10-04): the pg-side dependencies are published-library use as-is (tokio-postgres 0.7.18 + deadpool-postgres 0.14.2: clean, zero conflicts, actively maintained); postgres-notify 0.3.8 evaluated in-probe and passed over (derive-not-adopt — lazy reconnect, no connect_script on initial connect, unquoted identifier LISTENs, single-maintainer posture; the hand-rolled ~90-line forwarder with test-pinned pitfalls is the preferred shape; fallback if upstream improves). This is the OQ-ST-06 calculus applied per-subsystem, recorded, not a Phase 0 ADR.

OQ-ST-07: SQLite-side scope — loadable extension, embedded rusqlite, or both?

Status: resolved (cut-only, 2026-10-04 via the inventory + POC #1). Honker ships as a loadable extension usable by any SQLite client, plus per-language bindings. This crate (a Rust library) does not need the loadable-extension surface: every identified consumer is in-process Rust attaching to its own connection (the honker-core/attach_honker_functions shape — the alknet-filesystem POC's actual usage), and POC #1 sealed the engine posture as library linkage on our own rusqlite. The loadable-extension option is out unless a consumer appears; nothing remains to extract or decide here.

OQ-ST-08: Multi-host / deployment posture

Status: resolved (2026-10-06, Phase 1 — promoted as OQ-08; ADR-016.)

Honker is explicitly single-machine (file-backed SQLite). Postgres is natively multi-host — POC #2 verified the pg engine side has no single-host assumption to remove (the property tests ran all-through-network over the docker bridge; the listener/wake machinery is connection-based, per-process). The open question is the trait-surface half: the unified surface must not pretend SQLite is multi-host, but where does the honest boundary live — per-engine capability flags? A documented deployment matrix? Does the trait need to expose engine capabilities at all?

Open; rides OQ-ST-04 (the trait's shape constrains where capability differences can surface).

POC register

Proposals run as standalone crates in the global workspace; findings land in docs/research/ here. Named per the OQ each feeds:

# POC Spec Findings
1 SQLite engine posture: honker-core-on-rusqlite vs honker-extension-over-sqlx (async seam, watcher, transactional contract, packaging, interop) poc-sqlite-posture-spec.md poc-sqlite-posture-findings.md — passed (verdict: Arm A; ran 2026-10-04)
2 Postgres engine posture: LISTEN/NOTIFY plumbing, tx-seam over the pool (caller-owned tx vs closure-scoped), wake-vs-poll claim latency, reconnect recovery poc-pg-posture-spec.md poc-pg-posture-findings.md — passed (verdict: tokio-postgres+deadpool, hand-rolled listener, caller-tx seam; ran 2026-10-04)

Phase 0 plan — final state

The expected sequence, with what actually happened:

  1. Consumer-driven scope inventory (OQ-ST-01) — done (2026-10-04): consumer-inventory.md, run against the paused consumers' documents (alkfs phase-0, alkgit architecture, alkblobs architecture + the alknet-filesystem POC). OQ-ST-01 answered down to named per-feature rows; OQ-ST-02/07 sharpened by it.
  2. Crate scope (OQ-ST-02) — resolved (2026-10-04, operator decision): reactive-core + engine crates. Reasoning and the superseded inventory lean recorded at the OQ and in the inventory.
  3. Driver + reactive-shape research rounds (OQ-ST-03/04) — resolved/de-risked by the two POCs:
    • POC #1 (SQLite driver posture) passed — posture 1 (honker-core on our rusqlite).
    • POC #2 (Postgres side: tokio-postgres LISTEN/pool/tx-seam validation) passed — OQ-ST-03 closed with per-engine drivers.
    • OQ-ST-04's remaining contract pinning is Phase 1 paper work over a complete evidence base (§OQ-ST-04).
  4. Ownership decisions (OQ-ST-05/06) — the option space narrowed with the POCs; per-subsystem votes recorded at the OQs. Remaining: the semantics-depth design inputs (OQ-ST-05) and the Phase-1 quality read (OQ-ST-06's fork trigger).
  5. Converge — done (§Convergence). Phase 1 opened 2026-10-04: docs/architecture/ now exists (README index, seven ADRs 001–007 carrying this register's resolved decisions, spec docs, and the promoted open-questions tracker).

References

  • honker — /workspace/honker (git checkout @ f4e53c6 of github.com/russellromney/honker; README + honker-core/src/ read 2026-10-03): the SQLite-side feature/wake template. The four guides (queues/streams/pubsub/scheduler on honker.dev) + packages/honker-rs/src/lib.rs (v0.5.0) are the interface prior art (§Interface finding).
  • pgboss-rs — /workspace/pgboss-rs (git checkout @ 98f7d9e of github.com/rustworthy/pgboss-rs, v0.1.0-rc6; read 2026-10-03, LISTEN/NOTIFY-absence verified): the Postgres queue-family reference.
  • honker's own prior-art section: pg_notify, pg-boss, Oban, Huey — the external lineage this crate inherits from both sides.
  • alkblobs — /workspace/@alkdev/alkblobs (spec+POCs, paused): the paused planning this crate unblocks; its POC findings (poc-postgres-kv-findings.md, poc-pglo-findings.md) are the tokio-postgres evidence base; its store-api.md pins the spawn_blocking engine-execution posture (family precedent with alktty REQ-TTY-01).
  • alkgit — /workspace/@alkdev/alkgit (paused mid-Phase-1, architecture reviewed): its backend.md trait seam and ADR set are consumer evidence for the inventory (queues/locks rows).
  • alkfs — /workspace/@alkdev/alkfs (Phase 0 drafted, 2026-09-23): its phase-0 OQs (OQ-FS-05/07/14/16/17) are consumer evidence for the inventory (notify/locks/queues rows).
  • alknet-filesystem POC — /workspace/@alkdev/alknet/docs/research/alknet-filesystem/ poc-summary.md: the ran-once evidence that the honker-coordination layer works (notify-on-commit test; named-locks and outbox usage identified).
  • consumer-inventory.md (docs/research/consumer-inventory.md) — the per-feature synthesis (2026-10-04) answering OQ-ST-01 from the above.
  • alkcall — /workspace/@alkdev/alkcall: the family substrate; referenced for the store-layer-isolation principle only.
  • alkstore-sqlite-posture-poc — /workspace/alkstore-sqlite-posture-poc (standalone POC crate, published-deps-only): POC #1's code — both arms end-to-end, property tests, seam/watcher probes.
  • alkstore-pg-posture-poc — /workspace/alkstore-pg-posture-poc (standalone POC crate, published-deps-only): POC #2's code — the pg engine posture end-to-end (engine + hand-rolled listener + postgres-notify wrapper), 11-test contract suite, seam/wake/ burst/claim/pollvlisten/pnlisten probes; harness server pglo-poc (postgres:16-alpine, :15432).
  • alktty — /workspace/@alkdev/alktty (architecture reviewed): REQ-TTY-01 (docs/architecture/tty-backend.md) — the async-facing-trait + sync-bridge posture ("backends are not required to be natively async"), the family precedent bearing on OQ-ST-03's async sub-question.