Files
alkstore/docs/research/phase-0.md
T

808 lines
44 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
status: draft
last_updated: 2026-10-04 (Phase 0 complete and the register promoted:
docs/architecture/ opened with ADR-001..007 carrying the resolved
decisions and the OQ tracker mirroring OQ-ST-01..08 1:1. Phase 0
research itself is closed; remaining work is in docs/architecture/
open-questions.md. One erratum fixed in OQ-ST-04's tx-friction line —
thread-affinity is SQLite-side, previously garbled as "pg-side".)
---
# alkstore — Phase 0 (Exploration)
This document captures Phase 0 (Exploration) for the `alkstore` crate:
vision, guiding principles, prior art, the open-question register
(OQ-ST-01..NN), the POC register, and the convergence the phase
objective asks for. Phase 0's objective per `docs/sdd_process.md`:
*capture vision and guiding principles; research options; validate
approaches; converge on a recommended approach.* All of that
objective is now met: the scope question (OQ-ST-01) is answered
per-feature from the consumers' documents, the crate-split (OQ-ST-02)
is decided, the driver question (OQ-ST-03) is closed POC-backed on
both engines, the reactive contract (OQ-ST-04) is de-risked down to
paper work, and the ownership postures (OQ-ST-05/06) have recorded
per-subsystem votes. §Convergence collects the recommendation; the
OQ register's open residue is Phase 1 architecture work, not further
research.
Context for why this crate starts now: **alkblobs**
(`/workspace/@alkdev/alkblobs` — spec + POCs only, paused mid-planning)
hit repeated circular hedging in its Phase 0, and a root cause was that
the storage substrate it deploys onto was itself disjoint and fuzzy — a
repo pattern with a default in-memory adapter across the alk* ecosystem,
cache-invalidation patches in hot paths, non-invalidated caches where
delay was tolerable, no single definition of "how does a change in the
database become visible to other processes/connections?" alkblobs paused
partly to let this crate answer that first. alkstore is the attempt to
make that substrate real once, so downstream stores don't re-derive it.
## Vision and guiding principles
**One sentence (draft):** one reactive store interface over SQLite and
Postgres — durable pub/sub notify, queues, streams, and the transactional
integration (write + enqueue in one transaction) that honker delivers on
SQLite — with the Postgres side building on the natively-available
machinery (`pg_notify`/`LISTEN`, and the pgboss job-queue schema family)
rather than emulating it.
**The honker relationship.** `/workspace/honker` (reference checkout,
not for direct use as a dependency; alpha-quality per its own README,
MIT/Apache-2.0 dual) adds
Postgres-style `NOTIFY`/`LISTEN` semantics to SQLite without a broker:
durable at-least-once queues with retries/delay/priority/visibility
timeouts/dead-letter, durable streams with per-consumer offsets,
cron/`@every` scheduling, named locks, rate limits, transactional
outbox — all as INSERTs inside the caller's transaction, with the
cross-process wake delivered by a shared watcher that polls `PRAGMA
data_version` (default 1 ms → single-digit-ms delivery) and re-reads
indexed state after every wake. Its own Prior Art section names the
lineage: `pg_notify`, pg-boss, Oban, Huey.
**The goal is not "port honker to Postgres." The goal is the
interface:** one store API whose consumer code (queues, streams, notify)
looks the same whether the backing engine is SQLite or Postgres, while
each engine uses its own native wake/delivery story under the hood. The
honker docs recommend pgboss + `pg_notify` for the Postgres equivalent —
that recommendation is the design brief for this crate's Postgres engine.
### The interface finding
*(2026-10-03, from the four honker.dev guides +
`packages/honker-rs/src/lib.rs`, the Rust binding, v0.5.0.)* Honker's
*surface* is already the unified-interface candidate. Its Rust binding
exposes exactly the surface this crate wants, engine-clean:
- `db.queue(name, QueueOpts)` → `enqueue / enqueue_tx / claim_one /
claim_batch / ack_batch / cancel / get_job / sweep_expired /
claim_waker`, with `job.ack / retry / fail / heartbeat` and
`EnqueueOpts {delay, priority, max_attempts, expires, ...}` —
semantically the pg-boss model (visibility timeouts, retries,
dead-letter via move-to-`_honker_dead`), not a LISTEN-emulation.
- `db.stream(name)` → `publish / publish_tx / publish_with_key /
read_since / read_from_consumer / save_offset(_tx) / get_offset /
subscribe(consumer)` — offsets are explicit, transaction-aware
(`save_offset_tx`) for the exactly-once-within-a-business-tx shape,
and replay-on-reconnect is the default.
- `db.notify(channel, payload)` / `notify_tx` / `db.listen(channel)` —
the `pg_notify`-analogue fire-and-forget signal layer ("fire-and-forget,
no replay, no guarantees" — streams are the durable cousin), listener
starts from `MAX(id)` at attach, no historical replay.
- `db.scheduler()` → `add/pause/resume/update/list/remove/tick/run` —
cron + `@every` enqueueing into named queues, leader-elected via
advisory lock with TTL heartbeat, missed-boundary catch-up.
- `db.outbox(name)` — the transactional outbox helper (enqueue +
`run_once` delivery worker).
- `db.try_lock / try_rate_limit / save_result / get_result /
sweep_results` — the coordination/adjacent-tools surface.
The implication flips the framing of the unified-API work: it is not
"invent a shape both engines fit" — it is **"this shape both engines
can fit"** (pg-boss's queue model is *already* the native Postgres
tooling model; streams/offsets and notify have direct Postgres
counterparts) **and the work is pinning which parts of the shape are
the crate's contract** — the delivery-guarantee differences honker's
own guide documents per-binding (auto-checkpoint cadence vs manual
offset save; the processing-guarantees table) are exactly the seams a
single-crate version must clean up. Honker-rs is the concrete prior
art for that pinning exercise.
Postgres side: `pg_notify` gives fast triggers with no retry or
visibility semantics; pg-boss/Oban are the durable-layer gold standards
— "If you already run Postgres, use the Postgres tools."
**Consumer shape (from the paused alkblobs planning):** any crate that
today uses the ecosystem's repo-pattern + in-memory adapter should be
able to swap in an alkstore-backed engine and get durability + true
cross-process/cross-instance reactivity. That means the reactive surface
must compose with *client-side* caching: a subscriber that also holds a
cache can invalidate on notification instead of re-querying or re-polling
— the hot-path pattern the ecosystem already uses, given a real
invalidation source.
Guiding principles:
1. **One interface, two engines, native underneath.** The abstraction
layer unifies the consumer-visible features; the engines stay
dialects, not two emulations of one dialect. SQLite follows honker's
design (queue in the same file, same transaction, watcher-based
wake); Postgres follows pg-boss' design (schema-based job tables +
`pg_notify`-driven wake). A "lowest common denominator" unification
(both sides polling, both sides emulating LISTEN) is explicitly the
failure mode to avoid — it would re-create the fuzziness this crate
exists to remove.
2. **Transactional local-adjacency is the load-bearing property.**
Honker's core claim: enqueue/publish/notify in the same transaction
as the business write; rollback drops both. The unified surface must
preserve this on any engine, because the ecosystem's repo pattern
assumes it (a business write that loses its side-effect notification
is the dual-write problem honker names). Both POCs verified this
property holds natively on their engines — it is now evidence, not
aspiration.
3. **Ownership of the whole stack.** honker is third-party; pgboss-rs
is third-party. Whether any of them are adopted, forked, or used as
schema/design reference only is a deliberate per-question decision —
not inherited by adjacency. The workspace precedent is the targeted
fork (alksocks' fast-socks5 extraction: adopt the design, own the
code, port to our conventions).
4. **Substrate-agnostic consumer API, engine-specific setup.** A
consumer opens a `Store` from a connection string / file path and
gets the same trait surface. Which engine is behind what can vary
(per-deployment config), but consumer code must not branch on
engine type.
5. **No panics, `tokio`, `thiserror`, lean base crate, feature-gated
optional engines** — family-standard, pre-committed (AGENTS.md).
## The driver conflict (Phase 0's central tension — now resolved)
The immediate design fork, flagged by the user early in Phase 0:
a single crate with both engines means a driver decision, and the
reactivity story is entangled with it. The facts that made the tension
real:
- **The alkblobs POCs used `tokio-postgres` + `deadpool-postgres`**
(validated in `poc-postgres-kv-findings.md`, `poc-pglo-findings.md` —
including Large Objects work), while **pgboss-rs**
(`/workspace/pgboss-rs` @ 98f7d9e) uses `sqlx`, and **honker-core
uses `rusqlite`.**
- **pgboss-rs currently has no `LISTEN`/`NOTIFY` at all** (verified
2026-10-03 against the checkout — `src/` contains no LISTEN/NOTIFY
usage; consumption is `fetch_job` polling). So *even* "use pgboss
for the queue" does not deliver reactivity — LISTEN/NOTIFY wiring
would be new work either way, and the driver choice determines
*whose* LISTEN plumbing.
- **Honker's reactivity on SQLite is a watcher polling `PRAGMA
data_version`** — a fundamentally different mechanism from
LISTEN/NOTIFY. The unified reactive trait must abstract over both
without collapsing to the polling behavior of the weaker side.
- **honker-rs is sync** (`std` threads + blocking iterators;
parking_lot + rusqlite, no tokio) while the family standard is
tokio-async — so even the SQLite side looked like a port-and-adapt,
and the async question was entangled with whether rusqlite-in-a-pool
or a native-async driver (sqlx sqlite) is the right shape.
Two corrections sharpened the framing before the POCs landed: the
honker-rs *interface* is largely driver-independent, with the
transactional seam (`*_tx` methods assuming a live transaction handle
from the caller's driver) as the one genuinely driver-coupled design
point; and the alktty REQ-TTY-01 family precedent ("backends are not
required to be natively async" — bridge-at-the-seam is a supported
posture, not a workaround) already blesses the sync-machinery +
async-facing-trait shape, weakening native-async's main differentiator.
The tension dissolved with the POCs (details and measurements at
OQ-ST-03): **per-engine drivers under OQ-ST-02's crate split**, with
the bridged-rusqlite seam measurably *faster* than sqlx's native async
on SQLite, and tokio-postgres natively async-native on Postgres. What
pgboss-rs genuinely offers (schema DDL, job states, retry semantics,
the node-compatible API) is design reference regardless of driver;
the push channel is this crate's own work either way (OQ-ST-05).
## Prior art
Notes below are from reading the checkouts on 2026-10-03; both external
projects are reference checkouts — read freely, but not for direct use
as a dependency. We use the published version of anything that lives in
the global workspace unless we vendor or fork it (the alksocks
fast-socks5 precedent); if adoption ever requires a fork, forking is
normal work we own, not an exception. Provenance/licensing gets
recorded per AGENTS.md §3 when code is adopted, not while only reading.
### honker — the SQLite-side template
`/workspace/honker` (checkout @ f4e53c6; SQLite extension +
bindings). What matters for this crate:
- **The full feature set to match on Postgres** (its §What It Does):
notify/listen across processes, durable at-least-once queues
(retries, delayed jobs, priority, visibility timeouts, dead-letter
rows, result storage), durable streams with per-consumer offsets,
cron/`@every` scheduling, named locks, rate limits, transactional
outbox helpers. Deliberately excluded there: workflow DAGs, task
chains/chords, multi-writer replication, cross-machine locking —
scope line likely inherited, to be confirmed.
- **The wake mechanism** — `PRAGMA data_version` polling watcher
(default 1 ms; raise for idle CPU), re-read indexed state after
wake, overtriggering on purpose ("one indexed SELECT is cheap; a
missed wake is a correctness bug."). Optional kernel-events and WAL
shared-memory backends exist in source builds.
- **Single-machine honesty** — file-backed, one host; NFS-two-writers
explicitly not supported. This posture needs an explicit Postgres
counterpart (multi-host is Postgres' normal case, so the interface
must not bake SQLite's single-host assumption into the shared
surface).
- **The transactional enqueue shape** — every feature is an INSERT
inside the caller's transaction. This is the pattern the unified API
must keep visible and cheap.
- **The honker-rs binding is the concrete interface prior art** (v0.5.0,
`packages/honker-rs`, read 2026-10-03): the full surface per
§Interface finding. Notable honest limitations documented by its own
guides — the per-binding processing-guarantees table (auto-checkpoint
cadence vs manual offset save; several bindings "may persist an
offset on a cadence... without knowing whether downstream application
work committed"), the Node reverse-order consumer-checkpoint bug,
per-binding feature gaps (JVM missing cancel/get_job, Go/Bun/C++
missing typed pruning) — are exactly the seams a single-crate version
designed-for-the-contract from day one can clean up. Its sync-only
shape (std threads, blocking iterators, no tokio) is a
port-and-adapt constraint, not an adopt candidate as-is.
### pgboss-rs — the Postgres queue family reference
`/workspace/pgboss-rs` (checkout @ 98f7d9e; v0.1.0-rc6, MIT/Apache-2.0
dual). Ported from node pg-boss: builder-based queue/job API,
retry/delay/priority/singleton/dead-letter concepts, `sqlx` 0.8,
schema-scoped DDL.
- Verified gap (2026-10-03): **no LISTEN/NOTIFY** anywhere in `src/`
— consumption is polling `fetch_job`. Any push-reactivity is new
work, not an adoption freebie. This is the substantive difference
between pgboss-rs and what alkstore needs: regardless of fork vs
re-derive, reactivity is ours to build on the Postgres side either
way.
- Its value as reference: the pg-boss schema family (job states,
maintenance/dead-letter behavior) is battle-tested against real
Postgres semantics — worth borrowing *as design*, independent of the
driver decision.
- The node original (pg-boss) is the upstream of record for semantics
the port may have dropped; compare against it when adopting queue
semantics.
### Honker's Postgres-side recommendation
The honker README's own posture: if you run Postgres, use the Postgres
tools. `pg_notify` + pgboss is the recommended assembly. The design
brief: the queue machinery from the pg-boss family, the push semantics
from LISTEN/NOTIFY, the unified API shape from honker's Rust binding.
### The alk* repo pattern — what this crate replaces
The ecosystem's current shape: a repository trait with a default
in-memory adapter; cache-invalidation wiring in hot paths; uninvalidated
(non-reactive) caches where delay was acceptable; each project
composing these slightly differently. No persistence-backed reactive
substrate exists in the family — alkblobs was the first project to try
to plan against one, found it missing, and paused. This crate's reason
to exist is precisely that that substrate should exist once, well,
instead of per-project approximations.
### alkcall — the substrate (not a dependency of the store layer)
`/workspace/@alkdev/alkcall` (pure protocol crate, no transport). The
alk* crates (alktty, alktunnels, alksocks) are its consumers; a future
alkstore ops/protocol surface (if this crate ever exposes store access
over alkcall channels) rides the same substrate. Like the alkblobs
split (store layer stays substrate-free), the store layer here stays
alkcall-free; any networked surface is an ops module/sibling concern
and a separate decision.
## Convergence
Phase 0's objective — converge on a recommended approach — is met.
The recommendation, assembled from the OQ resolutions below (each
carries its own evidence):
**Shape (OQ-ST-02):** a reactive-core crate carrying the trait
surface/types, plus per-engine crates implementing it (SQLite,
Postgres; a mem-shaped test engine as a third impl if useful). The
split makes the engines' real asymmetry of work structural: the
SQLite engine rides honker's existing machinery; the Postgres engine
is the build-heavy side; any future engine is additive. It also keeps
each engine binary single-driver, which the dependency constraints
below effectively require.
**Engines (OQ-ST-03):**
- *SQLite engine:* rusqlite + published `honker-core` 0.5.0
(POC posture 1) — we own connection/schema/watcher wiring per the
alknet-filesystem POC's shape; the async seam is the bridge-at-the-seam
posture (writer-slot + `spawn_blocking`, REQ-TTY-01 precedent),
measured ~2× faster at p50 than sqlx's native async; honker-core's
`SharedUpdateWatcher` is inherited for wake (p50 ≈ 1.4 ms,
battle-tested failure handling). No `.so` runtime artifact, no
vendored patches.
- *Postgres engine:* tokio-postgres + deadpool-postgres (POC #2) —
pooled connections for queries/claims (`FOR UPDATE SKIP LOCKED`),
a dedicated non-pooled listener connection with a hand-rolled ~90-line
LISTEN forwarder (immediate reconnect, synthetic reconnect-wake on a
reserved channel closing the no-replay hole, identifier quoting),
LISTEN-driven claim beats poll 5–16× at p50. `postgres-notify`
evaluated and passed over (derive-not-adopt).
**Contract starting shape (OQ-ST-04):** the honker-rs surface
(§Interface finding), scoped to the inventory-confirmed features, is
the starting artifact for contract pinning. The load-bearing pieces
hold identically on both engines, POC-verified: the opaque-wake +
re-read listener contract; notify = fire-and-forget commit-atomic (no
replay) vs streams = durable with explicit per-consumer offsets; and
the caller-held tx handle (`*_tx` on the handle) whose per-engine
difference is bridging mechanism, not trait shape.
**Ownership (OQ-ST-05/06):** published-library dependencies as the
default posture — honker-core (SQLite), tokio-postgres +
deadpool-postgres (Postgres); the pg queue machinery is re-derived on
our driver with the pg-boss schema family as design reference; the
hand-rolled listener forwarder replaces `postgres-notify`. Named fork
triggers remain: a Phase 1 quality read of honker-core's
watcher/transactional core, or a needed change upstream won't take.
**Scope (OQ-ST-01):** notify/listen, named locks, queues, outbox,
scheduler, streams are in (with the inventory's per-row evidence
grades); rate limits and result storage are cut-flags; the loadable
extension surface is out (OQ-ST-07); honker's exclusion lines (DAGs,
task chains/chords, multi-writer replication, distributed locking)
stay out.
**Phase 1 inherits, as architecture work over complete evidence:**
the contract-pinning itself (OQ-ST-04's remainder — which surface
parts become contract, per the inventory rows, and the per-engine
capability surface, OQ-ST-08); queue semantics depth
(retry/backoff/dead-letter/sweep design, OQ-ST-05); the honker-core
quality read (OQ-ST-06's fork trigger); and the deployment-matrix /
capability-flags decision (OQ-ST-08). No further Phase 0 research is
required. The known deployment constraints to design around:
honker-core 0.5.0 pins rusqlite ^0.40.1 whose rustc requirement
(≥1.99) is a deployment note; pooled connections cannot carry LISTEN
(deadpool#360) so the listener connection is a per-process budget line
outside the pool; notify payloads are ≤ 8000 bytes (large payloads
ride a table row with the id in the notification — the honker-outbox
shape); the reserved reconnect-wake channel name needs a namespace
convention in the contract.
## Open Questions
Register in `docs/research/phase-0.md`; IDs OQ-ST-NN (stable, append
only). Promotion target: Phase 1 `docs/architecture/open-questions.md`
— **promoted 2026-10-04**: OQ-ST-01..08 mirror 1:1 to OQ-01..08 there
(their statuses and resolutions carried; resolved decisions carried
into ADRs 001–007), and new Phase 1 questions append from OQ-09.
Status conventions: `resolved` (evidence recorded here); `open —
<work-type>` where the work-type names the remaining work and the
entry is de-risked (the remainder is Phase 1 architecture work, not
further research); `open` (genuinely open).
### OQ-ST-01: Scope boundary — which honker features are in-scope?
**Status: resolved (2026-10-04).** Answered by the consumer inventory —
`docs/research/consumer-inventory.md`; scope votes shrink to named
rows, not the whole feature list. The original framing ("blocked on a
consumer-driven inventory pass") was circular hedging: the consumers
are paused, so the input would never arrive — but their *documents*
are stable evidence, and the inventory walks them per feature.
- In scope, first-class: **notify/listen** (pinned: alkfs path-tree
invalidation), **streams** (operator-authority record —
type-filtered event watching from several places, e.g. repo-change
subscriptions; a reactivity requirement notify cannot serve honestly,
being fire-and-forget).
- In scope: **named locks** (pinned: alkblobs fleet sweeper lock;
documented: alkfs OQ-FS-05 writer coordination), **queues + the
outbox helper** (documented: alkfs sync/fetch-on-miss outbox;
alkblobs embedder-owned maintenance cadence), **scheduler**
(documented-thin: the family-wide "who sweeps/renews/reaps" problem,
possibly collapsing into queues — watch at OQ-ST-04).
- Cut-flags (no named consumer; carried per the
keep-until-implementation posture, cut later rather than silently
included): **rate limits** (alkgit enforces budgets in its own wire
layer — an in-crate alternative exists), **result storage**.
- Out: honker's exclusion lines (DAGs, task chains/chords,
multi-writer replication, distributed locking) — no consumer names
these either; they stay out unless a consumer document grows one.
New consumers (alksftp, the alknet rewrite) add a row to the inventory
*before* being assumed into scope.
### OQ-ST-02: Crate scope — one store crate, or reactive-core + engines?
**Status: resolved (2026-10-04, operator decision): reactive-core +
engine crates** — a core crate carrying the trait surface/types,
per-engine crates implementing it (sqlite, postgres; mem-shaped test
engine as a third impl if useful). Options considered: single crate
with feature-gated engines (the alk* feature-gate pattern); a core
trait crate + per-engine crates; engine crates consuming a thin core.
The reasoning, recorded because it overrode the inventory's lean: the
split isolates the engines' real asymmetry of work — the SQLite engine
rides honker's existing machinery as the baseline (port-and-adapt
sync→async), the Postgres engine is the build-heavy side
(LISTEN/NOTIFY wiring + pg-boss-family schema work, OQ-ST-05) — and it
makes any future engine (alkfs's in-tree needs, an ops-surface engine)
additive rather than a feature-graph edit to one crate.
Base-crate-lean becomes structural rather than a feature-discipline.
It also keeps each engine binary single-driver, which the
libsqlite3-sys link-collision constraint (OQ-ST-03) effectively
requires.
The inventory's uniform-feature-family fact still stands, not
contradicted: it reads as "the core contract can stay small — one
feature family, both engines," not as an argument for the crates to
merge (its single-crate lean was inductive from that fact; the
structural reasoning here supersedes it — correction recorded in
consumer-inventory.md too).
### OQ-ST-03: Driver story — sqlx, tokio-postgres, or per-engine drivers?
**Status: resolved (2026-10-04, both halves, POC-backed): per-engine
drivers — rusqlite + honker-core 0.5.0 (SQLite engine); tokio-postgres
0.7.18 + deadpool-postgres 0.14.2 (Postgres engine) — under OQ-ST-02's
per-engine-crate split.**
Original option list (the decision as first framed): one driver across
engines (sqlx: both engines native, one API — but the alkblobs POC
evidence is tokio-postgres); per-engine drivers under a unified trait
(tokio-postgres + deadpool-postgres, POC-validated in alkblobs
findings, + rusqlite, honker's choice); adopt/fork pgboss-rs (brings
sqlx along where the queue lives). Honest unknowns at framing time:
does sqlx support SQLite `data_version`/extension-style machinery
equally well; does a unified trait over `(tokio-postgres, rusqlite)`
pay more trait-fitting cost than sqlx's single-API convenience costs
elsewhere; extension loading under sqlx vs rusqlite; and the async
question honker-rs's sync shape forces — rusqlite-in-a-pool with a
bridge, or a native-async driver?
The SQLite option space, named explicitly (2026-10-04, operator +
verified against the checkout @ f4e53c6) — three distinct postures, not
one "rusqlite vs sqlx" axis:
1. **honker-core on our own rusqlite connection** (the
`attach_honker_functions` shape — the alknet-filesystem POC's actual
usage): we own the connection, the schema bootstrap, and the
watcher; honker supplies the SQL-function machinery.
2. **honker-rs as the crate's SQLite substrate** (`Database::open`,
typed Queue/Stream/Transaction primitives): maximum reuse, least
control — honker-rs opens and holds its own connections, its
`Database` wraps a connection mutex (transactions pin the mutex;
same-thread `*_tx` methods only), and the whole engine is sync
under OUR async core (bridge at every seam).
3. **raw SQL over sqlx-sqlite with the honker loadable extension**
(`SqliteConnectOptions::extension(ext)` + `SELECT
honker_bootstrap()`, then every feature is plain SQL callable
through `SqliteExecutor<'e>` — pool, connection, AND Transaction
alike). Verified in-harness: the pattern is CI-proven in the honker
checkout itself (scripts/proof/orm/rust — async business-write +
`honker_enqueue` inside a `conn.begin()` tx, commit-visibility +
rollback-drops-job asserted).
Family precedent bearing on the async sub-question (2026-10-04): the
async-facing-trait + sync-bridge posture is family-standard, twice
over — alktty REQ-TTY-01 ("backends are not required to be natively
async": blocking work on dedicated threads or `spawn_blocking` feeding
tokio channels is a documented, supported implementation strategy, not
a workaround; the wezterm/portable_pty pattern), and alkblobs
store-api.md (blocking file work in `spawn_blocking` inside engine
impls). This reframed the sync→async port: less "rewrite onto a
native-async driver," more "keep the sync machinery and bridge at the
trait seam" — weakening sqlx's main differentiator for the SQLite side.
**Resolution, SQLite half — POC #1 (2026-10-04; findings:
`poc-sqlite-posture-findings.md`): posture 1 — honker-core on our
rusqlite.** All three gate conditions fired in A's favor, mildly: the
bridged path is ~2× sqlx's native-async at p50 (0.354 vs 0.707 ms on
the tx-enqueue workload; B's premise measured false), honker-core's
inherited watcher is tighter than a re-derived one (p50 1.40 vs 2.15
ms, max 29 vs 172 ms, with battle-tested failure handling), and the
`.so` runtime dependency is packaging cost A doesn't pay for no
compensating advantage. The transactional contract holds identically
on both (it is SQLite's property, not the posture's). Constraints
recorded for the engine crate regardless:
honker-core=0.5.0 pins rusqlite ^0.40.1 whose rustc requirement
(≥1.99) is a deployment note, and mixed rusqlite+sqlx binaries
currently need a vendored one-line libsqlite3-sys patch — OQ-ST-02's
per-engine-crate split is what keeps the engine binary single-driver.
**Resolution, Postgres half — POC #2 (2026-10-04; findings:
`poc-pg-posture-findings.md`): tokio-postgres + deadpool-postgres.**
All three gate conditions held: the transactional property is native
(in-tx NOTIFY delivers only at commit; rollback drops job row +
business row + notification), exactly-once claim via `FOR UPDATE SKIP
LOCKED`, and the LISTEN wake layer is push (p50 1.1 ms, 300/300;
claim latency 3–6 ms vs poll-only 32–50 ms — 5–16×, the push channel
pgboss-rs lacks, measured). Structural findings the contract must
absorb: pooled connections cannot carry LISTEN (deadpool#360 — pinned
as our own test; the dedicated listener connection is a per-process
budget line outside the pool), and the `*_tx` seam differs from
SQLite's exactly where expected (pg's client is Send+Sync — the tx
handle is held directly across awaits; SQLite's is a bridged
writer-slot lease — same seam shape, different bridging). Listener
substrate vote: hand-rolled ~90-line forwarder (immediate reconnect,
quoting control, no dependency posture) over `postgres-notify` 0.3.8
(lazy reconnect, connect_script skipped at initial connect,
unquoted-identifier LISTEN — derive-not-adopt; fallback if upstream
improves). The sqlx `PgListener` fallback retired unfired.
### OQ-ST-04: The reactive abstraction — what does the unified notify surface look like?
**Status: open — contract work (de-risked; both engine sides
POC-verified). The remainder is contract-pinning paper work over a
complete evidence base — Phase 1, not further research.**
Background: the two engines' wake mechanisms are structurally
different — SQLite = watcher polling `PRAGMA data_version` (deliver
on commit; no server-side push exists), Postgres = LISTEN/NOTIFY
(server push, connection-bound, no retry/visibility semantics). The
reactive trait must have a shape both implement without one emulating
the other's weaknesses.
The honker-rs surface (§Interface finding) is the concrete starting
point — the work decomposes into contract-pinning rather than
shape-invention:
- Which parts of the honker-rs surface become the crate's *contract*:
the `notify`/`listen` pair, the stream/offset/consumer model (in
scope per the inventory — subscriptions are the durable reactivity
half notify can't serve), the queue claim/ack/visibility model,
locks, outbox, scheduler — all have named consumers now; rate-limits
have an in-crate alternative mechanism (alkgit's wire layer) —
subset, renamed/regrouped, decided against the inventory
rows rather than against the whole honker menu.
- What is the delivery-guarantee contract, per mechanism (honker's
own guide table shows how easily per-binding auto-checkpoint vs
manual-save ambiguity produces *different* guarantees under one
function name — the single-crate version must pick one answer, not
inherit the table)?
- Listener semantics: honker starts from `MAX(id)` and replays
nothing; Postgres LISTEN has no replay either but delivers via a
dedicated connection with its own lifecycle. Does `listen()`
abstract over both honestly (opaque wake + re-read contract) or
promise durability it only has on one engine (that's what streams
are for)?
- The transactional seam (`enqueue_tx`/`publish_tx`/`save_offset_tx`)
across two transaction models — the driver-coupled point (§The
driver conflict).
- How a caching subscriber receives sufficient invalidation
information (keys? table/channel names? opaque wake + re-read
contract?) — rides the same contract decision.
Evidence now in hand (both POCs, 2026-10-04):
- **The opaque-wake + re-read contract holds on both engines,
unchanged.** SQLite: same `data_version` mechanism under either
posture (wake coalescing verified correct by design, p50 1.4–2.2 ms
at the default 1 ms cadence, missed-wake stress passes with correct
re-reads on both). Postgres: LISTEN delivers push (~1.1 ms p50,
300/300 isolated, no coalescing needed), no replay, and the
reconnect gap is made recoverable by the listener broadcasting a
synthetic reconnect-wake on a reserved channel (verified through a
killed-connection recovery — subscribers wake and re-read state
completely despite the in-gap notification never being delivered).
The two engines now share the *same* wake contract.
- **notify is commit-atomic natively on both** (in-tx NOTIFY delivers
only at commit; rollback drops it — the exact analogue of honker's
notify-in-tx property). Delivery-guarantee split is measurable and
native: notify = fire-and-forget (commit-atomic, at-most-once per
listener session, no replay); streams = durable with explicit
offsets. The trait must NOT promise replay under `listen()`.
- **The `*_tx` seam resolves to the same shape both engines:**
caller-held tx handle (`*_tx` methods on the handle); pg's instance
is async-native (tokio-postgres Client is Send+Sync — the handle
holds the pooled connection directly, no spawn_blocking), SQLite's
is a bridged writer-slot lease. The core-crate `TxHandle` trait
from POC #1's sketch stands unchanged; the per-engine difference is
bridging mechanism, not trait shape. Phase 1 starts from that shape
plus its two recorded frictions (the `as_any_mut` downcast and the
thread-affinity of rusqlite tx ops — the latter SQLite-side only;
erratum 2026-10-04, this line previously said "pg-side only," which
garbled POC #2's finding that the affinity friction does not carry
over to pg).
### OQ-ST-05: Queue semantics — adopt, fork, or re-derive?
**Status: open — design work (posture resolved; the remainder is
semantics-depth design on measured ground — Phase 1, not further
research).**
If queues land in scope (OQ-ST-01: they do, documented need), the
pg-boss schema family is the Postgres-side incumbent and honker's
queue design is the SQLite-side one. Options: adopt pgboss-rs as a
dependency (new feature-gated option); targeted-fork the relevant
subsystem (alksocks precedent, ported to our conventions);
schema/design-reference only (re-derive on our driver). Fork-vs-derive
depends on how much of pgboss-rs is queue-machinery vs driver-wiring
(the sqlx coupling — OQ-ST-03), on our tolerance for the alpha-state
rc port, and on the verified gap (§Prior art → pgboss-rs): the
push-reactivity half has to be built on top of any choice, so the
queue-machinery reuse value is the honest comparison point, not the
whole.
**Posture evidence (both POCs, 2026-10-04):**
- SQLite side: the adopt question dissolved — honker-core is consumed
as a published-library dependency; the inventory-confirmed feature
rows ride its machinery (fork-vs-reference for that consumption is
OQ-ST-06's calculus).
- Postgres side: the re-derive posture is strengthened — the minimal
queue table + `FOR UPDATE SKIP LOCKED` claim + LISTEN wake is ~40
lines of SQL over the pool (all claim/atomicity properties pass in
POC #2's suite); reactivity is built by this crate either way (the
verified pgboss-rs LISTEN/NOTIFY gap stands). The pg engine's
default consumption posture is LISTEN-driven claim with a re-poll
safety net; poll-only remains the fallback (measured: p50 3–6 ms
vs 32–50 ms).
Remaining question: the *semantics depth* — retry/backoff/dead-letter/
sweep design on that ground, with pgboss-rs (and the node original)
as schema/design reference.
### OQ-ST-06: Honker relationship — reference, fork, or vendor?
**Status: open — quality-read gate (default posture evidenced;
the remainder is the Phase 1 fork-trigger assessment).**
Options: design-reference only (read, don't copy); targeted fork of
honker-core's engine machinery; vendor the extension. Honker is
alpha-quality per its own README, MIT/Apache-2.0 dual-licensed, and
covers only the SQLite side — but it embodies exactly the watcher/
transactional design this crate wants on SQLite, and honker-rs
demonstrates the interface shape is sound.
Refinements from the 2026-10-03 reading:
- honker-rs is **sync-only** (std threads, blocking iterators) — the
tokio port is required work under any fork posture, which changes
the fork-vs-reference calculus (a fork is already a serious port).
- The crate likely needs only the core engine machinery (honker-core
minus the extension C surface — see OQ-ST-07), a smaller extraction
than the whole project.
- Honker's own documented per-binding inconsistencies (the
processing-guarantees table, OQ-ST-04) suggest extracting *design+
semantics* with our contract pinned, rather than preserving its
behavior verbatim — closer to the alkblobs "borrow conclusions, not
wire surface" principle than to alksocks' verbatim extraction.
**Default posture, evidenced by POC #1 (2026-10-04):** depend on the
published crate. Under the resolved SQLite posture, honker-core is
consumed as a *published library* (Writer/Readers/SharedUpdateWatcher/
attach_* — nearly all its surface minus the experimental backends),
not vendored or forked to ship; 0.5.0 is published with clean deps and
the reference-usage posture works as-is. The fork trigger is now
specifically: the Phase-1 quality read of honker-core's
watcher/transactional core, or a needed change upstream won't take.
**Per-subsystem dependency votes (POC #2, 2026-10-04):** the pg-side
dependencies are published-library use as-is (tokio-postgres 0.7.18 +
deadpool-postgres 0.14.2: clean, zero conflicts, actively maintained);
`postgres-notify` 0.3.8 evaluated in-probe and passed over
(derive-not-adopt — lazy reconnect, no connect_script on initial
connect, unquoted identifier LISTENs, single-maintainer posture; the
hand-rolled ~90-line forwarder with test-pinned pitfalls is the
preferred shape; fallback if upstream improves). This is the OQ-ST-06
calculus applied per-subsystem, recorded, not a Phase 0 ADR.
### OQ-ST-07: SQLite-side scope — loadable extension, embedded rusqlite, or both?
**Status: resolved (cut-only, 2026-10-04 via the inventory + POC #1).**
Honker ships as a loadable extension usable by *any* SQLite client,
plus per-language bindings. This crate (a Rust library) does not need
the loadable-extension surface: every identified consumer is in-process
Rust attaching to its own connection (the
`honker-core`/`attach_honker_functions` shape — the alknet-filesystem
POC's actual usage), and POC #1 sealed the engine posture as library
linkage on our own rusqlite. The loadable-extension option is out
unless a consumer appears; nothing remains to extract or decide here.
### OQ-ST-08: Multi-host / deployment posture
**Status: resolved (2026-10-06, Phase 1 — promoted as OQ-08;
[ADR-016](../architecture/decisions/016-deployment-honesty.md).)**
Honker is explicitly single-machine (file-backed SQLite). Postgres is
natively multi-host — POC #2 verified the pg engine side has no
single-host assumption to remove (the property tests ran
all-through-network over the docker bridge; the listener/wake
machinery is connection-based, per-process). The open question is the
*trait-surface* half: the unified surface must not pretend SQLite is
multi-host, but where does the honest boundary live — per-engine
capability flags? A documented deployment matrix? Does the trait need
to expose engine capabilities at all?
Open; rides OQ-ST-04 (the trait's shape constrains where capability
differences can surface).
## POC register
Proposals run as standalone crates in the global workspace; findings
land in `docs/research/` here. Named per the OQ each feeds:
| # | POC | Spec | Findings |
|---|---|---|---|
| 1 | SQLite engine posture: honker-core-on-rusqlite vs honker-extension-over-sqlx (async seam, watcher, transactional contract, packaging, interop) | [poc-sqlite-posture-spec.md](poc-sqlite-posture-spec.md) | [poc-sqlite-posture-findings.md](poc-sqlite-posture-findings.md) — **passed** (verdict: Arm A; ran 2026-10-04) |
| 2 | Postgres engine posture: LISTEN/NOTIFY plumbing, tx-seam over the pool (caller-owned tx vs closure-scoped), wake-vs-poll claim latency, reconnect recovery | [poc-pg-posture-spec.md](poc-pg-posture-spec.md) | [poc-pg-posture-findings.md](poc-pg-posture-findings.md) — **passed** (verdict: tokio-postgres+deadpool, hand-rolled listener, caller-tx seam; ran 2026-10-04) |
## Phase 0 plan — final state
The expected sequence, with what actually happened:
1. **Consumer-driven scope inventory (OQ-ST-01)** — done
(2026-10-04): `consumer-inventory.md`, run against the paused
consumers' documents (alkfs phase-0, alkgit architecture, alkblobs
architecture + the alknet-filesystem POC). OQ-ST-01 answered down
to named per-feature rows; OQ-ST-02/07 sharpened by it.
2. **Crate scope (OQ-ST-02)** — resolved (2026-10-04, operator
decision): reactive-core + engine crates. Reasoning and the
superseded inventory lean recorded at the OQ and in the inventory.
3. **Driver + reactive-shape research rounds (OQ-ST-03/04)** —
resolved/de-risked by the two POCs:
- POC #1 (SQLite driver posture) passed — posture 1 (honker-core
on our rusqlite).
- POC #2 (Postgres side: tokio-postgres LISTEN/pool/tx-seam
validation) passed — OQ-ST-03 closed with per-engine drivers.
- OQ-ST-04's remaining contract pinning is Phase 1 paper work
over a complete evidence base (§OQ-ST-04).
4. **Ownership decisions (OQ-ST-05/06)** — the option space narrowed
with the POCs; per-subsystem votes recorded at the OQs. Remaining:
the semantics-depth design inputs (OQ-ST-05) and the Phase-1
quality read (OQ-ST-06's fork trigger).
5. **Converge** — done (§Convergence). Phase 1 opened 2026-10-04:
`docs/architecture/` now exists (README index, seven ADRs
001–007 carrying this register's resolved decisions, spec docs,
and the promoted open-questions tracker).
## References
- honker — `/workspace/honker` (git checkout @ f4e53c6 of
github.com/russellromney/honker; README + `honker-core/src/` read
2026-10-03): the SQLite-side feature/wake template. The four guides
(queues/streams/pubsub/scheduler on honker.dev) +
`packages/honker-rs/src/lib.rs` (v0.5.0) are the interface prior
art (§Interface finding).
- pgboss-rs — `/workspace/pgboss-rs` (git checkout @ 98f7d9e of
github.com/rustworthy/pgboss-rs, v0.1.0-rc6; read 2026-10-03,
LISTEN/NOTIFY-absence verified): the Postgres queue-family
reference.
- honker's own prior-art section: pg_notify, pg-boss, Oban, Huey —
the external lineage this crate inherits from both sides.
- alkblobs — `/workspace/@alkdev/alkblobs` (spec+POCs, paused): the
paused planning this crate unblocks; its POC findings
(`poc-postgres-kv-findings.md`, `poc-pglo-findings.md`) are the
tokio-postgres evidence base; its store-api.md pins the
spawn_blocking engine-execution posture (family precedent with
alktty REQ-TTY-01).
- alkgit — `/workspace/@alkdev/alkgit` (paused mid-Phase-1,
architecture reviewed): its backend.md trait seam and ADR set are
consumer evidence for the inventory (queues/locks rows).
- alkfs — `/workspace/@alkdev/alkfs` (Phase 0 drafted, 2026-09-23):
its phase-0 OQs (OQ-FS-05/07/14/16/17) are consumer evidence for
the inventory (notify/locks/queues rows).
- alknet-filesystem POC —
`/workspace/@alkdev/alknet/docs/research/alknet-filesystem/
poc-summary.md`: the ran-once evidence that the honker-coordination
layer works (notify-on-commit test; named-locks and outbox usage
identified).
- consumer-inventory.md (`docs/research/consumer-inventory.md`) — the
per-feature synthesis (2026-10-04) answering OQ-ST-01 from the
above.
- alkcall — `/workspace/@alkdev/alkcall`: the family substrate;
referenced for the store-layer-isolation principle only.
- alkstore-sqlite-posture-poc — `/workspace/alkstore-sqlite-posture-poc`
(standalone POC crate, published-deps-only): POC #1's code — both
arms end-to-end, property tests, seam/watcher probes.
- alkstore-pg-posture-poc — `/workspace/alkstore-pg-posture-poc`
(standalone POC crate, published-deps-only): POC #2's code — the
pg engine posture end-to-end (engine + hand-rolled listener +
postgres-notify wrapper), 11-test contract suite, seam/wake/
burst/claim/pollvlisten/pnlisten probes; harness server `pglo-poc`
(postgres:16-alpine, :15432).
- alktty — `/workspace/@alkdev/alktty` (architecture reviewed):
REQ-TTY-01 (`docs/architecture/tty-backend.md`) — the
async-facing-trait + sync-bridge posture ("backends are not
required to be natively async"), the family precedent bearing on
OQ-ST-03's async sub-question.