docs: open Phase 1 — architecture spec set over the Phase 0 evidence

docs/architecture/ now exists: README index, overview, five component
specs (core-contract, engine-sqlite, engine-postgres, queues,
deployment), ADR-001..007 carrying the Phase 0 resolved decisions
(crate split, feature scope, per-engine drivers, dependency
ownership, wake contract, tx seam), and the centralized
open-questions tracker promotion: OQ-ST-01..08 mirror to OQ-01..08
one-to-one with statuses/resolutions carried; new Phase 1 questions
append (OQ-09 scheduler collapse, OQ-10 contract versioning).
Open Phase 1 work: OQ-04 contract pinning (high), OQ-05 queue
semantics depth (high), OQ-06 honker-core quality read (high;
fork-trigger gate), OQ-08 capability surface, OQ-09, OQ-10.

Erratum fixed in phase-0 OQ-ST-04 (thread-affinity friction is
SQLite-side, previously garbled as pg-side) and a stale scheduler-
boundary pointer corrected in consumer-inventory.md. Two review
passes run (findings: OQ-promotion numbering faithfulness, ADR
back-reference sync) — all critical/warning findings resolved.
This commit is contained in:
glm-5.3-flash committed 2026-10-04 18:13:10 +00:00
1 parent db73678090
commit 4391f6e879
17 files changed
+1812 -14

No files matched your search

@@ -0,0 +1,88 @@
# ADR-001: Reactive-core crate + per-engine crates
## Status
Accepted
## Context
The store must present one reactive interface (notify, streams, queues,
locks, scheduler, outbox) over two backing engines, SQLite and Postgres.
The engines differ sharply in how much of that surface is natively
theirs:
- The **SQLite** engine rides honker's existing, battle-tested machinery
(published `honker-core` — a port-and-adapt of a sync library to the
family's async-facing-trait posture).
- The **Postgres** engine is the build-heavy side: LISTEN/NOTIFY wiring
and the pg-boss-family queue schema are built by this crate.
A single crate with feature-gated engines would force both engines'
dependency graphs into every downstream binary that enables both
features. That is not hypothetical: rusqlite 0.40's `libsqlite3-sys`
link collision with sqlx's sqlite driver is a real constraint
(see [ADR-003](003-sqlite-driver.md)); single-driver binaries are the
only clean escape.
Scope evidence lives in `docs/research/consumer-inventory.md`; the
operator decision was recorded 2026-10-04 against
OQ-ST-02 (`docs/research/phase-0.md`).
## Decision
The project ships as a **family of crates**:
1. **`alkstore` (core)** — the crate carrying the unified trait
surface, types, error model, capability flags, and the contract
documentation. No driver dependencies. Compile-lean by construction:
the base crate has no engine machinery to keep out.
2. **`alkstore-sqlite`** — the SQLite engine implementing the core
surface. Single driver: rusqlite + honker-core ([ADR-003]).
3. **`alkstore-postgres`** — the Postgres engine implementing the core
surface. Single driver: tokio-postgres + deadpool-postgres
([ADR-004]).
4. **A mem-shaped engine** (in-tree or separate test crate) may exist as
a third implementation for tests and doctests, if the test story
turns out to want it. Treated as an implementation convenience, not
a contract artifact — decided at implementation time, not now.
Downstream consumers depend on `alkstore` plus exactly one engine
crate. A consumer binary links at most one driver.
## Consequences
**Positive**
- The engines' asymmetry of work is structural: the SQLite engine has a
small, mostly-wiring job; the Postgres engine carries the build-heavy
LISTEN/queue work. Neither is burdened by the other's dependencies.
- Any future engine (alkfs's in-tree needs, an ops-surface engine) is
additive — a new crate implementing the core traits — rather than a
feature-graph edit to one crate.
- Core-lean is structural, not a feature-discipline to be enforced by
review.
- The libsqlite3-sys link-collision class of problems is eliminated by
construction (single-driver binaries).
- Per-engine compilation/test gates are independent.
**Negative**
- A version-coordination duty: the core's trait surface is a contract
the engine crates must track. Version bumps in core must be adopted
by engines in lockstep when the contract changes (minor/major
discipline; no trait-default drift).
- Slightly more crate plumbing; naming/publishing overhead.
- Consumers who want *both* engines in one binary (rare, unsupported-by
design) cannot get a both-features build today.
## References
- OQ-ST-02 (`docs/research/phase-0.md`) — the decision record with
options considered.
- `docs/research/consumer-inventory.md` — the inventory whose lean
(single crate) was superseded, with the correction recorded there.
- [ADR-003](003-sqlite-driver.md) — the SQLite driver choice whose
link-collision constraint motivates this split.
- [ADR-004](004-postgres-driver.md) — the Postgres driver choice.
- OQ-10 (`docs/architecture/open-questions.md`) — trait-versioning
duties created here.
@@ -0,0 +1,99 @@
# ADR-002: Feature scope — what the store surface includes
## Status
Accepted
## Context
Honker's full surface (the interface prior art) covers nine feature
families: notify/listen, streams, queues, scheduler, outbox, named
locks, rate limits, result storage, and loadable-extension serving.
The crate must decide which of these it ships, and the evidence-first
answer (`docs/research/consumer-inventory.md`, 2026-10-04) grades each
row by whether any consumer document actually names a need. This ADR
fixes that scope so contract work isn't done against features no
consumer wants.
## Decision
**In scope, first-class** (a named consumer is pinned or the operator
has recorded the need):
- **notify / listen** — the transactional fire-and-forget signal layer.
Pinned (alkfs path-tree invalidation; the `watch_fires_on_commit`
POC evidence) and documented (alkblobs fleet visibility).
- **streams** — durable pub/sub with per-consumer offsets and
replay-on-attach. Operator-authority record (2026-10-04,
`consumer-inventory.md`): type-filtered event watching / repo-change
subscriptions that notify cannot honestly serve (no replay, no
durability).
**In scope** (documented need, thinner):
- **named locks** — TTL coordination locks. Pinned (alkblobs fleet
sweeper) and documented (alkfs writer coordination); alkgit's CAS has
an alternative design a lock is allowed to replace but needn't.
- **queues** — durable at-least-once work with the transactional
enqueue shape. Documented (alkfs sync/fetch-on-miss outbox; alkblobs
maintenance cadence).
- **outbox helper** — a thin helper *over* queues (enqueue inside the
business transaction + a delivery/consumption worker shape), not a
separately-needed feature. Same evidence as queues.
- **scheduler** — cron/`@every` enqueueing into named queues,
leader-elected. Documented-thin; the family-wide
"who sweeps/renews/reaps" need. May collapse into queues if design
shows scheduler = queues + tick (tracked in OQ-09).
**Cut-flag** (no named consumer; not silently included):
- **rate limits** — alkgit enforces budgets in its own wire layer; a
store-level rate limit is not the proven mechanism. The row stays
listed with its flag visible; cut at implementation time rather than
carried on momentum. If a distributed (cross-process) rate-limit need
materializes, it gets a consumer-inventory row first.
- **result storage** (`save_result`/`get_result`/`sweep_results`) —
adjacent-tooling shape, useful but named by no consumer. Same
cut-flag treatment.
**Out:**
- The **loadable-extension surface** — no consumer needs it; every
identified consumer is in-process Rust attaching to its own
connection (OQ-ST-07, resolved cut-only). Cut unless a consumer
appears.
- Honker's exclusion lines, inherited as out: workflow DAGs, task
chains/chords, multi-writer replication, distributed (cross-machine)
locking. They stay out unless a consumer document grows a row.
New consumers add a row to `docs/research/consumer-inventory.md`
*before* being assumed into scope.
## Consequences
**Positive**
- The contract surface is small and every part of it has a nameable
consumer — contract-pinning work isn't spent on speculative surface.
- Scope honesty is mechanical: the inventory is the ledger, and the
cut-flag rows keep their flags visible instead of being silently
dropped or silently shipped.
**Negative**
- Rate limits and result storage, if ever needed, require the consumer
inventory + contract process first — no ad-hoc additions.
- Cutting the extension surface narrows honker compatibility; any
future non-Rust consumer of these SQLite files cannot use honker's
extension (acceptable: none exists).
## References
- `docs/research/consumer-inventory.md` — per-feature evidence grades.
- OQ-ST-01, OQ-ST-07 (`docs/research/phase-0.md`) — resolved by this
evidence; promoted as OQ-01 and OQ-07 in
`docs/architecture/open-questions.md`.
- [core-contract.md](../core-contract.md) — the trait surface this
scope defines.
- [ADR-007](007-transactional-seam.md) — the transactional property all
in-scope enqueue/publish/notify features lean on.
@@ -0,0 +1,104 @@
# ADR-003: SQLite engine — rusqlite + published honker-core, bridged at the trait seam
## Status
Accepted
## Context
The SQLite engine's options were never a single "rusqlite vs sqlx"
axis (OQ-ST-03). Three distinct postures existed:
1. **honker-core on our own rusqlite connection** — we own
connection/schema/watcher wiring; honker supplies the SQL-function
machinery (`attach_notify`, `attach_honker_functions`,
`bootstrap_honker_schema`).
2. **honker-rs as the SQLite substrate** (`Database::open` — honker
holds its own connections; sync-only; transactions pin the
connection mutex).
3. **raw SQL over sqlx-sqlite with the honker loadable extension**
(`.so` built from published source, loaded per pool connection).
POC #1 (`docs/research/poc-sqlite-posture-findings.md`, ran 2026-10-04;
POC #2 is its pg twin) measured postures 1
and 3 end-to-end and settled the choice. Constraint facts that outlive
the comparison:
- rusqlite 0.40.x (what honker-core 0.5.0 pins) needs rustc ≥ 1.99
(`cfg_select!`) — a deployment note for any binary linking it.
- sqlx's `libsqlite3-sys` range collides with rusqlite 0.40's — mixed
rusqlite+sqlx binaries need a vendored one-line patch today. The
per-engine-crate split ([ADR-001]) keeps engine binaries
single-driver, dissolving this.
- The async-facing-trait + sync-bridge posture is family-standard
(alktty REQ-TTY-01; alkblobs store-api.md): bridge-at-the-seam is a
supported posture, not a workaround — which removes native-async's
main differentiator before any measurement.
## Decision
The SQLite engine uses **posture 1**: published `honker-core = 0.5`
as a library dependency over the crate's own rusqlite connections
(`bundled-sqlite` for hermetic builds).
- **Driver**: rusqlite, riding honker-core's pinned version.
- **Async seam**: bridged. The writer slot (`Writer`) +
reader pool (`Readers`) + every engine call in `spawn_blocking`.
A dedicated std-thread bridge (one thread owning the writer conn,
ops over mpsc) is the recorded optimization path if seam throughput
ever demands it — not the default.
- **Wake**: honker-core's `SharedUpdateWatcher` inherited (p50 ≈ 1.4 ms
at the default 1 ms `PRAGMA data_version` cadence, with
battle-tested failure handling — subscriber senders close on watcher
death, never silent-hang). A self-built thin watcher was benchmarked
only to confirm honker-core's is tighter (p50 1.40 vs 2.15 ms; max
29 vs 172 ms); no hybrid is warranted.
- **Transactions**: the engine opens `BEGIN IMMEDIATE` on the writer
slot and hands back a caller-held tx handle ([ADR-007]). The slot
lease models WAL's single-writer honestly: a long transaction parks
the writer — that is SQLite's shape, not an emulation artifact.
- **No `.so` runtime artifact, no vendored patches** in the engine
crate.
Rejected alternatives, briefly, for the record: posture 3 is *viable*
but pays ~2× p50 seam cost, a runtime dlopen artifact, wider
dependency tree, and sqlx 0.9 ergonomic warts (unsafe `extension()`,
`SqlSafeStr`) for no compensating advantage. Posture 2 (honker-rs)
was never measured because posture 1 dominates it on control with
comparable reuse.
## Consequences
**Positive**
- ~0.35 ms p50 transactional seam (measured), ~1.4 ms wake latency,
and honker-core's watcher failure-handling inherited for free.
- Static linkage: "open a path, get a store" needs no runtime
artifacts.
- The engine's job shrinks to wiring + the SQL surfaces honker doesn't
cover (queue semantics depth, OQ-ST-05's SQLite side rides honker's
machinery directly) + the async bridge.
**Negative**
- honker-core pins rusqlite ^0.40.1; version movement in honker-core
moves our rusqlite. A honker-core quality read is the recorded fork
trigger ([ADR-005], OQ-06).
- rustc ≥ 1.99 required by rusqlite 0.40.x — binaries linking this
engine carry that toolchain floor (deployment-matrix row,
[deployment.md](../deployment.md)).
- The sync bridge means engine ops pay a `spawn_blocking` hop — mostly
invisible next to SQLite write costs, but it shapes the tx-handle
design ([ADR-007]).
## References
- `docs/research/poc-sqlite-posture-findings.md` — the measurements.
- OQ-ST-03 (`docs/research/phase-0.md`) — resolution record.
- [ADR-001](001-crate-split.md) — why the engine crate is
single-driver.
- [ADR-004](004-postgres-driver.md) — the Postgres counterpart.
- [ADR-007](007-transactional-seam.md) — the tx-handle shape both
engines implement.
- [engine-sqlite.md](../engine-sqlite.md) — the engine spec this
decision defines.
@@ -0,0 +1,104 @@
# ADR-004: Postgres engine — tokio-postgres + deadpool-postgres, hand-rolled LISTEN forwarder
## Status
Accepted
## Context
The Postgres engine's driver question (OQ-ST-03, Postgres half) had
three candidate families: sqlx (single API across engines, and the
driver pgboss-rs carries), tokio-postgres + deadpool-postgres (the
alkblobs POC evidence base, natively async), and adopt/fork pgboss-rs
(bringing sqlx where the queue lives). POC #2
(`docs/research/poc-pg-posture-findings.md`, ran 2026-10-04) validated
the tokio-postgres posture end-to-end and measured the pieces the
contract depends on. The sqlite-arm comparison facts from
[ADR-003] apply on this side too: pgboss-rs brings sqlx (a second
driver per binary), has no LISTEN/NOTIFY at all (verified against the
checkout — consumption is `fetch_job` polling), so the push-reactivity
half is this crate's work regardless of fork-or-adopt.
## Decision
The Postgres engine uses:
- **tokio-postgres 0.7.x + deadpool-postgres 0.14.x** — pooled
connections for queries/claims (exactly-once via
`FOR UPDATE SKIP LOCKED`), natively async (client is `Send + Sync` —
no bridge, no spawn_blocking, the tx handle holds the pooled object
directly).
- **A hand-rolled LISTEN forwarder** (the POC's ~90-line shape) — a
dedicated non-pooled listener connection per process, a poll-message
loop fanning out to a bounded broadcast channel, immediate reconnect
with exponential backoff (50 ms → 2 s cap), re-LISTEN from the
channel list after every reconnect, and a **synthetic reconnect-wake
on a reserved channel** closing the no-replay hole
([ADR-006]). `postgres-notify` 0.3.8 was evaluated in-probe and
passed over (derive-not-adopt: lazy reconnect, connect_script
skipped at initial connect, unquoted-identifier LISTENs,
single-maintainer posture; it stays a recorded fallback if upstream
improves).
- **Queue machinery re-derived on this driver** with the pg-boss
schema family as design reference — not adopted/forked as a
dependency ([ADR-005]; semantics depth is [queues.md]'s work,
OQ-ST-05). The minimal-queue ground is measured: enqueue/claim/ack
with `SKIP LOCKED` is ~40 lines of SQL and passed the full property
suite.
- **Default consumption posture: LISTEN-driven claim with a re-poll
safety net**; poll-only remains the fallback (interval-tunable) when
a LISTEN connection is unavailable or unwanted. Measured: claim
latency p50 3–6 ms LISTEN-driven vs 32–50 ms at a 50 ms poll.
Structural constraints the engine owns (all test-pinned in POC #2):
- Pooled connections **cannot carry LISTEN** (deadpool#360 —
registers server-side, never delivers). The listener connection is a
per-process budget line outside the pool (`max_size + 1`).
- `pg_notify` payloads are ≤ 8000 bytes; the engine checks the limit
client-side and returns a typed error before the round-trip. Large
payloads ride a table row with the id in the notification (the
outbox shape).
- A query-vs-poll starvation deadlock exists if the listener's poll
loop isn't running before the first client query on the connection;
and dropping the `Client` closes the server session even if the
`Connection` task survives. Both are hand-rolled-forwarder failure
modes, owned and test-pinned by the engine.
## Consequences
**Positive**
- Single-driver engine, natively async, zero version conflicts, no
vendoring.
- The push channel pgboss-rs lacks, this engine gets — measured at
5–16× claim-latency improvement over even a 50 ms poll.
- Multi-host is native (no single-host assumption anywhere; property
tests ran all-through-network).
- The forwarder's two deadlock pitfalls are *learned territory* with
pinned tests, not unknowns.
**Negative**
- The listener forwarder is ours to maintain (~90 lines + the two
pinned failure modes) — the cost of not adopting `postgres-notify`.
- Per-process connection budget grows by one listener connection per
process that listens (a deployment-matrix row; consumers sizing
pools must account for it).
- The pg queue schema is re-derived work; the pg-boss family is
reference, not free.
## References
- `docs/research/poc-pg-posture-findings.md` — measurements and the
pinned property suite.
- OQ-ST-03, OQ-ST-05's posture half (`docs/research/phase-0.md`).
- [ADR-001](001-crate-split.md) — single-driver binaries.
- [ADR-005](005-dependency-ownership.md) — the ownership calculus this
decision applies per-subsystem.
- [ADR-006](006-wake-and-delivery-contract.md) — the wake contract the
forwarder implements on this engine.
- [engine-postgres.md](../engine-postgres.md) — the engine spec.
@@ -0,0 +1,70 @@
# ADR-005: Dependency ownership — published libraries by default, named fork triggers
## Status
Accepted (posture); fork triggers are live assessments tracked in OQ-06
## Context
The crate's reactive machinery sits on third-party work: honker
(SQLite side) and, conceptually, pgboss-rs (Postgres queue family).
Neither is adopted-by-adjacency — per guiding principle 3, ownership of
each subsystem is a deliberate per-question decision. The workspace
precedent is the targeted fork (alksocks' fast-socks5 extraction), but
a fork is expensive (especially when the forked code is sync and the
family standard is tokio — a fork of honker-rs is already a serious
port).
Two POCs (2026-10-04) measured the relevant postures per subsystem;
each carries its own record. This ADR fixes the *default calculus* and
the per-subsystem resolutions so Phase 1 design work starts from
stable ground.
## Decision
**Default posture: consume published libraries.** Fork or vendor only
when a named trigger fires. Per subsystem:
| Subsystem | Posture | Notes |
|---|---|---|
| honker-core (SQLite engine machinery) | published library, `honker-core = 0.5` | Fork trigger: the Phase 1 quality read of the watcher/transactional core finds a defect, OR a needed change upstream won't take. OQ-06 tracks the read. |
| rusqlite | published, riding honker-core's pin | Pin movement is honker-core's; we ride it ([ADR-003]). |
| tokio-postgres + deadpool-postgres | published library, as-is | Clean, zero conflicts, actively maintained (POC #2). |
| postgres-notify | **not adopted** (derive-not-adopt) | Lazy reconnect, connect_script skipped at initial connect, unquoted-identifier LISTENs, single maintainer. Hand-rolled forwarder instead ([ADR-004]). Fallback if upstream improves materially. |
| pg-boss queue machinery (pgboss-rs / node pg-boss) | **re-derived** on our driver; schema family as *design reference* only | pgboss-rs brings sqlx (second driver per binary) and has no LISTEN/NOTIFY (verified); the push half is ours either way, so queue-machinery reuse value is the honest comparison point and it loses on that arithmetic. Semantics depth: [queues.md], OQ-05. |
| honker-rs interface shape | **design reference**, not a dependency | Sync-only (no tokio); the surface is prior art for contract pinning ([ADR-006]), not code to ship. |
If any fork fires, forking is normal work we own (license/provenance
recorded per AGENTS.md §3 at adoption time) — it is not an exception
to this posture.
## Consequences
**Positive**
- Zero vendored code in the tree by default; upgrades are
`cargo update` work, not patch-management work.
- The fork triggers are concrete and named *before* the quality read,
so the read produces a decision, not a debate.
- Per-subsystem votes are recorded with evidence, so no future
discussion re-litigates from scratch.
**Negative**
- honker-core 0.5 remains alpha-quality software per its own README in
the link graph; the quality read (OQ-06) is the mitigation gate.
- The pg queue machinery is written by us — the pg-boss family's
battle-tested edge cases must be re-earned by design + tests
([queues.md]).
## References
- `docs/research/poc-sqlite-posture-findings.md` §"What feeds where"
(posture votes), `docs/research/poc-pg-posture-findings.md`
§"Dependency postures".
- OQ-ST-05, OQ-ST-06 (`docs/research/phase-0.md`) — the resolution
records.
- [ADR-003](003-sqlite-driver.md), [ADR-004](004-postgres-driver.md) —
the per-engine driver decisions applying this posture.
- OQ-06 (`docs/architecture/open-questions.md`) — the quality-read
tracker.
@@ -0,0 +1,128 @@
# ADR-006: Wake contract — opaque wake + re-read, and the notify-vs-streams delivery split
## Status
Accepted (the contract skeleton the evidence supports); remaining
surface pinning is OQ-04's contract work against this ADR
## Context
The two engines' wake mechanisms are structurally different — SQLite
wakes by a watcher polling `PRAGMA data_version` (deliver-on-commit,
no server push), Postgres by LISTEN/NOTIFY (server push,
connection-bound, no replay). A naive unification either collapses to
polling behavior on the strong side or promises semantics only one
engine has. Both POCs (2026-10-04) measured the pieces and — the
load-bearing finding — verified that **the two engines can share the
same wake contract unchanged**: wake is opaque, delivery is not
guaranteed to be precise, and consumers re-read state.
Honker's own processing-guarantees table is the cautionary prior art:
per-binding auto-checkpoint-vs-manual-save ambiguity produces
*different* guarantees under one function name. A single-crate version
must pick one answer per mechanism, not inherit the table.
## Decision
The contract fixes three things:
### 1. The unified wake contract: opaque wake + re-read state
`listen()` (and stream/queue subscription wake) delivers an **opaque
wake signal: "something changed; re-read the state you care about."**
The contract never carries semantic content, ordering promises, or
change descriptions. Consumers that need the actual data re-read it
(the hot-path pattern the ecosystem already uses: a subscriber holding
a cache invalidates on wake and re-reads, instead of re-polling).
- SQLite: `data_version` watcher fires on commit; wakes **coalesce**
under bursts (overtriggering on purpose — "one indexed SELECT is
cheap; a missed wake is a correctness bug").
- Postgres: LISTEN delivers per-notification (no coalescing), and the
forwarder's synthetic reconnect-wake covers connection gaps.
- Both: consumer code must be correct if wakes repeat, coalesce, or
arrive in any order, with at-least-once wake delivery. Exactly-once
*processing* semantics belong to queues/streams (below), never to
the wake layer.
Measured ground: wake latency p50 ≈ 1.1–2.2 ms on both engines at
default cadence; per-payload burst delivery verified (30/30 on pg,
coalesced-but-complete re-read on SQLite).
### 2. Delivery guarantees, pinned per mechanism (no per-binding table)
| Mechanism | Durability | Replay | Atomicity | Guarantee |
|---|---|---|---|---|
| `notify` / listen | none | never | commit-atomic (delivers at commit; rollback drops) | fire-and-forget, at-most-once per listener session |
| streams | durable table row | yes, per-consumer offset, replay-on-attach default | publish is commit-atomic | every committed event readable by every consumer that hasn't passed its offset |
| queues | durable row | claim/ack model | enqueue is commit-atomic | at-least-once *work* with visibility timeouts |
- The trait **does not promise replay under `listen()`** — on either
engine. Durability+replay needs are what streams are for; this split
is the contract's load-bearing line.
- Listener sessions start "from now" (SQLite: the watcher's current
state; Postgres: `MAX(id)`-equivalent at attach). History belongs to
streams.
- Wake *delivery* failures surface, not silence: watcher death closes
subscriber channels (SQLite, honker-core's
`WatcherDeathGuard` behavior); Postgres connection gaps surface as
the synthetic reconnect-wake on a reserved channel
([ADR-004], [engine-postgres.md]).
### 3. Reserved names are a contract surface
The Postgres forwarder's synthetic reconnect-wake needs a reserved
channel namespace no consumer channel may collide with. The naming
convention for reserved/meta channels (and any internal queue/stream
names) is pinned in the core contract — exact strings are OQ-04
contract work; the *existence of a reserved namespace* is decided
here.
### The caching-subscriber pattern rides the same contract
A subscriber that also holds a cache gets invalidation from the
opaque wake + re-read: the contract deliberately does not carry
key-level payloads, so cache clients key their invalidation on their
own read-set, informed by channel names (the channel name is the one
piece of semantic content `listen()` carries). Key-payload design,
if any consumer needs richer invalidation, is a
consumer-inventory-row-gated extension — not assumed now.
## Consequences
**Positive**
- One wake contract, verified identically on both engines — the
"how does a change become visible" question alkstore exists to
answer, answered once.
- The delivery-guarantee table is pinned per mechanism, killing the
per-binding ambiguity class honker's docs exhibit.
- Consumer code branches on *mechanism choice* (notify vs streams vs
queues), never on engine type (guiding principle 4; the honest
single/multi-host line is [deployment.md]'s and OQ-08's, not this
contract's).
**Negative**
- The opaque wake is deliberately less informative than
key-payload notify systems; consumers needing fine-grained
invalidation pay a re-read or need streams.
- Wake overtriggering on SQLite means consumers must be idempotent on
wake (a documented consumer obligation, part of the contract text).
## References
- `docs/research/poc-sqlite-posture-findings.md` §Probe 2,
§"What feeds where"; `docs/research/poc-pg-posture-findings.md`
§Sub-module W.
- OQ-ST-04 (`docs/research/phase-0.md`) — de-risked by these
measurements; the contract remainder is OQ-04 +
`core-contract.md`.
- [ADR-002](002-feature-scope.md) — notify/listen and streams are
first-class; the guarantee split is why both exist.
- [ADR-004](004-postgres-driver.md) — the forwarder owning the
Postgres side of this contract.
- [ADR-007](007-transactional-seam.md) — commit-atomicity's mechanism
(the `*_tx` seam).
- [core-contract.md](../core-contract.md) — the spec carrying this
contract's full surface.
@@ -0,0 +1,102 @@
# ADR-007: Transactional seam — caller-held tx handle with `*_tx` methods
## Status
Accepted (the seam shape both POCs verified); handle-type ergonomics
(generic vs downcast) is OQ-04 contract work
## Context
The load-bearing property of the whole store (guiding principle 2) is
**transactional local-adjacency**: enqueue/publish/notify inside the
same transaction as the caller's business write; rollback drops both.
Serving it through an async trait over two transaction models —
rusqlite's sync, thread-bound connection vs tokio-postgres's
`Send + Sync` client — is the one genuinely driver-coupled design
point (OQ-ST-03's framing). Both POCs (2026-10-04) built the candidates
and measured them.
## Decision
The core contract's transactional seam is a **caller-held transaction
handle**:
```text
store.begin_tx() -> TxHandle
handle.enqueue_tx(name, opts, payload) -> job id
handle.publish_tx(name, payload) -> event id
handle.notify_tx(channel, payload)
handle.save_offset_tx(stream, consumer, offset)
handle.commit() / handle.rollback()
```
- The handle is owned by the caller across `await` points; every
`*_tx` operation lands in *the caller's transaction*; commit/rollback
are explicit and owned by the caller. Non-`_tx` operations are the
auto-commit convenience counterparts, each atomic alone.
- **Postgres** ([ADR-004]): the handle holds the pooled connection
object directly (tokio-postgres `Client` is `Send + Sync` — verified
by a compile-time probe). Straight `.await`s, no bridge, no hop
cost. Commit/rollback return the object to the pool.
- **SQLite** ([ADR-003]): the handle is a **writer-slot lease** —
`begin_tx` acquires the `Writer` slot and opens `BEGIN IMMEDIATE`;
every handle op round-trips `spawn_blocking` to that connection
(rusqlite is not `Send`-across-await / not `Sync`). Holding the slot
across `await` points *is* how the lease works — a long transaction
parks the single WAL writer, which is SQLite's own shape.
- **Closure-scoped transactions** (`with_tx`) are *available as a
wrapper over* the handle shape, not instead of it (the POC verified
the reverse composition doesn't work: a closure cannot outlive
itself, so caching-subscriber state can't escape it).
- Both engines deliver the property **natively** — no emulation:
in-tx notify/Notify delivers only at commit; rollback drops job rows,
business rows, and notifications together (the POC property tests on
both engines green: no ghosts).
The POC's trait sketch carried two frictions, both now Phase 1
contract-pinning input rather than open risk:
1. The `as_any_mut` downcast per engine for `*_tx` methods on a
`dyn TxHandle` — small, but a generic or enum-favored handle may be
cleaner. OQ-04 decides with the full contract.
2. Thread-affinity of rusqlite tx ops (SQLite handle ops must
round-trip the same blocking thread) — inherent to [ADR-003]'s
bridge, documented as a contract note ("SQLite tx ops are
serialized by the writer slot"), not a defect.
## Consequences
**Positive**
- The load-bearing property is one seam, both engines, POC-verified
end-to-end — the dual-write problem is structurally prevented, not
consumer-disciplined.
- Long transactions compose naturally (checkout-then-work shape);
nothing about the seam forbids multi-statement business
transactions with interleaved reads.
- Per-engine bridging differences are invisible to consumer code: the
same `*_tx` calls, the same commit/rollback ownership.
**Negative**
- The SQLite handle parks the writer for its duration — a slow
consumer transaction throttles all writers on the file. This is
honest (WAL single-writer), but contract docs must say it so
consumers budget transactions accordingly.
- Per-op `spawn_blocking` hop on SQLite tx ops (measured fine at
~0.35 ms p50; the dedicated-thread bridge is the recorded
optimization).
- `with_tx` wrapping is a convenience surface we must ship and test
so it doesn't accrete ad-hoc in consumer code.
## References
- `docs/research/poc-sqlite-posture-findings.md` §Arm A, §Probe 1/3;
`docs/research/poc-pg-posture-findings.md` §Sub-module T.
- OQ-ST-04 (`docs/research/phase-0.md`), the `*_tx` seam bullet.
- [ADR-003](003-sqlite-driver.md), [ADR-004](004-postgres-driver.md) —
the per-engine handle mechanics.
- [ADR-006](006-wake-and-delivery-contract.md) — commit-atomicity is
the `*_tx` property this seam carries into the notify contract.
- [core-contract.md](../core-contract.md) — the spec whose transaction
section this ADR defines.