Files
alkstore/docs/architecture/core-contract.md
T

29 KiB

status, last_updated
status last_updated
draft 2026-10-06

Core contract

The unified, engine-agnostic surface a Store exposes. This document specifies WHAT the contract is; the per-engine specs map it onto their machinery; ADRs carry the WHY. The starting artifact is the honker-rs surface (the Phase 0 §Interface finding) scoped to the inventory-confirmed features — ADR-002 — and pinned against it. The base surface is pinned by ADR-008 (contract v1 partition, TxHandle representation, wake type, reserved strings, error taxonomy, config split — amended in place, pre-implementation, by ADR-014: outbox_enqueue_tx joins the TxHandle trait, and by ADR-015: streams depth — key semantics, StreamEvent shape, the ordering row, trim_to, and publish_with_key_tx); ADR-009/ADR-010 add the first post-v1 contract extensions (scheduler collapse surface, QueueOpts depth — versioning discipline for such extensions is OQ-10's); ADR-016 decides the parked capability-surface question: none, by default ever — engine differences surface at compile time (engine-crate identity) and in the deployment matrix, never as a runtime descriptor; the obligations are this document.

Concepts

  • Store — what a consumer opens from a connection string or file path (ADR-001): one handle, the engine behind it chosen at open time. Consumer code never branches on engine type (guiding principle 4).
  • Mechanism — one of the surface's coordination families: notify, streams, queues (+outbox), locks, scheduler. Delivery guarantees are pinned per mechanism in ADR-006's table (notify/streams/queues), extended with the locks row by ADR-008, the scheduler row and collapse by ADR-009. (Guiding principles are defined in docs/research/phase-0.md §Vision and cited by number throughout this directory.)
  • Wake — the opaque "something changed, re-read" signal the wake contract delivers (ADR-006). The contract's central abstraction.
  • TxHandle — the caller-held transaction object whose *_tx methods make side effects commit-atomic with a business write (ADR-007). Contract shape: the *_tx methods live on the TxHandle trait itself — engines implement it for their concrete handle, no downcast (ADR-008 §2).

The seam

Per ADR-008 §6, the trait is constructed by engine crates (durability knobs, pool sizing, watcher cadence are engine options — deployment.md carries the facts); the contract is the trait surface it returns:

Store::begin_tx() -> Box<dyn TxHandle + Send>

trait TxHandle {
    enqueue_tx(name, opts, payload) -> job_id
    publish_tx(stream, payload) -> offset
    publish_with_key_tx(stream, key, payload) -> offset
    notify_tx(channel, payload)
    save_offset_tx(stream, consumer, offset)
    outbox_enqueue_tx(outbox, opts, payload) -> job_id
    commit(self: Box<Self>) -> Result<()>   // or rollback
}

outbox_enqueue_tx takes the outbox name and derives the backing queue engine-side (the reserved prefix makes the derived name unreachable by enqueue_tx by design — ADR-014). publish_with_key_tx validates the key like the non-tx form (a present key must be non-empty → InvalidName) and is publish_tx's keyed twin — commit-atomic keyed publishes are the point (ADR-015 §2).

Mechanism handles come off the store (or, for transactional variants, off the handle):

store.notify(channel, payload)          handle.notify_tx(channel, payload)
store.stream(name) -> Stream            handle.publish_tx / publish_with_key_tx
                                        handle.save_offset_tx
                                        (the Stream handle also carries
                                        trim_to — stream-side, outside the
                                        tx seam)
store.queue(name, opts) -> Queue        handle.enqueue_tx
store.outbox(name) -> Outbox            handle.outbox_enqueue_tx(outbox, opts, payload)
store.try_lock(name, owner, ttl) -> Option<Lock>
store.listen(channel) -> Box<dyn WakeReceiver>
store.schedule(name, spec, queue, payload, opts) -> Result<Schedule>
store.unschedule(name) -> bool
store.run_schedules(stop) -> Result<()>

Payloads cross the trait as core value types (non-generic trait methods — object safety of the boxed handles, ADR-008 §2); payload_as<T> decodes on concrete returned values (Job, StreamEvent). The with_tx closure wrapper ships over the handle shape (the POC-verified composition direction, ADR-007). Mechanism handles (Stream, Queue, Outbox, Lock) are core-owned trait objects like the tx handle — engine types never appear in consumer signatures; their trait methods pin at implementation, mirroring the TxHandle pattern (ADR-008 §2's rationale applies identically).

Mechanism contracts

notify / listen

Fire-and-forget signals, commit-atomic when sent in a transaction (ADR-007). No durability, no replay, no per-listener retry; a listener attached after a commit never sees it.

  • notify(channel, payload) — payload ≤ 8000 bytes on Postgres (client-side checked, typed PayloadTooLarge before the round trip, verified by POC #2); no limit on SQLite. The error variant is contract-wide (callers match it identically on both engines), the occurrence is the documented engine asymmetry (ADR-008 §5) — with no capability surface (ADR-016), this variant is the one runtime carriage of an engine asymmetry; its occurrence asymmetry is pinned in the verification backlog below.
  • listen(channel) -> Box<dyn WakeReceiver> — starts from "now"; delivers opaque Wake { channel } signals (ADR-006), never payloads or ids (ADR-008 §3). Wakes are best-effort hints, not per-commit delivery promises: possibly coalesced (SQLite) or per-notify (Postgres), possibly repeated after reconnect, possibly absent entirely (SQLite burst coalescing; the pg no-replay hole) — recovery is the consumer's re-read plus the engine's reconnect surface, never a guaranteed deliver. Consumers must be idempotent on wake.
  • Failure surfaces as channel events, never silence: watcher death closes the receiver (recv() -> None, SQLite); the synthetic reconnect-wake (a reserved channel (ADR-004)) covers Postgres connection gaps.
  • Channel name is the one piece of semantic content a wake carries — the invalidation key for caching subscribers.

streams

Durable pub/sub with per-consumer offsets (ADR-006). The durable cousin of notify: publish is commit-atomic; every committed event is readable by every consumer whose offset hasn't passed it; replay-on-attach is the default; offsets are explicit and transaction-aware (save_offset_tx gives exactly-once-within-a- business-tx shape). Depth pinned by ADR-015.

  • publish / publish_tx / publish_with_key / publish_with_key_tx — append to the stream's durable log; return the assigned offset. Key semantics (ADR-015 §1): the key is carried metadata — stored on the event row, round-tripped on every read, Option<String>-shaped on StreamEvent — with no engine-enforced behavioral role: the ordering guarantee stays global FIFO by offset within the stream; per-key in-order reading is the documented emergent pattern (filter by key, read in offset order), not a server promise. A present key must be non-empty (InvalidName). Server-enforced per-key ordering re-enters only via a consumer-inventory row naming it.
  • StreamEvent shape (ADR-015 §3): { offset: i64, stream: String, key: Option<String>, payload: Vec<u8>, created_at: i64 } — stream renames honker's topic (one term everywhere, the ADR-008 §4 kinds list); offset is engine-assigned, monotone per stream, immutable, never renumbered (gaps after a trim are legal); created_at is unix-seconds-at- publish, informational, not an ordering field (offset is the only one); payload_as<T> decodes (the error is Codec).
  • Ordering guarantee row (ADR-015 §4, extending ADR-006 §2's streams row): read_since/subscribe yield offset ASC — global FIFO per stream; committed-event visibility and replay-on-attach as already pinned; offsets immutable. The contract suite pins cross-engine equivalence against this row (same publish sequence → same offset sequence → same read order).
  • read_since / read_from_consumer(offset) — cursor-based reads; get_offset(consumer) — checkpoint inspection. A read from a trimmed-away offset region (below) resumes at the trim horizon's first remaining row.
  • save_offset / save_offset_tx — consumer checkpoint, explicit; the contract's save is always explicit (no auto-checkpoint cadence — honker's per-binding ambiguity is not inherited). The subscription handle exposes no save_every / auto-save-on-drop (ADR-008 §8). Saves are monotone; a saved offset below the trim horizon (ADR-015 §5) stays a valid position marker.
  • trim_to(horizon) — delete events with offset <= horizon (ADR-015 §5). The stream-side bounded-growth op, consumer-invoked, no engine-default retention and no ambient sweeper (the ADR-010 §6 posture; the replay-forever default is the mechanism's purpose — growth is documented with the tool in-contract to bound it). Trim emits no dedicated wake and no notify — on SQLite the data_version watcher may still fire (any committed write bumps it; a spurious hint, contract-legal); surviving events keep their offsets.
  • subscribe(consumer) -> Box<dyn EventReceiver> — durable consumption: attach, read to current tail, resume after restart from the stored offset; explicit save_offset on the receiver (shape in ADR-008 §8). Consumption is wake-driven with table re-read — the same mechanism split as queues (durable row, LISTEN/watcher wake, ADR-006); the receiver's engine-side trigger is engine-internal, but events never require polling to become visible (the queues row's pg re-poll safety net is an admission of LISTEN loss, not the posture).

queues

Durable at-least-once work (ADR-002). Depth pinned by ADR-010; contract-level obligations:

  • enqueue / enqueue_tx — commit-atomic; EnqueueOpts { delay, run_at, priority, max_attempts, expires }; queue-level QueueOpts { visibility_timeout_s, max_attempts, backoff_base_s, dead_letter_retention_s } (ADR-010 §3/§4).
  • claim_one / claim_batch — exactly-once handout under concurrency (POC-pinned on both engines); claim ordering: priority DESC, then ready-time, then enqueue order (FIFO under equal priority).
  • Job handle: ack / retry / fail / heartbeat. ack deletes the row; heartbeat(extend) is renewal — an absolute reset of the claim deadline (extend is the new full deadline from now, not additive to elapsed time); late heartbeat refused; a reclaim consumes an attempt; retry(err, None) computes the queue's equal-jitter exponential delay (range pinned by ADR-010 §3), retry(err, Some(d)) overrides; fail = immediate dead-letter. Handle-op validity (uniform predicate): an op succeeds only while the row is processing and the caller's claim deadline is unexpired — deadline lapse refuses all handle ops (heartbeats, acks included) and a reclaim does too; the dual-execution window's worker may still complete its work but its ack will not land (at-least-once, as documented); false, not error (ADR-010 §2). Dead letters: move-to-dead storage, get_job sees dead rows (with last_error/died_at), retention via dead_letter_retention_s (default forever), no redrive API (ADR-010 §1–§4). QueueOpts stamp onto the job row at enqueue (§3a) — queues are names, not config owners. The remaining v1-skeleton ops (ack_batch, cancel — unconditional delete, not an interrupt) pin in ADR-010 §1.
  • sweep_expired(queue) — moves every past-expiry row (any state) to dead + enforces dead-letter retention — the no-stranded-rows property (ADR-010 §5); cadence recipe in queues.md (no ambient sweeper).

named locks

TTL-bounded coordination locks. Lock rows live in the same database as the caller's business data — coordination and data co-locate — but lock ops are not part of the tx seam: there is no lock_tx; a lock held across a business transaction is held by explicit acquisition before/inside it, not transaction-scoped (acquisition is a separate auto-commit op; rollback of the business tx does not release it — the TTL discipline governs).

  • try_lock(name, owner, ttl) — acquire or fail (Option<Lock> — no-work is a value, not an error); renew and release on the lock handle. Release on explicit unlock or TTL expiry. Re-acquirable after expiry (POC-pinned on Postgres; on SQLite it rests on the forked substrate's lock machinery — the SQLite-side pin is in the verification backlog below).
  • Guarantee row (ADR-008 §7): mutual exclusion bounded by TTL + renewal — after TTL expiry exclusion lapses silently (no revocation event); holders must renew within TTL; expiry is a loss of exclusivity, not an error.
  • Lock names are shared-namespace (reserved-prefix rules below, ADR-008 §4).

outbox

A helper over queues, not a separate mechanism: enqueue inside the business transaction + the delivery/consumption worker entry points. Same evidence base and guarantee as queues (ADR-002). Worker semantics (honker's run_once shape, inherited as the pinned posture): run_once(worker_id, delivery) is a pull op — claim one job, run the delivery closure, ack on Ok, retry(err, None) (the queue's curve) on Err; the consumer calls it in its own loop. No heartbeat inside delivery (honker parity, inherited deliberately): delivery slower than the backing queue's stamped visibility timeout (default 60 s for outbox-backed queues) can have its claim expire mid-delivery and be redelivered — the dual-execution window; idempotent delivery is the consumer's obligation, same as queues.

The transactional enqueue side is outbox_enqueue_tx(outbox, opts, payload) on the TxHandle trait (ADR-014): it takes the outbox name (validated like store.outbox(name) — empty → InvalidName, reserved-prefixed → ReservedName), derives the backing queue name engine-side, and lands the job row in the caller's transaction with EnqueueOpts stamped per [ADR-010](decisions/010-queue-semantics- depth.md) §3a over the backing queue's derived QueueOpts. Commit makes the job visible to run_once exactly when the business write commits; rollback drops both (the no-ghosts property, ADR-007). The derived backing queue name is not expressible through enqueue_tx (the reserved prefix is rejected on directly-supplied names) — outbox_enqueue_tx is the only transactional path into it, which is itself the guarantee that the derivation cannot be collided with.

scheduler

Collapsed into queues per ADR-009: schedule(name, spec, queue, payload, opts) (upsert by name), unschedule(name), and the opt-in run_schedules(stop) runner (leader-elected where the engine has peers — the leadership lock is the reserved name __alkstore_scheduler). Schedules never fire without a runner (no ambient timers). v1 spec grammar: @every <n><unit> only (s|m|h|d); cron strings are rejected (InvalidSpec) pending a consumer-inventory row naming wall-clock cron. Each boundary fire enqueues ordinary work stamped with ScheduleOpts { priority, max_attempts, expires } (ADR-009 §3). Boundary guarantee: ADR-009 §4 (at-least-once per elapsed boundary, bounded catch-up with skip-forward past the cap).

Cross-cutting contracts

Delivery-guarantee table

The per-mechanism table in ADR-006 is the contract of record for notify/streams/queues; the locks row is ADR-008 §7's and the scheduler row is ADR-009 §4's. This spec inherits them and adds the consumer obligations:

  • wake idempotence (notify's best-effort, coalescing behavior);
  • explicit offset saves (streams);
  • visibility-timeout budgeting — heartbeat inside the deadline for long work; the dual-execution window and reclaim-eats-attempt rules are documented consumer obligations (ADR-010 §2);
  • running run_schedules for schedules to fire at all, and scheduling sweep cadences via the collapse recipe — timeliness is the consumer's, correctness-of-transition the engine's (ADR-010 §6);
  • *_tx operations are only durable after the caller's commit — rollback drops job rows, event rows (keyed and unkeyed alike), notifications, and offset saves together (the no-ghosts property, POC-pinned on both engines).

Errors

thiserror-typed, per the family standard. Pinned by ADR-008 §5: one top-level Error, with v1 variants PayloadTooLarge (universal; produced pg-side, contract-wide matchable), ReservedName, InvalidName, Closed, Codec, and the opaque Database fallback (engine detail preserved via the source chain). Post-v1 additions: InvalidSpec (schedule spec grammar) and LeadershipLost (run_schedules return on leadership loss) — both from ADR-009 §6, the only taxonomy deltas so far. Pinning rule: a variant exists only when callers can act differently on it. Queue claim "no work" and job/lock boolean results are values, not errors; dead-letter is observable state (get_job), not an error.

Capability surface

None — decided by ADR-016 (2026-10-06): Store carries no capabilities accessor, in v1 and by default ever. The honest single-host/multi-host boundary lives at compile time (the engine-crate dependency is the deployment statement — ADR-001's single-driver binaries) and in deployment.md's matrix. Runtime carriage of the one caller-actionable engine asymmetry is the contract-wide matchable PayloadTooLarge variant (Errors above); re-entry requires a consumer-inventory row naming a runtime-adapt need.

Naming / reserved namespace

Consumer-visible names (channels, streams, queues, locks, and schedule names) share engine-visible namespaces on Postgres (LISTEN channel names are server-global per database). Pinned by ADR-008 §4 (schedule names added by ADR-009 §1):

  • Reserved prefix __alkstore_, engine-independent, applies across all name kinds; the one v1-reserved string is __alkstore_listener_reconnected__ (the Postgres reconnect-wake channel, ADR-004).
  • Reserved-prefix names are rejected at every entry point — the name-bearing methods and their *_tx counterparts — with the typed ReservedName error; empty names with InvalidName (ADR-008 §4). Stream-consumer names are consumer-local identifiers, not a reserved-namespace kind.
  • Engine-derived names in consumer namespaces carry the reserved prefix (the outbox's backing queue is derived under the prefix — honker's _outbox:{name} scheme is not inherited verbatim).
  • SQLite: the substrate's __alkstore_* internal table family is storage-internal (not consumer namespace) — re-owned from upstream's _honker_* by the fork (ADR-011, designed per ADR-012). Postgres: engine tables are schema-scoped — one engine-owned schema (ADR-010 §8) — the channel namespace is this section's.

Verification backlog

Contract properties POC-pinned on one engine only (or sketched rather than surface-verified) — the contract test suite must pin both engines before the engine specs are called stable:

  • Named-lock TTL/expiry re-acquisition on SQLite — pinned on Postgres (pg POC's lock probe); on SQLite it rests on the forked substrate's lock machinery (its lock_renew verified in the quality read; the re-acquire-does-not-refresh-TTL behavior is upstream's, inherited deliberately) — pin it in the contract suite.
  • Wake semantics under WakeReceiver shapes — POCs verified wake delivery/coalescing through their own probe types; the pinned Wake { channel } / recv forms (§3 of ADR-008) need the contract suite's own property tests on both engines.
  • save_offset_tx exactly-once-within-a-business-tx shape — both POCs verified it through their sketch implementations (SQLite's offset-save was a plain SQL upsert, not the forked substrate's full surface; the pg side likewise through its probe), so the property must be re-pinned against the real engines' save_offset_tx in the contract suite.
  • Concurrent try_lock loser/error behavior on SQLite (the pg side returns cleanly; the substrate's busy-path under lock contention is the thing to pin).
  • Scheduler + depth properties on Postgres — the SQLite side's tick/leader/catch-up machinery inherits the forked substrate's test-pinned implementation; the pg engine's re-derived tick (boundary advance, 64-cap skip-forward, leadership-loss discipline) and the ADR-010 depth properties (visibility reclaim consuming attempts, dead-letter moves, the no-stranded-rows sweep) pin in the contract suite at implementation.
  • Backoff-curve and stamp-resolution equivalence across engines — the equal-jitter curve (ADR-010 §3) and the opts-stamping resolution (§3a) are computed engine-side per ADR-012 §2's contract-blind boundary; the contract suite must pin both engines' arithmetic to identical outputs.
  • Job-handle validity predicate on both engines — the uniform processing-state + unexpired-deadline rule (ADR-010 §2): the late-heartbeat boundary (refused exactly when the deadline has lapsed), the post-lapse-ack-refusal (at-least-once), and the ack-vs-reclaim race (one wins atomically); both engines' SQL pin identical outcomes in the contract suite.
  • outbox_enqueue_tx commit-atomicity on both engines (ADR-014) — rollback drops the backing-queue job row together with the business write (no ghost job); commit makes it claimable by run_once; reserved/empty outbox names rejected identically on the tx path (ReservedName/InvalidName); stamped opts visible via get_job equal for both engines.
  • Cross-engine stream equivalence (ADR-015 §3/§4) — publish sequence → identical offset sequence → identical read_since output order on both engines, direct and subscriber reads alike; keyed/unkeyed interleavings preserve global FIFO; key round-trips exactly (None vs Some), stream/created_at fields equal.
  • publish_with_key_tx commit-atomicity on both engines (ADR-015 §2) — rollback drops the keyed event row with the business write (no ghost event); commit makes it visible to read_since/subscribe; empty-Some-key InvalidName on both engines' tx paths.
  • PayloadTooLarge occurrence asymmetry (ADR-016 §5) — with no capability surface, PayloadTooLarge is the one runtime carriage of an engine asymmetry, so its matchability pins in the suite: an oversized notify/notify_tx payload is rejected client-side by the Postgres engine before any round trip (the limit in the variant), and the SQLite engine never produces the variant at any size — engine-agnostic code writes the same match on both engines and the non-occurring arm simply never fires.
  • trim_to semantics on both engines (ADR-015 §5) — exact-boundary trim (<=), surviving rows keep their offsets (gaps legal, never renumbered), reads resume at the trim horizon's first remaining row, saved offsets below the horizon stay valid, trim wakes nothing, and a concurrent subscriber never loses its place.

Design Decisions

ADR Decision Summary
001 Crate split core + per-engine crates; single-driver binaries
002 Feature scope inventory-confirmed features only; cut-flags explicit
006 Wake contract opaque wake + re-read; notify-vs-streams guarantee split
007 Tx seam caller-held handle, *_tx methods, native commit-atomicity
008 Contract v1 pinning surface partition, TxHandle trait shape, Wake type, reserved strings, error taxonomy, config split, locks guarantee row
009 Scheduler collapse (post-v1 extension) queues + schedule()/run_schedules, @every-only v1, boundary guarantee row
010 Queue depth (post-v1 extension) visibility/renewal, opts stamping, backoff curve, dead-letter, sweep, layout
011 Substrate fork SQLite substrate owned (__alkstore_* naming); queue ops re-derived on contract v1
012 Fork design contract-blind substrate boundary; engine-side formula arithmetic pinned equivalent by the contract suite
014 Outbox tx enqueue (amends 008) outbox_enqueue_tx on TxHandle; outbox-name validation; derived backing queue reached only through the outbox surface
015 Streams depth (amends 006/008) key = carried metadata, global-FIFO ordering row, StreamEvent shape, publish_with_key_tx, trim_to
016 Deployment honesty (decides 008's parked question) no runtime capability surface — compile-time engine identity + documented matrix; PayloadTooLarge occurrence asymmetry pinned

Open Questions

Open questions are tracked in open-questions.md. Key questions affecting this document:

  • OQ-10: contract versioning discipline across engine crates (open)

Resolved on this document's surface: OQ-09 (scheduler collapse — ADR-009), OQ-05 (queue semantics depth — ADR-010), OQ-13 (transactional outbox enqueue shape — ADR-014), OQ-12 (streams depth — ADR-015), and OQ-08 (capability-surface shape — none, by default ever; ADR-016), 2026-10-05/06.