phase 0 setup: agent defs cleaned, AGENTS.md, initial phase-0.md draft

- sdd_process.md + coordinator.md: stale @alkdev/alkblobs name from the
  copy fixed to @alkdev/alkstore
- implementation-specialist.md: copied alkcall-family conventions
  (OperationEnv, vendored core types, BAST/wire formats) replaced with
  crate-neutral rules + an ADR-escalation rule
- code-reviewer.md: stale tls/iroh/acme feature list removed; anyhow
  posture corrected; db-specific review checks added
- AGENTS.md: phase-0 posture (no crate yet, no README, POC/discipline
  conventions, OQ-ST-NN register, reference checkouts)
- docs/research/phase-0.md: initial draft — vision, prior art (honker,
  pgboss-rs read incl. verified pgboss-rs LISTEN/NOTIFY absence), open
  questions OQ-ST-01..08, phase 0 plan
This commit is contained in:
glm-5.3-flash committed 2026-10-03 16:36:46 +00:00
1 parent 5bcd1b7a2f
commit f6531b5532
6 files changed
+504 -56

No files matched your search

+359
View File
@@ -0,0 +1,359 @@
---
status: draft
last_updated: 2026-10-03 (initial draft from the setup discussion)
---
# alkstore — Phase 0 (Exploration)
This document captures Phase 0 (Exploration) for the `alkstore` crate:
vision, guiding principles, prior art, and the open-question register
(OQ-ST-01..NN). Phase 0's objective per `docs/sdd_process.md`: *capture
vision and guiding principles; research options; validate approaches;
converge on a recommended approach.* This is the initial draft — none of
the questions below are settled, and there is no POC register yet.
Context for why this crate starts now: **alkblobs**
(`/workspace/@alkdev/alkblobs` — spec + POCs only, paused mid-planning)
hit repeated circular hedging in its Phase 0, and a root cause was that
the storage substrate it deploys onto was itself disjoint and fuzzy — a
repo pattern with a default in-memory adapter across the alk* ecosystem,
cache-invalidation patches in hot paths, non-invalidated caches where
delay was tolerable, no single definition of "how does a change in the
database become visible to other processes/connections?" alkblobs paused
partly to let this crate answer that first. alkstore is the attempt to
make that substrate real once, so downstream stores don't re-derive it.
## Vision and guiding principles
**One sentence (draft):** one reactive store interface over SQLite and
Postgres — durable pub/sub notify, queues, streams, and the transactional
integration (write + enqueue in one transaction) that honker delivers on
SQLite — with the Postgres side building on the natively-available
machinery (`pg_notify`/`LISTEN`, and the pgboss job-queue schema family)
rather than emulating it.
**The honker relationship.** `/workspace/honker` (read-only reference;
alpha-quality per its own README, MIT/Apache-2.0 dual) adds
Postgres-style `NOTIFY`/`LISTEN` semantics to SQLite without a broker:
durable at-least-once queues with retries/delay/priority/visibility
timeouts/dead-letter, durable streams with per-consumer offsets,
cron/`@every` scheduling, named locks, rate limits, transactional
outbox — all as INSERTs inside the caller's transaction, with the
cross-process wake delivered by a shared watcher that polls `PRAGMA
data_version` (default 1 ms → single-digit-ms delivery) and re-reads
indexed state after every wake. Its own Prior Art section names the
Postgres side: `pg_notify` gives fast triggers with no retry or
visibility semantics; pg-boss/Oban are the durable-layer gold standards
— "If you already run Postgres, use the Postgres tools."
The goal is **not** "port honker to Postgres." The goal is the interface:
one store API whose consumer code (queues, streams, notify) looks the
same whether the backing engine is SQLite or Postgres, while each engine
uses its own native wake/delivery story under the hood. The honker docs
recommend pgboss + `pg_notify` for the Postgres equivalent — that
recommendation is the design brief for this crate's Postgres engine.
**Consumer shape (from the paused alkblobs planning):** any crate that
today uses the ecosystem's repo-pattern + in-memory adapter should be
able to swap in an alkstore-backed engine and get durability + true
cross-process/cross-instance reactivity. That means the reactive surface
must compose with *client-side* caching: a subscriber that also holds a
cache can invalidate on notification instead of re-querying or re-polling
— the hot-path pattern the ecosystem already uses, given a real
invalidation source.
Guiding principles:
1. **One interface, two engines, native underneath.** The abstraction
layer unifies the consumer-visible features; the engines stay
dialects, not two emulations of one dialect. SQLite follows honker's
design (queue in the same file, same transaction, watcher-based
wake); Postgres follows pg-boss' design (schema-based job tables +
`pg_notify`-driven wake). A "lowest common denominator" unification
(both sides polling, both sides emulating LISTEN) is explicitly the
failure mode to avoid — it would re-create the fuzziness this crate
exists to remove.
2. **Transactional local-adjacency is the load-bearing property.**
Honker's core claim: enqueue/publish/notify in the same transaction
as the business write; rollback drops both. The unified surface must
preserve this on any engine, because the ecosystem's repo pattern
assumes it (a business write that loses its side-effect notification
is the dual-write problem honker names).
3. **Ownership of the whole stack.** honker is third-party; pgboss-rs
is third-party. Whether any of them are adopted, forked, or used as
schema/design reference only is a deliberate per-question decision —
not inherited by adjacency. The workspace precedent is the targeted
fork (alksocks' fast-socks5 extraction: adopt the design, own the
code, port to our conventions).
4. **Substrate-agnostic consumer API, engine-specific setup.** A
consumer opens a `Store` from a connection string / file path and
gets the same trait surface. Which engine is behind what can vary
(per-deployment config), but consumer code must not branch on
engine type.
5. **No panics, `tokio`, `thiserror`, lean base crate, feature-gated
optional engines** — family-standard, pre-committed (AGENTS.md).
## The driver conflict (Phase 0's central tension)
The immediate design fork, flagged by the user:
- **The alkblobs POCs used `tokio-postgres` + `deadpool-postgres`**
(validated in `poc-postgres-kv-findings.md`, `poc-pglo-findings.md` —
including Large Objects work).
- **pgboss-rs (`/workspace/pgboss-rs`) uses `sqlx`** (`sqlx Postgres
runtime-tokio`).
- honker-core uses `rusqlite`.
A single crate with both engines means a driver decision, and the
reactivity story is entangled with it:
- **pgboss-rs currently has no `LISTEN`/`NOTIFY` at all** (verified
2026-10-03 against the checkout — `src/` contains no LISTEN/NOTIFY
usage; consumption is `fetch_job` polling). The node original relies
on `pg-boss`'s own maintenance/polling; the port did not pick up a
push channel. So *even* "use pgboss for the queue" does not deliver
reactivity — LISTEN/NOTIFY wiring would be new work either way, and
the driver choice determines *whose* LISTEN plumbing (sqlx's
`PgListener` is built-in; tokio-postgres uses its `Connection`
notifications).
- **Honker's reactivity on SQLite is a watcher polling `PRAGMA
data_version`** — a fundamentally different mechanism from
LISTEN/NOTIFY. The unified reactive trait must abstract over both
without collapsing to the polling behavior of the weaker side.
This tension is OQ-ST-04 below. It is *not* resolved by "pgboss is well
written so start there" — that is exactly the inherited-assumption
shape the SDD process flags. What pgboss-rs genuinely offers (schema
DDL, job states, retry semantics, the node-compatible API) is design
reference regardless of driver.
## Prior art
Notes below are from reading the checkouts on 2026-10-03; both external
projects are read-only references. Provenance/licensing gets recorded
per AGENTS.md §3 when code is adopted, not while only reading.
### honker — the SQLite-side template
`/workspace/honker` (SQLite extension + bindings). What matters for
this crate:
- **The full feature set to match on Postgres** (its §What It Does):
notify/listen across processes, durable at-least-once queues
(retries, delayed jobs, priority, visibility timeouts, dead-letter
rows, result storage), durable streams with per-consumer offsets,
cron/`@every` scheduling, named locks, rate limits, transactional
outbox helpers. Deliberately excluded there: workflow DAGs, task
chains/chords, multi-writer replication, cross-machine locking —
scope line likely inherited, to be confirmed.
- **The wake mechanism** — `PRAGMA data_version` polling watcher
(default 1 ms; raise for idle CPU), re-read indexed state after
wake, overtriggering on purpose ("one indexed SELECT is cheap; a
missed wake is a correctness bug."). Optional kernel-events and WAL
shared-memory backends exist in source builds.
- **Single-machine honesty** — file-backed, one host; NFS-two-writers
explicitly not supported. This posture needs an explicit Postgres
counterpart (multi-host is Postgres' normal case, so the interface
must not bake SQLite's single-host assumption into the shared
surface).
- **The transactional enqueue shape** — every feature is an INSERT
inside the caller's transaction. This is the pattern the unified API
must keep visible and cheap.
### pgboss-rs — the Postgres queue family reference
`/workspace/pgboss-rs` (v0.1.0-rc6, MIT/Apache-2.0 dual). Ported from
node pg-boss: builder-based queue/job API, retry/delay/priority/
singleton/dead-letter concepts, `sqlx` 0.8, schema-scoped DDL.
- Verified gap (2026-10-03): **no LISTEN/NOTIFY** anywhere in `src/`
— consumption is polling `fetch_job`. Any push-reactivity is new
work, not an adoption freebie.
- Its value as reference: the pg-boss schema family (job states,
maintenance/dead-letter behavior) is battle-tested against real
Postgres semantics — worth borrowing *as design*, independent of the
driver decision.
- The node original (pg-boss) is the upstream of record for semantics
the port may have dropped; compare against it when adopting queue
semantics.
### Honker's Postgres-side recommendation
The honker README's own posture: if you run Postgres, use the Postgres
tools. `pg_notify` + pgboss is the recommended assembly. The design
brief: the queue machinery from the pg-boss family, the push semantics
from LISTEN/NOTIFY, the unified API shape from honker's Rust binding.
### The alk* repo pattern — what this crate replaces
The ecosystem's current shape: a repository trait with a default
in-memory adapter; cache-invalidation wiring in hot paths; uninvalidated
(non-reactive) caches where delay was acceptable; each project
composing these slightly differently. No persistence-backed reactive
substrate exists in the family — alkblobs was the first project to try
to plan against one, found it missing, and paused. This crate's reason
to exist is precisely that that substrate should exist once, well,
instead of per-project approximations.
### alkcall — the substrate (not a dependency of the store layer)
`/workspace/@alkdev/alkcall` (pure protocol crate, no transport). The
alk* crates (alktty, alktunnels, alksocks) are its consumers; a future
alkstore ops/protocol surface (if this crate ever exposes store access
over alkcall channels) rides the same substrate. Like the alkblobs
split (store layer stays substrate-free), the store layer here stays
alkcall-free; any networked surface is an ops module/sibling concern
and a separate decision.
## Open Questions
Register in `docs/research/phase-0.md`; IDs OQ-ST-NN (stable, append
only). Promotion target: Phase 1 `docs/architecture/open-questions.md`.
### OQ-ST-01: Scope boundary — which honker features are in-scope?
Honker's feature list (queues, streams, notify, scheduling, locks, rate
limits, outbox helpers) is large; per-feature scope decisions don't
exist yet. Also undetermined: the exclusion lines (honker deliberately
excludes DAGs, task chains/chords, multi-writer replication,
distributed locking).
Not yet decidable without working through the concrete consumers (the
paused alkblobs crates and alkfs planning) — per-feature need hasn't
been articulated. Blocked on a consumer-driven inventory pass.
### OQ-ST-02: Crate scope — one store crate, or reactive-core + engines?
Options include: single crate with feature-gated engines (the alk*
feature-gate pattern); a core trait crate + per-engine crates; engine
crates consuming a thin core. The answer constrains the driver decision
(OQ-ST-04) and the base-crate-lean invariant.
The use case isn't concrete yet (no engine code written, no consumer
wired). Deferred(scope) — concrete use case: the first engine
implementation would force this shape.
### OQ-ST-03: Driver story — sqlx, tokio-postgres, or per-engine drivers?
The named tension (§The driver conflict), restated as the decision:
- **One driver across engines**: sqlx (both sqlite + postgres native
support, one API — but the alkblobs POC evidence is tokio-postgres)
vs tokio-postgres per-engine (sqlite story unclear — tokio-postgres
is pg-only; rusqlite is the sqlite native).
- **Per-engine drivers under a unified trait**: tokio-postgres +
deadpool-postgres (POC-validated in alkblobs findings) + rusqlite
(honker's choice, so honker's SQLite machinery ports cleanly).
- **Adopt/fork pgboss-rs**: brings sqlx along where the queue lives.
Honest unknowns worth surfacing: does sqlx support SQLite
`data_version`/extension-style machinery equally well? Does a unified
trait over `(tokio-postgres, rusqlite)` pay more trait-fitting cost
than sqlx's single-API convenience costs elsewhere? What does
SQLITE_ENABLE/extension loading look like under sqlx vs rusqlite?
Genuinely open; needs research rounds (library capabilities vs the
unified-trait shape) and possibly a POC. Not deferred — this is the
central Phase 0 research question.
### OQ-ST-04: The reactive abstraction — what does the unified notify surface look like?
The two engines' mechanisms are structurally different: SQLite =
watcher polling `PRAGMA data_version` (deliver on commit; no
server-side push exists), Postgres = LISTEN/NOTIFY (server push,
connection-bound, no retry/visibility semantics). The reactive trait
must have a shape both implement without one emulating the other's
weaknesses:
- What is the subscription type (`channel`? stream of envelopes?)
- What is the delivery guarantee contract on each engine (honker's
wake-on-commit + re-read is *not* exactly-once — what does the trait
promise?)
- Does the trait absorb the enqueue+notify-in-one-transaction shape
(honker's core) — and how does that compose with Postgres
transaction-scoped LISTEN semantics?
- How does a caching subscriber (the ecosystem's hot-path pattern)
receive sufficient invalidation information (keys? table/channel
names? opaque wake + re-read contract?)
Open; this is the second central research question, coupled to OQ-ST-03
(the driver determines what LISTEN plumbing exists).
### OQ-ST-05: Queue semantics — adopt, fork, or re-derive?
If queues land in scope (OQ-ST-01), the pg-boss schema family is the
Postgres-side incumbent and honker's queue design is the SQLite-side
one. Options: adopt pgboss-rs as a dependency (new feature-gated
option); targeted-fork the relevant subsystem (alksocks precedent,
ported to our conventions); schema/design-reference only (re-derive on
our driver). Fork-vs-derive depends on how much of pgboss-rs is
queue-machinery vs driver-wiring (the sqlx coupling — OQ-ST-03) and on
our tolerance for the alpha-state rc port.
Open; inputs: the OQ-ST-01 inventory + OQ-ST-03 resolution.
### OQ-ST-06: Honker relationship — reference, fork, or vendor?
Design-reference only (read, don't copy), targeted fork of
honker-core's engine machinery, or vendor the extension? Honker is
alpha-quality per its own README, MIT/Apache-2.0 dual-licensed, and
covers only the SQLite side — but it embodies exactly the watcher/
transactional design this crate wants on SQLite. Fork-postures in this
workspace have precedent (alksocks' extraction) but have been for
*owning* a needed subset, not for adopting an alpha wholesale.
Open; needs the license/provenance check (AGENTS.md §3) and a quality
assessment honker's watcher/transactional core.
### OQ-ST-07: SQLite-side scope — loadable extension, embedded rusqlite, or both?
Honker ships as a loadable extension usable by *any* SQLite client,
plus per-language bindings. This crate (a Rust library) may not need
the loadable-extension surface at all — embedding the engine machinery
in-process may be the whole story (the honker-core shape minus the
extension/binding packaging). Determines how much of honker is even
candidate material.
Open; rides the first consumer-driven scope pass (OQ-ST-01).
### OQ-ST-08: Multi-host / deployment posture
Honker is explicitly single-machine (file-backed SQLite). Postgres is
natively multi-host. The unified surface must not pretend SQLite is
multi-host, but where does the honest boundary live — per-engine
capability flags? A documented deployment matrix? Does the trait need
to expose engine capabilities at all?
Open; partially rides OQ-ST-04 (the trait's shape constrains where
capability differences can surface).
## Phase 0 plan
Iteration expected; this register grows as the consumer inventory and
research rounds land. Expected sequence (deliberately rough):
1. Consumer-driven scope inventory (OQ-ST-01) — what the paused
alkblobs consumers and alkfs planning actually need from the reactive
store; walks back the per-feature scope decisions.
2. Research rounds on OQ-ST-03/04 (drivers + reactive shape) — library
capability matrices, then a POC if the unified-trait shape needs
validation (likely, given the structural mismatch noted in OQ-ST-04).
3. Ownership decisions (OQ-ST-05/06) — adopt/fork/derive per subsystem,
after the driver and shape questions narrow the option space.
4. Converge; Phase 1 opens with the ADR backlog this register becomes.
## References
- honker — `/workspace/honker` (git checkout; README + `honker-core/src/`
read 2026-10-03): the SQLite-side feature/wake template.
- pgboss-rs — `/workspace/pgboss-rs` (git checkout of
github.com/rustworthy/pgboss-rs, v0.1.0-rc6; read 2026-10-03,
LISTEN/NOTIFY-absence verified): the Postgres queue-family reference.
- honker's own prior-art section: pg_notify, pg-boss, Oban, Huey — the
external lineage this crate inherits from both sides.
- alkblobs — `/workspace/@alkdev/alkblobs` (spec+POCs, paused): the
paused planning this crate unblocks; its POC findings
(`poc-postgres-kv-findings.md`, `poc-pglo-findings.md`) are the
tokio-postgres evidence base.
- alkcall — `/workspace/@alkdev/alkcall`: the family substrate;
referenced for the store-layer-isolation principle only.