Files
alkstore/docs/research/quality-read-honker-core.md
T
glm-5.3-flash befbe2e714 OQ-06 resolved: honker-core quality read fires the fork trigger (ADR-011)
Quality read of honker-core's watcher/transactional core cross-checked
against the published crates.io artifact: the core itself is clean
(Writer/Readers, polling-watcher failure handling, WatcherDeathGuard
all verified), but published 0.5.0 predates upstream's unreleased fix
train carrying the issue-#133 savepoint hardening (silent job loss in
the dead-letter paths) and five .ok() error swallows — and ADR-010's
queue depth requires engine-owned queue SQL in any posture. Resolution:
fork honker-core at the reference revision, inherit the clean machinery
and test suites, re-derive queue ops on contract v1, rename tables to
__alkstore_*.

- docs/research/quality-read-honker-core.md — full evidence
- docs/architecture/decisions/011-sqlite-substrate-fork.md — decision
- OQ-06 resolved in open-questions.md; ADR-003/005/008, engine-sqlite,
  queues, README annotated for consistency
2026-10-05 03:39:46 +00:00

19 KiB

status, last_updated
status last_updated
draft 2026-10-05

Quality read: honker-core — OQ-06 (the SQLite dependency gate)

OQ-06's mandate (ADR-005): a Phase 1 quality read of honker-core 0.5's watcher/transactional core (Writer / Readers / SharedUpdateWatcher / attach_*), looking for defects the POCs wouldn't surface, unsafe assumptions in the watcher failure-handling, and schema-migration brittleness — and a weighing of the two concrete fork candidates ADR-010 handed it (the expired- processing-row zombie fix, dead-row get_job visibility) as complement-over-machinery vs fork.

Verdict: the fork trigger fires — ADR-011. The core the read was originally for is clean, but the read surfaced that the published artifact is materially behind the reference revision and carries confirmed silent-job-loss defects in exactly the queue operations the ride posture points at, while contract v1's queue depth (ADR-010) requires re-deriving that surface in any posture. The fork is normal work per ADR-005. This document is the evidence; the ADR is the decision.

1. Read scope and the published-artifact fact

Read in full: honker-core at the reference checkout (/workspace/honker @ f4e53c6, 2026-10-02) — all five source files (8,833 lines): lib.rs (PRAGMA/WAL handling, Writer, Readers, UpdateWatcher + run_poll_loop, SharedUpdateWatcher + WatcherDeathGuard, bootstrap + migrations), honker_ops.rs (savepoint machinery, every queue / lock / stream / scheduler / notifications function), kernel_watcher.rs, shm_watcher.rs, cron.rs (skimmed — contract-rejected per ADR-009).

Cross-checked against crates.io: honker-core's newest published version is 0.5.0 — published 2026-08-23 (commit 78dec24 = tag v0.5.0; same commit as rust-v0.5.0), single maintainer, 8 versions total, no yanks. The reference checkout HEAD is ~6 weeks ahead and contains a body of correctness work no published release carries (the v0.6.0 tag is a package-versioning tag; the honker-core crate at even v0.6.0 is still 0.5.0 and still lacks the fixes).

What 0.5.0 lacks (the upstream "Unreleased" fix train, all honker-core, verified by diffing the tag against HEAD):

  1. Issue #133's savepoint hardening (in_savepoint/UnwindUndo, honker_ops.rs:118-254 at f4e53c6) — absent from 0.5.0 entirely (grep: 0 occurrences). Five dead-letter/sweep paths in 0.5.0 are bare DELETE … RETURNING + decode + INSERT sequences with no savepoint (commits 3ab43aa, 753f6d0).
  2. Five .ok() error-swallows in 0.5.0's retry, fail, get_job, lock_acquire, result_get — every SQLite error (I/O, corruption, disk full) maps to "no row"/"not our claim"/ "lock held", not just QueryReturnedNoRows (commit 4881f27).
  3. claimed_at (schema column + migration) — absent from 0.5.0's BOOTSTRAP_HONKER_SQL; added at HEAD only (6780cae/PR #140).

Present in 0.5.0 already (i.e. the watch-side hardening predates the release): issue #80's -shm-descriptor registry and the kqueue lock-bearing-path guards (verified in the tag's shm_watcher.rs / kernel_watcher.rs), and the full WatcherDeathGuard machinery.

Consequence of the artifact fact: riding published 0.5.0 means shipping the confirmed job-loss windows of §3 below inside the exact functions ADR-010's ride posture points at; the fixes exist upstream, are documented in honker's own CHANGELOG as real defects, and are stranded unreleased with no announced date.

2. Watcher / transactional core — the read's original mandate

Verdict: clean. Per component (line cites at f4e53c6 unless prefixed v0.5.0:, which cite the published tag):

  • Writer (lib.rs:549-617; identical in v0.5.0:lib.rs): single-connection write slot over Mutex<Option<Connection>> + condvar + closed AtomicBool. Correct — the acquire wait loop re-checks closed after each wakeup (close wakes blocked acquirers with None), release after close drops the connection instead of pooling it, close is idempotent. The explicit-close design (binding-GC pressure relief) is sound for our shape too.
  • Readers (lib.rs:629-716): the subtle race handling is right — capacity slot is re-checked after open_conn returns (a close-race drops the brand-new connection and releases the slot it reserved), and the failed-open path decrements outstanding so transient open failures cannot permanently shrink the pool. The closed sentinel (SQLITE_MISUSE + "Database is closed") is a hack, but a documented and harmless one.
  • Polling watcher (run_poll_loop, lib.rs:832-937): the three-layer defense is genuinely good and is the quality POC #1 measured. (a) PRAGMA data_version fast path (~3.5 µs/poll). (b) Transient-vs-fatal error classification: SQLITE_BUSY/LOCKED are retried in place — critically not treated as connection death, because a reconnect would silently re-baseline last_version and skip pending wakes; fatal errors drop the connection, fire one conservative on_change(), and reconnect (re-baselining again — the conservative wake covers the gap). (c) The dead-man's switch: stat identity (dev, ino) ~every 100 ms; a replaced file (atomic rename, litestream restore, remount) panics the watcher thread with a precise message rather than watching stale data. All present in 0.5.0.
  • SharedUpdateWatcher (lib.rs:1063-1176; death guard at :1053-1061, in 0.5.0): one poll thread, N subscribers, capacity-1 channels (bursts coalesce), disconnected-subscriber pruning on TrySendError::Disconnected, and WatcherDeathGuard — thread exit or panic → closure drops → guard's Drop clears every sender → each subscriber's next recv() returns Err. The failure-handling property ADR-006 pins ("watcher death closes receivers, never a silent hang") is mechanically verified in source, present in 0.5.0.
  • Experimental backends (kernel_watcher.rs, shm_watcher.rs): separate opt-in Cargo features with explicitly documented weaker contracts (missed wakes possible; init failure = stderr + a backend that produces no wakes). Their issue-#80 handling — never close a descriptor on a lock-bearing inode (-shm/main db under kqueue), with the reasoning stated at kernel_watcher.rs:393-423 — is exemplary systems work. We never enable these features; in the fork they are dropped (§6).
  • attach_* (honker_ops.rs:259+): standard rusqlite scalar-function registration; documented as not idempotent ("call exactly once per connection") — the engine's open path must respect that (it already does, per ADR-003's wiring).

Minor findings — none trigger anything, recorded as fork-port notes:

  • W-1 (reconnect backoff): the reconnect loop retries one open attempt + one eprintln per poll tick (1 ms). A db file that disappears mid-flight produces ~1000 log lines/sec until closure. In our engine the watcher closes with subscribers on db loss anyway, but the fork's port should add bounded backoff.
  • W-2 (thread-build panic): spawn_with_config ends in .expect("spawn update-poll thread") — a panic in library code (thread-budget exhaustion). Our crate rules ban panics outside tests; the fork's port makes watcher spawn fallible. Trivial.
  • W-3 (data_version u32 wrap): at sustained max commit rates the counter can wrap and alias the last seen value — one missed wake per 2³² commits. Hint-signal semantics absorb it (consumers re-read; subsequent commits re-wake). Not actionable.
  • W-4 (watcher connection opens RW): run_poll_loop opens with READ_WRITE; a data_version read is readable RO. Harmless (the watcher legitimately needs write access on non-WAL journals' busy paths? unproven) — noted so the fork's port makes a deliberate choice rather than inheriting one.

3. The published-0.5.0 defect register (what the fork fires on)

Cites prefixed v0.5.0: are into the published tag's sources (read directly from the tag); HEAD cites are into the reference checkout. "Would-be engine impact" assumes the ADR-003/ADR-010 ride posture (consuming these functions through the tx seam).

  • D-1 — fail() destroys a job and reports "not our claim." (v0.5.0:honker-core/src/honker_ops.rs:911-948) fail runs DELETE … RETURNING and then swallows the read with .ok() (:934): a decode/mapper failure is converted to Ok(0) after the DELETE has executed inside that statement — the row is gone from _honker_live, the _honker_dead INSERT never runs, the caller is told the claim wasn't theirs. Silent job loss on any decode failure or I/O error mid-statement. Upstream: 4881f27 (error propagation) + 3ab43aa (savepoint) — fixed at HEAD, unreleased. The upstream CHANGELOG states it outright: "a failed fail() leaves the job claimable instead of destroyed."
  • D-2 — dead-letter moves are not atomic. (v0.5.0:honker_ops.rs:607-650 dead_letter_exhausted_claimable, :1079-1119 sweep_expired, and retry's dead-letter branch :890-907) DELETE-then-INSERT with no savepoint; a mid-loop INSERT failure leaves rows deleted from _honker_live whose _honker_dead INSERT never came — stranded between tables. (The pre-claim path in claim_batch :663 uses it too.) Upstream fixed all four sites with the in_savepoint train.
  • D-3 — retry swallow. (v0.5.0:honker_ops.rs:866) The claim read ends .ok(): a SQLite error maps to Ok(0) ("not our claim") while the row sits processing with a dead worker — silently extending the dual-execution window until a reclaimer arrives.
  • D-4 — lock_acquire swallow. (v0.5.0:honker_ops.rs:1146) The owner read-back ends .ok(): a broken database reports the lock as not acquired — if leadership acquisition rides it, the leader silently never starts (a stall, not an error). The scheduler's run_schedules would have inherited this via the leadership-lock path.
  • D-5 — get_job swallow. (v0.5.0:honker_ops.rs:1016) Error reported as job-miss. Diagnosis-by-API under a broken stored value lies.
  • D-6 — no zombie reachability (both 0.5.0 and HEAD): an expired processing row whose worker died is unreachable by the claim predicate (HEAD :878), by pre-claim dead-lettering (HEAD :792), and by sweep_expired (pending-only, HEAD :1346) — the ADR-010 §5 no-stranded-rows property fails upstream of anything we control. (This is fork candidate #1.)
  • D-7 — schema gaps vs contract v1 (0.5.0; mostly HEAD too): no claimed_at in 0.5.0's schema; no per-job stamps anywhere (visibility/backoff/retention columns of ADR-010 §3a); get_job reads live-rows only (HEAD :1237-1240 — fork candidate #2) and exposes neither stamps nor claimed_at in its JSON (HEAD's #136 tracks claimed_at exposure as open).

The POCs measured happy paths, commit/rollback atomicity, concurrency, and latency; D-1..D-5 are failure-interrupted windows and corrupted-value paths — exactly the class POC measurement does not surface.

4. Schema / migration brittleness — verdict: minor

  • Bootstrap (lib.rs:441-510): CREATE TABLE IF NOT EXISTS + ALTER TABLE ADD COLUMN migrations guarded by pragma_table_info. The concurrent-bootstrap "duplicate column" race is handled honestly (SQLite serializes writes file-wide; the loser swallows that specific error) — but the swallow keys on a lowercase error-string match (.contains("duplicate column")), brittle to SQLite message rewording. A rewording turns a benign race into a loud bootstrap error (recoverable by retry) — not a correctness hole.
  • No schema_version table; migrations are append-column-only, per-table, in-code. Adequate at this scale (three columns added over the project's life).
  • The co-migration exposure: realizing contract v1's claimed_at (get_job, ADR-010 §1) and opt stamps (§3a) over honker's tables means the engine runs its own migrations over _honker_* — two codebases (ours and honker-core's) migrating one table family, with only an exact version pin holding the seam stable. Either posture must resolve this; the fork resolves it by unifying ownership.

5. The fork calculus — the two candidates, plus the §3a stamps

Framing per OQ-06: complement-over-machinery vs fork, on the concrete candidates.

  • Zombie fix (D-6): complement is genuinely small — one engine-side SQL function over honker's own tables (move expired-processing rows to dead; the pending half exists, the processing half doesn't). Achievable over the machinery.
  • Dead-visible get_job (D-7): complement is small nominally — but contract get_job also carries claimed_at (absent from 0.5.0's schema → engine ALTER TABLE over honker's table, or a sidecar) and the stamps. Achievable, but it crosses schema ownership.
  • Stamps (ADR-010 §3a): per-row visibility deadlines honored in a single claim statement (the pinned deadline semantics) require rows carrying their stamps and a claim statement that reads them — i.e. engine-owned enqueue + claim SQL in any posture. The over-machinery alternative (claim with the uniform queue timeout, then re-stamp per row post-claim) is a two-statement bridge that bends the pinned "deadline = the job's stamp" semantics for one bounded window — the kind of contract-honesty fudge ADR-010 explicitly worked to avoid.

The pivot: once §3a and the contract get_job land, the engine owns enqueue / claim / retry / fail / sweep / get_job — the majority of the queue-op surface — in both postures. The ride then covers ack/heartbeat/cancel/locks/streams/notify/scheduler plus plumbing (watcher, Writer/Readers, pragmas, bootstrap). The fork question reduces to: own the remaining half, or keep depending on it?

The artifact fact (§1) decides it: the half you'd keep riding carries confirmed silent-job-loss windows in its published form (D-1..D-5), its fixes are stranded unreleased on a bindings' cadence, its crate description says "not intended for direct use" (its intended consumers are three language bindings, not direct library consumption — no upstream support contract for our posture), and the co-migration exposure keeps the schema dual-owned in the complement posture. The fork also inherits ~4,000 lines of upstream tests (savepoint, watcher, multiprocess pressure suites) — a large share of the fork cost is pre-paid.

Rejected postures, recorded for revisit:

  • Wait for upstream 0.5.1+: unbounded timing; the read's facts stand for the state actually measured.
  • Git-dependency pinned to upstream main: pins an unreleased branch as a production substrate, still lacks every contract delta and the zombie fix (both present at f4e53c6 — verified), still dual-owned schema; adds workspace/monorepo resolution gymnastics. Rejected.
  • Complement v2 (engine-owned queue SQL over honker's tables, keep honker for the rest): the strongest alternative — recorded with its genuine merits (zero vendored code; the defect functions are simply never called) — and rejected on the dual-owner schema, dead-weight dependency, release-cadence hostage, and direct-use-unnovation factors above.

6. Fork scope (feeds ADR-011)

  • Fork at the f4e53c6 state (the full fix train included), not the published tag — provenance: honker-core, MIT OR Apache-2.0, upstream commit f4e53c6 (package node-v0.5.1-10-gf4e53c6), recorded per AGENTS.md §3.
  • Keep as-ported: the PRAGMA block + set_journal_mode_wal retry logic; Writer; Readers; the polling watcher + SharedUpdateWatcher
    • WatcherDeathGuard + stat_identity dead-man's switch (W-1 backoff, W-2 fallible spawn applied in port); the in_savepoint / UnwindUndo machinery as the mutation-transition discipline; the REAL-coercion arg helpers; the notify() scalar + notifications table (renamed, plus an at-attach pruning cap implementing ADR-010 §6's engine-internal hygiene); stream functions; lock functions (with the same-owner-TTL-not-refreshed behavior documented as-is — contract pinning already covers it via the locks guarantee row).
  • Re-derive on contract v1 (ADR-010): enqueue + stamping; single-statement claim with per-row visibility from stamps; savepoint-guarded retry/fail/dead-letter/sweep; both-states no-stranded-rows sweep_expired (+ dead_letter_retention_s deletion); dead-visible get_job (stamps + claimed_at + last_error/died_at); scheduler tick over the new enqueue with @every next-boundary math (numeric, a few lines) — cron boundary machinery not ported.
  • Drop, do not port: cron.rs (ADR-009 rejects cron strings); the kernel-watcher / shm-fast-path features and their optional deps (notify, memmap2, libc); the rate-limit and result tables (cut-flags, ADR-002); the superseded queue functions.
  • Table naming: _honker_* → __alkstore_* across the fork's storage surface — ADR-010 §8 pre-authorized the fork re-owning names ("nothing consumer-visible changes"); storage-internal per ADR-008 §4.
  • Tests: inherit honker-core's suites as the floor (PRAGMA/WAL, watcher lifecycle + failure handling, savepoint + multiprocess pressure) and add the contract-property tests (no-stranded-rows, dead-visible get_job, stamps immutability, wake-latency floor) — the core-contract verification backlog's SQLite column.

7. Follow-through

  • ADR-011 — the fork decision and packaging.
  • ADR-005 — posture row for honker-core annotated (trigger fired); the calculus is applied, not renegotiated.
  • ADR-003 — "published honker-core" is superseded on ownership only; driver, seam, and watcher architecture are unchanged (the fork inherits them verbatim where kept).
  • engine-sqlite.md — living spec updated to the forked substrate.
  • OQ-06 marked resolved in open-questions.md with a short form + reference to this document.
  • Revisit note for whoever picks this up later: upstream fixed the defect class once pressed (issues #80/#133 → commits at HEAD). If a future upstream release ships the train and moves toward contract shapes we re-derived, re-adoption is conceivable — but by then the fork is owned, tested, and ahead on contract deltas; re-adoption would have to beat it on maintenance, not on novelty.