Quality read of honker-core's watcher/transactional core cross-checked against the published crates.io artifact: the core itself is clean (Writer/Readers, polling-watcher failure handling, WatcherDeathGuard all verified), but published 0.5.0 predates upstream's unreleased fix train carrying the issue-#133 savepoint hardening (silent job loss in the dead-letter paths) and five .ok() error swallows — and ADR-010's queue depth requires engine-owned queue SQL in any posture. Resolution: fork honker-core at the reference revision, inherit the clean machinery and test suites, re-derive queue ops on contract v1, rename tables to __alkstore_*. - docs/research/quality-read-honker-core.md — full evidence - docs/architecture/decisions/011-sqlite-substrate-fork.md — decision - OQ-06 resolved in open-questions.md; ADR-003/005/008, engine-sqlite, queues, README annotated for consistency
19 KiB
status, last_updated
| status | last_updated |
|---|---|
| draft | 2026-10-05 |
Quality read: honker-core — OQ-06 (the SQLite dependency gate)
OQ-06's mandate (ADR-005):
a Phase 1 quality read of honker-core 0.5's watcher/transactional core
(Writer / Readers / SharedUpdateWatcher / attach_*), looking for defects
the POCs wouldn't surface, unsafe assumptions in the watcher
failure-handling, and schema-migration brittleness — and a weighing of
the two concrete fork candidates ADR-010 handed it (the expired-
processing-row zombie fix, dead-row get_job visibility) as
complement-over-machinery vs fork.
Verdict: the fork trigger fires — ADR-011. The core the read was originally for is clean, but the read surfaced that the published artifact is materially behind the reference revision and carries confirmed silent-job-loss defects in exactly the queue operations the ride posture points at, while contract v1's queue depth (ADR-010) requires re-deriving that surface in any posture. The fork is normal work per ADR-005. This document is the evidence; the ADR is the decision.
1. Read scope and the published-artifact fact
Read in full: honker-core at the reference checkout
(/workspace/honker @ f4e53c6, 2026-10-02) — all five source files
(8,833 lines): lib.rs (PRAGMA/WAL handling, Writer, Readers,
UpdateWatcher + run_poll_loop, SharedUpdateWatcher +
WatcherDeathGuard, bootstrap + migrations), honker_ops.rs
(savepoint machinery, every queue / lock / stream / scheduler /
notifications function), kernel_watcher.rs, shm_watcher.rs,
cron.rs (skimmed — contract-rejected per ADR-009).
Cross-checked against crates.io: honker-core's newest published
version is 0.5.0 — published 2026-08-23 (commit 78dec24 = tag
v0.5.0; same commit as rust-v0.5.0), single maintainer, 8 versions
total, no yanks. The reference checkout HEAD is ~6 weeks ahead and
contains a body of correctness work no published release carries
(the v0.6.0 tag is a package-versioning tag; the honker-core crate at
even v0.6.0 is still 0.5.0 and still lacks the fixes).
What 0.5.0 lacks (the upstream "Unreleased" fix train, all honker-core, verified by diffing the tag against HEAD):
- Issue #133's savepoint hardening (
in_savepoint/UnwindUndo,honker_ops.rs:118-254at f4e53c6) — absent from 0.5.0 entirely (grep: 0 occurrences). Five dead-letter/sweep paths in 0.5.0 are bareDELETE … RETURNING+ decode + INSERT sequences with no savepoint (commits3ab43aa,753f6d0). - Five
.ok()error-swallows in 0.5.0'sretry,fail,get_job,lock_acquire,result_get— every SQLite error (I/O, corruption, disk full) maps to "no row"/"not our claim"/ "lock held", not justQueryReturnedNoRows(commit4881f27). claimed_at(schema column + migration) — absent from 0.5.0'sBOOTSTRAP_HONKER_SQL; added at HEAD only (6780cae/PR #140).
Present in 0.5.0 already (i.e. the watch-side hardening predates the
release): issue #80's -shm-descriptor registry and the kqueue
lock-bearing-path guards (verified in the tag's shm_watcher.rs /
kernel_watcher.rs), and the full WatcherDeathGuard machinery.
Consequence of the artifact fact: riding published 0.5.0 means shipping the confirmed job-loss windows of §3 below inside the exact functions ADR-010's ride posture points at; the fixes exist upstream, are documented in honker's own CHANGELOG as real defects, and are stranded unreleased with no announced date.
2. Watcher / transactional core — the read's original mandate
Verdict: clean. Per component (line cites at f4e53c6 unless
prefixed v0.5.0:, which cite the published tag):
Writer(lib.rs:549-617; identical inv0.5.0:lib.rs): single-connection write slot overMutex<Option<Connection>>+ condvar +closedAtomicBool. Correct — theacquirewait loop re-checksclosedafter each wakeup (close wakes blocked acquirers withNone),releaseafterclosedrops the connection instead of pooling it,closeis idempotent. The explicit-closedesign (binding-GC pressure relief) is sound for our shape too.Readers(lib.rs:629-716): the subtle race handling is right — capacity slot is re-checked afteropen_connreturns (a close-race drops the brand-new connection and releases the slot it reserved), and the failed-open path decrementsoutstandingso transient open failures cannot permanently shrink the pool. The closed sentinel (SQLITE_MISUSE+ "Database is closed") is a hack, but a documented and harmless one.- Polling watcher (
run_poll_loop,lib.rs:832-937): the three-layer defense is genuinely good and is the quality POC #1 measured. (a)PRAGMA data_versionfast path (~3.5 µs/poll). (b) Transient-vs-fatal error classification:SQLITE_BUSY/LOCKEDare retried in place — critically not treated as connection death, because a reconnect would silently re-baselinelast_versionand skip pending wakes; fatal errors drop the connection, fire one conservativeon_change(), and reconnect (re-baselining again — the conservative wake covers the gap). (c) The dead-man's switch:statidentity(dev, ino)~every 100 ms; a replaced file (atomic rename, litestream restore, remount) panics the watcher thread with a precise message rather than watching stale data. All present in 0.5.0. SharedUpdateWatcher(lib.rs:1063-1176; death guard at:1053-1061, in 0.5.0): one poll thread, N subscribers, capacity-1 channels (bursts coalesce), disconnected-subscriber pruning onTrySendError::Disconnected, andWatcherDeathGuard— thread exit or panic → closure drops → guard'sDropclears every sender → each subscriber's nextrecv()returnsErr. The failure-handling property ADR-006 pins ("watcher death closes receivers, never a silent hang") is mechanically verified in source, present in 0.5.0.- Experimental backends (
kernel_watcher.rs,shm_watcher.rs): separate opt-in Cargo features with explicitly documented weaker contracts (missed wakes possible; init failure = stderr + a backend that produces no wakes). Their issue-#80 handling — never close a descriptor on a lock-bearing inode (-shm/main db under kqueue), with the reasoning stated atkernel_watcher.rs:393-423— is exemplary systems work. We never enable these features; in the fork they are dropped (§6). attach_*(honker_ops.rs:259+): standard rusqlite scalar-function registration; documented as not idempotent ("call exactly once per connection") — the engine's open path must respect that (it already does, per ADR-003's wiring).
Minor findings — none trigger anything, recorded as fork-port notes:
- W-1 (reconnect backoff): the reconnect loop retries one open
attempt + one
eprintlnper poll tick (1 ms). A db file that disappears mid-flight produces ~1000 log lines/sec until closure. In our engine the watcher closes with subscribers on db loss anyway, but the fork's port should add bounded backoff. - W-2 (thread-build panic):
spawn_with_configends in.expect("spawn update-poll thread")— a panic in library code (thread-budget exhaustion). Our crate rules ban panics outside tests; the fork's port makes watcher spawn fallible. Trivial. - W-3 (
data_versionu32 wrap): at sustained max commit rates the counter can wrap and alias the last seen value — one missed wake per 2³² commits. Hint-signal semantics absorb it (consumers re-read; subsequent commits re-wake). Not actionable. - W-4 (watcher connection opens RW):
run_poll_loopopens withREAD_WRITE; adata_versionread is readable RO. Harmless (the watcher legitimately needs write access on non-WAL journals' busy paths? unproven) — noted so the fork's port makes a deliberate choice rather than inheriting one.
3. The published-0.5.0 defect register (what the fork fires on)
Cites prefixed v0.5.0: are into the published tag's sources (read
directly from the tag); HEAD cites are into the reference checkout.
"Would-be engine impact" assumes the ADR-003/ADR-010 ride posture
(consuming these functions through the tx seam).
- D-1 —
fail()destroys a job and reports "not our claim." (v0.5.0:honker-core/src/honker_ops.rs:911-948)failrunsDELETE … RETURNINGand then swallows the read with.ok()(:934): a decode/mapper failure is converted toOk(0)after the DELETE has executed inside that statement — the row is gone from_honker_live, the_honker_deadINSERT never runs, the caller is told the claim wasn't theirs. Silent job loss on any decode failure or I/O error mid-statement. Upstream:4881f27(error propagation) +3ab43aa(savepoint) — fixed at HEAD, unreleased. The upstream CHANGELOG states it outright: "a failedfail()leaves the job claimable instead of destroyed." - D-2 — dead-letter moves are not atomic.
(
v0.5.0:honker_ops.rs:607-650dead_letter_exhausted_claimable,:1079-1119sweep_expired, and retry's dead-letter branch:890-907) DELETE-then-INSERT with no savepoint; a mid-loop INSERT failure leaves rows deleted from_honker_livewhose_honker_deadINSERT never came — stranded between tables. (The pre-claim path inclaim_batch:663 uses it too.) Upstream fixed all four sites with thein_savepointtrain. - D-3 —
retryswallow. (v0.5.0:honker_ops.rs:866) The claim read ends.ok(): a SQLite error maps toOk(0)("not our claim") while the row sitsprocessingwith a dead worker — silently extending the dual-execution window until a reclaimer arrives. - D-4 —
lock_acquireswallow. (v0.5.0:honker_ops.rs:1146) The owner read-back ends.ok(): a broken database reports the lock as not acquired — if leadership acquisition rides it, the leader silently never starts (a stall, not an error). The scheduler'srun_scheduleswould have inherited this via the leadership-lock path. - D-5 —
get_jobswallow. (v0.5.0:honker_ops.rs:1016) Error reported as job-miss. Diagnosis-by-API under a broken stored value lies. - D-6 — no zombie reachability (both 0.5.0 and HEAD): an expired
processing row whose worker died is unreachable by the claim
predicate (HEAD :878), by pre-claim dead-lettering (HEAD :792), and by
sweep_expired(pending-only, HEAD :1346) — the ADR-010 §5 no-stranded-rows property fails upstream of anything we control. (This is fork candidate #1.) - D-7 — schema gaps vs contract v1 (0.5.0; mostly HEAD too): no
claimed_atin 0.5.0's schema; no per-job stamps anywhere (visibility/backoff/retention columns of ADR-010 §3a);get_jobreads live-rows only (HEAD :1237-1240 — fork candidate #2) and exposes neither stamps norclaimed_atin its JSON (HEAD's #136 tracks claimed_at exposure as open).
The POCs measured happy paths, commit/rollback atomicity, concurrency, and latency; D-1..D-5 are failure-interrupted windows and corrupted-value paths — exactly the class POC measurement does not surface.
4. Schema / migration brittleness — verdict: minor
- Bootstrap (
lib.rs:441-510):CREATE TABLE IF NOT EXISTS+ALTER TABLE ADD COLUMNmigrations guarded bypragma_table_info. The concurrent-bootstrap "duplicate column" race is handled honestly (SQLite serializes writes file-wide; the loser swallows that specific error) — but the swallow keys on a lowercase error-string match (.contains("duplicate column")), brittle to SQLite message rewording. A rewording turns a benign race into a loud bootstrap error (recoverable by retry) — not a correctness hole. - No
schema_versiontable; migrations are append-column-only, per-table, in-code. Adequate at this scale (three columns added over the project's life). - The co-migration exposure: realizing contract v1's
claimed_at(get_job, ADR-010 §1) and opt stamps (§3a) over honker's tables means the engine runs its own migrations over_honker_*— two codebases (ours and honker-core's) migrating one table family, with only an exact version pin holding the seam stable. Either posture must resolve this; the fork resolves it by unifying ownership.
5. The fork calculus — the two candidates, plus the §3a stamps
Framing per OQ-06: complement-over-machinery vs fork, on the concrete candidates.
- Zombie fix (D-6): complement is genuinely small — one engine-side SQL function over honker's own tables (move expired-processing rows to dead; the pending half exists, the processing half doesn't). Achievable over the machinery.
- Dead-visible
get_job(D-7): complement is small nominally — but contractget_jobalso carriesclaimed_at(absent from 0.5.0's schema → engineALTER TABLEover honker's table, or a sidecar) and the stamps. Achievable, but it crosses schema ownership. - Stamps (ADR-010 §3a): per-row visibility deadlines honored in a single claim statement (the pinned deadline semantics) require rows carrying their stamps and a claim statement that reads them — i.e. engine-owned enqueue + claim SQL in any posture. The over-machinery alternative (claim with the uniform queue timeout, then re-stamp per row post-claim) is a two-statement bridge that bends the pinned "deadline = the job's stamp" semantics for one bounded window — the kind of contract-honesty fudge ADR-010 explicitly worked to avoid.
The pivot: once §3a and the contract get_job land, the engine
owns enqueue / claim / retry / fail / sweep / get_job — the majority of
the queue-op surface — in both postures. The ride then covers
ack/heartbeat/cancel/locks/streams/notify/scheduler plus plumbing
(watcher, Writer/Readers, pragmas, bootstrap). The fork question
reduces to: own the remaining half, or keep depending on it?
The artifact fact (§1) decides it: the half you'd keep riding carries confirmed silent-job-loss windows in its published form (D-1..D-5), its fixes are stranded unreleased on a bindings' cadence, its crate description says "not intended for direct use" (its intended consumers are three language bindings, not direct library consumption — no upstream support contract for our posture), and the co-migration exposure keeps the schema dual-owned in the complement posture. The fork also inherits ~4,000 lines of upstream tests (savepoint, watcher, multiprocess pressure suites) — a large share of the fork cost is pre-paid.
Rejected postures, recorded for revisit:
- Wait for upstream 0.5.1+: unbounded timing; the read's facts stand for the state actually measured.
- Git-dependency pinned to upstream main: pins an unreleased branch as a production substrate, still lacks every contract delta and the zombie fix (both present at f4e53c6 — verified), still dual-owned schema; adds workspace/monorepo resolution gymnastics. Rejected.
- Complement v2 (engine-owned queue SQL over honker's tables, keep honker for the rest): the strongest alternative — recorded with its genuine merits (zero vendored code; the defect functions are simply never called) — and rejected on the dual-owner schema, dead-weight dependency, release-cadence hostage, and direct-use-unnovation factors above.
6. Fork scope (feeds ADR-011)
- Fork at the f4e53c6 state (the full fix train included), not the
published tag — provenance: honker-core, MIT OR Apache-2.0, upstream
commit
f4e53c6(packagenode-v0.5.1-10-gf4e53c6), recorded per AGENTS.md §3. - Keep as-ported: the PRAGMA block +
set_journal_mode_walretry logic;Writer;Readers; the polling watcher +SharedUpdateWatcherWatcherDeathGuard+stat_identitydead-man's switch (W-1 backoff, W-2 fallible spawn applied in port); thein_savepoint/UnwindUndomachinery as the mutation-transition discipline; the REAL-coercion arg helpers; thenotify()scalar + notifications table (renamed, plus an at-attach pruning cap implementing ADR-010 §6's engine-internal hygiene); stream functions; lock functions (with the same-owner-TTL-not-refreshed behavior documented as-is — contract pinning already covers it via the locks guarantee row).
- Re-derive on contract v1 (ADR-010): enqueue + stamping;
single-statement claim with per-row visibility from stamps;
savepoint-guarded retry/fail/dead-letter/sweep; both-states
no-stranded-rows
sweep_expired(+dead_letter_retention_sdeletion); dead-visibleget_job(stamps +claimed_at+last_error/died_at); scheduler tick over the new enqueue with@everynext-boundary math (numeric, a few lines) — cron boundary machinery not ported. - Drop, do not port:
cron.rs(ADR-009 rejects cron strings); thekernel-watcher/shm-fast-pathfeatures and their optional deps (notify,memmap2,libc); the rate-limit and result tables (cut-flags, ADR-002); the superseded queue functions. - Table naming:
_honker_*→__alkstore_*across the fork's storage surface — ADR-010 §8 pre-authorized the fork re-owning names ("nothing consumer-visible changes"); storage-internal per ADR-008 §4. - Tests: inherit honker-core's suites as the floor (PRAGMA/WAL,
watcher lifecycle + failure handling, savepoint + multiprocess
pressure) and add the contract-property tests (no-stranded-rows,
dead-visible
get_job, stamps immutability, wake-latency floor) — the core-contract verification backlog's SQLite column.
7. Follow-through
- ADR-011 — the fork decision and packaging.
- ADR-005 — posture row for honker-core annotated (trigger fired); the calculus is applied, not renegotiated.
- ADR-003 — "published honker-core" is superseded on ownership only; driver, seam, and watcher architecture are unchanged (the fork inherits them verbatim where kept).
- engine-sqlite.md — living spec updated to the forked substrate.
- OQ-06 marked resolved in open-questions.md with a short form + reference to this document.
- Revisit note for whoever picks this up later: upstream fixed the defect class once pressed (issues #80/#133 → commits at HEAD). If a future upstream release ships the train and moves toward contract shapes we re-derived, re-adoption is conceivable — but by then the fork is owned, tested, and ahead on contract deltas; re-adoption would have to beat it on maintenance, not on novelty.