SQLite commit-error-arm coverage (task sqlite-commit-error-arm): the wave-3 review gate's deferred test landed — failed_commit_replenishes_the_writer_slot drives a failing COMMIT through the engine's commit path and pins the review's code-read fix end-to-end: the error surfaces as the opaque Database carrying the SQLITE_FULL-shaped source chain, the failed connection is dropped and the writer slot replenished via the handle's reopen closure (the next begin_tx proceeds within a bounded timeout, no parking — the store-wide-livelock posture), no partial-commit residue (the dropped connection's uncommitted writes read back None), the post-failure commit is a real clean commit (the fault disarms on consumption), and auto-commit notify works afterward. Injection is a cfg(test) commit-fault seam in seam.rs — a per-store Arc<AtomicBool> arm (born disarmed, armed via arm_commit_fault, take() disarms on first consumption so exactly one commit faults) whose fabricated rusqlite SqliteFailure feeds the production commit error arm rather than replicating it; the PRAGMA max_page_count route was probed live against both WAL and DELETE journal modes first and rejected: SQLite checks the page-count limit at page-allocation time, so the squeeze always fails the growth statement (SQLITE_FULL/DiskFull on the first INSERT) and leaves no transaction active for COMMIT to fail — the arm is unreachable through PRAGMA-space (also probed: the pragma is per-connection, so pre-begin arming on the writer conn would have ridden into the tx conn; the route failed on error placement, not delivery). Mechanism choice and probes documented in the seam doc comment and the task Notes. Cross-test safety is per-store scoping; parallel stores never see the arm. Replay-proofed live: with the error arm's writer_reopen replenish temporarily removed the test fails (begin_tx parks past the 5 s timeout — the stranding the review identified), reverted it passes. Plumbing follows the pg-fix-forwarder-reconnect cfg(test) precedent: fields on SqliteStore/SqliteTxHandle and a begin param are cfg-gated, production builds compile the plain path. The waves-1-2 review's optional watcher reconnect-success add rides here (taken — recorded in Notes): reconnect_success_resumes_wake_delivery drives run_poll_loop through its existing open_conn_fn seam (same instrument as the W-1 failure test), with the db file present throughout because the vanished-file route cannot reach the success body (file reappearance trips the dead-man's identity switch first): initial open + first two reconnects fail by injection, the third reconnect succeeds, and a subsequent commit wakes on_change — the success arm's data_version re-baseline and restored delivery pinned. Watcher shape untouched. Verified: cargo test -p alkstore-sqlite green server-less (191 lib + 25 suite), workspace cargo test 399/0, clippy -D warnings, fmt clean

This commit is contained in:
glm-5.3-flash committed 2026-10-10 07:46:40 +00:00
1 parent 36023914b2
commit 0c7744977c
6 files changed
+320 -8

No files matched your search

+39
View File
@@ -160,3 +160,42 @@ pub(crate) fn is_closed_err(e: &rusqlite::Error) -> bool {
pub(crate) fn closed_store_error() -> Error {
Error::database(std::io::Error::other("the store is closed"))
}
/// The `cfg(test)` commit-fault seam (`sqlite-commit-error-arm`): a
/// per-store arm that, when set, substitutes a `SQLITE_FULL`-shaped
/// driver error for one `COMMIT`, feeding the real commit error arm
/// (drop + `reopen` replenish) instead of fabricating its behavior.
/// Chosen over the `PRAGMA max_page_count` squeeze after live probing
/// both WAL and DELETE journal modes: SQLite checks the page-count
/// limit at page-allocation time, so the squeeze always fails the
/// growth statement first and leaves no transaction for `COMMIT` to
/// fail on — the arm is unreachable through PRAGMA-space. The arm
/// holds a flag per [`crate::SqliteStore`] (born disarmed), so
/// parallel test stores cannot interfere; `take` disarms on read —
/// exactly one commit faults, and the post-failure commits are real.
#[cfg(test)]
pub(crate) mod commit_fault {
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
pub(crate) type Flag = Arc<AtomicBool>;
pub(crate) fn disarmed() -> Flag {
Arc::new(AtomicBool::new(false))
}
pub(crate) fn arm(flag: &Flag) {
flag.store(true, Ordering::Release);
}
pub(crate) fn take(flag: &Flag) -> bool {
flag.swap(false, Ordering::AcqRel)
}
pub(crate) fn failure() -> rusqlite::Error {
rusqlite::Error::SqliteFailure(
rusqlite::ffi::Error::new(rusqlite::ffi::SQLITE_FULL),
Some("database or disk is full".to_string()),
)
}
}
+23 -1
View File
@@ -50,6 +50,12 @@ pub struct SqliteStore {
watcher: Arc<SharedUpdateWatcher>,
db_path: PathBuf,
closed: Arc<AtomicBool>,
/// The store's commit-fault arm (`sqlite-commit-error-arm`) — born
/// disarmed; a test arms it through [`SqliteStore::arm_commit_fault`]
/// so the next handle commit fails through the real commit error
/// arm. Dead in production builds.
#[cfg(test)]
commit_fault: crate::seam::commit_fault::Flag,
}
impl std::fmt::Debug for SqliteStore {
@@ -71,6 +77,16 @@ impl SqliteStore {
self.readers.close();
self.writer.close();
}
/// Arm the commit-fault seam (`sqlite-commit-error-arm`): the
/// next `commit` through a handle of this store fails with a
/// `SQLITE_FULL`-shaped driver error and replenishes the writer
/// slot through the real commit error arm. One-store scoping keeps
/// parallel test stores unaffected.
#[cfg(test)]
pub(crate) fn arm_commit_fault(&self) {
crate::seam::commit_fault::arm(&self.commit_fault)
}
}
impl Drop for SqliteStore {
@@ -109,6 +125,8 @@ pub(crate) fn open_store(path: &str, opts: SqliteOpts) -> alkstore::Result<Sqlit
watcher: Arc::new(watcher),
db_path: PathBuf::from(path),
closed: Arc::new(AtomicBool::new(false)),
#[cfg(test)]
commit_fault: crate::seam::commit_fault::disarmed(),
})
}
@@ -121,7 +139,11 @@ impl Store for SqliteStore {
let reopen: std::sync::Arc<
dyn Fn() -> alkstore::Result<rusqlite::Connection> + Send + Sync,
> = std::sync::Arc::new(move || crate::seam::open_writer_connection(&db_path));
crate::tx::SqliteTxHandle::begin(writer, reopen)
#[cfg(test)]
let fut = crate::tx::SqliteTxHandle::begin(writer, reopen, self.commit_fault.clone());
#[cfg(not(test))]
let fut = crate::tx::SqliteTxHandle::begin(writer, reopen);
fut
}
fn notify<'a>(
+68
View File
@@ -1199,6 +1199,74 @@ async fn cancelling_a_tx_future_rolls_back() {
cleanup(&dir);
}
/// Acceptance: a failing `COMMIT` surfaces as the opaque `Database`
/// (`sqlite-commit-error-arm`, the wave-3 review's deferred coverage —
/// the replenish fix pinned by code-read there, now test-exercised),
/// does not strand the writer slot — the failed connection is dropped
/// (its state unknowable) and the slot replenished via the handle's
/// `reopen` closure — and the store stays fully usable: the next
/// `begin_tx` proceeds, finds no partial-commit residue (the failed
/// tx's writes rolled back with the dropped connection), commits
/// cleanly, and auto-commit ops work.
///
/// Injection is the engine's `cfg(test)` commit-fault seam
/// (seam.rs): a `SQLITE_FULL`-shaped driver error substituted for one
/// `COMMIT`, feeding the production error arm (drop + `reopen`
/// replenish) rather than fabricating its behavior. The seam over the
/// `PRAGMA max_page_count` squeeze: live probes of both WAL and
/// DELETE journal modes showed SQLite checks the page-count limit at
/// page-allocation time — the squeeze always fails the growth
/// statement (first `INSERT`) and leaves no transaction for `COMMIT`
/// to fail on, so the PRAGMA route cannot reach this arm
/// deterministically. The seam carries no cross-test risk: the flag
/// is per-store and disarms on its first consumption.
#[tokio::test(flavor = "multi_thread")]
async fn failed_commit_replenishes_the_writer_slot() {
let dir = temp_dir("commit-err");
let concrete = open_store(dir.join("store.db").to_str().unwrap(), Default::default()).unwrap();
concrete.arm_commit_fault();
let mut tx = concrete.begin_tx().await.unwrap();
let id = tx
.enqueue_tx("q", EnqueueOpts::default(), serde_json::json!({"n": 1}))
.await
.unwrap();
assert!(id > 0, "the in-tx write landed before the failed commit");
let err = tx.commit().await.unwrap_err();
match err {
Error::Database(source) => {
assert!(
source.to_string().contains("full"),
"the commit error carries the SQLITE_FULL-shaped source chain: {source}"
);
}
other => panic!("a failed commit errs opaque Database, got {other:?}"),
}
let next = tokio::time::timeout(std::time::Duration::from_secs(5), concrete.begin_tx())
.await
.expect("the slot must be grantable after the failed commit (replenished via reopen)")
.unwrap();
let mut next = next;
let residue = next.get_job_tx("q", id).await.unwrap();
assert_eq!(
residue, None,
"the failed commit's writes left no partial-commit residue"
);
next.commit()
.await
.expect("the post-failure commit must be a real, clean commit (fault disarmed)");
concrete
.notify("probe", serde_json::json!({"after": "commit-error"}))
.await
.expect("auto-commit ops work after the failed commit");
drop(concrete);
cleanup(&dir);
}
/// Acceptance: a tokio task cancelled *mid-op* (its future dropped
/// while an op's blocking round trip is in flight) cannot strand the
/// writer slot — the op's RAII connection lease replenishes the slot
+88
View File
@@ -820,6 +820,94 @@ mod watcher_tests {
assert_eq!(idx, 0);
}
/// The reconnect loop's *success* body (the coverage doc's open
/// line 187–202 — `shared_update_watcher`'s W-1 test drives only
/// the failure path, and a vanished-file drive cannot test the
/// success body: after the file reappears the identity check reads
/// a fresh `(dev, ino)` and kills the watcher through the
/// dead-man's switch). Driven here without the vanished-file trick
/// and without touching the engine's watcher shape: the db file
/// exists for the whole run (identity stable), and the
/// `run_poll_loop`'s own `open_conn_fn` seam — the same mechanism
/// the W-1 test instruments — fails the initial open and the first
/// two reconnects, with the third reconnect succeeding. A commit
/// from a separate writer connection after that success must wake
/// `on_change` again, proving the success arm re-baselines
/// `data_version` and restores wake delivery (and that injected
/// failed opens do not poison the re-established connection).
#[test]
fn reconnect_success_resumes_wake_delivery() {
static OPENS: std::sync::atomic::AtomicU64 = std::sync::atomic::AtomicU64::new(0);
const FAIL_FIRST: u64 = 3;
fn instrumented_open(path: &Path) -> rusqlite::Result<Connection> {
let n = OPENS.fetch_add(1, Ordering::Relaxed) + 1;
if n <= FAIL_FIRST {
return Err(rusqlite::Error::InvalidPath(path.to_path_buf()));
}
open_watcher_conn(path)
}
let tmp = temp_db("reconnect-success");
wal_db(&tmp);
let writer = Connection::open(&tmp).unwrap();
writer
.execute("CREATE TABLE IF NOT EXISTS _test_reconnect(x INTEGER)", [])
.unwrap();
let wakes = Arc::new(AtomicU64::new(0));
let wakes_t = wakes.clone();
let stop = Arc::new(AtomicBool::new(false));
let stop_t = stop.clone();
let (ready_tx, ready_rx) = std::sync::mpsc::sync_channel::<()>(1);
let db_path = tmp.clone();
let handle = std::thread::spawn(move || {
run_poll_loop(
db_path,
move || {
wakes_t.fetch_add(1, Ordering::Relaxed);
},
stop_t,
ready_tx,
Duration::from_millis(5),
instrumented_open,
)
});
ready_rx
.recv_timeout(Duration::from_secs(2))
.expect("watcher loop baseline capture");
// The reconnect schedule retries every MAX_RECONNECT_TICKS
// ticks — 500 ms apart at the 5 ms cadence. Let the first two
// reconnect attempts fail (opens 2 and 3) and the third
// reconnect succeed (open 4).
std::thread::sleep(Duration::from_millis(1_800));
let opens = OPENS.load(Ordering::Relaxed);
assert!(
opens > FAIL_FIRST,
"a reconnect open must have succeeded after the injected failures, saw {opens}"
);
writer
.execute("INSERT INTO _test_reconnect(x) VALUES (1)", [])
.unwrap();
let deadline = std::time::Instant::now() + Duration::from_secs(3);
while wakes.load(Ordering::Relaxed) == 0 && std::time::Instant::now() < deadline {
std::thread::sleep(Duration::from_millis(5));
}
assert!(
wakes.load(Ordering::Relaxed) > 0,
"no wake delivered after the successful reconnect — the success \
body must re-baseline data_version and restore delivery"
);
stop.store(true, Ordering::Release);
handle.join().unwrap();
let _ = std::fs::remove_file(&tmp);
let _ = std::fs::remove_file(format!("{}-wal", tmp.display()));
let _ = std::fs::remove_file(format!("{}-shm", tmp.display()));
}
/// W-1: with the watcher connection down and the db file vanished,
/// the reconnect path is throttled — not one open attempt per
/// poll tick. Driven directly: a missing db path fails every open,
+22 -4
View File
@@ -47,6 +47,11 @@ pub struct SqliteTxHandle {
conn: Option<rusqlite::Connection>,
writer: Arc<Writer>,
reopen: Arc<dyn Fn() -> Result<rusqlite::Connection> + Send + Sync>,
/// This store's commit-fault arm (`sqlite-commit-error-arm`) —
/// read once per commit by the test-only fault branch; dead in
/// production builds.
#[cfg(test)]
commit_fault: crate::seam::commit_fault::Flag,
}
impl std::fmt::Debug for SqliteTxHandle {
@@ -61,6 +66,7 @@ impl SqliteTxHandle {
pub(crate) fn begin(
writer: Arc<Writer>,
reopen: Arc<dyn Fn() -> Result<rusqlite::Connection> + Send + Sync>,
#[cfg(test)] commit_fault: crate::seam::commit_fault::Flag,
) -> BoxedFuture<'static, Result<Box<dyn TxHandle + Send>>> {
Box::pin(async move {
let writer2 = writer.clone();
@@ -83,6 +89,8 @@ impl SqliteTxHandle {
conn: Some(conn),
writer,
reopen,
#[cfg(test)]
commit_fault,
}) as _)
})
}
@@ -392,9 +400,19 @@ impl TxHandle for SqliteTxHandle {
};
let writer = self.writer.clone();
let writer_reopen = self.reopen.clone();
#[cfg(test)]
let commit_fault = self.commit_fault.clone();
Box::pin(async move {
blocking(
move || match conn.execute_batch("COMMIT").map_err(sqlite_error) {
blocking(move || {
#[cfg(test)]
let committed = if crate::seam::commit_fault::take(&commit_fault) {
Err(crate::seam::commit_fault::failure())
} else {
conn.execute_batch("COMMIT")
};
#[cfg(not(test))]
let committed = conn.execute_batch("COMMIT");
match committed.map_err(sqlite_error) {
Ok(()) => {
writer.release(conn);
Ok(())
@@ -406,8 +424,8 @@ impl TxHandle for SqliteTxHandle {
}
Err(e)
}
},
)
}
})
.await
})
}