SQLite commit-error-arm coverage (task sqlite-commit-error-arm): the wave-3 review gate's deferred test landed — failed_commit_replenishes_the_writer_slot drives a failing COMMIT through the engine's commit path and pins the review's code-read fix end-to-end: the error surfaces as the opaque Database carrying the SQLITE_FULL-shaped source chain, the failed connection is dropped and the writer slot replenished via the handle's reopen closure (the next begin_tx proceeds within a bounded timeout, no parking — the store-wide-livelock posture), no partial-commit residue (the dropped connection's uncommitted writes read back None), the post-failure commit is a real clean commit (the fault disarms on consumption), and auto-commit notify works afterward. Injection is a cfg(test) commit-fault seam in seam.rs — a per-store Arc<AtomicBool> arm (born disarmed, armed via arm_commit_fault, take() disarms on first consumption so exactly one commit faults) whose fabricated rusqlite SqliteFailure feeds the production commit error arm rather than replicating it; the PRAGMA max_page_count route was probed live against both WAL and DELETE journal modes first and rejected: SQLite checks the page-count limit at page-allocation time, so the squeeze always fails the growth statement (SQLITE_FULL/DiskFull on the first INSERT) and leaves no transaction active for COMMIT to fail — the arm is unreachable through PRAGMA-space (also probed: the pragma is per-connection, so pre-begin arming on the writer conn would have ridden into the tx conn; the route failed on error placement, not delivery). Mechanism choice and probes documented in the seam doc comment and the task Notes. Cross-test safety is per-store scoping; parallel stores never see the arm. Replay-proofed live: with the error arm's writer_reopen replenish temporarily removed the test fails (begin_tx parks past the 5 s timeout — the stranding the review identified), reverted it passes. Plumbing follows the pg-fix-forwarder-reconnect cfg(test) precedent: fields on SqliteStore/SqliteTxHandle and a begin param are cfg-gated, production builds compile the plain path. The waves-1-2 review's optional watcher reconnect-success add rides here (taken — recorded in Notes): reconnect_success_resumes_wake_delivery drives run_poll_loop through its existing open_conn_fn seam (same instrument as the W-1 failure test), with the db file present throughout because the vanished-file route cannot reach the success body (file reappearance trips the dead-man's identity switch first): initial open + first two reconnects fail by injection, the third reconnect succeeds, and a subsequent commit wakes on_change — the success arm's data_version re-baseline and restored delivery pinned. Watcher shape untouched. Verified: cargo test -p alkstore-sqlite green server-less (191 lib + 25 suite), workspace cargo test 399/0, clippy -D warnings, fmt clean
This commit is contained in:
1 parent
36023914b2
commit
0c7744977c
6 files changed
+320
-8
No files matched your search
@@ -160,3 +160,42 @@ pub(crate) fn is_closed_err(e: &rusqlite::Error) -> bool {
|
||||
pub(crate) fn closed_store_error() -> Error {
|
||||
Error::database(std::io::Error::other("the store is closed"))
|
||||
}
|
||||
|
||||
/// The `cfg(test)` commit-fault seam (`sqlite-commit-error-arm`): a
|
||||
/// per-store arm that, when set, substitutes a `SQLITE_FULL`-shaped
|
||||
/// driver error for one `COMMIT`, feeding the real commit error arm
|
||||
/// (drop + `reopen` replenish) instead of fabricating its behavior.
|
||||
/// Chosen over the `PRAGMA max_page_count` squeeze after live probing
|
||||
/// both WAL and DELETE journal modes: SQLite checks the page-count
|
||||
/// limit at page-allocation time, so the squeeze always fails the
|
||||
/// growth statement first and leaves no transaction for `COMMIT` to
|
||||
/// fail on — the arm is unreachable through PRAGMA-space. The arm
|
||||
/// holds a flag per [`crate::SqliteStore`] (born disarmed), so
|
||||
/// parallel test stores cannot interfere; `take` disarms on read —
|
||||
/// exactly one commit faults, and the post-failure commits are real.
|
||||
#[cfg(test)]
|
||||
pub(crate) mod commit_fault {
|
||||
use std::sync::Arc;
|
||||
use std::sync::atomic::{AtomicBool, Ordering};
|
||||
|
||||
pub(crate) type Flag = Arc<AtomicBool>;
|
||||
|
||||
pub(crate) fn disarmed() -> Flag {
|
||||
Arc::new(AtomicBool::new(false))
|
||||
}
|
||||
|
||||
pub(crate) fn arm(flag: &Flag) {
|
||||
flag.store(true, Ordering::Release);
|
||||
}
|
||||
|
||||
pub(crate) fn take(flag: &Flag) -> bool {
|
||||
flag.swap(false, Ordering::AcqRel)
|
||||
}
|
||||
|
||||
pub(crate) fn failure() -> rusqlite::Error {
|
||||
rusqlite::Error::SqliteFailure(
|
||||
rusqlite::ffi::Error::new(rusqlite::ffi::SQLITE_FULL),
|
||||
Some("database or disk is full".to_string()),
|
||||
)
|
||||
}
|
||||
}
|
||||
@@ -50,6 +50,12 @@ pub struct SqliteStore {
|
||||
watcher: Arc<SharedUpdateWatcher>,
|
||||
db_path: PathBuf,
|
||||
closed: Arc<AtomicBool>,
|
||||
/// The store's commit-fault arm (`sqlite-commit-error-arm`) — born
|
||||
/// disarmed; a test arms it through [`SqliteStore::arm_commit_fault`]
|
||||
/// so the next handle commit fails through the real commit error
|
||||
/// arm. Dead in production builds.
|
||||
#[cfg(test)]
|
||||
commit_fault: crate::seam::commit_fault::Flag,
|
||||
}
|
||||
|
||||
impl std::fmt::Debug for SqliteStore {
|
||||
@@ -71,6 +77,16 @@ impl SqliteStore {
|
||||
self.readers.close();
|
||||
self.writer.close();
|
||||
}
|
||||
|
||||
/// Arm the commit-fault seam (`sqlite-commit-error-arm`): the
|
||||
/// next `commit` through a handle of this store fails with a
|
||||
/// `SQLITE_FULL`-shaped driver error and replenishes the writer
|
||||
/// slot through the real commit error arm. One-store scoping keeps
|
||||
/// parallel test stores unaffected.
|
||||
#[cfg(test)]
|
||||
pub(crate) fn arm_commit_fault(&self) {
|
||||
crate::seam::commit_fault::arm(&self.commit_fault)
|
||||
}
|
||||
}
|
||||
|
||||
impl Drop for SqliteStore {
|
||||
@@ -109,6 +125,8 @@ pub(crate) fn open_store(path: &str, opts: SqliteOpts) -> alkstore::Result<Sqlit
|
||||
watcher: Arc::new(watcher),
|
||||
db_path: PathBuf::from(path),
|
||||
closed: Arc::new(AtomicBool::new(false)),
|
||||
#[cfg(test)]
|
||||
commit_fault: crate::seam::commit_fault::disarmed(),
|
||||
})
|
||||
}
|
||||
|
||||
@@ -121,7 +139,11 @@ impl Store for SqliteStore {
|
||||
let reopen: std::sync::Arc<
|
||||
dyn Fn() -> alkstore::Result<rusqlite::Connection> + Send + Sync,
|
||||
> = std::sync::Arc::new(move || crate::seam::open_writer_connection(&db_path));
|
||||
crate::tx::SqliteTxHandle::begin(writer, reopen)
|
||||
#[cfg(test)]
|
||||
let fut = crate::tx::SqliteTxHandle::begin(writer, reopen, self.commit_fault.clone());
|
||||
#[cfg(not(test))]
|
||||
let fut = crate::tx::SqliteTxHandle::begin(writer, reopen);
|
||||
fut
|
||||
}
|
||||
|
||||
fn notify<'a>(
|
||||
|
||||
@@ -1199,6 +1199,74 @@ async fn cancelling_a_tx_future_rolls_back() {
|
||||
cleanup(&dir);
|
||||
}
|
||||
|
||||
/// Acceptance: a failing `COMMIT` surfaces as the opaque `Database`
|
||||
/// (`sqlite-commit-error-arm`, the wave-3 review's deferred coverage —
|
||||
/// the replenish fix pinned by code-read there, now test-exercised),
|
||||
/// does not strand the writer slot — the failed connection is dropped
|
||||
/// (its state unknowable) and the slot replenished via the handle's
|
||||
/// `reopen` closure — and the store stays fully usable: the next
|
||||
/// `begin_tx` proceeds, finds no partial-commit residue (the failed
|
||||
/// tx's writes rolled back with the dropped connection), commits
|
||||
/// cleanly, and auto-commit ops work.
|
||||
///
|
||||
/// Injection is the engine's `cfg(test)` commit-fault seam
|
||||
/// (seam.rs): a `SQLITE_FULL`-shaped driver error substituted for one
|
||||
/// `COMMIT`, feeding the production error arm (drop + `reopen`
|
||||
/// replenish) rather than fabricating its behavior. The seam over the
|
||||
/// `PRAGMA max_page_count` squeeze: live probes of both WAL and
|
||||
/// DELETE journal modes showed SQLite checks the page-count limit at
|
||||
/// page-allocation time — the squeeze always fails the growth
|
||||
/// statement (first `INSERT`) and leaves no transaction for `COMMIT`
|
||||
/// to fail on, so the PRAGMA route cannot reach this arm
|
||||
/// deterministically. The seam carries no cross-test risk: the flag
|
||||
/// is per-store and disarms on its first consumption.
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn failed_commit_replenishes_the_writer_slot() {
|
||||
let dir = temp_dir("commit-err");
|
||||
let concrete = open_store(dir.join("store.db").to_str().unwrap(), Default::default()).unwrap();
|
||||
|
||||
concrete.arm_commit_fault();
|
||||
|
||||
let mut tx = concrete.begin_tx().await.unwrap();
|
||||
let id = tx
|
||||
.enqueue_tx("q", EnqueueOpts::default(), serde_json::json!({"n": 1}))
|
||||
.await
|
||||
.unwrap();
|
||||
assert!(id > 0, "the in-tx write landed before the failed commit");
|
||||
|
||||
let err = tx.commit().await.unwrap_err();
|
||||
match err {
|
||||
Error::Database(source) => {
|
||||
assert!(
|
||||
source.to_string().contains("full"),
|
||||
"the commit error carries the SQLITE_FULL-shaped source chain: {source}"
|
||||
);
|
||||
}
|
||||
other => panic!("a failed commit errs opaque Database, got {other:?}"),
|
||||
}
|
||||
|
||||
let next = tokio::time::timeout(std::time::Duration::from_secs(5), concrete.begin_tx())
|
||||
.await
|
||||
.expect("the slot must be grantable after the failed commit (replenished via reopen)")
|
||||
.unwrap();
|
||||
let mut next = next;
|
||||
let residue = next.get_job_tx("q", id).await.unwrap();
|
||||
assert_eq!(
|
||||
residue, None,
|
||||
"the failed commit's writes left no partial-commit residue"
|
||||
);
|
||||
next.commit()
|
||||
.await
|
||||
.expect("the post-failure commit must be a real, clean commit (fault disarmed)");
|
||||
|
||||
concrete
|
||||
.notify("probe", serde_json::json!({"after": "commit-error"}))
|
||||
.await
|
||||
.expect("auto-commit ops work after the failed commit");
|
||||
drop(concrete);
|
||||
cleanup(&dir);
|
||||
}
|
||||
|
||||
/// Acceptance: a tokio task cancelled *mid-op* (its future dropped
|
||||
/// while an op's blocking round trip is in flight) cannot strand the
|
||||
/// writer slot — the op's RAII connection lease replenishes the slot
|
||||
|
||||
@@ -820,6 +820,94 @@ mod watcher_tests {
|
||||
assert_eq!(idx, 0);
|
||||
}
|
||||
|
||||
/// The reconnect loop's *success* body (the coverage doc's open
|
||||
/// line 187–202 — `shared_update_watcher`'s W-1 test drives only
|
||||
/// the failure path, and a vanished-file drive cannot test the
|
||||
/// success body: after the file reappears the identity check reads
|
||||
/// a fresh `(dev, ino)` and kills the watcher through the
|
||||
/// dead-man's switch). Driven here without the vanished-file trick
|
||||
/// and without touching the engine's watcher shape: the db file
|
||||
/// exists for the whole run (identity stable), and the
|
||||
/// `run_poll_loop`'s own `open_conn_fn` seam — the same mechanism
|
||||
/// the W-1 test instruments — fails the initial open and the first
|
||||
/// two reconnects, with the third reconnect succeeding. A commit
|
||||
/// from a separate writer connection after that success must wake
|
||||
/// `on_change` again, proving the success arm re-baselines
|
||||
/// `data_version` and restores wake delivery (and that injected
|
||||
/// failed opens do not poison the re-established connection).
|
||||
#[test]
|
||||
fn reconnect_success_resumes_wake_delivery() {
|
||||
static OPENS: std::sync::atomic::AtomicU64 = std::sync::atomic::AtomicU64::new(0);
|
||||
const FAIL_FIRST: u64 = 3;
|
||||
fn instrumented_open(path: &Path) -> rusqlite::Result<Connection> {
|
||||
let n = OPENS.fetch_add(1, Ordering::Relaxed) + 1;
|
||||
if n <= FAIL_FIRST {
|
||||
return Err(rusqlite::Error::InvalidPath(path.to_path_buf()));
|
||||
}
|
||||
open_watcher_conn(path)
|
||||
}
|
||||
|
||||
let tmp = temp_db("reconnect-success");
|
||||
wal_db(&tmp);
|
||||
let writer = Connection::open(&tmp).unwrap();
|
||||
writer
|
||||
.execute("CREATE TABLE IF NOT EXISTS _test_reconnect(x INTEGER)", [])
|
||||
.unwrap();
|
||||
|
||||
let wakes = Arc::new(AtomicU64::new(0));
|
||||
let wakes_t = wakes.clone();
|
||||
let stop = Arc::new(AtomicBool::new(false));
|
||||
let stop_t = stop.clone();
|
||||
let (ready_tx, ready_rx) = std::sync::mpsc::sync_channel::<()>(1);
|
||||
let db_path = tmp.clone();
|
||||
let handle = std::thread::spawn(move || {
|
||||
run_poll_loop(
|
||||
db_path,
|
||||
move || {
|
||||
wakes_t.fetch_add(1, Ordering::Relaxed);
|
||||
},
|
||||
stop_t,
|
||||
ready_tx,
|
||||
Duration::from_millis(5),
|
||||
instrumented_open,
|
||||
)
|
||||
});
|
||||
ready_rx
|
||||
.recv_timeout(Duration::from_secs(2))
|
||||
.expect("watcher loop baseline capture");
|
||||
|
||||
// The reconnect schedule retries every MAX_RECONNECT_TICKS
|
||||
// ticks — 500 ms apart at the 5 ms cadence. Let the first two
|
||||
// reconnect attempts fail (opens 2 and 3) and the third
|
||||
// reconnect succeed (open 4).
|
||||
std::thread::sleep(Duration::from_millis(1_800));
|
||||
let opens = OPENS.load(Ordering::Relaxed);
|
||||
assert!(
|
||||
opens > FAIL_FIRST,
|
||||
"a reconnect open must have succeeded after the injected failures, saw {opens}"
|
||||
);
|
||||
|
||||
writer
|
||||
.execute("INSERT INTO _test_reconnect(x) VALUES (1)", [])
|
||||
.unwrap();
|
||||
|
||||
let deadline = std::time::Instant::now() + Duration::from_secs(3);
|
||||
while wakes.load(Ordering::Relaxed) == 0 && std::time::Instant::now() < deadline {
|
||||
std::thread::sleep(Duration::from_millis(5));
|
||||
}
|
||||
assert!(
|
||||
wakes.load(Ordering::Relaxed) > 0,
|
||||
"no wake delivered after the successful reconnect — the success \
|
||||
body must re-baseline data_version and restore delivery"
|
||||
);
|
||||
|
||||
stop.store(true, Ordering::Release);
|
||||
handle.join().unwrap();
|
||||
let _ = std::fs::remove_file(&tmp);
|
||||
let _ = std::fs::remove_file(format!("{}-wal", tmp.display()));
|
||||
let _ = std::fs::remove_file(format!("{}-shm", tmp.display()));
|
||||
}
|
||||
|
||||
/// W-1: with the watcher connection down and the db file vanished,
|
||||
/// the reconnect path is throttled — not one open attempt per
|
||||
/// poll tick. Driven directly: a missing db path fails every open,
|
||||
|
||||
@@ -47,6 +47,11 @@ pub struct SqliteTxHandle {
|
||||
conn: Option<rusqlite::Connection>,
|
||||
writer: Arc<Writer>,
|
||||
reopen: Arc<dyn Fn() -> Result<rusqlite::Connection> + Send + Sync>,
|
||||
/// This store's commit-fault arm (`sqlite-commit-error-arm`) —
|
||||
/// read once per commit by the test-only fault branch; dead in
|
||||
/// production builds.
|
||||
#[cfg(test)]
|
||||
commit_fault: crate::seam::commit_fault::Flag,
|
||||
}
|
||||
|
||||
impl std::fmt::Debug for SqliteTxHandle {
|
||||
@@ -61,6 +66,7 @@ impl SqliteTxHandle {
|
||||
pub(crate) fn begin(
|
||||
writer: Arc<Writer>,
|
||||
reopen: Arc<dyn Fn() -> Result<rusqlite::Connection> + Send + Sync>,
|
||||
#[cfg(test)] commit_fault: crate::seam::commit_fault::Flag,
|
||||
) -> BoxedFuture<'static, Result<Box<dyn TxHandle + Send>>> {
|
||||
Box::pin(async move {
|
||||
let writer2 = writer.clone();
|
||||
@@ -83,6 +89,8 @@ impl SqliteTxHandle {
|
||||
conn: Some(conn),
|
||||
writer,
|
||||
reopen,
|
||||
#[cfg(test)]
|
||||
commit_fault,
|
||||
}) as _)
|
||||
})
|
||||
}
|
||||
@@ -392,9 +400,19 @@ impl TxHandle for SqliteTxHandle {
|
||||
};
|
||||
let writer = self.writer.clone();
|
||||
let writer_reopen = self.reopen.clone();
|
||||
#[cfg(test)]
|
||||
let commit_fault = self.commit_fault.clone();
|
||||
Box::pin(async move {
|
||||
blocking(
|
||||
move || match conn.execute_batch("COMMIT").map_err(sqlite_error) {
|
||||
blocking(move || {
|
||||
#[cfg(test)]
|
||||
let committed = if crate::seam::commit_fault::take(&commit_fault) {
|
||||
Err(crate::seam::commit_fault::failure())
|
||||
} else {
|
||||
conn.execute_batch("COMMIT")
|
||||
};
|
||||
#[cfg(not(test))]
|
||||
let committed = conn.execute_batch("COMMIT");
|
||||
match committed.map_err(sqlite_error) {
|
||||
Ok(()) => {
|
||||
writer.release(conn);
|
||||
Ok(())
|
||||
@@ -406,8 +424,8 @@ impl TxHandle for SqliteTxHandle {
|
||||
}
|
||||
Err(e)
|
||||
}
|
||||
},
|
||||
)
|
||||
}
|
||||
})
|
||||
.await
|
||||
})
|
||||
}
|
||||
|
||||
Reference in new issue
Block a user