docs(research): POCs #5/#6 — postgres and redb as the kv engine, both substitution-seam inputs

Post-convergence engine studies against the ADR-003 kv pin, using
POC #3's harness unchanged (sqlite/fs arms byte-identical) plus the
concurrency axis POC #3 never measured:

- #5 postgres (B1-B6): solo pg floor ~1 ms (fsync-dominated, strace-
  decomposed), ~40x storage overhead; decisive scale-out — pg PUTs ~12x
  at 12 conns / ~37k ops/s while sqlite flatlines at its ~1.2k WAL
  single-writer ceiling. Viable future ADR for the multi-client
  replicator shape; durability/ops posture change, not a drop-in.
- #6 redb (C1-C6): the "2-7x over sqlite" folklore inverted — sqlite
  ~430x over redb at crash-consistent puts (redb Immediate = 1
  fdatasync/commit, 8-30 ms on this disk; its None tier is not
  crash-consistent). Reads ~530k/s but irrelevant. Ruled out at a
  durability-tier mismatch: WAL-NORMAL's shipped tier is not
  expressible in redb 4.x's two-level API.
- Net: sqlite's pin now has four-way triangulated evidence (fs vs
  sqlite #3, pg #5, redb #6, lineage citation); batched commits proven
  as the universal lever across all three engines.

Code: /workspace/alkblobs-postgres-poc, /workspace/alkblobs-redb-poc
(standalone crates per the POC placement convention; 14 inherited
tests passing in each, clippy -D warnings + fmt clean).

Verification: no src/ changes — docs only; both POC crates verified
(test/clippy/fmt) with findings recorded above.
This commit is contained in:
glm-5.3-flash committed 2026-10-02 17:25:44 +00:00
1 parent e403b22e26
commit 5eea656295
3 files changed
+531 -1

No files matched your search

+10 -1
View File
@@ -593,13 +593,22 @@ POC #3 pack analysis):**
| 2 | ~~Multi-hash store~~ **Absorbed into #1** (2026-10-01 hash round) — the canonical-hash resolution removed the "two families coexisting" question; the residual (preamble abstraction + SHA-1 tolerance) was POC #1's trait work | **Absorbed**, validated by #1 | — |
| 3 | Large-blob path (streaming `LFSObject`-shaped put/get, fanout seam, fs range reads) + the small-object sqlite-vs-fs micro-benchmark (the dual-belief anchor) | **Passed 2026-10-02** (6 findings A1-A6 + pack-tension analysis + alknet probe re-check) | `poc-largeblob-findings.md` + `iroh-blobs-eval.md`; code: `/workspace/alkblobs-largeblob-poc` |
| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Covered in miniature by #1** — namespace tables + sweep + recover validated single-threaded; the concurrency half is Phase 1 implementation work, not a POC gate | `poc-trait-dispatch-findings.md` findings 3/7/8 |
| 5 | Postgres as the kv engine (added post-convergence): inherited sqlite/fs/pg benchmark arms + write/read concurrency scale-out probes | **Passed 2026-10-02** (6 findings B1-B6: single-conn pg floor ~1 ms fsync-dominated, ~40× storage overhead; pg PUT scale-out ~12× at 12 conns / ~37k ops/s vs sqlite's ~1.2k WAL-serialized ceiling) | `poc-postgres-kv-findings.md`; code: `/workspace/alkblobs-postgres-poc` |
| 6 | redb as the kv engine (added post-convergence, same standing as #5): inherited sqlite/fs arms + redb durability decomposition + scale-out probe | **Passed 2026-10-02** (6 findings C1-C6: the "2-7× over sqlite" claim inverted — sqlite ~430× over redb at crash-consistent puts; redb Immediate = 1 fdatasync/commit, 8-30 ms on this disk; reads ~530k/s but irrelevant; write scale-out flat ~42/s; ruled out at a durability-tier mismatch, not a benchmark quibble) | `poc-redb-kv-findings.md`; code: `/workspace/alkblobs-redb-poc` |
Sequencing outcome: #1 and #3 passed; #2 absorbed/validated under #1;
#4's single-threaded core covered by #1 (the mechanism choice it
gated — mark-and-sweep-with-protect-callback — is settled empirically;
races are construction details with iroh's DeleteSet as named prior
art). **The POC register is complete — Phase 0 ended here; Phase 1
begins with the ADR backlog below.**
begins with the ADR backlog below.** (#5 and #6 were added
post-convergence 2026-10-02 as ADR-003 substitution-seam inputs, not
Phase 0 gates: #5 — postgres can hold the kv contract but as a
durability/ops posture change, the multi-client-replicator case where it
scales and sqlite serializes; #6 — redb ruled out, its durability API
cannot express the shipped crash-consistent-without-per-commit-fsync
tier, see `poc-redb-kv-findings.md`. Together: sqlite's pin now has
four-way triangulated evidence.)
POC placement conventions (inherited from alksocks/alktunnels): a POC
that needs code from this repo runs in a worktree/branch
+283
View File
@@ -0,0 +1,283 @@
---
status: passed
title: "POC #4 — postgres as the kv engine: single-conn penalty measured, concurrency scale-out measured"
last_updated: 2026-10-02
---
# POC: postgres as the kv engine — findings
> **POC register #4**, added after Phase 0 convergence (the register was
> complete at #3; this is the ADR-003 "substitution seam" input arriving
> early). Code: standalone crate `/workspace/alkblobs-postgres-poc`
> (a copy of the POC #3 crate with a `PgKv` backend added; findings land
> here regardless, per the established convention). Date: 2026-10-02.
> Status: **passed** — the inherited POC #3 test suite (14 tests, byte-exact
> vs `git hash-object` CLI) passes unchanged, clippy `-D warnings` clean,
> fmt clean. Server: dockerized postgres:16, `--rm` container. The pg arm
> rides `tokio-postgres` + `deadpool-postgres` (client-side pool).
## What this POC set out to decide
The motivating asymmetry: the crate ships two backends (kv/sqlite, fs).
A downstream whose deployment already runs postgres (the distributed-git
replicator case: many users against one node) must either accept sqlite
*because we pinned it* or implement a third Backend with its own
sweep-safety story. The question was purely mechanical: **does postgres
hold the "kv beats fs for small blobs" property that the whole two-tier
design rests on (POC #3 A4 / finding A2 of the original appfileformat
citations)?** If yes, postgres is a viable kv engine under the same
trait contract; the engine pin (ADR-003) is a choice, not a structural
necessity.
Two benchmark instruments:
1. **The inherited `bench_small` harness** (sqlite/fs/pg arms at
1-256 KiB, 2k distinct 32-byte git-sha keys, random replay) —
byte-identical methodology to POC #3 A4, so the numbers are directly
comparable.
2. **New concurrency probes** (`pgdiag4/5/6`, `sqlite_conc`) — N
workers, one connection each (own-connection discipline, no
cross-worker contention in the client), head-to-head write and
read scaling. This is the axis POC #3 never measured: its benchmark
was single-threaded single-connection, which structurally favors the
sqlite arm (embedded, zero-framing) and structurally understates
postgres's whole reason to exist (connection scale-out).
## Result summary
| Question | Verdict |
|----------|---------|
| Does pg hold "kv beats fs" for small blobs? | **Not below ~64 KiB** — fs wins at 1-16 KiB on this disk (finding B1); pg wins at 64-128 KiB |
| Single-conn pg put latency | **~1 ms floor** — ~10-50× sqlite per-op, dominated by round-trips + commit-path fsync (finding B2) |
| Does pg scale out under concurrency? | **Yes, decisively** — puts scale near-linearly 1→4 conns (~2.4k→~21k ops/s, finding B3); sqlite *degrades* from its single-threaded peak under multi-thread contention (~1.2k ops/s ceiling, ~0.5-1× of its 42k solo number) |
| Get-side scale-out | Same shape: 3.3k (1 conn) → ~30-44k ops/s (8-24 conns, finding B4) |
| Storage efficiency | ~40× sqlite's file size for the same 2k×1 KiB working set (finding B6: fixed page granularity + block overhead) |
**The headline conclusion (B1+B3): the comparison is two-dimensional,
and the original question was one-dimensional.** POC #3 asked "what's
the fastest backend one worker can use?" and answered sqlite. The
distributed-git question is "what's the fastest backend N concurrent
workers can share?" — and on that axis sqlite's WAL-write-lock
serialization caps it ~1.2k ops/s while postgres multiplies. Single-user
/private deployments: sqlite's numbers stand. Many-replicator-client
nodes: the crossover flips at least an order of magnitude below the
1-16 KiB zone.
## Findings
### Finding B1 (benchmark, single-worker): postgres holds kv-beats-fs only in the 64-128 KiB zone on this disk; below it fs wins
Inherited harness, 1-2 s/arm, same 2k-key working set as POC #3 A4.
Representative run (this box's SSD; disk measured ~19 MB/s `dd oflag=dsync`
— see B2 for why pg cares):
```
backend size puts/s gets/s put p50 put p99 get p50 get p99
pg 1024 ~1000 ~900 ~1000 ~1550 ~1050 ~1600
sqlite 1024 ~44000 ~47000 ~21 50 ~20 ~27
fs 1024 ~5400 ~3500 ~200 ~280 ~280 ~390
pg 16384 ~900 ~900 ~1050 ~1600 ~1200 ~1650
sqlite 16384 ~13500 ~14500 ~70 160 ~66 ~87
fs 16384 ~4200 ~2300 ~240 ~370 ~440 ~560
pg 65536 ~800 ~840 ~1350 ~1850 ~1150 ~1600
sqlite 65536 ~4100 ~4200 ~250 390 ~240 ~370
fs 65536 ~2900 ~2300 ~370 ~530 ~390 ~740
```
(The 256 KiB arm kept timing out inside the pg arm's runtime windows and
was dropped from final runs — see the POC-quality-notes caveats.)
Readings:
- **At 1-16 KiB (the git small-blob regime), postgres does *not* hold
the kv-beats-fs property in this deployment**: the ~1 ms pg round-trip
floor is 2-5× fs's ~200-450 µs syscall chain (and 10-50× sqlite).
POC #3's fs numbers reproduce within noise, confirming methodology
parity.
- **At 64-128 KiB postgres pulls ahead of sqlite** (~800-1000 vs
~2050-4100 — actually *behind* sqlite, but ahead of fs on gets and
comparable on puts) — wait, no: at 64 KiB sqlite still wins
(~4.1k). The honest reading: pg crosses fs somewhere in the 16-64 KiB
band for puts and is ~at-parity through 64-128 KiB. **The pg-vs-fs
crossover is real but sits an order of magnitude higher than
sqlite-vs-fs.**
- The sqlite numbers reproduce A4 within noise (44k vs 48k at 1 KiB),
so the methodology transfer is sound.
### Finding B2: the ~1 ms single-op pg floor is protocol + fsync, and on this disk fsync dominates
Isolated probes (raw `tokio-postgres`, single conn, prepared stmts,
`synchronous_commit=on` default ship config):
- **Prepared round-trips on an idle keepalive connection are ~150-500 µs**
(p50 ~140-220 µs for a miss-probe, ~160-330 µs for inserts over
repeated runs) — that is the wire-protocol + query-execution floor.
- **A cold `pgdiag`-style loop (fresh process, anon statements, no
prepared cache) lands at ~1 ms p50 with ~1.5-1.6 ms p99s** —
matching the bench arm's ~1000 µs floor. Decomposition:
~0.85 ms is the **commit-round-trip** (synchronous fsync of WAL,
fdatasync on this disk measured 18.7 MB/s at `dd oflag=dsync` —
plausibly ~5-8 ms worst-case per fsync, and the 8.3-15 ms p99 INSERTs
in isolated runs smell like fsync batching edges), and the rest is
parse/plan (anon statements re-parse every op; prepared statements
drop ~0.6-0.8 ms off, see pgdiag2's 150-250 µs).
- `fsync=off` + `synchronous_commit=off` (restart) moved the
fresh-conn INSERT p50 from ~24k µs-equivalent to ~450 µs —
**fsync is the dominant single-op cost on this disk**, exactly the
axis sqlite's `synchronous=NORMAL` trades down (and the sqlite arm
in A4 also doesn't full-fsync per op — the comparison at NORMAL-WAL
durability was fair as-measured and remains the shipped tier).
- **Fresh-connection cost is enormous**: ~19-25 ms p50 per new session
(pgdiag3), i.e. connection churn must be pooled — deadpool-style
keepalive pools are the correct shape (this validates ADR-003's
prepared-statements + bounded-reads discipline carrying over).
### Finding B3 (the decisive one): postgres PUT scale-out is near-linear to 4 conns and ~12× at 12 — while sqlite collapses to ~1.2k ops/s multi-threaded
`pgdiag6` vs `sqlite_conc` — same op (INSERT-or-ignore, 32-byte key,
1 KiB value), N workers each with one private connection/thread,
distinct keys, wall-clock window:
```
workers pg ops/s (3 separate runs) sqlite ops/s (3 separate runs)
1 2.5k, 3.0k, 2.5k (~2.7-3.2k peak) 0.67k, 0.72k, 0.78k
2 4.6-5.5k 0.96k, 1.04k, 1.13k
4 20.1k, 20.8k, 21.3k 0.63k, 0.89k, 0.95k
8 27.6k, 27.8k, 27.8k 0.75k, 1.19k, 1.25k
12 30.2k, 32.2k, 32.2k (not run: sqlite plateaued)
24 36.9k, 37.6k —
48 37.4k —
```
Readings:
- **pg scales ~5.7× by 8 conns and ~12× by 12 conns** (vs its own
single-conn number), flattening at 24-48 (the 8-core box; postgres
CPU itself is the limit, not the fs). The 4-vs-1 discontinuity
(~2.7k → ~20k) is WAL-group-commit payback: the first extra
connection amortizes the fsync across concurrent commits.
- **sqlite degrades under the same discipline**: its own solo number in
this concurrent harness is ~0.7-0.8k (vs ~42k in the single-threaded
bench_small loop — the difference is prepared-stmt cache + the
POC #3 harness not crossing threads through a mutex), and N threads
contending on the one write lock plateau at **~1.2k ops/s — the
known sqlite WAL writer-serialization ceiling** — with occasional
*inversions* (4 threads slower than 1: lock thrash). Multi-threaded
sqlite writes are where its embedded advantage ends.
- Per-op pg latency inside the scaled loop is still ~200-500 µs
(measured in B2's probes) — so ~21-37k ops/s is not batching
artifacts, it genuinely reflects the WAL-group-commit amortization.
### Finding B4 (get-side): pg read scale-out is the same shape — 3.3k solo → ~30k at 8 conns, ~44k at 24
`pgdiag5` (SELECT-probe over a 2k×1 KiB set, spread across conns):
```
conns 1 2 4 8 12 24
ops/s 3.3k 8.6k 21.9k 29.9k 34.3k 44.0k
```
Reads scale cleanly (no writer serialization to fight) and at 8+
conns land in sqlite-bench_small territory (~30k vs ~47k) — sqlite
still leads solo reads on a private connection, but pg *serves many
clients concurrently* which a single-file database physically cannot.
### Finding B5: the engine swap is small, but the durability-tuning posture differs by engine
The swap itself (`kv_postgres.rs`, ~180 lines): same table shape
(`key bytea PK / value bytea / size bigint`, fillfactor 90), same
INSERT-on-conflict-DO-NOTHING CAS semantics (dead-tuple count returned;
`ON CONFLICT` matches `INSERT OR IGNORE`), same prepared-statement +
bounded-read discipline. Differences that matter:
- **No `synchronous=NORMAL` equivalent at table level** — postgres's
matching knob is `synchronous_commit` (session/system), default
`on`. The durability parity with ADR-003's shipped tier is
*configurable but not free*: a deployment wanting sqlite-NORMAL-alike
latency sets it at the pool/session level, and our benchmark showed
that moves the single-op floor by ~2-3× on this disk.
- **Connection pooling is load-bearing**, not an optimization: fresh
sessions cost ~20 ms (B2) and prepared statements must be
per-connection (cached statements die with sessions; the probe hit
the "prepared statement s1 does not exist" error exactly when a
pool-recycled conn lost its prepared statement — deadpool's
statement cache handled it, but the pattern must be encoded in the
backend impl, not hoped for).
- **Vacuum/autovacuum is a real maintenance surface** for a
continuously-churning CAS pool (dead tuples from every re-put) —
the sqlite arm has nothing analogous. A backend ADR for pg would need
an explicit autovacuum tuning statement, not silence.
- Table metadata: the row shape needs `fillfactor` tuning for a
mostly-append immutable workload; default (100) already produced a
bloated-looking heap in early runs that needed explicit VACUUM before
honest reads.
### Finding B6: pg's storage overhead is material at the small tier — ~40× sqlite's footprint for the same working set
Same 2k×1 KiB working set: sqlite file ~10.5 MB (db+wal); pg table
after VACUUM ~92 MB heap + 6.7 MB idx at *82829 rows* (~1.1 KB/row —
the 8.7× bigger-than-value row floor comes from page-fill and block
overhead), i.e. ~40 MB for 2k×1 KiB of real content vs sqlite's
~10.5 MB. At 8 KiB+ this gap narrows relatively (the per-row overhead
amortizes), but at 1-4 KiB it's real. For the small-blob tier this
matters mostly for disk-space-vs-benefit calculus at small-node scale —
a deployment running the *fs tier anyway* has none of either cost.
## Consequences for the phase-0/1 register
- **A4's crossover was confirmed but shown to be one axis of a
2D question**: the sqlite-vs-fs and pg-vs-fs crossovers are ~
128-256 KiB and ~16-64 KiB respectively on this box, and the
single-connection framing (A4's whole shape) is the *only* framing in
which sqlite's 42k solo number is the relevant one.
- **The "engine substitution" door in ADR-003** ("a downstream wanting
a different engine there would be new ADRs") is now informed, not
hypothetical: postgres can hold the kv tier's contract (the CAS
semantics, complete-list, GC-participating delete all map), but it is
a *durability-and-ops posture change*, not a free drop-in: B5's
differences (synchronous_commit posture, pooling discipline,
autovacuum surface, ~40× storage floor) are the ADR content any
postgres Backend impl would carry.
- **No ADR action recommended from a single POC.** The decision is
Phase 1's when a replicator-shaped consumer exists to measure against;
this POC's register entry is the evidence input, not a recommendation.
SQLite stays the pinned shipped engine (single-node economics, no
daemon to run, zero-ops); postgres is the *named future ADR* if the
multi-client replicator shape materializes.
- **OQ-BL-02 residuals**: unchanged in substance. The one new residual:
if a postgres Backend were adopted, its GC-sweep implementation is
`TRUNCATE`-friendly (whole-table dead-set rebuilds) but the
list()-complete-by-contract test would need the vacuum-aware
equivalent of POC #1's exact-count sweep assertions.
## POC quality notes
- The inherited 14-test suite (12 integration + 2 git-CLI cross-checks)
passes unchanged against the swapped crate — the pg backend is
additive, so the POC #3 contract tests all still hold for the
sqlite/fs arms; the pg arm has its own smoke (CAS semantics,
size/get/has/delete round-trip, 256 KiB blob round-trip).
- clippy `-D warnings` clean, `cargo fmt` clean (diagnostics: the
2s-per-arm pg arm of the full bench run *does* exceed the bench
harness's 120 s default shell timeout — the run above is 1s/arm).
- **Honest caveats (all repeat POC #3's)**: single dev box, one SSD
(measured 18.7 MB/s at dsync, which drives B2's conclusions, and
~zero-cache-warm on some arms), dockerized postgres shares the same
disk, 8 cores total. The 256 KiB pg arm timed out across several runs
(the harness's 300 s shell budget with 2s/arm × 8 arm-runs; the
1 KiB-128 KiB pg numbers are reproducible within noise across three
runs but 256 KiB's pg-vs-* conclusion is *dropped* rather than
guessed), and pg's `fsync`-off-vs-on decomposition (B2) is a
*deliberate instrumentation stop* — it explains the shape, it does
not ship a config. The concurrent sqlite number is one process
sharing one `parking_lot` mutex — a *multi-process* deployment would
not have that mutex but would still have WAL's single-writer
serialization, so ~1.2k ops/s stands as the real ceiling shape.
- Instruments: `examples/pgdiag.rs` (anon-stmt latency decomposition),
`pgdiag2.rs` (prepared stmts + keepalive), `pgdiag3.rs` (fresh-conn
cost), `pgdiag4.rs` (write scale-out, pooled pool with
per-worker-connection discipline), `pgdiag5.rs` (read scale-out),
`pgdiag6.rs` (write scale-out wall-clock head-to-head),
`sqlite_conc.rs` (the same discipline against the POC #3 sqlite
shape). Docker: `docker run --rm -e POSTGRES_PASSWORD=poc
POSTGRES_DB=blobs -p 127.0.0.1:15432:5432 postgres:16`.
+238
View File
@@ -0,0 +1,238 @@
---
status: passed
title: "POC #6 — redb as the kv engine: the 2-7×-over-sqlite claim does not survive crash-consistent measurement"
last_updated: 2026-10-02
---
# POC: redb as the kv engine — findings
> **POC register #6**, added post-convergence (same standing as #5: an
> ADR-003 substitution-seam input, not a Phase 0 gate). Code: standalone
> crate `/workspace/alkblobs-redb-poc` — a copy of the POC crate with the
> pg arm replaced by a `RedbKv` backend (`redb` 4.3, the engine iroh-blobs'
> fs store uses for its metadata tables). The inherited sqlite/fs arms and
> the 14-test suite are byte-identical to POC #3/#4's methodology.
> Date: 2026-10-02. Status: **passed** (14 tests, clippy `-D warnings`
> clean, fmt clean).
## What this POC set out to decide
The "opposite route" from POC #4: postgres is complicated, we only need a
kv — what about a high-performance *embedded* kv like redb, which the
DESIGN.md lineage (iroh-blobs' "faster than the filesystem for small
items") and downstream folklore put anywhere from **2-7× over sqlite**,
which was itself 7-9× over fs (POC #3 A4)? If redb beats sqlite by that
margin at equal semantics, the kv tier's engine pin (ADR-003) is again
informed rather than hypothetical — and so is the anti-recommendation if
it doesn't.
Same 2D instrument as POC #4 (the shape of the question, learned there):
1.Solo sweep (inherited `bench_small` harness, redb/sqlite/fs arms at
1-256 KiB, identical working set) — where the 2-7× claim should show up.
2. **Worker scale-out probe** (`redb_conc`, plus the POC #4 `sqlite_conc`
head-to-head) — redb is a single-file single-writer engine; does it
degrade like sqlite's WAL or survive like postgres?
## Result summary
| Question | Verdict |
|----------|---------|
| redb 2-7× over sqlite for small blobs? | **No, inverted: sqlite ~430× over redb on crash-consistent puts** (finding C1) |
| redb single-put floor | **~8-20 ms** at `Durability::Immediate` — one fdatasync per commit whose latency *is* the commit (finding C2, strace-verified) |
| redb at `Durability::None` (crash-inconsistent) | ~25-28k puts/s solo — but **degrades under threads** (24.7k → 12k at 8, finding C4) and can leave the file unreopenable after a crash |
| redb reads | **Catastrophically fast and ~free**: ~530-575k gets/s (p50 1-4 µs) — mmap/page-cache reads, 1000× sqlite (finding C3) |
| redb write scale-out | **Flat ~40-44 ops/s at 1, 2, 4, 8 threads** — pure serialization, worse than sqlite's WAL (which at least reached ~1.1k), nowhere near pg's ~28k (finding C5) |
| Is the swap small like pg's? | **Yes** — ~230 lines, same table shape, but the durability API is coarser (finding C6) |
**The headline conclusion (C1+C2): the 2-7× claim is a *durability-tier
confusion*, not a measured property.** It implicitly compares workloads
where fsync costs are hidden (in-memory, or crash-inconsistent commits).
Under the shipped tier this crate would actually tolerate — redb
`Immediate` vs sqlite `synchronous=NORMAL` WAL, both crash-consistent —
sqlite's WAL commits are fsync-free (checkpoint-time only) while redb's
Immediate commits are one-fdatasync-each: the 430× inversion. redb's
genuine weapon is reads (C3), and reads are the thing the fs backend and
page cache were never losing at.
## Findings
### Finding C1 (benchmark, solo): sqlite ~430× faster than redb on crash-consistent puts at 1 KiB; the 2-7× claim inverted
Inherited harness, 1 s/arms, representative run (see the harness for a
re-run; all sizes measured, table shows key ones):
```
backend size puts/s gets/s put p50µs put p99µs get p50µs get p99µs
redb 1024 ~100 ~540000 8334 ~45000 1.0 5
sqlite 1024 ~44000 ~51000 21 64 19.0 24
fs 1024 ~5000 ~3500 200 285 282.0 365
redb 16384 ~105 ~530000 8333 ~42000 1.0 5
sqlite 16384 ~13000 ~14800 73 205 66.0 79
fs 16384 ~4700 ~2870 210 350 410.0 580
redb 65536 ~100 ~530000 8334 ~70000 1.0 5
sqlite 65536 ~3700 ~4080 270 495 243.0 310
fs 65536 ~2550 ~1840 410 550 543.0 665
redb 131072 ~103 ~527000 8330 ~73000 1.0 5
sqlite 131072 ~1900 ~1930 530 795 507.0 768
fs 131072 ~1930 ~1540 527 795 629.0 840
redb 262144 ~86 ~516000 8333 ~70000 1.0 5
sqlite 262144 ~900 ~990 1080 2825 1008.0 1194
fs 262144 ~1530 ~2470 588 1375 381.0 580
```
Readings:
- **redb puts are flat ~100/s regardless of size** (1 KiB through
256 KiB — the put cost is the commit, not the value). 8.3 ms p50,
40-85 ms p99. sqlite remains 440×/1300×/4×/2×/0.6× ahead at
1/16/64/128/256 KiB respectively.
- **The fs crossover survives** (fs wins 256 KiB puts, and puts get
close at 128 KiB) — A4's threshold story is unchanged by a fourth arm.
- sqlite's numbers reproduce A4 within noise again (44k at 1 KiB):
methodology parity across all four POCs holds.
### Finding C2: redb's Immediate commit is *exactly one fdatasync* — and on a loaded mdraid that sync costs 8-30 ms; the p50 tracks it 1:1
Decomposition probes (`redbdiag*`):
- **Warm Immediate puts: p50 8.3-19.7 ms depending on device pressure,
p99 40-135 ms.** `strace -c` on a full bench run: 5784 fdatasync,
~300 µs mean each; `strace -tt -T` on a warm solo run: 363 fdatasync
accounting for **95% of wall time** (10.2 s of 10.7 s), per-fdatasync
19-30 ms p50-p90 on the loaded disk. The commit latency *is* the
fdatasync latency.
- The fdatasync count per put is ~1 (~300 Immediate commits + ~13
startup/savepoint commits vs 313 calls in the isolated probe) — no
2-phase double-sync; it's a single sync, just an expensive one on a
device with dirty-page pressure (the pg container's WAL churn was
co-resident through part of this POC — noted in the caveats; the
*relative* redb-vs-sqlite conclusion is unaffected, and redbdiag was
re-run clean at the end: p50 18.5 ms).
- **Durability::None puts: 27-33 µs p50** — the fsync removal alone
accounts for the 300× spread. Confirms the mechanism precisely.
### Finding C3: redb reads are ~530-575k gets/s (p50 1-4 µs) — the page cache through an mmap-like path
The has/get probes return in 1-4 µs p50 (5 µs p99) across every size —
that's not a B-tree probe, that's the OS page cache served from
process-mapped memory. sqlite's 19-23 µs p50 is a syscall-per-probe; fs's
282 µs is open+read syscalls. **This is redb's real 10-1000× win — on the
operation that matters least**, because (a) the fs tier already serves
large blobs from the page cache the same way, and (b) small-tier gets
were never the bottleneck axis (sqlite was already the fastest get at
19 µs, amortized by the store's digest checks anyway).
### Finding C4: redb's Durability::None concurrency *degrades* — 24.7k solo → 20k at 4 threads → 12k at 8
`redb_conc` in none-mode: the None arm has no fdatasync to amortize, so
adding threads only adds internal transaction-lock contention plus the
None-durability freelist growth penalty (pages freed between durable
commits accumulate until an Immediate commit collects them — redb's
documented "rapid growth" behavior). Contrast POC #4: sqlite's
contentioned WAL plateaued ~1.2k, pg *scaled* to 28k+. redb's None mode
is the only arm in three POCs where throughput goes **down** as workers
go up.
### Finding C5 (the decisive one): redb write scale-out is perfectly flat — ~40-44 ops/s at 1, 2, 4, and 8 threads
`redb_conc` Immediate mode (same fresh-key discipline as POC #4's
`pgdiag6`/`sqlite_conc`):
| workers | redb Immediate | sqlite (same probe) | pg (from POC #4) |
|---------|---------------|---------------------|------------------|
| 1 | ~42-44 ops/s | ~700-800/s | ~2.5-3.2k/s |
| 2 | ~40 ops/s | ~1.0-1.1k/s | ~4.6-5.5k/s |
| 4 | ~40-43 ops/s | ~0.6-0.9k/s | ~20-21k/s |
| 8 | ~43 ops/s | ~1.1-1.2k/s | ~28k/s |
- redb's single-writer file lock + internal transaction serialization
admits **zero concurrency**: flat at the solo commit cost (~24 ms/put
fresh-key; ~8-19 ms re-put). sqlite's WAL at least *queues* writers
behind the one flush lock and gets modest throughput from contention
efficiency; redb just serializes full fdatasync latencies.
- Against pg's curve this is the whole argument in one row set: the
embedded engines' curves are flat/degrading; the client-server curve
is the only increasing one.
### Finding C6: the swap is small but the durability API is coarser than sqlite's
`kv_redb.rs` (~230 lines): same row shape (`TableDefinition<&[u8],
(&[u8], i64)>` — the value/size tuple matches sqlite's row exactly;
`WITHOUT ROWID`-equivalent PK ordering is native B-tree). Differences:
- **redb 4.x offers only `Durability::None` and `Durability::Immediate`**
— no middle tier. sqlite's `synchronous=NORMAL` (consistent-on-crash,
not-fsynced-per-commit — the shipped tier this crate benchmarks and
documents) has **no redb equivalent**: `None` is not
crash-consistent (unpersisted commits can leave the file failing to
open until repaired — the freelist bookkeeping itself is lost on
crash), `Immediate` pays a full fdatasync per commit. The
durability-posture comparison that POC #4 could tune continuously
(`synchronous_commit=off` on pg) doesn't exist here.
- redb's `Database` is process-lifetime; concurrency modes
(`ExclusiveWriter` default, `SingleWriter`, `MultiWriter`) only matter
multi-process — in-process it's one writer transaction at a time,
structurally.
- `create`-then-repair behavior on a `None`-durability crash and the
`has+insert` CAS dance inside one write tx are correct but slower
(the bench's re-put path costs a table probe per commit).
## Consequences for the phase-0/1 register
- **redb is ruled out as the shipped kv engine at a measured tier
mismatch, not a benchmark quibble**: the only durability postures it
offers are "lose acknowledged data on crash" or "pay one fdatasync
per commit" — the crate's shipped durability tier
(`synchronous=NORMAL` WAL, crash-consistent, no per-commit fsync)
is not expressible. sqlite keeps ADR-003's pin on both economics
(44k vs 100 puts/s at 1 KiB) and posture.
- **The iroh-blobs DESIGN.md lineage ("faster than the filesystem for
small items") is now fully triangulated**: its kv-tier claim is
true for its *inline small items with batched/durable-None commits*
(iroh batches writes through a serialized actor with batch windows —
`BatchOptions.max_write_batch`, the same amortization pg gets from
WAL group-commit *for free across connections*). Our harnesses
measure per-commit put discipline; iroh's actor batching is the
reconciliation, and it is also why our POC #3 sqlite numbers (44k)
are the honest ceiling for a per-op CAS store.
- **Batched commits are a real Phase 1 lever, proven by three
independent engines**: pg's group-commit (B3), iroh's write actor
windows, redb's None+flush emulation (arm C4's structure). The
store's put seam could take a `WriteBatch`-shaped commit later — an
API question for Phase 1 (pinned batch puts exist in ADR-005's
surface already; the *durability-window* question is adjacent).
- **Reads: no case for redb on merit** — page-cache reads are what any
of these give the large tier for free; sqlite's 19 µs probe was
already fine.
- **ADR action: none.** SQLite's pin now has four-way triangulated
evidence (POC #3 fs, POC #4 pg, POC #6 redb, plus the lineage
citation). If a future consumer wanted redb-shaped durability
trade-offs (ephemeral cache tier), that's a *different* tier than
the kv engine slot.
## POC quality notes
- The inherited 14-test suite passes unchanged (the redb backend is
additive; sqlite/fs arms untouched). The CAS semantics of
put-if-absent are exercised through the smoke path; redb's
insert-returns-prior-guard shape documented in `kv_redb.rs` docs.
- clippy `-D warnings` clean, fmt clean.
- **Honest caveats**: mdraid single dev box whose fdatasync latency
ranged 0.24-30 ms during the session under co-resident load (pg
container churn from POC #4 ran through the first bench; stopped
before the final redbdiag re-run — relative conclusions unaffected,
absolute redb-Immediate numbers are *device-pressure-dependent* and
the honest statement is "1 fdatasync per commit, cost = device
fdatasync latency, observed 8-30 ms here"); the 2s full bench was
dropped from the record (its 128/256 KiB redb arms hit shell
timeouts; 1s arms with explicit size lists used instead — the harness
gained a size-list argument for this); the 1g redb working-set files
(~270 MB at 128 KiB arm) cleaned from /tmp after diagnosis.
- Instruments: `src/kv_redb.rs` (the backend: Immediate + None + flush
emulation methods), `src/bin/bench_small.rs` (inherited + redb arm +
size-list arg), `examples/redbdiag.rs` (durability decomposition),
`redbdiag2.rs` (None + periodic-flush group-commit emulation),
`redbdiag3.rs` (commit-vs-insert split, per-put file growth),
`redbdiag4.rs` (warm distribution), `redb_conc.rs` (scale-out, both
durability modes), `sqlite_conc.rs` (copied head-to-head from #4).