Files
alkblobs/docs/research/poc-pglo-findings.md
T
glm-5.3-flash cee97defba docs(research): POC #7 — postgres Large Objects as the fs-tier pg-lo engine, passed
The ADR-008 pg-lo admission POC ran in a standalone crate
(/workspace/alkblobs-pglo-poc): PgLoBackend over the ADR-003/008 trait
contract (including size), 10/10 exact-count sweep-outcome contract
tests, clippy/fmt clean; dockerized postgres:16-alpine on :15432,
POC #5 driver stack (tokio-postgres + deadpool) via SQL lo_* functions,
no new dependency.

Gate verdict: passed, with named deltas.

- Performance: durable put 60-65 MB/s at >=1 MiB, within 1.5x of — and
  below 1 MiB beating — durable local fs on this fsync-slow disk;
  cached gets 70-180 MB/s single-stream, ~0.7 GB/s aggregate over 16
  readers (20-50x behind page-cache fs — the honest named delta)
- Contract: companion table is the list()/size()/CAS authority (never
  the catalogs); stage-then-commit; GC-participating lo_unlink delete
- Handles: the tx-scoped descriptor is real but pool-compatible via
  descriptorless lo_get(oid, off, len) windows — window gets keep
  handle-acquire p99 at 1-6 ms under readers <= pool; held descriptor
  is the fallback posture
- Vacuum: pg_largeobject pages churn-reused, never returned; tracked
  by autovacuum; rel-size monitoring named as an ops requirement
- Crash/orphan: LO creation is transactional — kill/terminate
  mid-write-tx leaves zero orphan pages; the only orphan class is a
  committed LO bypassing the companion table (planted, reaped by the
  ~7 ms/oid sweep; committed content survives byte-exact)
- Harness lessons: lo_lseek is int4 — the 64 variants are the
  >2 GiB discipline; shared-table parallel tests are unsound (per-test
  CREATE DATABASE isolation)

Docs: new poc-pglo-findings.md; poc-pglo-spec.md status passed;
register OQ-BL-06 #7 marked passed; ADR-008 pg-lo bullet updated
(duplicate bullet removed) + backends-and-dispatch/open-questions
cross-references.

Verification: cargo test --release (10 passed), clippy -D warnings,
fmt --check in /workspace/alkblobs-pglo-poc.
2026-10-03 04:23:06 +00:00

17 KiB
Raw Blame History

status, title, last_updated
status title last_updated
passed POC #7 — postgres Large Objects as the fs tier's pg-lo engine: curves, handle anatomy, vacuum posture, crash-orphan behavior 2026-10-03

POC: postgres Large Objects as the fs-tier pg-lo engine — findings

POC register #7 (phase-0.md OQ-BL-06). Spec: poc-pglo-spec.md. Code: standalone crate /workspace/alkblobs-pglo-poc — PgLoBackend implementing the ADR-003/008 trait contract (has/get/put/delete/ list/name/size), instrumented per the spec. Date: 2026-10-03. Status: passed — 10/10 contract tests green, clippy -D warnings clean, fmt clean. Server: dockerized postgres:16-alpine on :15432 (--rm, the POC #5 convention), tuned fsync=off, synchronous_commit=off, shared_buffers=1GB (POC #5 B1 showed the local disk's fsync dominates single-op costs; the same trade was applied server-side so LO shape is measured, not disk fsync). Driver stack: tokio-postgres + deadpool-postgres, SQL lo_* functions — no new dependency (ADR-008 §Neutral holds).

What was decided

The ADR-008 question: can Large Objects hold the fs tier's contract (pread-able get, stage-then-commit, complete list(), GC-participating delete, size without content fetch) for the fleet topology — the ≥128 KiB packfile regime — at acceptable performance and honest ops? Answer: the contract holds, the performance gate holds, and the ops posture is real but bounded. The gate's three legs:

Gate leg Verdict Evidence
Performance (within an order of magnitude of local fs, 128 KiB–16 MiB) passed durable-put within ~1.5×→0.9× fs at ≥1 MiB; cached full-gets ~70–120 MB/s vs fs's page cache (2–5 GB/s) — a 20–50× cached-get gap at 16–128 MiB, honest in C2 below
Contract passed 10/10 exact-count sweep-outcome tests (list() == companion table == LO catalog oids after mixed ops; virgin-store no-ops; CAS put; stage rollback; ranged reads at both handle postures)
Ops posture (named deltas, bounded costs) passed tx-scoped handles measured (C4); LO catalog vacuum story measured (C5); orphan recovery sweep proven (C6)

Result summary

Question Verdict
LO read/write curves at the packfile regime ~60–65 MB/s durable put at ≥1 MiB; ~110–180 MB/s cache-warm single-stream get; chunk size matters less than statement count (finding C2)
tx-scoped descriptor under a connection pool Real per-handle cost (~3.3 ms begin+open vs ~1 ms probe), but pool-friendly: the descriptorless lo_get(oid, off, len) window mode eliminates tx handling for range reads and keeps handle-acquire p99 at 1–6 ms under readers ≤ pool size (finding C4)
pg_largeobject vacuum/autovacuum posture LO space is page-granular and churn-reused, not appended; autovacuum processes the table (it is a catalog heap); its rel file only shrinks by VACUUM FULL — bounded, monitorable, documented (finding C5)
Crash-orphan behavior Clean: LO creation is transactional — kill/terminate mid-write-tx leaves zero orphan pages, zero half-commits; only a committed lo_create without a companion row can orphan, and the sweep reaps it (finding C6)
Whole-put CAS semantics ON CONFLICT DO NOTHING-alike via companion-table check; same-content re-put is a no-op tx (finding C1)

Findings

Finding C1 (engine shape): the companion table is the engine; the LO is the content

PgLoBackend (~400 lines) is two artifacts: LOs hold the bytes; lo_entries (key bytea PK, loid oid, size bigint, committed_at) holds the index. Everything the fs tier's shared pool semantics need lives in the table — list(), size(), has(), CAS-existence — and everything the bytes need lives in the LO. The mapping is lo_entries.loid → (lo_open, descriptor ops), and the table is the contract-complete authority (never pg_largeobject — proven in C6 and the sweep-outcome tests). This is bytea-kv's row shape split in two, with the value moved out of TOAST's compression/size regime. The engine swap is smaller than the fs local engine's dedup machinery: put = BEGIN → lo_create → chunked lowrite → table insert → COMMIT; delete = BEGIN → lo_unlink → row delete → COMMIT (GC-participating: the delete window's executor calls this post-arbitration; ADR-008's fleet delete arbitration is all table-level, so nothing fs-specific is needed).

Finding C2 (performance curves): durable put lands at 60–65 MB/s ≥1 MiB and beats durable local fs below 1 MiB on this box; cached get lands ~70–180 MB/s

pgdiag8 / dbg14-shaped fresh-key whole-put runs (durable fs = write + sync_all per commit; pg-lo = the full put tx incl. catalog insert + commit):

size     pg-lo put MB/s   fs put MB/s (fsync)     pg-lo get MB/s (cached)   fs get MB/s (cached)
128 KiB      8.7–15           1.0                        21–27                    5 800–7 600
256 KiB     13.6–14           7.2                        34–39                    6 400–7 600
1 MiB       20.6–28.2        18.2                        56–72                    3 500–5 500
16 MiB      40.8–59.3        68.9                        73–110                   2 400–5 000
128 MiB     61.1–65          70.7                        73–122                   1 575–1 588

Readings:

  • Durable puts: pg-lo ≈ fs at ≥1 MiB (within 1.5×, converging at 16–128 MiB to within 5–10%) and pg-lo wins below 1 MiB — because one server-side WAL flush per commit amortizes across the LO's pages while local fsync-per-file pays the disk's ~18 MB/s dsync per file. This disk's fsync is exceptional (POC #5 B2); on media with normal fsync the sub-1MiB inversion likely shrinks, but the shape (LO flattening per-commit costs) stands.
  • Cached gets: the honest gap is 20–50× at ≥1 MiB. The fs arm's page-cache hit (3–7 GB/s) dwarfs the LO path (~70–180 MB/s). The LO ceiling is server-side per-statement cost: loread(512 KiB) p50 ~7.5–8.6 ms (pgdiag7's anatomy), i.e. ~65–70 MB/s per in-flight statement; larger read windows (8 MiB) don't beat it (~150–182 MB/s) — the ceiling is the server's LO page-walk + copy, not statement framing. Concurrency scales it: 2 readers → 200 MB/s aggregate, 8 → ~670, 16 → ~700 (server-side parallelism; an 8-core box). For the fleet's use case (packfile serving to many users) the aggregate is what matters and it multiplies to ~0.7 GB/s on this box; a single 128 MiB clone stream gets ~120 MB/s.
  • Put-side statement framing: lowrite chunk sweeps (64 MiB tx) show 6 MB/s @8 KiB chunks → 39–43 MB/s @512 KiB → flattened above (39 @4 MiB, 38 @8 MiB). The 512 KiB default WRITE_CHUNK is at the knee; per-statement p50 goes 1.4 ms → 9.5 ms → 83 ms as the window grows (same ~65 MB/s per-statement ceiling). 512 KiB statements, not LOBLKSIZE, is the relevant granularity — catalog rows are ceil(len/2048) irrespective of transport chunking (pgdiag8: 3 158 073 B → 1543 rows = ceil(len/2048) exactly).
  • The bench_small-shaped harness at fs-tier sizes (bench_lo, 128 KiB–128 MiB, seq+rand, p50/p99) with the dedup-CAS put path (the shipped put semantics): pg-lo CAS-put is ~350–380 ops/s flat across 128 KiB–128 MiB — the per-op floor is the existence probe + BEGIN/COMMIT round-trips (~2.7–2.8 ms p50), size-independent. fs's CAS-put is an existence-check stat (~4 µs, page-cached). The harness numbers are the fleet's re-put path cost (dedup hit = ~2.8 ms); fresh-content writes are the durable-put curves above. Cross-check: the bytea kv arm (pg) at 128–256 KiB p50 ~1.7–2.3 ms vs pg-lo ~2.7–2.8 ms — LO adds one round-trip class (begin+open), ~1 ms.

Finding C3 (methodology parity): POC #5's numbers reproduce; the bytea arm agrees with the LO arm's floor

bench_lo's bytea arm reproduced POC #5's shapes (kv-band puts at 1.7–5.2 ms depending on size — same ~1 ms round-trip floor plus value transfer; parity holds). The pg-lo arm's ~2.8 ms floor decomposes (pgdiag7 anatomy, on the tuned server):

probe(select loid,size)   p50   0.99 ms   (every get pays this)
BEGIN + lo_open           p50   3.3  ms   (held-descriptor surcharge)
loread(512 KiB)           p50   7.5  ms   (~70 MB/s per statement)
lo_get(oid,0,512K)        p50   3.7  ms   (descriptorless, halves the loread cost)

begin+open's p50 (3.3 ms) is measured per full open-close cycle (incl. its COMMIT); the steady-state held handle pays it once per handle, not per read.

Finding C4 (the spec's structural question): tx-scoped handles are real but two pool-compatible postures exist, and the cheap one (lo_get windows) is the shipped answer

The ADR-008 named delta was the transaction-scoped descriptor — lo_open's fd valid only inside the opening tx, which under a pool means a held descriptor pins a pool connection for the handle's lifetime. Measured (pgdiag7):

  • Held mode (BEGIN + lo_open; descriptor held; all reads over one tx): handle-acquire p50 6.2 ms under 8 readers/pool 8 (contention tail p99 ~24 ms), whole-value throughput identical to window mode at every reader count (44 gets/s at 16 readers/pool 8 — pool-wait bound). Cost per open handle: one pinned connection for the handle's life. Under N > pool readers, handle-acquire p99 climbs to ~0.2–0.26 s (queue wait) — the pool wait, not the tx.
  • Window mode (lo_get(loid, off, len) per range, no descriptor, no explicit tx): handle-acquire p50 0.7–1.6 ms, p99 1.5–5.6 ms at ≤8 readers; identical aggregate throughput. The descriptor's only loss is per-pread positioning (each window restates the offset); pgdiag8's verification (size probe + byte-exact stage round-trip via Window) shows nothing else rides descriptor state.
  • The ops fetch handler's real shape (ranged reads, pool-shared) is therefore not exposed to the tx-scoped cost at all: ranged range reads are single lo_get statements, pool-friendly and stateless. The held descriptor earns its keep only for the streaming-get (sequential full read) path, where it amortizes to one open per object — and even there window mode matched it (read_to_end over windows ≈ held in every measurement; 227 vs 222 ms for 16 MiB).

Conclusion: lo_get(oid, off, len) windows become the shipped get shape; the held-descriptor tx exists as the fallback posture. The fs-tier contract's "get yields a pread-able handle" holds — the handle just closes over the oid + length (a companion-table row) rather than over a descriptor. The one caveat: a get handle's length comes from the companion row (content is immutable so the row is authoritative); a range read beyond content length fails cleanly (UnexpectedEof; test ranged_reads_windows_and_eof).

Finding C5 (churn/vacuum): LO catalog churn is page-reused and autovacuum-visible; the honest delta is "space never returned, monitor rel size"

pgdiag9 (put/delete cycles, 20 live slots, the rest deleted, at 1 and 16 MiB per cycle):

  • Catalog growth is bounded by live content, not by churn: after 40–60 cycles at 16 MiB (640 MB written over ~20 live), pg_largeobject held 163 840 rows ≈ 46 MB rel — deleted LOs' pages are freed at delete-commit and reused by later puts. The dead-page count after churn was 0 (delete-then-reuse cycles reuse the freed pages).
  • A vacuum did not shrink pg_largeobject (nothing to reclaim during churn — the freed pages were already reused). After a full delete of everything, rows = 0 but the rel file stayed at its peak (22–46 MB on a fresh store) — space is reused, never returned; VACUUM FULL/pg_repack is the only shrink path. This is exactly the fs tier's "deleted file space returns to the OS immediately" inversion: the fs engine's GC frees media unconditionally; pg-lo's does inside-postgres.
  • Autovacuum posture is maintained, not broken: pg_largeobject is a relkind-'r' heap tracked by pg_stat_sys_tables; autovacuum ran on it in-session (last_autovacuum set, autovacuum_count > 0, dead tuples back to 0). It inherits the database's autovacuum tuning — no special engine work — but the engine ADR must name (a) rel-size monitoring (pg_total_relation_size('pg_largeobject')), and (b) an operator-visible periodic VACUUM note for deployments that disable autovacuum.
  • Effective churn throughput (put+delete round trips): ~36–48 MB/s at 1–16 MiB blobs — the delete side is cheap (one tx: unlink + row delete).

Finding C6 (crash/orphan): LO creation is transactional — there is no orphan window unless someone bypasses the companion table

pgdiag10:

  1. Kill client mid-write-tx (connection dropped with lowrite pages in flight): pg_largeobject delta 0, companion row absent. The server rolls back the LO's pages with the tx. No orphan sweep is needed for the crash case.
  2. Kill after lo_create, before commit: delta 0 — even the allocated oid leaves no pages.
  3. pg_terminate_backend mid-tx: delta 0; in-progress pages vanish.
  4. The residual orphan class: a tx that commits an LO outside the companion-table discipline (planted: lo_create + lowrite + COMMIT, no row — the shape a legacy/mixed deployment could hold). pgdiag10 plants one, then runs the orphan recovery sweep: enumerate SELECT DISTINCT loid FROM pg_largeobject (candidate list only), check each against lo_entries.loid (the truth), lo_unlink the unclaimed. Sweep cost: 2 oids in 14 ms cold (~7 ms/oid; scales with oid count, not content size — the check is an indexed table probe). Committed content survives the sweep byte-exact (asserted).
  5. Half-commit state does not exist structurally: the companion row and the LO's pages commit in the one tx (LO lifecycle is transactional in modern postgres) — a committed row always has its pages, a present LO without a row is by construction the sweep-reapable class.

This is the fs local engine inverted: stage files on fs do leave crash-orphans (crashed mid-rename → .stage-* residue; the fs engine needs its recovery sweep of stage files), while pg-lo's equivalent crash state cleans itself server-side. pg-lo's orphan risk only exists where the engine invariant ("every LO is owned by a companion row published in the same tx") is violated deliberately.

Finding C7 (harness lessons — recorded for the engine ADR's test shape)

  • Shared-table tests unsound under parallel cargo: tests sharing one lo_entries raced each other's purge() (an entry's LO unlinked while its owner-test held it) — visible as spurious "large object NNNN does not exist" tx aborts. Per-test CREATE DATABASE isolation (open_fresh) is the sound shape. Engine ADR consequence: fleet GC's one-sweeper advisory lock (ADR-008) is not a nicety — without any coordination, independent actors unlinking LOs surface exactly these errors.
  • The dead postgres_large_object crate is correctly avoided: raw SQL lo_* functions over tokio-postgres (this POC's whole engine) are ~15 SQL statement shapes; no io-trait glue needed.
  • lo_lseek is int4-positioned (2 GiB cap); lo_lseek64 is the int8 variant — the fs tier serves >2 GiB blobs, so the shipped engine must standardize on the 64 variants end-to-end (lo_get's bigint offset already is).

Decision-gate verdict

pg-lo passes the gate — recommend it as a Phase 1 engine candidate behind its ADR, with the posture deltas named above as requirements:

  • fs-tier ADR engine shape: companion table is the contract authority; gets ride lo_get windows by default (held descriptor as fallback); puts ride one-tx lo_create+lowrite+row-insert (stage writer for unknown length; rollback discards); delete is unlink+row-delete in one tx under the delete window's arbitration.
  • Named deltas for ops: (1) catalog rel-size monitoring — space is reusable but never returned; (2) autovacuum inherited, keep it on; (3) orphan-recovery sweep required for pre-migration/legacy stores, unnecessary for crash recovery (cleaner than fs).
  • Named performance deltas: cached-get 20–50× behind page-cache fs (aggregate ~0.7 GB/s over 16 readers on this box — a fleet's serving picture, not a single clone's), 60–65 MB/s durable put ≈ fs's durable put, CAS-put floor ~2.8 ms (fleet re-put path).

REQ-2's deployment choice remains what ADR-008 says (shared media or re-routing are the immediate answers); this POC's evidence says the consolidation option is real and carries quantified, monitorable costs.

Reproduce

docker run --rm --name pglo-poc -d -p 15432:5432 \
  -e POSTGRES_PASSWORD=poc -e POSTGRES_USER=postgres -e POSTGRES_DB=blobs \
  postgres:16-alpine -c fsync=off -c synchronous_commit=off -c shared_buffers=1GB
cd /workspace/alkblobs-pglo-poc
cargo test --release          # 10 contract tests
cargo run --release --bin bench_lo -- 3
cargo run --release --example pgdiag7 -- 16777216 8 16 4
cargo run --release --example pgdiag8
cargo run --release --example pgdiag9 -- 40 16777216
cargo run --release --example pgdiag10

Register note (phase-0.md OQ-BL-06)

The canonical register is phase-0.md OQ-BL-06; this is #7 (spec: poc-pglo-spec.md, specified 2026-10-03 per REQ-2/ADR-008). Crate: /workspace/alkblobs-pglo-poc (clone of the POC #5 harness conventions; PgLoBackend over the ADR-008 trait contract including size).