Files
alkblobs/docs/research/poc-pglo-findings.md
T
glm-5.3-flash cee97defba docs(research): POC #7 — postgres Large Objects as the fs-tier pg-lo engine, passed
The ADR-008 pg-lo admission POC ran in a standalone crate
(/workspace/alkblobs-pglo-poc): PgLoBackend over the ADR-003/008 trait
contract (including size), 10/10 exact-count sweep-outcome contract
tests, clippy/fmt clean; dockerized postgres:16-alpine on :15432,
POC #5 driver stack (tokio-postgres + deadpool) via SQL lo_* functions,
no new dependency.

Gate verdict: passed, with named deltas.

- Performance: durable put 60-65 MB/s at >=1 MiB, within 1.5x of — and
  below 1 MiB beating — durable local fs on this fsync-slow disk;
  cached gets 70-180 MB/s single-stream, ~0.7 GB/s aggregate over 16
  readers (20-50x behind page-cache fs — the honest named delta)
- Contract: companion table is the list()/size()/CAS authority (never
  the catalogs); stage-then-commit; GC-participating lo_unlink delete
- Handles: the tx-scoped descriptor is real but pool-compatible via
  descriptorless lo_get(oid, off, len) windows — window gets keep
  handle-acquire p99 at 1-6 ms under readers <= pool; held descriptor
  is the fallback posture
- Vacuum: pg_largeobject pages churn-reused, never returned; tracked
  by autovacuum; rel-size monitoring named as an ops requirement
- Crash/orphan: LO creation is transactional — kill/terminate
  mid-write-tx leaves zero orphan pages; the only orphan class is a
  committed LO bypassing the companion table (planted, reaped by the
  ~7 ms/oid sweep; committed content survives byte-exact)
- Harness lessons: lo_lseek is int4 — the 64 variants are the
  >2 GiB discipline; shared-table parallel tests are unsound (per-test
  CREATE DATABASE isolation)

Docs: new poc-pglo-findings.md; poc-pglo-spec.md status passed;
register OQ-BL-06 #7 marked passed; ADR-008 pg-lo bullet updated
(duplicate bullet removed) + backends-and-dispatch/open-questions
cross-references.

Verification: cargo test --release (10 passed), clippy -D warnings,
fmt --check in /workspace/alkblobs-pglo-poc.
2026-10-03 04:23:06 +00:00

300 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
status: passed
title: "POC #7 — postgres Large Objects as the fs tier's pg-lo engine: curves, handle anatomy, vacuum posture, crash-orphan behavior"
last_updated: 2026-10-03
---
# POC: postgres Large Objects as the fs-tier pg-lo engine — findings
> **POC register #7** (phase-0.md OQ-BL-06). Spec: `poc-pglo-spec.md`.
> Code: standalone crate `/workspace/alkblobs-pglo-poc` — `PgLoBackend`
> implementing the ADR-003/008 trait contract (has/get/put/delete/
> list/name/size), instrumented per the spec. Date: 2026-10-03.
> Status: **passed** — 10/10 contract tests green, clippy `-D warnings`
> clean, fmt clean. Server: dockerized postgres:16-alpine on :15432
> (`--rm`, the POC #5 convention), tuned `fsync=off`,
> `synchronous_commit=off`, `shared_buffers=1GB` (POC #5 B1 showed the
> local disk's fsync dominates single-op costs; the same trade was
> applied server-side so LO *shape* is measured, not disk fsync).
> Driver stack: `tokio-postgres` + `deadpool-postgres`, SQL `lo_*`
> functions — no new dependency (ADR-008 §Neutral holds).
## What was decided
The ADR-008 question: can Large Objects hold the fs tier's contract
(pread-able get, stage-then-commit, complete `list()`, GC-participating
`delete`, `size` without content fetch) for the fleet topology — the
≥128 KiB packfile regime — at acceptable performance and honest ops?
Answer: **the contract holds, the performance gate holds, and the ops
posture is real but bounded.** The gate's three legs:
| Gate leg | Verdict | Evidence |
|---|---|---|
| Performance (within an order of magnitude of local fs, 128 KiB–16 MiB) | **passed** | durable-put within ~1.5×→0.9× fs at ≥1 MiB; cached full-gets ~70–120 MB/s vs fs's page cache (2–5 GB/s) — a 20–50× *cached-get* gap at 16–128 MiB, honest in C2 below |
| Contract | **passed** | 10/10 exact-count sweep-outcome tests (`list()` == companion table == LO catalog oids after mixed ops; virgin-store no-ops; CAS put; stage rollback; ranged reads at both handle postures) |
| Ops posture (named deltas, bounded costs) | **passed** | tx-scoped handles measured (C4); LO catalog vacuum story measured (C5); orphan recovery sweep proven (C6) |
## Result summary
| Question | Verdict |
|----------|---------|
| LO read/write curves at the packfile regime | ~60–65 MB/s durable put at ≥1 MiB; ~110–180 MB/s cache-warm single-stream get; chunk size matters less than statement count (finding C2) |
| tx-scoped descriptor under a connection pool | Real per-handle cost (~3.3 ms begin+open vs ~1 ms probe), but **pool-friendly**: the descriptorless `lo_get(oid, off, len)` window mode eliminates tx handling for range reads and keeps handle-acquire p99 at 1–6 ms under readers ≤ pool size (finding C4) |
| `pg_largeobject` vacuum/autovacuum posture | LO space is page-granular and churn-reused, not appended; autovacuum processes the table (it is a catalog heap); its rel file only shrinks by `VACUUM FULL` — bounded, monitorable, documented (finding C5) |
| Crash-orphan behavior | Clean: LO creation is transactional — kill/terminate mid-write-tx leaves **zero** orphan pages, zero half-commits; only a committed lo_create without a companion row can orphan, and the sweep reaps it (finding C6) |
| Whole-put CAS semantics | `ON CONFLICT DO NOTHING`-alike via companion-table check; same-content re-put is a no-op tx (finding C1) |
## Findings
### Finding C1 (engine shape): the companion table is the engine; the LO is the content
`PgLoBackend` (~400 lines) is two artifacts: LOs hold the bytes;
`lo_entries (key bytea PK, loid oid, size bigint, committed_at)` holds
the index. Everything the fs tier's *shared pool* semantics need lives
in the table — `list()`, `size()`, `has()`, CAS-existence — and
everything the *bytes* need lives in the LO. The mapping is
`lo_entries.loid → (lo_open, descriptor ops)`, and the table is the
contract-complete authority (never `pg_largeobject` — proven in C6 and
the sweep-outcome tests). This is bytea-kv's row shape split in two,
with the value moved out of TOAST's compression/size regime. The engine
swap is smaller than the fs `local` engine's dedup machinery: put =
BEGIN → `lo_create` → chunked `lowrite` → table insert → COMMIT;
delete = BEGIN → `lo_unlink` → row delete → COMMIT (GC-participating:
the delete window's executor calls this post-arbitration; ADR-008's
fleet delete arbitration is all table-level, so nothing fs-specific is
needed).
### Finding C2 (performance curves): durable put lands at 60–65 MB/s ≥1 MiB and *beats* durable local fs below 1 MiB on this box; cached get lands ~70–180 MB/s
`pgdiag8` / `dbg14`-shaped fresh-key whole-put runs (durable fs =
write + `sync_all` per commit; pg-lo = the full put tx incl. catalog
insert + commit):
```
size pg-lo put MB/s fs put MB/s (fsync) pg-lo get MB/s (cached) fs get MB/s (cached)
128 KiB 8.7–15 1.0 21–27 5 800–7 600
256 KiB 13.6–14 7.2 34–39 6 400–7 600
1 MiB 20.6–28.2 18.2 56–72 3 500–5 500
16 MiB 40.8–59.3 68.9 73–110 2 400–5 000
128 MiB 61.1–65 70.7 73–122 1 575–1 588
```
Readings:
- **Durable puts: pg-lo ≈ fs at ≥1 MiB (within 1.5×, converging at
16–128 MiB to within 5–10%) and pg-lo *wins* below 1 MiB** — because
one server-side WAL flush per commit amortizes across the LO's pages
while local fsync-per-file pays the disk's ~18 MB/s dsync per file.
This disk's fsync is exceptional (POC #5 B2); on media with normal
fsync the sub-1MiB inversion likely shrinks, but the shape (LO
flattening per-commit costs) stands.
- **Cached gets: the honest gap is 20–50× at ≥1 MiB.** The fs arm's
page-cache hit (3–7 GB/s) dwarfs the LO path (~70–180 MB/s). The LO
ceiling is *server-side per-statement cost*: `loread(512 KiB)` p50
~7.5–8.6 ms (pgdiag7's anatomy), i.e. ~65–70 MB/s per in-flight
statement; larger read windows (8 MiB) don't beat it (~150–182
MB/s) — the ceiling is the server's LO page-walk + copy, not
statement framing. Concurrency scales it: 2 readers → 200 MB/s
aggregate, 8 → ~670, 16 → ~700 (server-side parallelism; an
8-core box). For the fleet's use case (packfile serving to many
users) the aggregate is what matters and it multiplies to ~0.7 GB/s
on this box; a single 128 MiB clone stream gets ~120 MB/s.
- **Put-side statement framing**: `lowrite` chunk sweeps (64 MiB tx)
show 6 MB/s @8 KiB chunks → 39–43 MB/s @512 KiB → flattened above
(39 @4 MiB, 38 @8 MiB). The 512 KiB default `WRITE_CHUNK` is at the
knee; per-statement p50 goes 1.4 ms → 9.5 ms → 83 ms as the window
grows (same ~65 MB/s per-statement ceiling). **512 KiB statements,
not LOBLKSIZE, is the relevant granularity** — catalog rows are
ceil(len/2048) irrespective of transport chunking (`pgdiag8`:
3 158 073 B → 1543 rows = ceil(len/2048) exactly).
- **The `bench_small`-shaped harness at fs-tier sizes** (`bench_lo`,
128 KiB–128 MiB, seq+rand, p50/p99) with the *dedup-CAS* put path
(the shipped put semantics): pg-lo CAS-put is ~350–380 ops/s flat
across 128 KiB–128 MiB — the per-op floor is the existence probe +
BEGIN/COMMIT round-trips (~2.7–2.8 ms p50), size-independent. fs's
CAS-put is an existence-check stat (~4 µs, page-cached). The harness
numbers *are* the fleet's re-put path cost (dedup hit = ~2.8 ms);
fresh-content writes are the durable-put curves above. Cross-check:
the bytea kv arm (pg) at 128–256 KiB p50 ~1.7–2.3 ms vs pg-lo
~2.7–2.8 ms — LO adds one round-trip class (begin+open), ~1 ms.
### Finding C3 (methodology parity): POC #5's numbers reproduce; the bytea arm agrees with the LO arm's floor
`bench_lo`'s bytea arm reproduced POC #5's shapes (kv-band puts at
1.7–5.2 ms depending on size — same ~1 ms round-trip floor plus value
transfer; parity holds). The pg-lo arm's ~2.8 ms floor decomposes
(pgdiag7 anatomy, on the tuned server):
```
probe(select loid,size) p50 0.99 ms (every get pays this)
BEGIN + lo_open p50 3.3 ms (held-descriptor surcharge)
loread(512 KiB) p50 7.5 ms (~70 MB/s per statement)
lo_get(oid,0,512K) p50 3.7 ms (descriptorless, halves the loread cost)
```
`begin+open`'s p50 (3.3 ms) is measured *per full open-close cycle*
(incl. its COMMIT); the steady-state held handle pays it once per
handle, not per read.
### Finding C4 (the spec's structural question): tx-scoped handles are real but two pool-compatible postures exist, and the cheap one (`lo_get` windows) is the shipped answer
The ADR-008 named delta was the transaction-scoped descriptor —
`lo_open`'s fd valid only inside the opening tx, which under a pool
means a held descriptor pins a pool connection for the handle's
lifetime. Measured (`pgdiag7`):
- **Held mode** (BEGIN + lo_open; descriptor held; all reads over one
tx): handle-acquire p50 6.2 ms under 8 readers/pool 8 (contention
tail p99 ~24 ms), whole-value throughput identical to window mode at
every reader count (44 gets/s at 16 readers/pool 8 — pool-wait
bound). Cost per *open handle*: one pinned connection for the
handle's life. Under N > pool readers, handle-acquire p99 climbs to
~0.2–0.26 s (queue wait) — *the pool wait, not the tx*.
- **Window mode** (`lo_get(loid, off, len)` per range, no descriptor,
no explicit tx): handle-acquire p50 0.7–1.6 ms, p99 1.5–5.6 ms at ≤8
readers; identical aggregate throughput. The descriptor's only loss
is per-pread positioning (each window restates the offset); pgdiag8's
verification (`size` probe + byte-exact stage round-trip via Window)
shows nothing else rides descriptor state.
- **The ops fetch handler's real shape** (ranged reads, pool-shared) is
therefore *not* exposed to the tx-scoped cost at all: ranged range
reads are single `lo_get` statements, pool-friendly and stateless.
The held descriptor earns its keep only for the streaming-get
(sequential full read) path, where it amortizes to one open per
object — and *even there window mode matched it* (read_to_end over
windows ≈ held in every measurement; 227 vs 222 ms for 16 MiB).
**Conclusion: `lo_get(oid, off, len)` windows become the shipped get
shape; the held-descriptor tx exists as the fallback posture.** The
fs-tier contract's "get yields a pread-able handle" holds — the handle
just closes over the oid + length (a companion-table row) rather than
over a descriptor. The one caveat: a get handle's length comes from the
companion row (content is immutable so the row is authoritative); a
range read beyond content length fails cleanly (`UnexpectedEof`; test
`ranged_reads_windows_and_eof`).
### Finding C5 (churn/vacuum): LO catalog churn is page-reused and autovacuum-visible; the honest delta is "space never returned, monitor rel size"
`pgdiag9` (put/delete cycles, 20 live slots, the rest deleted, at 1 and
16 MiB per cycle):
- **Catalog growth is bounded by live content, not by churn**: after
40–60 cycles at 16 MiB (640 MB written over ~20 live), `pg_largeobject`
held 163 840 rows ≈ 46 MB rel — deleted LOs' pages are freed *at
delete-commit* and reused by later puts. The dead-page count after
churn was 0 (delete-then-reuse cycles reuse the freed pages).
- **A vacuum did not shrink** `pg_largeobject` (nothing to reclaim
during churn — the freed pages were already reused). After a full
delete of everything, rows = 0 but the rel file stayed at its peak
(22–46 MB on a fresh store) — **space is reused, never returned**;
`VACUUM FULL`/`pg_repack` is the only shrink path. This is exactly
the fs tier's "deleted file space returns to the OS immediately"
inversion: the fs engine's GC frees media unconditionally; pg-lo's
does inside-postgres.
- **Autovacuum posture is maintained, not broken**: `pg_largeobject`
is a relkind-'r' heap tracked by `pg_stat_sys_tables`; autovacuum ran
on it in-session (`last_autovacuum` set, `autovacuum_count` > 0,
dead tuples back to 0). It inherits the database's autovacuum tuning
— no special engine work — but the engine ADR must name (a) rel-size
monitoring (`pg_total_relation_size('pg_largeobject')`), and (b) an
operator-visible periodic `VACUUM` note for deployments that disable
autovacuum.
- Effective churn throughput (put+delete round trips): ~36–48 MB/s at
1–16 MiB blobs — the delete side is cheap (one tx: unlink + row
delete).
### Finding C6 (crash/orphan): LO creation is transactional — there is no orphan window *unless* someone bypasses the companion table
`pgdiag10`:
1. **Kill client mid-write-tx** (connection dropped with lowrite pages
in flight): pg_largeobject delta 0, companion row absent. The
server rolls back the LO's pages with the tx. **No orphan sweep is
needed for the crash case.**
2. **Kill after lo_create, before commit**: delta 0 — even the
allocated oid leaves no pages.
3. **`pg_terminate_backend` mid-tx**: delta 0; in-progress pages vanish.
4. **The residual orphan class**: a tx that commits an LO *outside* the
companion-table discipline (planted: lo_create + lowrite + COMMIT,
no row — the shape a legacy/mixed deployment could hold). `pgdiag10`
plants one, then runs the **orphan recovery sweep**: enumerate
`SELECT DISTINCT loid FROM pg_largeobject` (candidate list only),
check each against `lo_entries.loid` (the truth), `lo_unlink` the
unclaimed. Sweep cost: 2 oids in 14 ms cold (~7 ms/oid; scales with
oid count, not content size — the check is an indexed table probe).
Committed content survives the sweep byte-exact (asserted).
5. **Half-commit state does not exist structurally**: the companion row
and the LO's pages commit in the one tx (LO lifecycle is
transactional in modern postgres) — a committed row always has its
pages, a present LO without a row is by construction the
sweep-reapable class.
This is the fs `local` engine inverted: stage files on fs *do* leave
crash-orphans (crashed mid-rename → `.stage-*` residue; the fs engine
needs its recovery sweep of stage files), while pg-lo's equivalent
crash state cleans itself server-side. pg-lo's orphan risk only exists
where the engine invariant ("every LO is owned by a companion row
published in the same tx") is violated deliberately.
### Finding C7 (harness lessons — recorded for the engine ADR's test shape)
- **Shared-table tests unsound under parallel cargo**: tests sharing
one `lo_entries` raced each other's `purge()` (an entry's LO unlinked
while its owner-test held it) — visible as spurious
"large object NNNN does not exist" tx aborts. Per-test `CREATE
DATABASE` isolation (`open_fresh`) is the sound shape. Engine ADR
consequence: fleet GC's one-sweeper advisory lock (ADR-008) is not a
nicety — without any coordination, independent actors unlinking LOs
surface exactly these errors.
- The dead `postgres_large_object` crate is correctly avoided: raw SQL
`lo_*` functions over tokio-postgres (this POC's whole engine) are
~15 SQL statement shapes; no io-trait glue needed.
- `lo_lseek` is int4-positioned (2 GiB cap); **`lo_lseek64`** is the
int8 variant — the fs tier serves >2 GiB blobs, so the shipped
engine must standardize on the 64 variants end-to-end (`lo_get`'s
bigint offset already is).
## Decision-gate verdict
**pg-lo passes the gate — recommend it as a Phase 1 engine candidate
behind its ADR**, with the posture deltas named above as requirements:
- fs-tier ADR engine shape: companion table is the contract authority;
gets ride `lo_get` windows by default (held descriptor as fallback);
puts ride one-tx lo_create+lowrite+row-insert (stage writer for
unknown length; rollback discards); delete is unlink+row-delete in
one tx under the delete window's arbitration.
- Named deltas for ops: (1) catalog rel-size monitoring — space is
reusable but never returned; (2) autovacuum inherited, keep it on;
(3) orphan-recovery sweep required for pre-migration/legacy stores,
unnecessary for crash recovery (cleaner than fs).
- Named performance deltas: cached-get 20–50× behind page-cache fs
(aggregate ~0.7 GB/s over 16 readers on this box — a fleet's
serving picture, not a single clone's), 60–65 MB/s durable put
≈ fs's durable put, CAS-put floor ~2.8 ms (fleet re-put path).
REQ-2's deployment choice remains what ADR-008 says (shared media or
re-routing are the *immediate* answers); this POC's evidence says the
consolidation option is real and carries quantified, monitorable costs.
## Reproduce
```sh
docker run --rm --name pglo-poc -d -p 15432:5432 \
-e POSTGRES_PASSWORD=poc -e POSTGRES_USER=postgres -e POSTGRES_DB=blobs \
postgres:16-alpine -c fsync=off -c synchronous_commit=off -c shared_buffers=1GB
cd /workspace/alkblobs-pglo-poc
cargo test --release # 10 contract tests
cargo run --release --bin bench_lo -- 3
cargo run --release --example pgdiag7 -- 16777216 8 16 4
cargo run --release --example pgdiag8
cargo run --release --example pgdiag9 -- 40 16777216
cargo run --release --example pgdiag10
```
## Register note (phase-0.md OQ-BL-06)
The canonical register is phase-0.md OQ-BL-06; this is **#7** (spec:
`poc-pglo-spec.md`, specified 2026-10-03 per REQ-2/ADR-008). Crate:
`/workspace/alkblobs-pglo-poc` (clone of the POC #5 harness conventions;
PgLoBackend over the ADR-008 trait contract including `size`).