docs(research): POC #7 — postgres Large Objects as the fs-tier pg-lo engine, passed
The ADR-008 pg-lo admission POC ran in a standalone crate (/workspace/alkblobs-pglo-poc): PgLoBackend over the ADR-003/008 trait contract (including size), 10/10 exact-count sweep-outcome contract tests, clippy/fmt clean; dockerized postgres:16-alpine on :15432, POC #5 driver stack (tokio-postgres + deadpool) via SQL lo_* functions, no new dependency. Gate verdict: passed, with named deltas. - Performance: durable put 60-65 MB/s at >=1 MiB, within 1.5x of — and below 1 MiB beating — durable local fs on this fsync-slow disk; cached gets 70-180 MB/s single-stream, ~0.7 GB/s aggregate over 16 readers (20-50x behind page-cache fs — the honest named delta) - Contract: companion table is the list()/size()/CAS authority (never the catalogs); stage-then-commit; GC-participating lo_unlink delete - Handles: the tx-scoped descriptor is real but pool-compatible via descriptorless lo_get(oid, off, len) windows — window gets keep handle-acquire p99 at 1-6 ms under readers <= pool; held descriptor is the fallback posture - Vacuum: pg_largeobject pages churn-reused, never returned; tracked by autovacuum; rel-size monitoring named as an ops requirement - Crash/orphan: LO creation is transactional — kill/terminate mid-write-tx leaves zero orphan pages; the only orphan class is a committed LO bypassing the companion table (planted, reaped by the ~7 ms/oid sweep; committed content survives byte-exact) - Harness lessons: lo_lseek is int4 — the 64 variants are the >2 GiB discipline; shared-table parallel tests are unsound (per-test CREATE DATABASE isolation) Docs: new poc-pglo-findings.md; poc-pglo-spec.md status passed; register OQ-BL-06 #7 marked passed; ADR-008 pg-lo bullet updated (duplicate bullet removed) + backends-and-dispatch/open-questions cross-references. Verification: cargo test --release (10 passed), clippy -D warnings, fmt --check in /workspace/alkblobs-pglo-poc.
This commit is contained in:
1 parent
281c37876e
commit
cee97defba
6 files changed
+329
-19
No files matched your search
@@ -145,7 +145,10 @@ for fleets):
|
|||||||
- **`pg-lo` engine** (candidate, *not shipped*; ADR-008): postgres
|
- **`pg-lo` engine** (candidate, *not shipped*; ADR-008): postgres
|
||||||
Large Objects as the fs tier's storage — the same engine-behind-
|
Large Objects as the fs tier's storage — the same engine-behind-
|
||||||
one-trait move as ADR-007, one tier over. Admission is gated on
|
one-trait move as ADR-007, one tier over. Admission is gated on
|
||||||
POC #7's measured evidence and its own engine ADR (see
|
POC #7's measured evidence (passed 2026-10-03:
|
||||||
|
`docs/research/poc-pglo-findings.md` — durable put ≈ fs's durable
|
||||||
|
put, `lo_get` window gets, companion-table authority, clean
|
||||||
|
crash-orphan behavior) and its own engine ADR (see
|
||||||
`docs/research/poc-pglo-spec.md`).
|
`docs/research/poc-pglo-spec.md`).
|
||||||
- **Fleet locality contract (ADR-008):** over one shared pool, the fs
|
- **Fleet locality contract (ADR-008):** over one shared pool, the fs
|
||||||
tier is either shared media (every node's `local` root on the same
|
tier is either shared media (every node's `local` root on the same
|
||||||
|
|||||||
@@ -171,7 +171,14 @@ valid:
|
|||||||
pooling, `pg_largeobject`/vacuum posture under churn, crash-orphan
|
pooling, `pg_largeobject`/vacuum posture under churn, crash-orphan
|
||||||
behavior) and its own engine ADR. **Not shipped speculatively;**
|
behavior) and its own engine ADR. **Not shipped speculatively;**
|
||||||
REQ-2's immediate fleet answers are shared media or re-routing;
|
REQ-2's immediate fleet answers are shared media or re-routing;
|
||||||
pg-lo is the consolidation option if those are unacceptable.
|
pg-lo is the consolidation option if those are unacceptable. *Update
|
||||||
|
2026-10-03: POC #7 ran and passed its gate —
|
||||||
|
`docs/research/poc-pglo-findings.md` (companion-table authority,
|
||||||
|
`lo_get` window gets, durable put ≈ fs's durable put, cached gets
|
||||||
|
behind page-cache fs by a measured 20-50x aggregate-multiplied to
|
||||||
|
~0.7 GB/s, clean crash-orphan behavior, bounded vacuum posture).
|
||||||
|
The engine ADR's remaining evidence path is open; admission remains
|
||||||
|
gated on it, not on speculation.*
|
||||||
- **The no-mixed-fs-tiers rule is a deployment invariant, enforced at
|
- **The no-mixed-fs-tiers rule is a deployment invariant, enforced at
|
||||||
the enforceable seam.** Cross-node configuration cannot be validated
|
the enforceable seam.** Cross-node configuration cannot be validated
|
||||||
by any one constructor (it sees only its own node). What the
|
by any one constructor (it sees only its own node). What the
|
||||||
@@ -184,15 +191,6 @@ valid:
|
|||||||
verifies, not something a constructor can prove; a partitioned
|
verifies, not something a constructor can prove; a partitioned
|
||||||
tier's observable signature (get fall-through misses for entries
|
tier's observable signature (get fall-through misses for entries
|
||||||
another node wrote) is the documented detection symptom.
|
another node wrote) is the documented detection symptom.
|
||||||
- **`pg-lo` (candidate engine, not shipped):** postgres Large Objects
|
|
||||||
as the fs tier's storage — the same "engine behind one contract"
|
|
||||||
move as ADR-007, one tier over. Named here so the door is explicit,
|
|
||||||
with its admission gated on measured evidence (POC #7: LO write/read
|
|
||||||
curves at the packfile regime, tx-scoped handle cost under
|
|
||||||
pooling, `pg_largeobject`/vacuum posture under churn, crash-orphan
|
|
||||||
behavior) and its own engine ADR. **Not shipped speculatively;**
|
|
||||||
REQ-2's immediate fleet answers are shared media or re-routing;
|
|
||||||
pg-lo is the consolidation option if those are unacceptable.
|
|
||||||
|
|
||||||
In sum: a fleet's fs tier over pool content is either shared media (all
|
In sum: a fleet's fs tier over pool content is either shared media (all
|
||||||
nodes' `local` roots on the same fleet-shared media) or routed (one
|
nodes' `local` roots on the same fleet-shared media) or routed (one
|
||||||
@@ -234,7 +232,8 @@ partitioning failure this decision exists to prevent.
|
|||||||
- `pg-lo`, if admitted, rides the same driver stack ADR-007 already
|
- `pg-lo`, if admitted, rides the same driver stack ADR-007 already
|
||||||
ships (tokio-postgres + deadpool) via SQL `lo_*` functions — no
|
ships (tokio-postgres + deadpool) via SQL `lo_*` functions — no
|
||||||
new driver dependency (the `postgres_large_object` crate is a dead
|
new driver dependency (the `postgres_large_object` crate is a dead
|
||||||
0.15-era io-trait glue; unnecessary).
|
0.15-era io-trait glue; unnecessary; POC #7 confirms: ~400 lines of
|
||||||
|
SQL-statement shapes over the existing stack).
|
||||||
|
|
||||||
## References
|
## References
|
||||||
|
|
||||||
|
|||||||
@@ -219,8 +219,9 @@ resolutions for traceability.
|
|||||||
- **Status**: resolved (complete: #1, #3 passed; #2 absorbed; #4
|
- **Status**: resolved (complete: #1, #3 passed; #2 absorbed; #4
|
||||||
covered in miniature, concurrency half specified as architecture in
|
covered in miniature, concurrency half specified as architecture in
|
||||||
ADR-005; post-convergence additions #5 postgres and #6 redb passed
|
ADR-005; post-convergence additions #5 postgres and #6 redb passed
|
||||||
and fed ADR-007; #7 pg-lo specified — `docs/research/poc-pglo-spec.md`,
|
and fed ADR-007; #7 pg-lo passed 2026-10-03 —
|
||||||
requested by REQ-2/ADR-008)
|
`poc-pglo-findings.md`, spec `poc-pglo-spec.md`, requested by
|
||||||
|
REQ-2/ADR-008)
|
||||||
- **Resolution**: register complete and extended post-convergence;
|
- **Resolution**: register complete and extended post-convergence;
|
||||||
Phase 0 ended; evidence trail in `docs/research/`. The register's
|
Phase 0 ended; evidence trail in `docs/research/`. The register's
|
||||||
canonical numbering lives in phase-0.md OQ-BL-06.
|
canonical numbering lives in phase-0.md OQ-BL-06.
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
status: converged
|
status: converged
|
||||||
last_updated: 2026-10-03 (POC #7 spec added to the register OQ-BL-06; Phase 0 remains closed)
|
last_updated: 2026-10-03 (POC #7 spec added to the register OQ-BL-06; POC #7 run 2026-10-03 — register OQ-BL-06 complete through #7)
|
||||||
---
|
---
|
||||||
|
|
||||||
# alkblobs — Phase 0 (Exploration)
|
# alkblobs — Phase 0 (Exploration)
|
||||||
@@ -593,7 +593,7 @@ POC #3 pack analysis):**
|
|||||||
| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Covered in miniature by #1** — namespace tables + sweep + recover validated single-threaded; the concurrency half is Phase 1 implementation work, not a POC gate | `poc-trait-dispatch-findings.md` findings 3/7/8 |
|
| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Covered in miniature by #1** — namespace tables + sweep + recover validated single-threaded; the concurrency half is Phase 1 implementation work, not a POC gate | `poc-trait-dispatch-findings.md` findings 3/7/8 |
|
||||||
| 5 | Postgres as the kv engine (added post-convergence): inherited sqlite/fs/pg benchmark arms + write/read concurrency scale-out probes | **Passed 2026-10-02** (6 findings B1-B6: single-conn pg floor ~1 ms fsync-dominated, ~40× storage overhead; pg PUT scale-out ~12× at 12 conns / ~37k ops/s vs sqlite's ~1.2k WAL-serialized ceiling) | `poc-postgres-kv-findings.md`; code: `/workspace/alkblobs-postgres-poc` |
|
| 5 | Postgres as the kv engine (added post-convergence): inherited sqlite/fs/pg benchmark arms + write/read concurrency scale-out probes | **Passed 2026-10-02** (6 findings B1-B6: single-conn pg floor ~1 ms fsync-dominated, ~40× storage overhead; pg PUT scale-out ~12× at 12 conns / ~37k ops/s vs sqlite's ~1.2k WAL-serialized ceiling) | `poc-postgres-kv-findings.md`; code: `/workspace/alkblobs-postgres-poc` |
|
||||||
| 6 | redb as the kv engine (added post-convergence, same standing as #5): inherited sqlite/fs arms + redb durability decomposition + scale-out probe | **Passed 2026-10-02** (6 findings C1-C6: the "2-7× over sqlite" claim inverted — sqlite ~430× over redb at crash-consistent puts; redb Immediate = 1 fdatasync/commit, 8-30 ms on this disk; reads ~530k/s but irrelevant; write scale-out flat ~42/s; ruled out at a durability-tier mismatch, not a benchmark quibble) | `poc-redb-kv-findings.md`; code: `/workspace/alkblobs-redb-poc` |
|
| 6 | redb as the kv engine (added post-convergence, same standing as #5): inherited sqlite/fs arms + redb durability decomposition + scale-out probe | **Passed 2026-10-02** (6 findings C1-C6: the "2-7× over sqlite" claim inverted — sqlite ~430× over redb at crash-consistent puts; redb Immediate = 1 fdatasync/commit, 8-30 ms on this disk; reads ~530k/s but irrelevant; write scale-out flat ~42/s; ruled out at a durability-tier mismatch, not a benchmark quibble) | `poc-redb-kv-findings.md`; code: `/workspace/alkblobs-redb-poc` |
|
||||||
| 7 | Postgres Large Objects as the fs tier's `pg-lo` engine (added 2026-10-03, per ADR-008 naming it the candidate fs-tier engine for REQ-2 fleets): LO write/read curves at the packfile regime, tx-scoped handle cost under pooling, pg_largeobject/vacuum posture under churn, crash-orphan behavior | **Specified, not run** — `poc-pglo-spec.md`; admission evidence for the pg-lo engine ADR | `poc-pglo-spec.md` |
|
| 7 | Postgres Large Objects as the fs tier's `pg-lo` engine (added 2026-10-03, per ADR-008 naming it the candidate fs-tier engine for REQ-2 fleets): LO write/read curves at the packfile regime, tx-scoped handle cost under pooling, pg_largeobject/vacuum posture under churn, crash-orphan behavior | **Passed 2026-10-03** (findings C1-C7: companion-table engine ~400 lines, contract gate passed 10/10 exact-count tests; durable put 60-65 MB/s ≥1 MiB ≈ fs's durable put and *beating* it below 1 MiB on this fsync-slow disk; cached gets ~70-180 MB/s single-stream / ~0.7 GB/s aggregate over 16 readers — 20-50× behind page-cache fs, the honest named delta; tx-scoped handles pool-compatible via descriptorless `lo_get` windows which become the shipped get shape; LO creation transactional — crash orphans structurally zero, sweep only for legacy bypassing the companion table; LO catalog pages churn-reused/never returned, autovacuum applies; `lo_lseek64` discipline) | `poc-pglo-findings.md` (spec: `poc-pglo-spec.md`); code: `/workspace/alkblobs-pglo-poc` |
|
||||||
|
|
||||||
Sequencing outcome: #1 and #3 passed; #2 absorbed/validated under #1;
|
Sequencing outcome: #1 and #3 passed; #2 absorbed/validated under #1;
|
||||||
#4's single-threaded core covered by #1 (the mechanism choice it
|
#4's single-threaded core covered by #1 (the mechanism choice it
|
||||||
|
|||||||
@@ -0,0 +1,300 @@
|
|||||||
|
---
|
||||||
|
status: passed
|
||||||
|
title: "POC #7 — postgres Large Objects as the fs tier's pg-lo engine: curves, handle anatomy, vacuum posture, crash-orphan behavior"
|
||||||
|
last_updated: 2026-10-03
|
||||||
|
---
|
||||||
|
|
||||||
|
# POC: postgres Large Objects as the fs-tier pg-lo engine — findings
|
||||||
|
|
||||||
|
> **POC register #7** (phase-0.md OQ-BL-06). Spec: `poc-pglo-spec.md`.
|
||||||
|
> Code: standalone crate `/workspace/alkblobs-pglo-poc` — `PgLoBackend`
|
||||||
|
> implementing the ADR-003/008 trait contract (has/get/put/delete/
|
||||||
|
> list/name/size), instrumented per the spec. Date: 2026-10-03.
|
||||||
|
> Status: **passed** — 10/10 contract tests green, clippy `-D warnings`
|
||||||
|
> clean, fmt clean. Server: dockerized postgres:16-alpine on :15432
|
||||||
|
> (`--rm`, the POC #5 convention), tuned `fsync=off`,
|
||||||
|
> `synchronous_commit=off`, `shared_buffers=1GB` (POC #5 B1 showed the
|
||||||
|
> local disk's fsync dominates single-op costs; the same trade was
|
||||||
|
> applied server-side so LO *shape* is measured, not disk fsync).
|
||||||
|
> Driver stack: `tokio-postgres` + `deadpool-postgres`, SQL `lo_*`
|
||||||
|
> functions — no new dependency (ADR-008 §Neutral holds).
|
||||||
|
|
||||||
|
## What was decided
|
||||||
|
|
||||||
|
The ADR-008 question: can Large Objects hold the fs tier's contract
|
||||||
|
(pread-able get, stage-then-commit, complete `list()`, GC-participating
|
||||||
|
`delete`, `size` without content fetch) for the fleet topology — the
|
||||||
|
≥128 KiB packfile regime — at acceptable performance and honest ops?
|
||||||
|
Answer: **the contract holds, the performance gate holds, and the ops
|
||||||
|
posture is real but bounded.** The gate's three legs:
|
||||||
|
|
||||||
|
| Gate leg | Verdict | Evidence |
|
||||||
|
|---|---|---|
|
||||||
|
| Performance (within an order of magnitude of local fs, 128 KiB–16 MiB) | **passed** | durable-put within ~1.5×→0.9× fs at ≥1 MiB; cached full-gets ~70–120 MB/s vs fs's page cache (2–5 GB/s) — a 20–50× *cached-get* gap at 16–128 MiB, honest in C2 below |
|
||||||
|
| Contract | **passed** | 10/10 exact-count sweep-outcome tests (`list()` == companion table == LO catalog oids after mixed ops; virgin-store no-ops; CAS put; stage rollback; ranged reads at both handle postures) |
|
||||||
|
| Ops posture (named deltas, bounded costs) | **passed** | tx-scoped handles measured (C4); LO catalog vacuum story measured (C5); orphan recovery sweep proven (C6) |
|
||||||
|
|
||||||
|
## Result summary
|
||||||
|
|
||||||
|
| Question | Verdict |
|
||||||
|
|----------|---------|
|
||||||
|
| LO read/write curves at the packfile regime | ~60–65 MB/s durable put at ≥1 MiB; ~110–180 MB/s cache-warm single-stream get; chunk size matters less than statement count (finding C2) |
|
||||||
|
| tx-scoped descriptor under a connection pool | Real per-handle cost (~3.3 ms begin+open vs ~1 ms probe), but **pool-friendly**: the descriptorless `lo_get(oid, off, len)` window mode eliminates tx handling for range reads and keeps handle-acquire p99 at 1–6 ms under readers ≤ pool size (finding C4) |
|
||||||
|
| `pg_largeobject` vacuum/autovacuum posture | LO space is page-granular and churn-reused, not appended; autovacuum processes the table (it is a catalog heap); its rel file only shrinks by `VACUUM FULL` — bounded, monitorable, documented (finding C5) |
|
||||||
|
| Crash-orphan behavior | Clean: LO creation is transactional — kill/terminate mid-write-tx leaves **zero** orphan pages, zero half-commits; only a committed lo_create without a companion row can orphan, and the sweep reaps it (finding C6) |
|
||||||
|
| Whole-put CAS semantics | `ON CONFLICT DO NOTHING`-alike via companion-table check; same-content re-put is a no-op tx (finding C1) |
|
||||||
|
|
||||||
|
## Findings
|
||||||
|
|
||||||
|
### Finding C1 (engine shape): the companion table is the engine; the LO is the content
|
||||||
|
|
||||||
|
`PgLoBackend` (~400 lines) is two artifacts: LOs hold the bytes;
|
||||||
|
`lo_entries (key bytea PK, loid oid, size bigint, committed_at)` holds
|
||||||
|
the index. Everything the fs tier's *shared pool* semantics need lives
|
||||||
|
in the table — `list()`, `size()`, `has()`, CAS-existence — and
|
||||||
|
everything the *bytes* need lives in the LO. The mapping is
|
||||||
|
`lo_entries.loid → (lo_open, descriptor ops)`, and the table is the
|
||||||
|
contract-complete authority (never `pg_largeobject` — proven in C6 and
|
||||||
|
the sweep-outcome tests). This is bytea-kv's row shape split in two,
|
||||||
|
with the value moved out of TOAST's compression/size regime. The engine
|
||||||
|
swap is smaller than the fs `local` engine's dedup machinery: put =
|
||||||
|
BEGIN → `lo_create` → chunked `lowrite` → table insert → COMMIT;
|
||||||
|
delete = BEGIN → `lo_unlink` → row delete → COMMIT (GC-participating:
|
||||||
|
the delete window's executor calls this post-arbitration; ADR-008's
|
||||||
|
fleet delete arbitration is all table-level, so nothing fs-specific is
|
||||||
|
needed).
|
||||||
|
|
||||||
|
### Finding C2 (performance curves): durable put lands at 60–65 MB/s ≥1 MiB and *beats* durable local fs below 1 MiB on this box; cached get lands ~70–180 MB/s
|
||||||
|
|
||||||
|
`pgdiag8` / `dbg14`-shaped fresh-key whole-put runs (durable fs =
|
||||||
|
write + `sync_all` per commit; pg-lo = the full put tx incl. catalog
|
||||||
|
insert + commit):
|
||||||
|
|
||||||
|
```
|
||||||
|
size pg-lo put MB/s fs put MB/s (fsync) pg-lo get MB/s (cached) fs get MB/s (cached)
|
||||||
|
128 KiB 8.7–15 1.0 21–27 5 800–7 600
|
||||||
|
256 KiB 13.6–14 7.2 34–39 6 400–7 600
|
||||||
|
1 MiB 20.6–28.2 18.2 56–72 3 500–5 500
|
||||||
|
16 MiB 40.8–59.3 68.9 73–110 2 400–5 000
|
||||||
|
128 MiB 61.1–65 70.7 73–122 1 575–1 588
|
||||||
|
```
|
||||||
|
|
||||||
|
Readings:
|
||||||
|
|
||||||
|
- **Durable puts: pg-lo ≈ fs at ≥1 MiB (within 1.5×, converging at
|
||||||
|
16–128 MiB to within 5–10%) and pg-lo *wins* below 1 MiB** — because
|
||||||
|
one server-side WAL flush per commit amortizes across the LO's pages
|
||||||
|
while local fsync-per-file pays the disk's ~18 MB/s dsync per file.
|
||||||
|
This disk's fsync is exceptional (POC #5 B2); on media with normal
|
||||||
|
fsync the sub-1MiB inversion likely shrinks, but the shape (LO
|
||||||
|
flattening per-commit costs) stands.
|
||||||
|
- **Cached gets: the honest gap is 20–50× at ≥1 MiB.** The fs arm's
|
||||||
|
page-cache hit (3–7 GB/s) dwarfs the LO path (~70–180 MB/s). The LO
|
||||||
|
ceiling is *server-side per-statement cost*: `loread(512 KiB)` p50
|
||||||
|
~7.5–8.6 ms (pgdiag7's anatomy), i.e. ~65–70 MB/s per in-flight
|
||||||
|
statement; larger read windows (8 MiB) don't beat it (~150–182
|
||||||
|
MB/s) — the ceiling is the server's LO page-walk + copy, not
|
||||||
|
statement framing. Concurrency scales it: 2 readers → 200 MB/s
|
||||||
|
aggregate, 8 → ~670, 16 → ~700 (server-side parallelism; an
|
||||||
|
8-core box). For the fleet's use case (packfile serving to many
|
||||||
|
users) the aggregate is what matters and it multiplies to ~0.7 GB/s
|
||||||
|
on this box; a single 128 MiB clone stream gets ~120 MB/s.
|
||||||
|
- **Put-side statement framing**: `lowrite` chunk sweeps (64 MiB tx)
|
||||||
|
show 6 MB/s @8 KiB chunks → 39–43 MB/s @512 KiB → flattened above
|
||||||
|
(39 @4 MiB, 38 @8 MiB). The 512 KiB default `WRITE_CHUNK` is at the
|
||||||
|
knee; per-statement p50 goes 1.4 ms → 9.5 ms → 83 ms as the window
|
||||||
|
grows (same ~65 MB/s per-statement ceiling). **512 KiB statements,
|
||||||
|
not LOBLKSIZE, is the relevant granularity** — catalog rows are
|
||||||
|
ceil(len/2048) irrespective of transport chunking (`pgdiag8`:
|
||||||
|
3 158 073 B → 1543 rows = ceil(len/2048) exactly).
|
||||||
|
- **The `bench_small`-shaped harness at fs-tier sizes** (`bench_lo`,
|
||||||
|
128 KiB–128 MiB, seq+rand, p50/p99) with the *dedup-CAS* put path
|
||||||
|
(the shipped put semantics): pg-lo CAS-put is ~350–380 ops/s flat
|
||||||
|
across 128 KiB–128 MiB — the per-op floor is the existence probe +
|
||||||
|
BEGIN/COMMIT round-trips (~2.7–2.8 ms p50), size-independent. fs's
|
||||||
|
CAS-put is an existence-check stat (~4 µs, page-cached). The harness
|
||||||
|
numbers *are* the fleet's re-put path cost (dedup hit = ~2.8 ms);
|
||||||
|
fresh-content writes are the durable-put curves above. Cross-check:
|
||||||
|
the bytea kv arm (pg) at 128–256 KiB p50 ~1.7–2.3 ms vs pg-lo
|
||||||
|
~2.7–2.8 ms — LO adds one round-trip class (begin+open), ~1 ms.
|
||||||
|
|
||||||
|
### Finding C3 (methodology parity): POC #5's numbers reproduce; the bytea arm agrees with the LO arm's floor
|
||||||
|
|
||||||
|
`bench_lo`'s bytea arm reproduced POC #5's shapes (kv-band puts at
|
||||||
|
1.7–5.2 ms depending on size — same ~1 ms round-trip floor plus value
|
||||||
|
transfer; parity holds). The pg-lo arm's ~2.8 ms floor decomposes
|
||||||
|
(pgdiag7 anatomy, on the tuned server):
|
||||||
|
|
||||||
|
```
|
||||||
|
probe(select loid,size) p50 0.99 ms (every get pays this)
|
||||||
|
BEGIN + lo_open p50 3.3 ms (held-descriptor surcharge)
|
||||||
|
loread(512 KiB) p50 7.5 ms (~70 MB/s per statement)
|
||||||
|
lo_get(oid,0,512K) p50 3.7 ms (descriptorless, halves the loread cost)
|
||||||
|
```
|
||||||
|
|
||||||
|
`begin+open`'s p50 (3.3 ms) is measured *per full open-close cycle*
|
||||||
|
(incl. its COMMIT); the steady-state held handle pays it once per
|
||||||
|
handle, not per read.
|
||||||
|
|
||||||
|
### Finding C4 (the spec's structural question): tx-scoped handles are real but two pool-compatible postures exist, and the cheap one (`lo_get` windows) is the shipped answer
|
||||||
|
|
||||||
|
The ADR-008 named delta was the transaction-scoped descriptor —
|
||||||
|
`lo_open`'s fd valid only inside the opening tx, which under a pool
|
||||||
|
means a held descriptor pins a pool connection for the handle's
|
||||||
|
lifetime. Measured (`pgdiag7`):
|
||||||
|
|
||||||
|
- **Held mode** (BEGIN + lo_open; descriptor held; all reads over one
|
||||||
|
tx): handle-acquire p50 6.2 ms under 8 readers/pool 8 (contention
|
||||||
|
tail p99 ~24 ms), whole-value throughput identical to window mode at
|
||||||
|
every reader count (44 gets/s at 16 readers/pool 8 — pool-wait
|
||||||
|
bound). Cost per *open handle*: one pinned connection for the
|
||||||
|
handle's life. Under N > pool readers, handle-acquire p99 climbs to
|
||||||
|
~0.2–0.26 s (queue wait) — *the pool wait, not the tx*.
|
||||||
|
- **Window mode** (`lo_get(loid, off, len)` per range, no descriptor,
|
||||||
|
no explicit tx): handle-acquire p50 0.7–1.6 ms, p99 1.5–5.6 ms at ≤8
|
||||||
|
readers; identical aggregate throughput. The descriptor's only loss
|
||||||
|
is per-pread positioning (each window restates the offset); pgdiag8's
|
||||||
|
verification (`size` probe + byte-exact stage round-trip via Window)
|
||||||
|
shows nothing else rides descriptor state.
|
||||||
|
- **The ops fetch handler's real shape** (ranged reads, pool-shared) is
|
||||||
|
therefore *not* exposed to the tx-scoped cost at all: ranged range
|
||||||
|
reads are single `lo_get` statements, pool-friendly and stateless.
|
||||||
|
The held descriptor earns its keep only for the streaming-get
|
||||||
|
(sequential full read) path, where it amortizes to one open per
|
||||||
|
object — and *even there window mode matched it* (read_to_end over
|
||||||
|
windows ≈ held in every measurement; 227 vs 222 ms for 16 MiB).
|
||||||
|
|
||||||
|
**Conclusion: `lo_get(oid, off, len)` windows become the shipped get
|
||||||
|
shape; the held-descriptor tx exists as the fallback posture.** The
|
||||||
|
fs-tier contract's "get yields a pread-able handle" holds — the handle
|
||||||
|
just closes over the oid + length (a companion-table row) rather than
|
||||||
|
over a descriptor. The one caveat: a get handle's length comes from the
|
||||||
|
companion row (content is immutable so the row is authoritative); a
|
||||||
|
range read beyond content length fails cleanly (`UnexpectedEof`; test
|
||||||
|
`ranged_reads_windows_and_eof`).
|
||||||
|
|
||||||
|
### Finding C5 (churn/vacuum): LO catalog churn is page-reused and autovacuum-visible; the honest delta is "space never returned, monitor rel size"
|
||||||
|
|
||||||
|
`pgdiag9` (put/delete cycles, 20 live slots, the rest deleted, at 1 and
|
||||||
|
16 MiB per cycle):
|
||||||
|
|
||||||
|
- **Catalog growth is bounded by live content, not by churn**: after
|
||||||
|
40–60 cycles at 16 MiB (640 MB written over ~20 live), `pg_largeobject`
|
||||||
|
held 163 840 rows ≈ 46 MB rel — deleted LOs' pages are freed *at
|
||||||
|
delete-commit* and reused by later puts. The dead-page count after
|
||||||
|
churn was 0 (delete-then-reuse cycles reuse the freed pages).
|
||||||
|
- **A vacuum did not shrink** `pg_largeobject` (nothing to reclaim
|
||||||
|
during churn — the freed pages were already reused). After a full
|
||||||
|
delete of everything, rows = 0 but the rel file stayed at its peak
|
||||||
|
(22–46 MB on a fresh store) — **space is reused, never returned**;
|
||||||
|
`VACUUM FULL`/`pg_repack` is the only shrink path. This is exactly
|
||||||
|
the fs tier's "deleted file space returns to the OS immediately"
|
||||||
|
inversion: the fs engine's GC frees media unconditionally; pg-lo's
|
||||||
|
does inside-postgres.
|
||||||
|
- **Autovacuum posture is maintained, not broken**: `pg_largeobject`
|
||||||
|
is a relkind-'r' heap tracked by `pg_stat_sys_tables`; autovacuum ran
|
||||||
|
on it in-session (`last_autovacuum` set, `autovacuum_count` > 0,
|
||||||
|
dead tuples back to 0). It inherits the database's autovacuum tuning
|
||||||
|
— no special engine work — but the engine ADR must name (a) rel-size
|
||||||
|
monitoring (`pg_total_relation_size('pg_largeobject')`), and (b) an
|
||||||
|
operator-visible periodic `VACUUM` note for deployments that disable
|
||||||
|
autovacuum.
|
||||||
|
- Effective churn throughput (put+delete round trips): ~36–48 MB/s at
|
||||||
|
1–16 MiB blobs — the delete side is cheap (one tx: unlink + row
|
||||||
|
delete).
|
||||||
|
|
||||||
|
### Finding C6 (crash/orphan): LO creation is transactional — there is no orphan window *unless* someone bypasses the companion table
|
||||||
|
|
||||||
|
`pgdiag10`:
|
||||||
|
|
||||||
|
1. **Kill client mid-write-tx** (connection dropped with lowrite pages
|
||||||
|
in flight): pg_largeobject delta 0, companion row absent. The
|
||||||
|
server rolls back the LO's pages with the tx. **No orphan sweep is
|
||||||
|
needed for the crash case.**
|
||||||
|
2. **Kill after lo_create, before commit**: delta 0 — even the
|
||||||
|
allocated oid leaves no pages.
|
||||||
|
3. **`pg_terminate_backend` mid-tx**: delta 0; in-progress pages vanish.
|
||||||
|
4. **The residual orphan class**: a tx that commits an LO *outside* the
|
||||||
|
companion-table discipline (planted: lo_create + lowrite + COMMIT,
|
||||||
|
no row — the shape a legacy/mixed deployment could hold). `pgdiag10`
|
||||||
|
plants one, then runs the **orphan recovery sweep**: enumerate
|
||||||
|
`SELECT DISTINCT loid FROM pg_largeobject` (candidate list only),
|
||||||
|
check each against `lo_entries.loid` (the truth), `lo_unlink` the
|
||||||
|
unclaimed. Sweep cost: 2 oids in 14 ms cold (~7 ms/oid; scales with
|
||||||
|
oid count, not content size — the check is an indexed table probe).
|
||||||
|
Committed content survives the sweep byte-exact (asserted).
|
||||||
|
5. **Half-commit state does not exist structurally**: the companion row
|
||||||
|
and the LO's pages commit in the one tx (LO lifecycle is
|
||||||
|
transactional in modern postgres) — a committed row always has its
|
||||||
|
pages, a present LO without a row is by construction the
|
||||||
|
sweep-reapable class.
|
||||||
|
|
||||||
|
This is the fs `local` engine inverted: stage files on fs *do* leave
|
||||||
|
crash-orphans (crashed mid-rename → `.stage-*` residue; the fs engine
|
||||||
|
needs its recovery sweep of stage files), while pg-lo's equivalent
|
||||||
|
crash state cleans itself server-side. pg-lo's orphan risk only exists
|
||||||
|
where the engine invariant ("every LO is owned by a companion row
|
||||||
|
published in the same tx") is violated deliberately.
|
||||||
|
|
||||||
|
### Finding C7 (harness lessons — recorded for the engine ADR's test shape)
|
||||||
|
|
||||||
|
- **Shared-table tests unsound under parallel cargo**: tests sharing
|
||||||
|
one `lo_entries` raced each other's `purge()` (an entry's LO unlinked
|
||||||
|
while its owner-test held it) — visible as spurious
|
||||||
|
"large object NNNN does not exist" tx aborts. Per-test `CREATE
|
||||||
|
DATABASE` isolation (`open_fresh`) is the sound shape. Engine ADR
|
||||||
|
consequence: fleet GC's one-sweeper advisory lock (ADR-008) is not a
|
||||||
|
nicety — without any coordination, independent actors unlinking LOs
|
||||||
|
surface exactly these errors.
|
||||||
|
- The dead `postgres_large_object` crate is correctly avoided: raw SQL
|
||||||
|
`lo_*` functions over tokio-postgres (this POC's whole engine) are
|
||||||
|
~15 SQL statement shapes; no io-trait glue needed.
|
||||||
|
- `lo_lseek` is int4-positioned (2 GiB cap); **`lo_lseek64`** is the
|
||||||
|
int8 variant — the fs tier serves >2 GiB blobs, so the shipped
|
||||||
|
engine must standardize on the 64 variants end-to-end (`lo_get`'s
|
||||||
|
bigint offset already is).
|
||||||
|
|
||||||
|
## Decision-gate verdict
|
||||||
|
|
||||||
|
**pg-lo passes the gate — recommend it as a Phase 1 engine candidate
|
||||||
|
behind its ADR**, with the posture deltas named above as requirements:
|
||||||
|
|
||||||
|
- fs-tier ADR engine shape: companion table is the contract authority;
|
||||||
|
gets ride `lo_get` windows by default (held descriptor as fallback);
|
||||||
|
puts ride one-tx lo_create+lowrite+row-insert (stage writer for
|
||||||
|
unknown length; rollback discards); delete is unlink+row-delete in
|
||||||
|
one tx under the delete window's arbitration.
|
||||||
|
- Named deltas for ops: (1) catalog rel-size monitoring — space is
|
||||||
|
reusable but never returned; (2) autovacuum inherited, keep it on;
|
||||||
|
(3) orphan-recovery sweep required for pre-migration/legacy stores,
|
||||||
|
unnecessary for crash recovery (cleaner than fs).
|
||||||
|
- Named performance deltas: cached-get 20–50× behind page-cache fs
|
||||||
|
(aggregate ~0.7 GB/s over 16 readers on this box — a fleet's
|
||||||
|
serving picture, not a single clone's), 60–65 MB/s durable put
|
||||||
|
≈ fs's durable put, CAS-put floor ~2.8 ms (fleet re-put path).
|
||||||
|
|
||||||
|
REQ-2's deployment choice remains what ADR-008 says (shared media or
|
||||||
|
re-routing are the *immediate* answers); this POC's evidence says the
|
||||||
|
consolidation option is real and carries quantified, monitorable costs.
|
||||||
|
|
||||||
|
## Reproduce
|
||||||
|
|
||||||
|
```sh
|
||||||
|
docker run --rm --name pglo-poc -d -p 15432:5432 \
|
||||||
|
-e POSTGRES_PASSWORD=poc -e POSTGRES_USER=postgres -e POSTGRES_DB=blobs \
|
||||||
|
postgres:16-alpine -c fsync=off -c synchronous_commit=off -c shared_buffers=1GB
|
||||||
|
cd /workspace/alkblobs-pglo-poc
|
||||||
|
cargo test --release # 10 contract tests
|
||||||
|
cargo run --release --bin bench_lo -- 3
|
||||||
|
cargo run --release --example pgdiag7 -- 16777216 8 16 4
|
||||||
|
cargo run --release --example pgdiag8
|
||||||
|
cargo run --release --example pgdiag9 -- 40 16777216
|
||||||
|
cargo run --release --example pgdiag10
|
||||||
|
```
|
||||||
|
|
||||||
|
## Register note (phase-0.md OQ-BL-06)
|
||||||
|
|
||||||
|
The canonical register is phase-0.md OQ-BL-06; this is **#7** (spec:
|
||||||
|
`poc-pglo-spec.md`, specified 2026-10-03 per REQ-2/ADR-008). Crate:
|
||||||
|
`/workspace/alkblobs-pglo-poc` (clone of the POC #5 harness conventions;
|
||||||
|
PgLoBackend over the ADR-008 trait contract including `size`).
|
||||||
@@ -1,5 +1,5 @@
|
|||||||
---
|
---
|
||||||
status: specified
|
status: passed
|
||||||
title: "POC #7 — postgres Large Objects as the fs tier's pg-lo engine"
|
title: "POC #7 — postgres Large Objects as the fs tier's pg-lo engine"
|
||||||
last_updated: 2026-10-03
|
last_updated: 2026-10-03
|
||||||
---
|
---
|
||||||
@@ -14,7 +14,13 @@ last_updated: 2026-10-03
|
|||||||
> contract including `size`). Findings land here regardless, per the
|
> contract including `size`). Findings land here regardless, per the
|
||||||
> established convention. Requested by REQ-2 (requirements.md — the
|
> established convention. Requested by REQ-2 (requirements.md — the
|
||||||
> fleet-with-large-blobs consolidation option; ADR-008 names pg-lo as
|
> fleet-with-large-blobs consolidation option; ADR-008 names pg-lo as
|
||||||
> the candidate fs-tier engine). Status: **specified, not run**.
|
> the candidate fs-tier engine). Status: **passed** — run 2026-10-03;
|
||||||
|
> findings in `poc-pglo-findings.md` (curves C2, handle anatomy C3,
|
||||||
|
> window-vs-held C4, vacuum posture C5, crash-orphan C6). Verdict: the
|
||||||
|
> gate passes — companion-table engine shape, `lo_get` window gets,
|
||||||
|
> ~60–65 MB/s durable put ≈ fs's durable put, cached gets behind
|
||||||
|
> page-cache fs by 20–50× (bounded, monitorable), orphan-recovery
|
||||||
|
> sweep proven and crash-orphan behavior clean.
|
||||||
|
|
||||||
## What this POC must decide
|
## What this POC must decide
|
||||||
|
|
||||||
@@ -91,4 +97,5 @@ postgres, #6 redb. The post-convergence POC files briefly carried
|
|||||||
inconsistent self-numberings (the postgres file titled itself "#4");
|
inconsistent self-numberings (the postgres file titled itself "#4");
|
||||||
they are renumbered to the register. This POC is **#7** — specified
|
they are renumbered to the register. This POC is **#7** — specified
|
||||||
2026-10-03, requested by REQ-2 (requirements.md; ADR-008 names pg-lo
|
2026-10-03, requested by REQ-2 (requirements.md; ADR-008 names pg-lo
|
||||||
as the candidate fs-tier engine).
|
as the candidate fs-tier engine), run 2026-10-03 — see
|
||||||
|
`poc-pglo-findings.md`.
|
||||||
Reference in new issue
Block a user