night 09-27/28: paperless 16 copy released on a real 18 dump (v0.276.0); demo-hp restore test refused for space -> R-701
gates / gates (push) Successful in 25s
gates / gates (push) Successful in 25s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,31 @@
|
||||
read at 2026-09-28T01:45:38Z
|
||||
conversion_copy null earlier null
|
||||
paperless-ngx_paperless_data
|
||||
paperless-ngx_paperless_postgres_data
|
||||
paperless-ngx_paperless_redis_data
|
||||
ls: cannot access '/mnt/sys_drive/felhom-data/backups/primary/paperless-ngx/db-dumps': No such file or directory
|
||||
*.sql:
|
||||
2026/09/28 01:30:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database)
|
||||
2026/09/28 01:30:45 dbdump.go:178: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=8e904f9d109a)
|
||||
2026/09/28 01:30:45 dbdump.go:896: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless")
|
||||
2026/09/28 01:30:45 dbdump.go:198: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless
|
||||
2026/09/28 01:30:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database)
|
||||
2026/09/28 01:35:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database)
|
||||
2026/09/28 01:35:45 dbdump.go:178: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=8e904f9d109a)
|
||||
2026/09/28 01:35:45 dbdump.go:896: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless")
|
||||
2026/09/28 01:35:45 dbdump.go:198: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless
|
||||
2026/09/28 01:35:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database)
|
||||
2026/09/28 01:40:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database)
|
||||
2026/09/28 01:40:45 dbdump.go:178: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=8e904f9d109a)
|
||||
2026/09/28 01:40:45 dbdump.go:896: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless")
|
||||
2026/09/28 01:40:45 dbdump.go:198: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless
|
||||
2026/09/28 01:40:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database)
|
||||
|
||||
2026/09/28 01:00:45 pgconvert.go:702: [INFO] [stacks] paperless-ngx: REMOVED the pre-conversion datadir copy paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z (PostgreSQL 16) — the converted app has a backup proven on 18: Tier 1 (own recovery unit) at 2026-09-28T00:30:00Z, its dump db-dumps/paperless-ngx-postgres.sql written 2026-09-28T00:30:00Z by map[paperless-postgres:postgres:18-alpine@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873 paperless-redis:redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 paperless-webserver:ghcr.io/paperless-ngx/paperless-ngx:2.20.15@sha256:6c86cad803970ea782683a8e80e7403444c5bf3cf70de63b4d3c8e87500db92f]
|
||||
unit=
|
||||
ls: cannot access '/db-dumps': No such file or directory
|
||||
*.sql:
|
||||
|
||||
/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql 2026-09-28 00:30:00.404026450 +0000 :: Dumped from database version 18.6
|
||||
/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx/db-dumps/pre-restore-20260927T094411Z-paperless-ngx-postgres.sql 2026-09-27 09:44:12.287984914 +0000 :: Dumped from database version 16.15
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
# Night 2026-09-27/28 — read at 01:45Z (03:45 CEST), read-only
|
||||
|
||||
**9202, paperless-ngx: the pre-conversion copy was released through v0.276.0's rewritten loop.** `N1-9202-paperless-release.txt`:
|
||||
- 00:30:00Z the nightly data run wrote `db-dumps/paperless-ngx-postgres.sql`; the file itself says "Dumped from database
|
||||
version 18.6" (the older `pre-restore-…` undo dump says 16.15 and was, correctly, not the one used).
|
||||
- 01:00:45Z the hourly job logged: „REMOVED the pre-conversion datadir copy
|
||||
paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z (PostgreSQL 16) — … Tier 1 (own recovery unit) at
|
||||
2026-09-28T00:30:00Z, its dump … written 2026-09-28T00:30:00Z by … postgres:18-alpine@sha256:77f5…".
|
||||
- After: the volume is gone from `docker volume ls`; `conversion_copy` and `earlier_conversion_copies` are both empty.
|
||||
|
||||
Before v0.275.0 this copy went at the first tick after a refresh, on a 16 dump (R-696). Here it waited for a real 18 dump.
|
||||
|
||||
**demo-hp restore test** (the earlier session's D1 item, read here): see `version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`
|
||||
— the pick is right (R-689 closed correctly), but the test is refused for space every cycle → R-701 filed.
|
||||
@@ -0,0 +1,12 @@
|
||||
# demo-hp restore-test cycles after agent 0.137.0 (restart 2026-09-27 14:13:17 CEST) — read 2026-09-28T01:46:09Z
|
||||
felhom-agent 0.137.0
|
||||
Sep 27 14:13:16 demo-hp felhom-agent[1319775]: time=2026-09-27T14:13:16.508+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled"
|
||||
Sep 27 14:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T14:13:17.506+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s
|
||||
Sep 27 20:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:17.836+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/felhom-golden-0.236.0.tar.zst size_bytes=654115664 reason="not a backup of a guest (the storage reports no vmid)"
|
||||
Sep 27 20:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:17.845+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst size_bytes=656970239 reason="guest 9100 does not exist on this node or is not one this agent manages"
|
||||
Sep 27 20:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:17.853+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different one)"
|
||||
Sep 27 20:13:19 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:19.466+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620981 avail_bytes=22657356320 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
Sep 27 20:13:19 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:19.467+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
Sep 28 02:13:17 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:17.870+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different one)"
|
||||
Sep 28 02:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:19.630+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=22657356320 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
Sep 28 02:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:19.630+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB"
|
||||
@@ -812,6 +812,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** |
|
||||
| **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** |
|
||||
| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |
|
||||
| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. | **OPEN — P3; owner: operator (which option), CC measures** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user