diff --git a/REPORT.md b/REPORT.md index d765e435..6cf9435f 100644 --- a/REPORT.md +++ b/REPORT.md @@ -3,7 +3,7 @@ Evidence: `documentation/audits/records-carried-2026-09-27/` (README = the session record). - `07` §6.6: new paragraph "What a restore keeps from the app it replaces" (and: a drive move is not a restore). -- Register: 336 → 336 rows (678,058 → 678,524 B). Opened R-700 (WATCHING, live proof open); closed R-697 (to +- Register: 336 → 337 rows. Opened R-700 (WATCHING, live proof open), R-701 (demo-hp restore test refused for space); closed R-697 (to `CLOSED-ITEMS.md`); R-691 annotated: (2) not built, and why. - Hub: no code change; global floor 0.276.0 (MinAgent 0.131.0) via `POST /configuration/global-floor`. - STATUS and CONTEXT updated. Capability map: not changed. diff --git a/STATUS.md b/STATUS.md index 70377b9b..1b13b8e5 100644 --- a/STATUS.md +++ b/STATUS.md @@ -17,6 +17,8 @@ **Evening follow-up (controller 0.276.0).** - **Fixed: moving an app to another drive made the box forget which version the app must run.** The app then took the newest version at its next start, and skipped the tested steps. Found by reading the code, not on a box. Tested in code only: no test box has two drives. - **Fixed: after a restore, the box forgot the old database copy, so it stayed on disk forever.** A restore now keeps that note, and also keeps whether you had the app switched on. +- **Seen in the night: the scratch box deleted paperless's old database copy by itself**, after a backup made by the new database version. This is correct. +- **Found in the night: the HP box's full-guest restore test can never run.** It now picks the right backup, but the disk has 21 GiB free and the test needs 31 GiB. It refuses safely every 6 hours. Filed, with three options; nothing decided. - **Not built: "Use my kept data" from the off-site copy.** It would be a new way to put data back, and no test box has an off-site copy to prove it on. **What broke, or is not done.** @@ -24,9 +26,8 @@ - **"Use my kept data" still cannot load from the off-site copy.** Filed. - **The file-browser group fix is tested in code only.** No nextcloud kept folder existed on the scratch box. - **A backup stores the name of each app version, not the app itself.** If a maker deletes an old version, a restore of it cannot start. Today all 42 versions the catalog names still exist. Filed with options; nothing decided. -- **Not watched yet:** the HP box converts paperless tonight; the scratch box releases its old paperless copy after tonight's backup. -**Rows.** 3 opened, 6 closed. The list went from 339 to 336. Evening: 1 opened, 1 closed; still 336. +**Rows.** 3 opened, 6 closed. The list went from 339 to 336. Evening and night: 2 opened, 1 closed; now 337. **Decision for you — D4: when the only backup holds data of an older app version, what does a restore bring back?** - **A — the older version with its own data; then the normal update climbs, one tested step at a time (I recommend this, and it is built).** Cost: after the restore the app runs an older version for a night or until someone presses Update. The page says so. diff --git a/documentation/audits/records-carried-2026-09-27/N/N1-9202-paperless-release.txt b/documentation/audits/records-carried-2026-09-27/N/N1-9202-paperless-release.txt new file mode 100644 index 00000000..95fb2c70 --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/N/N1-9202-paperless-release.txt @@ -0,0 +1,31 @@ +read at 2026-09-28T01:45:38Z +conversion_copy null earlier null +paperless-ngx_paperless_data +paperless-ngx_paperless_postgres_data +paperless-ngx_paperless_redis_data +ls: cannot access '/mnt/sys_drive/felhom-data/backups/primary/paperless-ngx/db-dumps': No such file or directory +*.sql: +2026/09/28 01:30:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database) +2026/09/28 01:30:45 dbdump.go:178: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=8e904f9d109a) +2026/09/28 01:30:45 dbdump.go:896: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless") +2026/09/28 01:30:45 dbdump.go:198: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless +2026/09/28 01:30:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database) +2026/09/28 01:35:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database) +2026/09/28 01:35:45 dbdump.go:178: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=8e904f9d109a) +2026/09/28 01:35:45 dbdump.go:896: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless") +2026/09/28 01:35:45 dbdump.go:198: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless +2026/09/28 01:35:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database) +2026/09/28 01:40:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-webserver (image=ghcr.io/paperless-ngx/paperless-ngx:2.20.15, not a database) +2026/09/28 01:40:45 dbdump.go:178: [DEBUG] DiscoverDatabases: found postgres container: paperless-postgres (id=8e904f9d109a) +2026/09/28 01:40:45 dbdump.go:896: [DEBUG] DiscoverDatabases: paperless-postgres → stack "paperless-ngx" from the compose project label (the container name would have given "paperless") +2026/09/28 01:40:45 dbdump.go:198: [DEBUG] DiscoverDatabases: paperless-postgres → stack=paperless-ngx, dbUser=paperless, dbName=paperless +2026/09/28 01:40:45 dbdump.go:171: [DEBUG] DiscoverDatabases: skipping container paperless-redis (image=redis:7-alpine, not a database) + +2026/09/28 01:00:45 pgconvert.go:702: [INFO] [stacks] paperless-ngx: REMOVED the pre-conversion datadir copy paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z (PostgreSQL 16) — the converted app has a backup proven on 18: Tier 1 (own recovery unit) at 2026-09-28T00:30:00Z, its dump db-dumps/paperless-ngx-postgres.sql written 2026-09-28T00:30:00Z by map[paperless-postgres:postgres:18-alpine@sha256:77f585114c32fbca283dc835b0596f4e52b51b4c6662d7810b2f4084f60a1873 paperless-redis:redis:7-alpine@sha256:858f009f9709ce576febc734aa78b8f6d624b82571f9ddb6bda4377c833b3499 paperless-webserver:ghcr.io/paperless-ngx/paperless-ngx:2.20.15@sha256:6c86cad803970ea782683a8e80e7403444c5bf3cf70de63b4d3c8e87500db92f] +unit= +ls: cannot access '/db-dumps': No such file or directory +*.sql: + +/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx/db-dumps/paperless-ngx-postgres.sql 2026-09-28 00:30:00.404026450 +0000 :: Dumped from database version 18.6 +/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx/db-dumps/pre-restore-20260927T094411Z-paperless-ngx-postgres.sql 2026-09-27 09:44:12.287984914 +0000 :: Dumped from database version 16.15 + diff --git a/documentation/audits/records-carried-2026-09-27/N/README.md b/documentation/audits/records-carried-2026-09-27/N/README.md new file mode 100644 index 00000000..07eb6c58 --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/N/README.md @@ -0,0 +1,14 @@ +# Night 2026-09-27/28 — read at 01:45Z (03:45 CEST), read-only + +**9202, paperless-ngx: the pre-conversion copy was released through v0.276.0's rewritten loop.** `N1-9202-paperless-release.txt`: +- 00:30:00Z the nightly data run wrote `db-dumps/paperless-ngx-postgres.sql`; the file itself says "Dumped from database + version 18.6" (the older `pre-restore-…` undo dump says 16.15 and was, correctly, not the one used). +- 01:00:45Z the hourly job logged: „REMOVED the pre-conversion datadir copy + paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z (PostgreSQL 16) — … Tier 1 (own recovery unit) at + 2026-09-28T00:30:00Z, its dump … written 2026-09-28T00:30:00Z by … postgres:18-alpine@sha256:77f5…". +- After: the volume is gone from `docker volume ls`; `conversion_copy` and `earlier_conversion_copies` are both empty. + +Before v0.275.0 this copy went at the first tick after a refresh, on a 16 dump (R-696). Here it waited for a real 18 dump. + +**demo-hp restore test** (the earlier session's D1 item, read here): see `version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt` +— the pick is right (R-689 closed correctly), but the test is refused for space every cycle → R-701 filed. diff --git a/documentation/audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt b/documentation/audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt new file mode 100644 index 00000000..9e68d9b6 --- /dev/null +++ b/documentation/audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt @@ -0,0 +1,12 @@ +# demo-hp restore-test cycles after agent 0.137.0 (restart 2026-09-27 14:13:17 CEST) — read 2026-09-28T01:46:09Z +felhom-agent 0.137.0 +Sep 27 14:13:16 demo-hp felhom-agent[1319775]: time=2026-09-27T14:13:16.508+02:00 level=INFO msg="backup: restore-test scheduler shutting down" reason="context canceled" +Sep 27 14:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T14:13:17.506+02:00 level=INFO msg="backup: restore-test scheduler starting (per-archive due-check)" eval_interval=6h0m0s settle=24h0m0s +Sep 27 20:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:17.836+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/felhom-golden-0.236.0.tar.zst size_bytes=654115664 reason="not a backup of a guest (the storage reports no vmid)" +Sep 27 20:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:17.845+02:00 level=INFO msg="backup: restore-test skips an entry that is not a backup of a guest" target=local volid=local:backup/vzdump-lxc-9100-2026_08_21-17_59_15.tar.zst size_bytes=656970239 reason="guest 9100 does not exist on this node or is not one this agent manages" +Sep 27 20:13:17 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:17.853+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different one)" +Sep 27 20:13:19 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:19.466+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620981 avail_bytes=22657356320 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB" +Sep 27 20:13:19 demo-hp felhom-agent[1370926]: time=2026-09-27T20:13:19.467+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB" +Sep 28 02:13:17 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:17.870+02:00 level=INFO msg="backup: restore-test tier is DUE (per-archive; oldest-proven first among due tiers)" target=felhom-pbs archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z landed=2026-09-24T20:06:25Z reason="newest settled archive (landed 2026-09-24T20:06:25Z) has not been proven (last proven archive was a different one)" +Sep 28 02:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:19.630+02:00 level=WARN msg="restore-test SKIPPED by the space preflight (R-672) — nothing was created" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z storage=local-lvm required_bytes=33252620982 avail_bytes=22657356320 reason="not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB" +Sep 28 02:13:19 demo-hp felhom-agent[1370926]: time=2026-09-28T02:13:19.630+02:00 level=ERROR msg="backup: scheduled restore-test FAILED" archive=felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z err="skipped: not enough space on local-lvm: restoring 21.6 GiB (pbs snapshot size) needs 31.0 GiB free, has 21.1 GiB" diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 61005d23..dd137081 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -812,6 +812,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** | | **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** | | **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** | +| **R-701** | **[P3-LOW] demo-hp's whole-guest restore test can never run: every 6 h it picks the right archive and the space preflight refuses it.** Read 2026-09-28 (agent 0.137.0, `audits/version-travel-2026-09-26/D1/D1-cycle-demo-hp.txt`): 20:13 and 02:13 CEST both skipped the golden file and the deleted guest 9100's archive (R-689 working), chose `felhom-pbs:backup/ct/9201/2026-09-24T20:06:25Z` (21.6 GiB), and were refused — "needs 31.0 GiB free, has 21.1 GiB" on `local-lvm` — logged `ERROR scheduled restore-test FAILED`. Same refusal first seen 2026-09-24 (R-672's delivery). So demo-hp's whole-guest tier is never proven, and the refusal is SAFE (nothing created). Not measured: whether each refusal reaches the hub or the operator as a failure. **Options (decide nothing yet):** (a) reclaim thin-pool space (`pct fstrim`, R-444) and see if 31 GiB frees; (b) restore-test into `nvme-scratch` instead of `local-lvm` — a config change on the host; (c) accept: demo-hp is a small box, record the tier as not testable there. | **OPEN — P3; owner: operator (which option), CC measures** |