From 06744dbaccfb0b6ec36fa9d816946fa195d07264 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 27 Sep 2026 18:04:00 +0200 Subject: [PATCH] =?UTF-8?q?records-carried:=20controller=20v0.276.0=20(R-6?= =?UTF-8?q?97=20closed,=20R-700=20filed),=2007=20=C2=A76.6,=20register=203?= =?UTF-8?q?36=20->=20336,=20STATUS?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 4 ++ REPORT.md | 16 +++---- STATUS.md | 10 ++-- .../architecture/07-backup-architecture.md | 2 + .../R/R1-9202-0.276.0.txt | 3 ++ .../records-carried-2026-09-27/R/R2-floor.txt | 8 ++++ .../R/R3-9202-paperless-record.txt | 12 +++++ .../records-carried-2026-09-27/README.md | 46 +++++++++++++++++++ .../RP1-restore-drops-conversion-copy.txt | 11 +++++ .../RP2-second-conversion-orphans-first.txt | 6 +++ .../redproofs/RP3-drive-move-unpins.txt | 6 +++ .../redproofs/RP4-wiring.txt | 13 ++++++ documentation/backlog/CLOSED-ITEMS.md | 1 + documentation/backlog/OPEN-ITEMS.md | 4 +- 14 files changed, 128 insertions(+), 14 deletions(-) create mode 100644 documentation/audits/records-carried-2026-09-27/R/R1-9202-0.276.0.txt create mode 100644 documentation/audits/records-carried-2026-09-27/R/R2-floor.txt create mode 100644 documentation/audits/records-carried-2026-09-27/R/R3-9202-paperless-record.txt create mode 100644 documentation/audits/records-carried-2026-09-27/README.md create mode 100644 documentation/audits/records-carried-2026-09-27/redproofs/RP1-restore-drops-conversion-copy.txt create mode 100644 documentation/audits/records-carried-2026-09-27/redproofs/RP2-second-conversion-orphans-first.txt create mode 100644 documentation/audits/records-carried-2026-09-27/redproofs/RP3-drive-move-unpins.txt create mode 100644 documentation/audits/records-carried-2026-09-27/redproofs/RP4-wiring.txt diff --git a/CONTEXT.md b/CONTEXT.md index 45e5a881..9ccf1978 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,10 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-27 evening — a restore and a drive move keep the app's records (controller v0.276.0) + +- **Controller v0.276.0** (floor 0.276.0, MinAgent 0.131.0): R-697 closed (the restore's write carries the life records; `earlier_conversion_copies`); R-700 filed + fixed (a drive move dropped the pin → the next start took the catalog's newest version) — live proof open, no Tier-0 guest has two drives. R-691 (2) deliberately not built (no off-site target on a box CC may touch). Rule added to `07` §6.6. Evidence `audits/records-carried-2026-09-27/`. + ## 2026-09-26/27 — a backup's data and its version travel together (controller v0.275.0), agent v0.135.0, catalog f1a7d6c - **Rulings recorded (`09` §3):** 40 (D3 A — the box never deletes kept data), 41 (adventurelog keeps its world-data diff --git a/REPORT.md b/REPORT.md index b041f2e3..d765e435 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,11 +1,9 @@ -# REPORT — 2026-09-26/27: version-travel brief (docs, register, evidence) +# REPORT — 2026-09-27 evening: a restore and a drive move keep the app's records -Record: `documentation/audits/DRILL-version-travel-2026-09-26.md`; evidence `documentation/audits/version-travel-2026-09-26/`. +Evidence: `documentation/audits/records-carried-2026-09-27/` (README = the session record). -- `09` §3: decisions 40, 41 (operator rulings), 42 (CC unattended). `07` §6.5 (kept view decision, decision 40) and new - §6.6 "Which version a restore brings back". STATUS rewritten (D4 in the decision section). CONTEXT updated. -- Register: 339 → 336 rows (682,159 → 678,058 B). Opened R-697, R-698, R-699; closed R-655, R-689, R-694, R-695, R-696, - R-699 (to `CLOSED-ITEMS.md`); narrowed R-691, R-463. -- Hub: no code change; global floor 0.275.0 (MinAgent 0.131.0) set via `POST /configuration/global-floor`. -- Capability map: not changed — no scenario's status moved (the restore routes existed; what changed is which version - they bring back, recorded in `07` §6.6). +- `07` §6.6: new paragraph "What a restore keeps from the app it replaces" (and: a drive move is not a restore). +- Register: 336 → 336 rows (678,058 → 678,524 B). Opened R-700 (WATCHING, live proof open); closed R-697 (to + `CLOSED-ITEMS.md`); R-691 annotated: (2) not built, and why. +- Hub: no code change; global floor 0.276.0 (MinAgent 0.131.0) via `POST /configuration/global-floor`. +- STATUS and CONTEXT updated. Capability map: not changed. diff --git a/STATUS.md b/STATUS.md index 07850f67..70377b9b 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,6 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-27. Both demo boxes run controller 0.275.0 and host agent 0.137.0. Hub 0.125.0.** +**Updated 2026-09-27 evening. Both demo boxes run controller 0.276.0 and host agent 0.137.0. Hub 0.125.0.** **Decisions I took on my own** (you may reverse each): 1. **paperless-ngx goes to PostgreSQL 18, tandoor to 17.** Each follows what its own makers ship. paperless's makers use 18. tandoor's use 16, and 17 is the newest its framework supports. @@ -14,15 +14,19 @@ - **The HP box's restore test now tests real backups of guests that exist**, not the golden template file and not an old backup of a deleted guest (two more small agent releases, each found when I checked the box). - **Small fixes:** the file browser no longer re-creates an empty kept folder; the "Kept data" name follows the box's language; after a load, the app page no longer shows a password that does not work. +**Evening follow-up (controller 0.276.0).** +- **Fixed: moving an app to another drive made the box forget which version the app must run.** The app then took the newest version at its next start, and skipped the tested steps. Found by reading the code, not on a box. Tested in code only: no test box has two drives. +- **Fixed: after a restore, the box forgot the old database copy, so it stayed on disk forever.** A restore now keeps that note, and also keeps whether you had the app switched on. +- **Not built: "Use my kept data" from the off-site copy.** It would be a new way to put data back, and no test box has an off-site copy to prove it on. + **What broke, or is not done.** - **Found and fixed today: an app installed minutes ago could be updated with no backup of its database.** The update trusted a backup that held only the app's settings. Now it backs up first. - **"Use my kept data" still cannot load from the off-site copy.** Filed. - **The file-browser group fix is tested in code only.** No nextcloud kept folder existed on the scratch box. - **A backup stores the name of each app version, not the app itself.** If a maker deletes an old version, a restore of it cannot start. Today all 42 versions the catalog names still exist. Filed with options; nothing decided. -- **After a restore, the box keeps an old database copy it can no longer release.** Disk only. Filed. - **Not watched yet:** the HP box converts paperless tonight; the scratch box releases its old paperless copy after tonight's backup. -**Rows.** 3 opened, 6 closed. The list went from 339 to 336. +**Rows.** 3 opened, 6 closed. The list went from 339 to 336. Evening: 1 opened, 1 closed; still 336. **Decision for you — D4: when the only backup holds data of an older app version, what does a restore bring back?** - **A — the older version with its own data; then the normal update climbs, one tested step at a time (I recommend this, and it is built).** Cost: after the restore the app runs an older version for a night or until someone presses Update. The page says so. diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index f73f9dbb..525c7ec3 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -636,6 +636,8 @@ both. Files written under different pins (a failed leg) make the unit `mixed`. | off-site (Tier 3) | present, the snapshot's version differs from what runs | the snapshot unit's definition is written into the stack dir (and pinned) right after the stop, before any file, volume or database is touched; the database service is resolved from it | | any | absent — a unit or snapshot written before v0.275.0, or with an unstamped file | restores **as before** (its captured definition; the off-site path the live one), with a WARN naming it | +**[DESIGN] What a restore keeps from the app it replaces (controller v0.276.0, R-697).** The restore writes a fresh `app.yaml` from the unit — env, locked fields, and the pin from the unit's definition — but keeps the records of the app's life on this box: the household's `desired_state`, the kept pre-conversion copies (`conversion_copy`, `earlier_conversion_copies` — a restore does not remove the volume, so it must not forget it), and the ladder history (`failed_update_step`, `last_update_undone`, `last_auto_update`). Not kept: `installed_images` (what ran before, maybe another version). **A drive move is not a restore:** it changes `HDD_PATH` and nothing else (R-700 — before v0.276.0 it dropped the pin, and the syncer then gave the app the catalog's newest version at its next start). Pinned by `internal/stacks/r700_records_carried_test.go`. + **[DESIGN] The jump guard.** An installed version that matches no ladder entry's `from` jumps to the catalog's current definition when a PERSON presses Update (`ladder.go`, measured live on vikunja 2.5.0, `09` §6.4 part 5); the AUTOMATIC leg never presses it (`LegSkipOlderThanLadder`, and a template without a ladder is diff --git a/documentation/audits/records-carried-2026-09-27/R/R1-9202-0.276.0.txt b/documentation/audits/records-carried-2026-09-27/R/R1-9202-0.276.0.txt new file mode 100644 index 00000000..41220534 --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/R/R1-9202-0.276.0.txt @@ -0,0 +1,3 @@ +# 9202 before: gitea.dooplex.hu/admin/felhom-controller:0.275.0 +# 9202 after: gitea.dooplex.hu/admin/felhom-controller:0.276.0 Up 25 seconds (healthy) +2026/09/27 16:00:45 main.go:337: [INFO] felhom-controller 0.276.0 starting (customer: demo-hp, domain: enkisfelhom.hu) diff --git a/documentation/audits/records-carried-2026-09-27/R/R2-floor.txt b/documentation/audits/records-carried-2026-09-27/R/R2-floor.txt new file mode 100644 index 00000000..1983ba70 --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/R/R2-floor.txt @@ -0,0 +1,8 @@ +# global floor 0.276.0 (declared MinAgent 0.131.0) — 2026-09-27T16:01:26Z +HTTP/1.1 303 See Other +Location: /configuration?flash=floor_set +2026-09-27T16:01:36Z +demo-hp: gitea.dooplex.hu/admin/felhom-controller:0.276.0 +demo-felhom: gitea.dooplex.hu/admin/felhom-controller:0.276.0 +gitea.dooplex.hu/admin/felhom-controller:0.276.0 Up 22 seconds (healthy) +gitea.dooplex.hu/admin/felhom-controller:0.276.0 Up 25 seconds (healthy) diff --git a/documentation/audits/records-carried-2026-09-27/R/R3-9202-paperless-record.txt b/documentation/audits/records-carried-2026-09-27/R/R3-9202-paperless-record.txt new file mode 100644 index 00000000..104b23be --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/R/R3-9202-paperless-record.txt @@ -0,0 +1,12 @@ +state running +pinned_images {"paperless-postgres": "postgres:18-alpine", "paperless-redis": "redis:7-alpine", "paperless-webserver": "ghcr.io/paperless-ngx/paperless-ngx:2.20.15"} +desired_state "running" +conversion_copy {"volume": "paperless-ngx_paperless_postgres_data", "copy": "paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z", "at": "2026-09-27T09:45:15Z", "from": 16, "to": 18, "service": "paperless-postgres"} +earlier_conversion_copies null +paperless-ngx_paperless_data +paperless-ngx_paperless_postgres_data +paperless-ngx_paperless_postgres_data.pre-update-20260927T094421Z +paperless-ngx_paperless_redis_data +2026/09/27 16:00:45 scheduler.go:102: [INFO] [scheduler] Registered periodic job: conversion-copy-release (every 1h0m0s) +2026/09/27 16:00:45 scheduler.go:67: [DEBUG] [scheduler] periodic job registered: name="conversion-copy-release" interval=1h0m0s totalJobs=10 + diff --git a/documentation/audits/records-carried-2026-09-27/README.md b/documentation/audits/records-carried-2026-09-27/README.md new file mode 100644 index 00000000..42851114 --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/README.md @@ -0,0 +1,46 @@ +# records-carried — 2026-09-27 (controller v0.276.0): a restore and a drive move keep the app's records + +Architecture read before any claim: `07-backup-architecture.md` §6.6 (now with "What a restore keeps"), `09` §6.4 part 10 +(the conversion copy). Rows: R-697 (closed), R-700 (new; fixed, live proof open), R-691 (2) (not built — why in the row). + +## Found + +- **R-697 (known):** `PersistUnitRedeployConfig` built a fresh `AppConfig`, so a restore dropped `conversion_copy` — the + kept 16 datadir volume was never released. Also dropped: `desired_state` (a dead app after a restore then reads as + "unknown intent" and is not alarmed), `failed_update_step`, `last_update_undone`, `last_auto_update`. +- **Second orphan (same row):** after a restore to the OLD major the ladder converts again; `recordConversionCopy` overwrote + the first copy's record with the second's. +- **R-700 (new, by reading, not seen on a box):** `doFlipRedeploy` (drive move) persisted through the same fresh write, so it + ALSO dropped `pinned_images`. `sync.renderSource`: deployed + unpinned → the catalog copied verbatim → the next `up` runs + the catalog's newest version. Pin adoption runs only at controller start. + +## Fixed (v0.276.0, controller `820e8ef`) + +`carryLifeRecords` in the restore write; `earlier_conversion_copies` + a release loop over every kept copy; +`persistDriveFlip` (load-then-save, `HDD_PATH` only) + `upFromAppConfig`. + +## Red-proofs (`redproofs/`, each seen failing, tree restored, suite green after) + +| | pre-fix shape | failed at | +|---|---|---| +| RP1 | no carry in the restore write | `conversion_copy = ` | +| RP2 | a newer record overwrites the older | `earlier []` | +| RP3 | the drive move persists through the restore write | `pinned_images = map[]` | +| RP4 | `doFlipRedeploy` back on `RedeployFromEnv` | the wiring test | + +## Live (endpoint level + guest reads) + +- `R/R1-9202-0.276.0.txt` — 9202 by hand, healthy. `R/R2-floor.txt` — global floor 0.276.0 (MinAgent 0.131.0), both demo + boxes on 0.276.0 within 10 s, healthy. +- `R/R3-9202-paperless-record.txt` — after the upgrade paperless-ngx keeps `conversion_copy` (16 → 18, at 09:45:15Z) and + its `.pre-update-20260927T094421Z` volume; the release job registered. +- **Night (N/):** the release on a real 18 dump — see `N/README.md` once written. + +## Not proven live + +- **R-700:** no Tier-0 guest has two drives (9202 has one: `scratch_hdd`). Unit-proven only; row stays WATCHING. +- **R-697's carry across a real restore:** restoring paperless now would bring it back at 16 (its unit's data is the + pre-conversion dump) and would spoil the night's release proof. Unit-proven only. + +Teardown: provisioned nothing. Machine: 9202 controller image changed 0.275.0 → 0.276.0 (kept). Host: nothing. Hub: global +floor 0.275.0 → 0.276.0. diff --git a/documentation/audits/records-carried-2026-09-27/redproofs/RP1-restore-drops-conversion-copy.txt b/documentation/audits/records-carried-2026-09-27/redproofs/RP1-restore-drops-conversion-copy.txt new file mode 100644 index 00000000..5ad5290c --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/redproofs/RP1-restore-drops-conversion-copy.txt @@ -0,0 +1,11 @@ +RP1: carryLifeRecords call removed from PersistUnitRedeployConfig (= v0.275.0 restore write) + controller/internal/stacks/deploy.go | 11 ++++ + controller/internal/stacks/migrate.go | 47 ++++++++++---- + controller/internal/stacks/pgconvert.go | 109 +++++++++++++++++++++++--------- + 3 files changed, 126 insertions(+), 41 deletions(-) +2026/09/27 17:49:58 [INFO] [stacks] SaveAppConfig: saved config for nextcloud +--- FAIL: TestR697_ARestoreKeepsTheConversionCopyRecordSoTheCopyIsReleased (0.01s) + r700_records_carried_test.go:43: after the restore conversion_copy = , want nextcloud_pgdata.pre-update-20260913T100000Z kept +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.014s +FAIL diff --git a/documentation/audits/records-carried-2026-09-27/redproofs/RP2-second-conversion-orphans-first.txt b/documentation/audits/records-carried-2026-09-27/redproofs/RP2-second-conversion-orphans-first.txt new file mode 100644 index 00000000..b05ed810 --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/redproofs/RP2-second-conversion-orphans-first.txt @@ -0,0 +1,6 @@ +RP2: recordConversionCopy overwrites the older record (= v0.275.0) +--- FAIL: TestR697_ASecondConversionDoesNotOrphanTheFirstCopy (0.00s) + r700_records_carried_test.go:74: records: current &{Volume:nextcloud_pgdata Copy:nextcloud_pgdata.pre-update-B At:2026-09-13T11:00:00Z From:16 To:18 Service:db}, earlier [] — want B current and A kept as earlier +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.010s +FAIL diff --git a/documentation/audits/records-carried-2026-09-27/redproofs/RP3-drive-move-unpins.txt b/documentation/audits/records-carried-2026-09-27/redproofs/RP3-drive-move-unpins.txt new file mode 100644 index 00000000..9957814f --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/redproofs/RP3-drive-move-unpins.txt @@ -0,0 +1,6 @@ +RP3: persistDriveFlip = PersistUnitRedeployConfig(env with HDD_PATH=target) (= v0.275.0's write, plus this release's carry) +--- FAIL: TestR700_ADriveMoveKeepsThePinAndTheRecords (0.00s) + r700_records_carried_test.go:109: pinned_images = map[] — the moved app is unpinned and will jump to the catalog's version +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.009s +FAIL diff --git a/documentation/audits/records-carried-2026-09-27/redproofs/RP4-wiring.txt b/documentation/audits/records-carried-2026-09-27/redproofs/RP4-wiring.txt new file mode 100644 index 00000000..443b2006 --- /dev/null +++ b/documentation/audits/records-carried-2026-09-27/redproofs/RP4-wiring.txt @@ -0,0 +1,13 @@ +RP4: doFlipRedeploy back on RedeployFromEnv + // the pin and every record. This used to be RedeployFromEnv, whose fresh app.yaml dropped the pin: the + // syncer then copied the catalog verbatim and the next start jumped the app past its ladder. + if err := m.RedeployFromEnv(name, nil); err != nil { // RP4 + return err + } + if !m.waitHealthy(name) { + return util.MsgError("err.stacks.az_alkalmazas_nem_indult_el_az") + } + return nil +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.008s +FAIL diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 5a3be786..6fcae516 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -419,3 +419,4 @@ Full original text of each row: `git show ^: | **R-694** | **A load/restore regenerated a withheld login and the page showed it as the password (P3).** Measured per app from each entrypoint: 6 of 7 keep the login in their data (code-server is the exception). v0.275.0: `restored_logins`; the page shows no value and says to use the password valid at the backup. | CLOSED 2026-09-27 — controller v0.275.0 | `audits/version-travel-2026-09-26/D4/` | | **R-689** | **demo-hp's restore test picked the golden template in `local:backup/` and failed every 6 h (P3).** Agent v0.135.0 + v0.136.0 + v0.137.0: only `vzdump--` files or PBS `ct/…` or `vm/…` snapshots with a reported vmid, of a guest that still EXISTS on the node (v0.136.0 — the first half, read live with `-selftest=restore-test-due`, fell to a leftover archive of a guest deleted in August; v0.137.0 — PVE answers 403, not "does not exist", for a guest outside the agent's pool, which 0.136.0 read as a lookup failure). Delivered to both demo hosts by signed `agent_update`. | CLOSED 2026-09-27 — agent v0.137.0 | `audits/version-travel-2026-09-26/D1/` | | **R-655** | **adventurelog v0.13.0 could not become healthy in the catalog's template (P2).** Operator ruling decision 41 (keep the download). Catalog `06ea7da`: the frontend override dropped; a cut-off world-data file set aside before start (a pending import crash-looped 10× without it). Proven on the bench and on 9202 (204 s under the default 5-min wait). | CLOSED 2026-09-27 — catalog `06ea7da` | `audits/version-travel-2026-09-26/C/` | +| **R-697** | **A restore dropped the `conversion_copy` record but not the kept pre-conversion volume — the copy was never released (P3).** Seen 2026-09-26 on 9202 (A1). v0.276.0: the restore's write carries the app's life records from the `app.yaml` it replaces (`carryLifeRecords`: conversion copies, `desired_state`, update history; not the pin, not `installed_images`); a second conversion after a restore to the old major keeps the first copy in `earlier_conversion_copies`, released by the same rule. Rule: a write that is not a new install is load-then-save, or it names every field it drops. Red-proofed RP1, RP2. | CLOSED 2026-09-27 — controller v0.276.0 | `audits/records-carried-2026-09-27/` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 53592c46..61005d23 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -808,10 +808,10 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | | **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). | **OPEN — P3, gaps (1)–(4) only; owner: CC** | | **R-688** | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** | -| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** | **OPEN — P3; owner: CC — NARROWED to (2)** | +| **R-691** | **[P3-LOW] Kept data (09 §3 decision 36): two gaps of the first build.** (1) **The read-only file-browser view cannot open a folder another user owns with mode 0770** — nextcloud's `appdata/nextcloud` is `www-data` `drwxrwx---` (measured on 9202 2026-09-25), FileBrowser runs as uid 1000, so „Megőrzött adatok" shows the folder and not its files; the files are still listed, sized, loadable and deletable. Fix direction needs a decision (a read-only ACL, or a helper that lists as root) — not a chmod of the household's data. (2) **„Use my kept data" / Load looks only at the own unit (Tier 1) and the second-drive mirror (Tier 2)**; an app whose only database copy is off-site gets "no backup". Controller `43e99d1`. `audits/night-2026-09-26/E/` **-- 2026-09-25 live proof:** (1) confirmed on 9202 — the view mounts nextcloud's kept folders `:ro` but its files are `www-data` 0770. Also seen: the source's name „Megőrzött adatok" is Hungarian on an English box (the file browser's config holds one name). **-- 2026-09-27 (controller v0.275.0): (1) FIXED: the view joins the kept folder's OWNING GROUP when it is group-readable (never root's, never its own), binds stay `:ro`, nothing on disk changes (CC-unattended decision, `07` §6.5); the source's name follows a language switch (the switch re-syncs the file browser). Red-proofed, `audits/version-travel-2026-09-26/D3/`. NOT live-proven with a real nextcloud kept folder. STILL OPEN: (2), the Use/Load choice does not look at the off-site copy.** **-- 2026-09-27 (second session): (2) NOT built on purpose** — it composes the unit-only off-site download (`RestoreOffboxScratch(full=false)`) with the unit restore into a new restore path on household data, and no box CC may touch has an off-site target to prove it on (9202 has none; 9201 on both demo hosts is fenced). Needs: a Tier-0 guest with an off-site target, or an operator word to use one.** | **OPEN — P3; owner: CC — NARROWED to (2)** | | **R-693** | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** | -| **R-697** | **[P3-LOW] A restore drops the `conversion_copy` record but not the kept pre-conversion volume — the copy is then never released.** Seen 2026-09-26 on 9202 (A1, controller 0.274.0): after docmost's 16 → 18 conversion `app.yaml` held `conversion_copy` (`docmost_docmost_postgres_data.pre-update-20260926T073003Z`); a restore from the own unit, then one from the second drive, both left `conversion_copy=None` while the volume stayed. `PersistUnitRedeployConfig` builds a fresh `AppConfig` (Deployed, DeployedAt, Env, LockedFields), so every record kept beside the env is dropped by a restore — the hourly release reads the record, so the volume is orphaned. **Harm: disk only** (the size of the old datadir), never data. **Fix direction:** carry `conversion_copy` across the restore's app.yaml rewrite (the release then removes it when a backup on the new major is proven), and say in the log when a restore supersedes one. `audits/version-travel-2026-09-26/A1/S4-restore-window-step2.txt`, `S5-restore-tier2-window.txt` | **OPEN — P3; owner: CC** | | **R-698** | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. | **OPEN — P3; owner: operator (a decision), CC measures** | +| **R-700** | **[P2] A drive move unpinned the app — its next start took the catalog's newest version, past the ladder.** Found 2026-09-27 reading the code for R-697 (not seen on a box): `doFlipRedeploy` (the per-app and whole-drive move) persisted through the restore's fresh `app.yaml` write, which drops `pinned_images`, `desired_state`, `installed_images`, the update records and the kept conversion copies. Unpinned, the catalog syncer copies the catalog's compose verbatim (`sync.renderSource`'s table) and the next `up` runs the newest version — for a PostgreSQL app past its conversion step, i.e. a new engine on an old datadir. Pin adoption repairs it only at a controller restart. **-- 2026-09-27 (controller v0.276.0): FIXED** — `persistDriveFlip` changes `HDD_PATH` and nothing else; red-proofed (`audits/records-carried-2026-09-27/redproofs/RP3`, `RP4`). **STILL OPEN: the live proof** — no Tier-0 guest has two drives (9202 has one); prove a move on a box with a second drive, reading `pinned_images` before and after and the running image after the next sync. | **WATCHING — P2; owner: CC (live proof)** |