diff --git a/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md index eecc677e..4778e52f 100644 --- a/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md @@ -86,3 +86,4 @@ cherry-picks onto `main`, writes CHANGELOG, closes rows, pushes and watches CI. | R-298 | moved to D — Needs a design ruling: is 'user-data AND backup target on one drive' supported (then the agent must | 10 | — | | R-756 | needs a live reading — A fix that walks up to a mounted ancestor would also loosen the boot reconciler's start gate (DriveL | 10 | — | | (live, 9202) | the live-proof helper was stopped at 23:30 after 90 min with no result; 9202's catalog pointer put back byte-identical; no proof run — R-776, R-613 stay held, R-763/R-764 ship as a hidden app | — | `live-9202-teardown.txt` | +| R-518, R-638, R-528 | group D: one-page design proposals written (no code), in this folder | 40 | this commit | diff --git a/documentation/audits/night-burndown-2026-10-05/design-R-518.md b/documentation/audits/night-burndown-2026-10-05/design-R-518.md new file mode 100644 index 00000000..cd6fa9bc --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-05/design-R-518.md @@ -0,0 +1,40 @@ +# R-518 — per-tier quiesce — design proposal (burn-down night 2026-10-05, no code) + +Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6.4. + +## 1. The problem +„Mentés most" stops every app until the slow local copy has fully finished, although the copy's snapshot is ready within seconds. Measured 2026-10-05 on demo-hp (9 apps): the agent reported `snapshotted` at the first sample, 12 s after the local job started (09:19:29 → 09:19:41), but the apps stayed down until 09:24:09 and the last one was back at 09:24:55 — 5 min 47 s (`audits/hub-safety-2026-10-05/partE/r518-watch.log`, `r518-measure.txt`). BIGNIGHT measured ≈ 7 min 45 s on 12 apps (`evidence-bignight-2026-09-14/phase4/guest-backup-quiesce-log.txt:3-45`). + +## 2. What the code does today (read in source) +- A manual press covers every tier in one window (`quiesce.go:428-431`, `allTiersForManualRun`). +- ONE stop for all due tiers, tiers run one after another (`quiesce.go:487-599`). +- Early resume at `snapshotted` happens ONLY on the last tier (`quiesce.go:630-636`). A non-last tier waits for `done` (`quiesce.go:643`), because the agent holds the guest lock until the upload ends. +- Order is primary (local) first, so the long local upload always runs with the apps down. +- This is a recorded R-82 choice: „ONE quiesce window for both due tiers (never two app outages for one night)" (`felhom.eu/CONTEXT.md:2709`). Two tests pin it: `TestBothTiersDue_ExactlyOneQuiesceWindow` (`tiers_test.go:140`), `TestNonLastTierSnapshot_DoesNotResumeApp` (`tiers_test.go:205`). +- After a successful primary copy the agent runs the OS leg under the same heavy-op gate (`felhom-agent/internal/localapi/server.go:904`). Inferred: that is why the PBS tier was BUSY at 09:24:09 on demo-hp. + +## 3. Options +**A. One window per tier.** Stop → start tier → resume at its `snapshotted` → let the upload finish with apps up → next tier gets its own stop. +- Changes: the loop body; two tests are rewritten to pin the new rule. +- Costs: two short outages instead of one long one when both tiers are due. Inferred per outage from the demo-hp parts: stop 21 s + snapshot ≤ 12 s + restart 46 s ≈ 80 s. +- Can go wrong: a second stop is wasted if the agent refuses tier 2 as BUSY (seen today). Mitigation: run tier 2 in a later cycle, not straight after tier 1. +- Measure first: the stop-to-`snapshotted` time on the PBS tier (never measured; BIGNIGHT's PBS run failed in 10 s, today's was refused). +- Every copy stays app-consistent (each tier is taken with the apps stopped). + +**B. Keep one window, resume at the first tier's snapshot.** The second tier then copies RUNNING apps. +- Costs: one line of logic. Loses app-consistency on the off-site (disaster) copy. That changes risk to customer data — not CC's call. + +**C. Do nothing more.** The page already states the measured minutes (controller v0.296.0). Every press still costs ≈ 6-8 min of no apps. + +## 4. The pick — PROPOSAL for the operator, not a decision +Option A, with tier 2 left to the next cycle. It keeps the reason for one window (every copy app-consistent) and drops the cost (4-5 minutes of upload with apps down). It does reverse the recorded R-82 choice „never two app outages for one night", so it needs the operator's word. At night, two outages of about 80 s each are less visible than one of 6-8 min. With the operator's word, the R-82 note and §6.4 change in the same commit. + +## 5. First slice and its proof +- Build: `quiesceAndPollTiers` runs only the FIRST due tier per window and resumes at its `snapshotted`; the other tiers stay due and the next cycle (5 min poll) picks them up. A manual press keeps covering all tiers, but each in its own window. +- Red test first (must FAIL on today's code): two tiers due; local reports `snapshotted, snapshotted, snapshotted, done`. Assert the stacks were STARTED before the second `snapshotted` poll answered, i.e. the apps run while local still uploads. Today it fails: start comes only after the last tier. +- Keep green: crash marker before any stop (`quiesce.go:490-494`); one unquiesce per window; the max-quiesce bound. +- Live proof on scratch 9202 (throwaway apps only): press the button; every 5 s sample (a) one throwaway app over HTTP, (b) the agent job phase. Positive observable: the app answers 200 while the phase reads `snapshotted`. Control from a different channel: container `StartedAt` from `docker inspect`, against the controller log. Evidence off the box before teardown. + +## 6. Open questions for the operator +1. Two short app stops in one night instead of one long one — acceptable? If you do nothing: every press keeps every app down for the whole local upload. +2. On a manual press, should the off-site tier still run (second short stop), or only the local one? diff --git a/documentation/audits/night-burndown-2026-10-05/design-R-528.md b/documentation/audits/night-burndown-2026-10-05/design-R-528.md new file mode 100644 index 00000000..37d518ba --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-05/design-R-528.md @@ -0,0 +1,39 @@ +# R-528 — OOM kills not reported — design proposal (burn-down night 2026-10-05, no code) + +Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `08-alarm-ladder.md` (OOM rung, lines 224-247). Memory note: `lxc-docker-oom-signals-unreliable`. + +## 1. The problem +The out-of-memory alarm starts from one Docker flag, and inside a Felhom guest that flag sometimes stays false after a real kill. Measured 2026-09-15 on 9202: Paperless at 128M restarted 11 times and a memory hog was killed (rc 137), with `OOMKilled=false` and no `oom` event (`audits/evidence-p1fixes-2026-09-15/E2-oom-signal-measure-9202.txt`); the same on the drill box 2026-09-16 (`audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`). Measured the other way: romm on demo-hp (2026-09-22) and on 9202 (2026-09-23) read `true`, and the alarm reached the operator's inbox. + +## 2. What the code does today (read in source) +- Every 30 s the box runs one `docker inspect` and keeps only containers with `OOMKilled=true` (`stacks/oom.go:55-66`, scheduled at `cmd/controller/main.go:929, 940`). +- The kernel's real kill counter (`memory.events` `oom_kill`, read by `docker exec … cat` inside the container) is read ONLY for those flagged containers (`oom.go:71-77`). +- So when the flag is false, nothing reads the counter: no `app_oom`, no `app_oom_storm`. `08` line 238 states this limit. +- Partly covered since controller v0.269.0: a crash loop (≥ 6 restarts in 10 min) stops the app and alarms as `app_stopped_unhealthy` (`08` line 246). It does not say "memory". +- Why the flag is set on some runs and not others is **unknown**. Searched: the two evidence files above, the memory note, `08`. + +## 3. Options +**A. Read the counter for every running container, not only flagged ones.** Same read, already proven live on 9202 (`oom_kill` 8 → 49). +- Changes: drop the flag gate in `oom.go`; read unflagged containers every 10th scan (5 min) to bound cost. +- Costs: one `docker exec` per container per read (≈ 20-40 on a full box; cost per exec not measured). +- Can go wrong: images with no `cat` return nothing (reads as unknown, not zero). Inferred: a container whose MAIN process is killed exits, its cgroup and counter go with it — that shape stays invisible here. +- Measure first: a memory hog killed inside a running container with the flag false — does the counter rise? + +**B. The agent reads the guest's cgroups from the Proxmox host.** A parent cgroup's counter survives a container's death (inferred from cgroup v2 rules), so it also sees main-process kills. +- Costs: a new agent read, a new local-API field, a controller client, a `MinAgent` raise. A mechanism nobody has measured. +- Measure first: the host-side cgroup path of a guest's Docker container, and whether the guest-wide counter rises on each kill. + +**C. Name the crash loop "probably memory".** The same `docker inspect` adds `ExitCode`; exit 137 with no stop from us → the crash-loop alarm text says "probably out of memory". +- Costs: small. Can go wrong: any other SIGKILL reads as memory too — the text must say "probably". + +## 4. The pick — PROPOSAL for the operator, not a decision +A first, then C. A uses a read the box already does and proved live; it closes the worker-kill case (the BIGNIGHT Paperless shape) when the flag lies. C closes the main-process case cheaply, on an alarm that already fires. B is the complete answer but is a new mechanism on the operator-tier agent; it should wait for a measurement that shows A + C miss real kills. + +## 5. First slice and its proof +- Red test first (must FAIL today): the fake `execCommand` answers `inspect` with `OOMKilled=false` and the container's `memory.events` with `oom_kill 3`. Assert `ScanOOMKilled` returns that container with `Kills=3`. Today it returns nothing (`oom.go:63`, the flag test). +- Second test: `oom_kill 0` on every container → nothing returned and no alarm (no false positive). +- Live proof on 9202, throwaway app only: run a memory hog in a child process of a running container under a low cap. Positive observable: `app_oom` in the hub's Events tab. Control from a different channel: `docker exec cat /sys/fs/cgroup/memory.events` read by hand, and the `OOMKilled` flag recorded beside it. Evidence off the box before teardown; teardown on box, host and hub. + +## 6. Open questions for the operator +1. Is one extra `docker exec` per container every 5 minutes acceptable on a small box? If you do nothing: kills the flag misses stay silent, as today. +2. Should the host-side agent read (B) be spiked now, or only after A + C have run a week? diff --git a/documentation/audits/night-burndown-2026-10-05/design-R-638.md b/documentation/audits/night-burndown-2026-10-05/design-R-638.md new file mode 100644 index 00000000..5d7a503c --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-05/design-R-638.md @@ -0,0 +1,40 @@ +# R-638 — restoring over a newer schema — design proposal (burn-down night 2026-10-05, no code) + +Baselines read: felhom-controller `ef199c5`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6 ("replay → rollback → hold", line 596). + +## 1. The problem +The database loader replays a copy on top of the live database, so it only removes what the copy knows about. Measured 2026-09-23 on 9202: after docmost 0.95 → 0.96 migrated, replaying the older copy FAILED on PostgreSQL (rc 3, a new table's foreign key blocked the drop); on MariaDB (romm 5.0 → 5.3) it "succeeded" and left 12 newer tables behind (`audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`). + +## 2. What the code does today (read in source) +- Loader: `psql -v ON_ERROR_STOP=1 --single-transaction` / plain `mariadb` over the live DB (`appbackup/dbdump.go:740-786`). Copies are made with `--clean --if-exists` (`dbdump.go:312`). +- **Unit restore** (the restore the hold sentence names): stop → volume tars REPLACE the named volumes (`volume rm -f` + create + untar, `backup/restore.go:154-178`) → definition from the unit, i.e. the data's own version (`restore_unit.go:362-383, 444`) → DB-only start → replay (`restore_unit.go:438-458`). +- **Off-site restore**: same order — volumes from the scratch unit, then replay (`offbox_reconstitute.go:856, 881`); a failed volume leg restarts and stops there (`:858-862`). +- Catalog: all 17 templates with a PostgreSQL or MariaDB data dir keep it in a NAMED volume (read in `app-catalog-felhom.eu/templates/*/docker-compose.yml`). +- Inferred from the three above: on the two main paths the replay meets the copy's OWN schema, so R-638 is probably moot there. **Not measured** — that is the measurement the row owes. +- Remaining exposures (inferred from source): + 1. **No-manifest fallback** `RestoreApp`: volumes back, then the WHOLE stack starts at the CURRENT definition (`restore.go:67, 76`) — a newer app can migrate the old data — then the replay runs (`restore.go:82`). This is the R-638 shape exactly. + 2. **Unit restore with a failed volume leg** still replays (`restore_unit.go:439-442` sets the error and continues to `:458`). + 3. **Rollback after a failed off-site replay** pours the NEWER pre-restore copy over the OLDER volume just put back (`offbox_reconstitute.go:898`, `:456-463`) — the reverse direction; tables the migration removed would stay. + +## 3. Options +**A. Measure, then close the three gaps by order, not by loader.** Fallback: start only DB services at the restored volume, replay, then start the app. Unit restore: do not replay when the DB's volume leg failed. +- Cost: small, in two files. Risk: low; no change to what the loader does. +- Measure first: the named unit restore after a real migration (docmost, romm) on 9202. + +**B. Make the loader rebuild instead of overlay.** PostgreSQL: `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the copy, in ONE transaction (measured working: rc 0, 1.38 s). MariaDB: drop every table first, then load. +- Cost: medium. Fixes every caller at once, the rollback included. +- Can go wrong: on MariaDB a failed load now leaves an EMPTY database (DDL is not transactional, `07` line 619) — the rollback must catch it. On PostgreSQL ≥ 15 the app user may not own `public` (speculative; must be measured per app). Objects in other schemas stay (immich-style extensions — unmeasured). + +**C. Refuse a replay when the copy's version is older than the live one.** Uses the unit's data version record (`restore_unit.go:362`). Cost: small. Leaves the household with no restore at all in that case. + +## 4. The pick — PROPOSAL for the operator, not a decision +Option A. Read in source, both shipped restore paths already put the copy's own database files back before they load the copy. So the loader is not the weak point; the order on three side paths is. A keeps the loader the drills have proven, and it adds no new delete step on customer data. B is the fallback if the measurement shows the main paths still fail. + +## 5. First slice and its proof +- Slice 0, measurement only (9202, throwaway docmost + romm): copy at the old version → update and migrate → unit restore. Positive observable: the replay log line `Imported DB dump` AND a read-back of a row written before the copy. Control from a different channel: `\dt` / `SHOW TABLES` counted against the copy's own table list (the 6 and 12 newer tables must be GONE). Evidence off the box before teardown. +- Slice 1 (fallback order): red test first — a fake stack provider records call order; assert no full `StartStack` happens before the replay in `RestoreApp`. Fails today (`restore.go:76` before `:82`). +- Slice 2: red test — a unit restore whose volume leg errors must NOT call the importer. Fails today. + +## 6. Open questions for the operator +1. If the measurement shows the main paths are safe, may the row close on A alone, with B kept as a note? If you do nothing: the three side paths stay as they are. +2. The rollback in the reverse direction (newer copy over an older volume): fix it now, or record it as a known limit? diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 231f5cec..e2b733db 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -179,8 +179,8 @@ stopping line that lies. | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC | | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | -| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** | — | — | CC | -| **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | — | — | CC | +| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. | — | — | CC | +| **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-638.md`** — for the operator. | — | — | CC | | **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | @@ -274,7 +274,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-243** | Monitoring & notifications | P2 | **A box in the R-241 state silently stops backing up off-site, and NO ALARM OF ANY KIND FIRES.** Found by the R-241 spike (2026-08-07) as a by-product; **not part of the walk's finding and not previously filed.** The R-241 state is self-locking in a second, worse way than the recovery-journey dead end: `escrow_state` is stuck `pending` forever (the auto-confirm flips only on a hash match, and the hash cannot match a key the box minted itself), and `runOffboxBackup` returns at the escrow gate (`offbox.go:743`) before touching anything. **So off-site backups never run again — and the hub never notices.** All three signals that could catch it are excluded, each for its own individually-correct reason, verified in the hub this session: `offsite_stale` — `isStale` (`monitor/offsite.go:135`) returns false unless `EscrowState == "escrowed"`, and its own comment reads *"Pending/disabled = normal onboarding, never stale"*, so the box is classified as **still being set up, forever**; `offsite_delivery_stuck` — `monitor/offsite_delivery.go:91` skips the `applied` shape, and delivery genuinely IS applied (the credential was consumed and the target is in every report); `backup_failed` — never fires, because nothing fails: the run returns `nil` before it starts. **Three correct exclusions leaving one state unobserved.** This is the same class as the workspace `CLAUDE.md` "presence is not success" rule, one level up: **the absence of a failure is being read as the presence of a working tier.** **Partly subsumed by R-241's fix** — a box that recovers leaves this state — but **not for a box that does not**, and the alarm gap is what makes "does not" survivable indefinitely. **Not fixed; no code written.** **⚠ UPDATED 2026-08-07 (v0.206.0) — the STATE this row describes can no longer be entered, but the ALARM GAP is untouched and the row stays open.** R-241's mint guard means a box no longer mints a key over a sealed package, so it no longer arrives in the "escrow stuck pending against a self-minted key" state by itself. **What replaces it is a state that is VISIBLE rather than silent:** the box declares `offsite.state=awaiting_recovery_key` and the customer is offered the recovery screen. **But the hub still raises nothing for it**, and for the same three reasons: `isStale` needs `escrowed`, the delivery checker skips the `applied` shape, and `backup_failed` needs a run that never happens. **So a box whose customer never acts still stops backing up off-site with no operator signal** — the difference is that the customer can now see it and act, where before nobody could. **The remaining work is an operator-side signal for a box held in `awaiting_recovery_key` past some age**, and it is deliberately not bundled into R-241's fix. **⚠ MEASURED ON A REBUILD, 2026-08-07 (fifth walk) — the gap is real for the state this row describes, and NOT for the state a rebuild produces.** 88 seconds after the walk5 guest was destroyed and rebuilt, the hub emitted `offsite_delivery_stuck` (**warning**) and wrote an **operator-channel** `notification_log` row recording `offsite_credential_restaged` / status **REFUSED** with an accurate reason — *"the credential was applied and worked; the target was lost afterwards … a guest rebuild does, R-193"*. So on the **regressed-apply** shape the operator IS told, promptly and correctly, and this row's *"skips the applied shape"* does not apply. The gap stands for a box that reaches the held state **without** a prior working tier in its report history. **Recorded so the row is not read wider than it measures.** | **READY** — owner Viktor | — | — | operator | -| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. | — | — | CC | +| **R-528** | Monitoring & notifications | P2 | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend **RE-MEASURED 2026-09-16 on the DRILL box (fresh install, nested VM 334, Docker in an LXC guest, controller 0.243.0), so the finding is not a property of one machine:** the Paperless webserver was capped at 128 M with `docker update --memory`; it restarted 9-10 times, and all three signals stayed silent - `OOMKilled=false` on every inspect, `docker events --filter event=oom` EMPTY for the whole window, the container's cgroup not visible from inside the guest, and `dmesg` unreadable there. Identical to scratch 9202. So the v0.243.0 OOM line cannot fire on ANY Felhom box as shipped, on either host. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`. | **READY — rank P2-MEDIUM; owner: CC** **2026-09-17 (chaos night): an OOM WAS detected on a fresh box, and named precisely.** On `tester-1-022354` (controller 0.245.0, guest 9201, 6 GB RAM) immich’s Postgres was killed by the memory limit during its reverse-geocoding import, and the controller pushed `app_oom` (warning, operator-only): „Alkalmazás memóriája elfogyott: immich (immich-postgres) — egy folyamatát a memóriakorlát leállította” — naming the app AND the exact container. The visible consequence was `write CONNECTION_CLOSED immich-postgres:5432` and twelve restarts of immich-server. So on THIS box the OOM scan works and was the fastest route to the diagnosis; recorded here rather than filed as a new row. Evidence: `audits/evidence-chaos-night-2026-09-17/round-2.txt`. **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-528.md`** — for the operator. | — | — | CC | | **R-79** | Monitoring & notifications | P3 | **`report.Issues` / `report.Warnings` are English on customer-facing surfaces** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the issues and warnings travel to the hub as sentences; changing them is the two-repo spike the row itself names. | — | **Whole-surface, not a one-off** (DIAG §6): every producer is English — `"SSD/HDD disk usage critical"`, `"Docker: %v"`, `"Protected container not running: %s"`, and all six `Warnings` strings. They render on the customer's Hungarian dashboard, and the `health_critical` path has reached the **customer** email channel three times historically. Deliberately NOT bundled into R-77: a copy sweep across every producer would have buried two safety fixes in string churn, and the seam is not obvious — translate at the producer, or at the render/notification boundary where operator-English and customer-Hungarian already diverge? Pick the seam in a spike; the strings are mechanical after. | CC | | **R-211** | Monitoring & notifications | P3 | **Prometheus has no config-reloader — a rules change reaches the pod and is never read** | **READY (S) — NEW 2026-08-05** | — | Found while verifying R-205 rather than by looking for it. The `mon-system/prometheus` Deployment runs **one** container (`prom/prometheus:v3.12.0`) with **no `configmap-reload`/`prometheus-config-reloader` sidecar**. After the ArgoCD sync the updated `node-housekeeping-alerts.yml` was present **inside the pod** (`grep -c "and on(instance)"` → 3 on the mounted symlink) while the Prometheus **rules API still served the old expression** — for **4+ minutes**, with no error anywhere. It only took effect after an explicit `POST /-/reload`. **The consequence is general, not specific to R-205: every rule edit in this repo since the stack was built has silently not applied until something happened to restart the pod** — so "committed and synced" has never meant "in force", and ArgoCD reporting `Synced/Healthy` is true and beside the point. `--web.enable-lifecycle` IS already set, so the fix is small: add a reloader sidecar watching the ConfigMap, or a `checksum/config` pod annotation so a rules change rolls the pod. **Same class as the four *built-but-never-wired* seams** — the control exists, nothing walks it | CC | | **R-271** | Monitoring & notifications | P3 | **The `agent_channel_unauthorized` alarm can never be closed, because its own prescribed remedy is what silences the recovery.** `channelhealth.Checker.Check`'s UP branch notifies only when `prev != "" && prev != "up"`; a controller restart resets `state` to `""`, so an unseeded→up transition is silent by construction. The alert text says *"token stale/rotated (**re-bootstrap**)"* — i.e. restart the controller — so **following the instruction guarantees no recovery event.** Observed live 2026-08-09: two `agent_channel_unauthorized` errors on the hub (one `sent`, one `suppressed` by the 1 h operator cooldown) and **nothing afterwards**, though the channel came up 3 minutes later and stayed up. The down side is deliberately asymmetric (F2: a born-down channel alerts on cycle 1); the up side never got the matching treatment. Customer dashboard is fine — `SetDashboard` reflects current state every cycle. It is the OPERATOR's trail that ends on "down" | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on controller `main`** (`f885100` — a down alert leaves a marker, so the first good probe after a restart reports the recovery; test + red-proof). Ships with the next controller release; close after delivery. | — | Notify on unseeded→up when the previous *persisted* state was down, or seed from the hub's last event | CC |