burn-down night: group D design proposals R-518, R-638, R-528 (no code)
gates / gates (push) Successful in 2m7s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 23:32:12 +02:00
parent b37902ce58
commit d549082be9
5 changed files with 123 additions and 3 deletions
@@ -86,3 +86,4 @@ cherry-picks onto `main`, writes CHANGELOG, closes rows, pushes and watches CI.
| R-298 | moved to D — Needs a design ruling: is 'user-data AND backup target on one drive' supported (then the agent must | 10 | — |
| R-756 | needs a live reading — A fix that walks up to a mounted ancestor would also loosen the boot reconciler's start gate (DriveL | 10 | — |
| (live, 9202) | the live-proof helper was stopped at 23:30 after 90 min with no result; 9202's catalog pointer put back byte-identical; no proof run — R-776, R-613 stay held, R-763/R-764 ship as a hidden app | — | `live-9202-teardown.txt` |
| R-518, R-638, R-528 | group D: one-page design proposals written (no code), in this folder | 40 | this commit |
@@ -0,0 +1,40 @@
# R-518 — per-tier quiesce — design proposal (burn-down night 2026-10-05, no code)
Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6.4.
## 1. The problem
„Mentés most" stops every app until the slow local copy has fully finished, although the copy's snapshot is ready within seconds. Measured 2026-10-05 on demo-hp (9 apps): the agent reported `snapshotted` at the first sample, 12 s after the local job started (09:19:29 → 09:19:41), but the apps stayed down until 09:24:09 and the last one was back at 09:24:55 — 5 min 47 s (`audits/hub-safety-2026-10-05/partE/r518-watch.log`, `r518-measure.txt`). BIGNIGHT measured ≈ 7 min 45 s on 12 apps (`evidence-bignight-2026-09-14/phase4/guest-backup-quiesce-log.txt:3-45`).
## 2. What the code does today (read in source)
- A manual press covers every tier in one window (`quiesce.go:428-431`, `allTiersForManualRun`).
- ONE stop for all due tiers, tiers run one after another (`quiesce.go:487-599`).
- Early resume at `snapshotted` happens ONLY on the last tier (`quiesce.go:630-636`). A non-last tier waits for `done` (`quiesce.go:643`), because the agent holds the guest lock until the upload ends.
- Order is primary (local) first, so the long local upload always runs with the apps down.
- This is a recorded R-82 choice: „ONE quiesce window for both due tiers (never two app outages for one night)" (`felhom.eu/CONTEXT.md:2709`). Two tests pin it: `TestBothTiersDue_ExactlyOneQuiesceWindow` (`tiers_test.go:140`), `TestNonLastTierSnapshot_DoesNotResumeApp` (`tiers_test.go:205`).
- After a successful primary copy the agent runs the OS leg under the same heavy-op gate (`felhom-agent/internal/localapi/server.go:904`). Inferred: that is why the PBS tier was BUSY at 09:24:09 on demo-hp.
## 3. Options
**A. One window per tier.** Stop → start tier → resume at its `snapshotted` → let the upload finish with apps up → next tier gets its own stop.
- Changes: the loop body; two tests are rewritten to pin the new rule.
- Costs: two short outages instead of one long one when both tiers are due. Inferred per outage from the demo-hp parts: stop 21 s + snapshot ≤ 12 s + restart 46 s ≈ 80 s.
- Can go wrong: a second stop is wasted if the agent refuses tier 2 as BUSY (seen today). Mitigation: run tier 2 in a later cycle, not straight after tier 1.
- Measure first: the stop-to-`snapshotted` time on the PBS tier (never measured; BIGNIGHT's PBS run failed in 10 s, today's was refused).
- Every copy stays app-consistent (each tier is taken with the apps stopped).
**B. Keep one window, resume at the first tier's snapshot.** The second tier then copies RUNNING apps.
- Costs: one line of logic. Loses app-consistency on the off-site (disaster) copy. That changes risk to customer data — not CC's call.
**C. Do nothing more.** The page already states the measured minutes (controller v0.296.0). Every press still costs ≈ 6-8 min of no apps.
## 4. The pick — PROPOSAL for the operator, not a decision
Option A, with tier 2 left to the next cycle. It keeps the reason for one window (every copy app-consistent) and drops the cost (4-5 minutes of upload with apps down). It does reverse the recorded R-82 choice „never two app outages for one night", so it needs the operator's word. At night, two outages of about 80 s each are less visible than one of 6-8 min. With the operator's word, the R-82 note and §6.4 change in the same commit.
## 5. First slice and its proof
- Build: `quiesceAndPollTiers` runs only the FIRST due tier per window and resumes at its `snapshotted`; the other tiers stay due and the next cycle (5 min poll) picks them up. A manual press keeps covering all tiers, but each in its own window.
- Red test first (must FAIL on today's code): two tiers due; local reports `snapshotted, snapshotted, snapshotted, done`. Assert the stacks were STARTED before the second `snapshotted` poll answered, i.e. the apps run while local still uploads. Today it fails: start comes only after the last tier.
- Keep green: crash marker before any stop (`quiesce.go:490-494`); one unquiesce per window; the max-quiesce bound.
- Live proof on scratch 9202 (throwaway apps only): press the button; every 5 s sample (a) one throwaway app over HTTP, (b) the agent job phase. Positive observable: the app answers 200 while the phase reads `snapshotted`. Control from a different channel: container `StartedAt` from `docker inspect`, against the controller log. Evidence off the box before teardown.
## 6. Open questions for the operator
1. Two short app stops in one night instead of one long one — acceptable? If you do nothing: every press keeps every app down for the whole local upload.
2. On a manual press, should the off-site tier still run (second short stop), or only the local one?
@@ -0,0 +1,39 @@
# R-528 — OOM kills not reported — design proposal (burn-down night 2026-10-05, no code)
Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `08-alarm-ladder.md` (OOM rung, lines 224-247). Memory note: `lxc-docker-oom-signals-unreliable`.
## 1. The problem
The out-of-memory alarm starts from one Docker flag, and inside a Felhom guest that flag sometimes stays false after a real kill. Measured 2026-09-15 on 9202: Paperless at 128M restarted 11 times and a memory hog was killed (rc 137), with `OOMKilled=false` and no `oom` event (`audits/evidence-p1fixes-2026-09-15/E2-oom-signal-measure-9202.txt`); the same on the drill box 2026-09-16 (`audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`). Measured the other way: romm on demo-hp (2026-09-22) and on 9202 (2026-09-23) read `true`, and the alarm reached the operator's inbox.
## 2. What the code does today (read in source)
- Every 30 s the box runs one `docker inspect` and keeps only containers with `OOMKilled=true` (`stacks/oom.go:55-66`, scheduled at `cmd/controller/main.go:929, 940`).
- The kernel's real kill counter (`memory.events` `oom_kill`, read by `docker exec … cat` inside the container) is read ONLY for those flagged containers (`oom.go:71-77`).
- So when the flag is false, nothing reads the counter: no `app_oom`, no `app_oom_storm`. `08` line 238 states this limit.
- Partly covered since controller v0.269.0: a crash loop (≥ 6 restarts in 10 min) stops the app and alarms as `app_stopped_unhealthy` (`08` line 246). It does not say "memory".
- Why the flag is set on some runs and not others is **unknown**. Searched: the two evidence files above, the memory note, `08`.
## 3. Options
**A. Read the counter for every running container, not only flagged ones.** Same read, already proven live on 9202 (`oom_kill` 8 → 49).
- Changes: drop the flag gate in `oom.go`; read unflagged containers every 10th scan (5 min) to bound cost.
- Costs: one `docker exec` per container per read (≈ 20-40 on a full box; cost per exec not measured).
- Can go wrong: images with no `cat` return nothing (reads as unknown, not zero). Inferred: a container whose MAIN process is killed exits, its cgroup and counter go with it — that shape stays invisible here.
- Measure first: a memory hog killed inside a running container with the flag false — does the counter rise?
**B. The agent reads the guest's cgroups from the Proxmox host.** A parent cgroup's counter survives a container's death (inferred from cgroup v2 rules), so it also sees main-process kills.
- Costs: a new agent read, a new local-API field, a controller client, a `MinAgent` raise. A mechanism nobody has measured.
- Measure first: the host-side cgroup path of a guest's Docker container, and whether the guest-wide counter rises on each kill.
**C. Name the crash loop "probably memory".** The same `docker inspect` adds `ExitCode`; exit 137 with no stop from us → the crash-loop alarm text says "probably out of memory".
- Costs: small. Can go wrong: any other SIGKILL reads as memory too — the text must say "probably".
## 4. The pick — PROPOSAL for the operator, not a decision
A first, then C. A uses a read the box already does and proved live; it closes the worker-kill case (the BIGNIGHT Paperless shape) when the flag lies. C closes the main-process case cheaply, on an alarm that already fires. B is the complete answer but is a new mechanism on the operator-tier agent; it should wait for a measurement that shows A + C miss real kills.
## 5. First slice and its proof
- Red test first (must FAIL today): the fake `execCommand` answers `inspect` with `OOMKilled=false` and the container's `memory.events` with `oom_kill 3`. Assert `ScanOOMKilled` returns that container with `Kills=3`. Today it returns nothing (`oom.go:63`, the flag test).
- Second test: `oom_kill 0` on every container → nothing returned and no alarm (no false positive).
- Live proof on 9202, throwaway app only: run a memory hog in a child process of a running container under a low cap. Positive observable: `app_oom` in the hub's Events tab. Control from a different channel: `docker exec <c> cat /sys/fs/cgroup/memory.events` read by hand, and the `OOMKilled` flag recorded beside it. Evidence off the box before teardown; teardown on box, host and hub.
## 6. Open questions for the operator
1. Is one extra `docker exec` per container every 5 minutes acceptable on a small box? If you do nothing: kills the flag misses stay silent, as today.
2. Should the host-side agent read (B) be spiked now, or only after A + C have run a week?
@@ -0,0 +1,40 @@
# R-638 — restoring over a newer schema — design proposal (burn-down night 2026-10-05, no code)
Baselines read: felhom-controller `ef199c5`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6 ("replay → rollback → hold", line 596).
## 1. The problem
The database loader replays a copy on top of the live database, so it only removes what the copy knows about. Measured 2026-09-23 on 9202: after docmost 0.95 → 0.96 migrated, replaying the older copy FAILED on PostgreSQL (rc 3, a new table's foreign key blocked the drop); on MariaDB (romm 5.0 → 5.3) it "succeeded" and left 12 newer tables behind (`audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`).
## 2. What the code does today (read in source)
- Loader: `psql -v ON_ERROR_STOP=1 --single-transaction` / plain `mariadb` over the live DB (`appbackup/dbdump.go:740-786`). Copies are made with `--clean --if-exists` (`dbdump.go:312`).
- **Unit restore** (the restore the hold sentence names): stop → volume tars REPLACE the named volumes (`volume rm -f` + create + untar, `backup/restore.go:154-178`) → definition from the unit, i.e. the data's own version (`restore_unit.go:362-383, 444`) → DB-only start → replay (`restore_unit.go:438-458`).
- **Off-site restore**: same order — volumes from the scratch unit, then replay (`offbox_reconstitute.go:856, 881`); a failed volume leg restarts and stops there (`:858-862`).
- Catalog: all 17 templates with a PostgreSQL or MariaDB data dir keep it in a NAMED volume (read in `app-catalog-felhom.eu/templates/*/docker-compose.yml`).
- Inferred from the three above: on the two main paths the replay meets the copy's OWN schema, so R-638 is probably moot there. **Not measured** — that is the measurement the row owes.
- Remaining exposures (inferred from source):
1. **No-manifest fallback** `RestoreApp`: volumes back, then the WHOLE stack starts at the CURRENT definition (`restore.go:67, 76`) — a newer app can migrate the old data — then the replay runs (`restore.go:82`). This is the R-638 shape exactly.
2. **Unit restore with a failed volume leg** still replays (`restore_unit.go:439-442` sets the error and continues to `:458`).
3. **Rollback after a failed off-site replay** pours the NEWER pre-restore copy over the OLDER volume just put back (`offbox_reconstitute.go:898`, `:456-463`) — the reverse direction; tables the migration removed would stay.
## 3. Options
**A. Measure, then close the three gaps by order, not by loader.** Fallback: start only DB services at the restored volume, replay, then start the app. Unit restore: do not replay when the DB's volume leg failed.
- Cost: small, in two files. Risk: low; no change to what the loader does.
- Measure first: the named unit restore after a real migration (docmost, romm) on 9202.
**B. Make the loader rebuild instead of overlay.** PostgreSQL: `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the copy, in ONE transaction (measured working: rc 0, 1.38 s). MariaDB: drop every table first, then load.
- Cost: medium. Fixes every caller at once, the rollback included.
- Can go wrong: on MariaDB a failed load now leaves an EMPTY database (DDL is not transactional, `07` line 619) — the rollback must catch it. On PostgreSQL ≥ 15 the app user may not own `public` (speculative; must be measured per app). Objects in other schemas stay (immich-style extensions — unmeasured).
**C. Refuse a replay when the copy's version is older than the live one.** Uses the unit's data version record (`restore_unit.go:362`). Cost: small. Leaves the household with no restore at all in that case.
## 4. The pick — PROPOSAL for the operator, not a decision
Option A. Read in source, both shipped restore paths already put the copy's own database files back before they load the copy. So the loader is not the weak point; the order on three side paths is. A keeps the loader the drills have proven, and it adds no new delete step on customer data. B is the fallback if the measurement shows the main paths still fail.
## 5. First slice and its proof
- Slice 0, measurement only (9202, throwaway docmost + romm): copy at the old version → update and migrate → unit restore. Positive observable: the replay log line `Imported DB dump` AND a read-back of a row written before the copy. Control from a different channel: `\dt` / `SHOW TABLES` counted against the copy's own table list (the 6 and 12 newer tables must be GONE). Evidence off the box before teardown.
- Slice 1 (fallback order): red test first — a fake stack provider records call order; assert no full `StartStack` happens before the replay in `RestoreApp`. Fails today (`restore.go:76` before `:82`).
- Slice 2: red test — a unit restore whose volume leg errors must NOT call the importer. Fails today.
## 6. Open questions for the operator
1. If the measurement shows the main paths are safe, may the row close on A alone, with B kept as a note? If you do nothing: the three side paths stay as they are.
2. The rollback in the reverse direction (newer copy over an older volume): fix it now, or record it as a known limit?