burn-down night: group D design proposals R-518, R-638, R-528 (no code)
gates / gates (push) Successful in 2m7s
gates / gates (push) Successful in 2m7s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -86,3 +86,4 @@ cherry-picks onto `main`, writes CHANGELOG, closes rows, pushes and watches CI.
|
||||
| R-298 | moved to D — Needs a design ruling: is 'user-data AND backup target on one drive' supported (then the agent must | 10 | — |
|
||||
| R-756 | needs a live reading — A fix that walks up to a mounted ancestor would also loosen the boot reconciler's start gate (DriveL | 10 | — |
|
||||
| (live, 9202) | the live-proof helper was stopped at 23:30 after 90 min with no result; 9202's catalog pointer put back byte-identical; no proof run — R-776, R-613 stay held, R-763/R-764 ship as a hidden app | — | `live-9202-teardown.txt` |
|
||||
| R-518, R-638, R-528 | group D: one-page design proposals written (no code), in this folder | 40 | this commit |
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
# R-518 — per-tier quiesce — design proposal (burn-down night 2026-10-05, no code)
|
||||
|
||||
Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6.4.
|
||||
|
||||
## 1. The problem
|
||||
„Mentés most" stops every app until the slow local copy has fully finished, although the copy's snapshot is ready within seconds. Measured 2026-10-05 on demo-hp (9 apps): the agent reported `snapshotted` at the first sample, 12 s after the local job started (09:19:29 → 09:19:41), but the apps stayed down until 09:24:09 and the last one was back at 09:24:55 — 5 min 47 s (`audits/hub-safety-2026-10-05/partE/r518-watch.log`, `r518-measure.txt`). BIGNIGHT measured ≈ 7 min 45 s on 12 apps (`evidence-bignight-2026-09-14/phase4/guest-backup-quiesce-log.txt:3-45`).
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
- A manual press covers every tier in one window (`quiesce.go:428-431`, `allTiersForManualRun`).
|
||||
- ONE stop for all due tiers, tiers run one after another (`quiesce.go:487-599`).
|
||||
- Early resume at `snapshotted` happens ONLY on the last tier (`quiesce.go:630-636`). A non-last tier waits for `done` (`quiesce.go:643`), because the agent holds the guest lock until the upload ends.
|
||||
- Order is primary (local) first, so the long local upload always runs with the apps down.
|
||||
- This is a recorded R-82 choice: „ONE quiesce window for both due tiers (never two app outages for one night)" (`felhom.eu/CONTEXT.md:2709`). Two tests pin it: `TestBothTiersDue_ExactlyOneQuiesceWindow` (`tiers_test.go:140`), `TestNonLastTierSnapshot_DoesNotResumeApp` (`tiers_test.go:205`).
|
||||
- After a successful primary copy the agent runs the OS leg under the same heavy-op gate (`felhom-agent/internal/localapi/server.go:904`). Inferred: that is why the PBS tier was BUSY at 09:24:09 on demo-hp.
|
||||
|
||||
## 3. Options
|
||||
**A. One window per tier.** Stop → start tier → resume at its `snapshotted` → let the upload finish with apps up → next tier gets its own stop.
|
||||
- Changes: the loop body; two tests are rewritten to pin the new rule.
|
||||
- Costs: two short outages instead of one long one when both tiers are due. Inferred per outage from the demo-hp parts: stop 21 s + snapshot ≤ 12 s + restart 46 s ≈ 80 s.
|
||||
- Can go wrong: a second stop is wasted if the agent refuses tier 2 as BUSY (seen today). Mitigation: run tier 2 in a later cycle, not straight after tier 1.
|
||||
- Measure first: the stop-to-`snapshotted` time on the PBS tier (never measured; BIGNIGHT's PBS run failed in 10 s, today's was refused).
|
||||
- Every copy stays app-consistent (each tier is taken with the apps stopped).
|
||||
|
||||
**B. Keep one window, resume at the first tier's snapshot.** The second tier then copies RUNNING apps.
|
||||
- Costs: one line of logic. Loses app-consistency on the off-site (disaster) copy. That changes risk to customer data — not CC's call.
|
||||
|
||||
**C. Do nothing more.** The page already states the measured minutes (controller v0.296.0). Every press still costs ≈ 6-8 min of no apps.
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
Option A, with tier 2 left to the next cycle. It keeps the reason for one window (every copy app-consistent) and drops the cost (4-5 minutes of upload with apps down). It does reverse the recorded R-82 choice „never two app outages for one night", so it needs the operator's word. At night, two outages of about 80 s each are less visible than one of 6-8 min. With the operator's word, the R-82 note and §6.4 change in the same commit.
|
||||
|
||||
## 5. First slice and its proof
|
||||
- Build: `quiesceAndPollTiers` runs only the FIRST due tier per window and resumes at its `snapshotted`; the other tiers stay due and the next cycle (5 min poll) picks them up. A manual press keeps covering all tiers, but each in its own window.
|
||||
- Red test first (must FAIL on today's code): two tiers due; local reports `snapshotted, snapshotted, snapshotted, done`. Assert the stacks were STARTED before the second `snapshotted` poll answered, i.e. the apps run while local still uploads. Today it fails: start comes only after the last tier.
|
||||
- Keep green: crash marker before any stop (`quiesce.go:490-494`); one unquiesce per window; the max-quiesce bound.
|
||||
- Live proof on scratch 9202 (throwaway apps only): press the button; every 5 s sample (a) one throwaway app over HTTP, (b) the agent job phase. Positive observable: the app answers 200 while the phase reads `snapshotted`. Control from a different channel: container `StartedAt` from `docker inspect`, against the controller log. Evidence off the box before teardown.
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. Two short app stops in one night instead of one long one — acceptable? If you do nothing: every press keeps every app down for the whole local upload.
|
||||
2. On a manual press, should the off-site tier still run (second short stop), or only the local one?
|
||||
@@ -0,0 +1,39 @@
|
||||
# R-528 — OOM kills not reported — design proposal (burn-down night 2026-10-05, no code)
|
||||
|
||||
Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `08-alarm-ladder.md` (OOM rung, lines 224-247). Memory note: `lxc-docker-oom-signals-unreliable`.
|
||||
|
||||
## 1. The problem
|
||||
The out-of-memory alarm starts from one Docker flag, and inside a Felhom guest that flag sometimes stays false after a real kill. Measured 2026-09-15 on 9202: Paperless at 128M restarted 11 times and a memory hog was killed (rc 137), with `OOMKilled=false` and no `oom` event (`audits/evidence-p1fixes-2026-09-15/E2-oom-signal-measure-9202.txt`); the same on the drill box 2026-09-16 (`audits/evidence-drill-0243-2026-09-16/phase2-m1-oom.txt`). Measured the other way: romm on demo-hp (2026-09-22) and on 9202 (2026-09-23) read `true`, and the alarm reached the operator's inbox.
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
- Every 30 s the box runs one `docker inspect` and keeps only containers with `OOMKilled=true` (`stacks/oom.go:55-66`, scheduled at `cmd/controller/main.go:929, 940`).
|
||||
- The kernel's real kill counter (`memory.events` `oom_kill`, read by `docker exec … cat` inside the container) is read ONLY for those flagged containers (`oom.go:71-77`).
|
||||
- So when the flag is false, nothing reads the counter: no `app_oom`, no `app_oom_storm`. `08` line 238 states this limit.
|
||||
- Partly covered since controller v0.269.0: a crash loop (≥ 6 restarts in 10 min) stops the app and alarms as `app_stopped_unhealthy` (`08` line 246). It does not say "memory".
|
||||
- Why the flag is set on some runs and not others is **unknown**. Searched: the two evidence files above, the memory note, `08`.
|
||||
|
||||
## 3. Options
|
||||
**A. Read the counter for every running container, not only flagged ones.** Same read, already proven live on 9202 (`oom_kill` 8 → 49).
|
||||
- Changes: drop the flag gate in `oom.go`; read unflagged containers every 10th scan (5 min) to bound cost.
|
||||
- Costs: one `docker exec` per container per read (≈ 20-40 on a full box; cost per exec not measured).
|
||||
- Can go wrong: images with no `cat` return nothing (reads as unknown, not zero). Inferred: a container whose MAIN process is killed exits, its cgroup and counter go with it — that shape stays invisible here.
|
||||
- Measure first: a memory hog killed inside a running container with the flag false — does the counter rise?
|
||||
|
||||
**B. The agent reads the guest's cgroups from the Proxmox host.** A parent cgroup's counter survives a container's death (inferred from cgroup v2 rules), so it also sees main-process kills.
|
||||
- Costs: a new agent read, a new local-API field, a controller client, a `MinAgent` raise. A mechanism nobody has measured.
|
||||
- Measure first: the host-side cgroup path of a guest's Docker container, and whether the guest-wide counter rises on each kill.
|
||||
|
||||
**C. Name the crash loop "probably memory".** The same `docker inspect` adds `ExitCode`; exit 137 with no stop from us → the crash-loop alarm text says "probably out of memory".
|
||||
- Costs: small. Can go wrong: any other SIGKILL reads as memory too — the text must say "probably".
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
A first, then C. A uses a read the box already does and proved live; it closes the worker-kill case (the BIGNIGHT Paperless shape) when the flag lies. C closes the main-process case cheaply, on an alarm that already fires. B is the complete answer but is a new mechanism on the operator-tier agent; it should wait for a measurement that shows A + C miss real kills.
|
||||
|
||||
## 5. First slice and its proof
|
||||
- Red test first (must FAIL today): the fake `execCommand` answers `inspect` with `OOMKilled=false` and the container's `memory.events` with `oom_kill 3`. Assert `ScanOOMKilled` returns that container with `Kills=3`. Today it returns nothing (`oom.go:63`, the flag test).
|
||||
- Second test: `oom_kill 0` on every container → nothing returned and no alarm (no false positive).
|
||||
- Live proof on 9202, throwaway app only: run a memory hog in a child process of a running container under a low cap. Positive observable: `app_oom` in the hub's Events tab. Control from a different channel: `docker exec <c> cat /sys/fs/cgroup/memory.events` read by hand, and the `OOMKilled` flag recorded beside it. Evidence off the box before teardown; teardown on box, host and hub.
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. Is one extra `docker exec` per container every 5 minutes acceptable on a small box? If you do nothing: kills the flag misses stay silent, as today.
|
||||
2. Should the host-side agent read (B) be spiked now, or only after A + C have run a week?
|
||||
@@ -0,0 +1,40 @@
|
||||
# R-638 — restoring over a newer schema — design proposal (burn-down night 2026-10-05, no code)
|
||||
|
||||
Baselines read: felhom-controller `ef199c5`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6 ("replay → rollback → hold", line 596).
|
||||
|
||||
## 1. The problem
|
||||
The database loader replays a copy on top of the live database, so it only removes what the copy knows about. Measured 2026-09-23 on 9202: after docmost 0.95 → 0.96 migrated, replaying the older copy FAILED on PostgreSQL (rc 3, a new table's foreign key blocked the drop); on MariaDB (romm 5.0 → 5.3) it "succeeded" and left 12 newer tables behind (`audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`).
|
||||
|
||||
## 2. What the code does today (read in source)
|
||||
- Loader: `psql -v ON_ERROR_STOP=1 --single-transaction` / plain `mariadb` over the live DB (`appbackup/dbdump.go:740-786`). Copies are made with `--clean --if-exists` (`dbdump.go:312`).
|
||||
- **Unit restore** (the restore the hold sentence names): stop → volume tars REPLACE the named volumes (`volume rm -f` + create + untar, `backup/restore.go:154-178`) → definition from the unit, i.e. the data's own version (`restore_unit.go:362-383, 444`) → DB-only start → replay (`restore_unit.go:438-458`).
|
||||
- **Off-site restore**: same order — volumes from the scratch unit, then replay (`offbox_reconstitute.go:856, 881`); a failed volume leg restarts and stops there (`:858-862`).
|
||||
- Catalog: all 17 templates with a PostgreSQL or MariaDB data dir keep it in a NAMED volume (read in `app-catalog-felhom.eu/templates/*/docker-compose.yml`).
|
||||
- Inferred from the three above: on the two main paths the replay meets the copy's OWN schema, so R-638 is probably moot there. **Not measured** — that is the measurement the row owes.
|
||||
- Remaining exposures (inferred from source):
|
||||
1. **No-manifest fallback** `RestoreApp`: volumes back, then the WHOLE stack starts at the CURRENT definition (`restore.go:67, 76`) — a newer app can migrate the old data — then the replay runs (`restore.go:82`). This is the R-638 shape exactly.
|
||||
2. **Unit restore with a failed volume leg** still replays (`restore_unit.go:439-442` sets the error and continues to `:458`).
|
||||
3. **Rollback after a failed off-site replay** pours the NEWER pre-restore copy over the OLDER volume just put back (`offbox_reconstitute.go:898`, `:456-463`) — the reverse direction; tables the migration removed would stay.
|
||||
|
||||
## 3. Options
|
||||
**A. Measure, then close the three gaps by order, not by loader.** Fallback: start only DB services at the restored volume, replay, then start the app. Unit restore: do not replay when the DB's volume leg failed.
|
||||
- Cost: small, in two files. Risk: low; no change to what the loader does.
|
||||
- Measure first: the named unit restore after a real migration (docmost, romm) on 9202.
|
||||
|
||||
**B. Make the loader rebuild instead of overlay.** PostgreSQL: `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the copy, in ONE transaction (measured working: rc 0, 1.38 s). MariaDB: drop every table first, then load.
|
||||
- Cost: medium. Fixes every caller at once, the rollback included.
|
||||
- Can go wrong: on MariaDB a failed load now leaves an EMPTY database (DDL is not transactional, `07` line 619) — the rollback must catch it. On PostgreSQL ≥ 15 the app user may not own `public` (speculative; must be measured per app). Objects in other schemas stay (immich-style extensions — unmeasured).
|
||||
|
||||
**C. Refuse a replay when the copy's version is older than the live one.** Uses the unit's data version record (`restore_unit.go:362`). Cost: small. Leaves the household with no restore at all in that case.
|
||||
|
||||
## 4. The pick — PROPOSAL for the operator, not a decision
|
||||
Option A. Read in source, both shipped restore paths already put the copy's own database files back before they load the copy. So the loader is not the weak point; the order on three side paths is. A keeps the loader the drills have proven, and it adds no new delete step on customer data. B is the fallback if the measurement shows the main paths still fail.
|
||||
|
||||
## 5. First slice and its proof
|
||||
- Slice 0, measurement only (9202, throwaway docmost + romm): copy at the old version → update and migrate → unit restore. Positive observable: the replay log line `Imported DB dump` AND a read-back of a row written before the copy. Control from a different channel: `\dt` / `SHOW TABLES` counted against the copy's own table list (the 6 and 12 newer tables must be GONE). Evidence off the box before teardown.
|
||||
- Slice 1 (fallback order): red test first — a fake stack provider records call order; assert no full `StartStack` happens before the replay in `RestoreApp`. Fails today (`restore.go:76` before `:82`).
|
||||
- Slice 2: red test — a unit restore whose volume leg errors must NOT call the importer. Fails today.
|
||||
|
||||
## 6. Open questions for the operator
|
||||
1. If the measurement shows the main paths are safe, may the row close on A alone, with B kept as a note? If you do nothing: the three side paths stay as they are.
|
||||
2. The rollback in the reverse direction (newer copy over an older volume): fix it now, or record it as a known limit?
|
||||
Reference in New Issue
Block a user