Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
4.8 KiB
R-518 — per-tier quiesce — design proposal (burn-down night 2026-10-05, no code)
Baselines read: felhom-controller ef199c5, felhom-agent 861d32a, felhom.eu b37902ce. Architecture: 07-backup-architecture.md §6.4.
1. The problem
„Mentés most" stops every app until the slow local copy has fully finished, although the copy's snapshot is ready within seconds. Measured 2026-10-05 on demo-hp (9 apps): the agent reported snapshotted at the first sample, 12 s after the local job started (09:19:29 → 09:19:41), but the apps stayed down until 09:24:09 and the last one was back at 09:24:55 — 5 min 47 s (audits/hub-safety-2026-10-05/partE/r518-watch.log, r518-measure.txt). BIGNIGHT measured ≈ 7 min 45 s on 12 apps (evidence-bignight-2026-09-14/phase4/guest-backup-quiesce-log.txt:3-45).
2. What the code does today (read in source)
- A manual press covers every tier in one window (
quiesce.go:428-431,allTiersForManualRun). - ONE stop for all due tiers, tiers run one after another (
quiesce.go:487-599). - Early resume at
snapshottedhappens ONLY on the last tier (quiesce.go:630-636). A non-last tier waits fordone(quiesce.go:643), because the agent holds the guest lock until the upload ends. - Order is primary (local) first, so the long local upload always runs with the apps down.
- This is a recorded R-82 choice: „ONE quiesce window for both due tiers (never two app outages for one night)" (
felhom.eu/CONTEXT.md:2709). Two tests pin it:TestBothTiersDue_ExactlyOneQuiesceWindow(tiers_test.go:140),TestNonLastTierSnapshot_DoesNotResumeApp(tiers_test.go:205). - After a successful primary copy the agent runs the OS leg under the same heavy-op gate (
felhom-agent/internal/localapi/server.go:904). Inferred: that is why the PBS tier was BUSY at 09:24:09 on demo-hp.
3. Options
A. One window per tier. Stop → start tier → resume at its snapshotted → let the upload finish with apps up → next tier gets its own stop.
- Changes: the loop body; two tests are rewritten to pin the new rule.
- Costs: two short outages instead of one long one when both tiers are due. Inferred per outage from the demo-hp parts: stop 21 s + snapshot ≤ 12 s + restart 46 s ≈ 80 s.
- Can go wrong: a second stop is wasted if the agent refuses tier 2 as BUSY (seen today). Mitigation: run tier 2 in a later cycle, not straight after tier 1.
- Measure first: the stop-to-
snapshottedtime on the PBS tier (never measured; BIGNIGHT's PBS run failed in 10 s, today's was refused). - Every copy stays app-consistent (each tier is taken with the apps stopped).
B. Keep one window, resume at the first tier's snapshot. The second tier then copies RUNNING apps.
- Costs: one line of logic. Loses app-consistency on the off-site (disaster) copy. That changes risk to customer data — not CC's call.
C. Do nothing more. The page already states the measured minutes (controller v0.296.0). Every press still costs ≈ 6-8 min of no apps.
4. The pick — PROPOSAL for the operator, not a decision
Option A, with tier 2 left to the next cycle. It keeps the reason for one window (every copy app-consistent) and drops the cost (4-5 minutes of upload with apps down). It does reverse the recorded R-82 choice „never two app outages for one night", so it needs the operator's word. At night, two outages of about 80 s each are less visible than one of 6-8 min. With the operator's word, the R-82 note and §6.4 change in the same commit.
5. First slice and its proof
- Build:
quiesceAndPollTiersruns only the FIRST due tier per window and resumes at itssnapshotted; the other tiers stay due and the next cycle (5 min poll) picks them up. A manual press keeps covering all tiers, but each in its own window. - Red test first (must FAIL on today's code): two tiers due; local reports
snapshotted, snapshotted, snapshotted, done. Assert the stacks were STARTED before the secondsnapshottedpoll answered, i.e. the apps run while local still uploads. Today it fails: start comes only after the last tier. - Keep green: crash marker before any stop (
quiesce.go:490-494); one unquiesce per window; the max-quiesce bound. - Live proof on scratch 9202 (throwaway apps only): press the button; every 5 s sample (a) one throwaway app over HTTP, (b) the agent job phase. Positive observable: the app answers 200 while the phase reads
snapshotted. Control from a different channel: containerStartedAtfromdocker inspect, against the controller log. Evidence off the box before teardown.
6. Open questions for the operator
- Two short app stops in one night instead of one long one — acceptable? If you do nothing: every press keeps every app down for the whole local upload.
- On a manual press, should the off-site tier still run (second short stop), or only the local one?