Files
felhom.eu/documentation/audits/night-burndown-2026-10-05/design-R-518.md
T
2026-10-05 23:32:12 +02:00

4.8 KiB

R-518 — per-tier quiesce — design proposal (burn-down night 2026-10-05, no code)

Baselines read: felhom-controller ef199c5, felhom-agent 861d32a, felhom.eu b37902ce. Architecture: 07-backup-architecture.md §6.4.

1. The problem

„Mentés most" stops every app until the slow local copy has fully finished, although the copy's snapshot is ready within seconds. Measured 2026-10-05 on demo-hp (9 apps): the agent reported snapshotted at the first sample, 12 s after the local job started (09:19:29 → 09:19:41), but the apps stayed down until 09:24:09 and the last one was back at 09:24:55 — 5 min 47 s (audits/hub-safety-2026-10-05/partE/r518-watch.log, r518-measure.txt). BIGNIGHT measured ≈ 7 min 45 s on 12 apps (evidence-bignight-2026-09-14/phase4/guest-backup-quiesce-log.txt:3-45).

2. What the code does today (read in source)

  • A manual press covers every tier in one window (quiesce.go:428-431, allTiersForManualRun).
  • ONE stop for all due tiers, tiers run one after another (quiesce.go:487-599).
  • Early resume at snapshotted happens ONLY on the last tier (quiesce.go:630-636). A non-last tier waits for done (quiesce.go:643), because the agent holds the guest lock until the upload ends.
  • Order is primary (local) first, so the long local upload always runs with the apps down.
  • This is a recorded R-82 choice: „ONE quiesce window for both due tiers (never two app outages for one night)" (felhom.eu/CONTEXT.md:2709). Two tests pin it: TestBothTiersDue_ExactlyOneQuiesceWindow (tiers_test.go:140), TestNonLastTierSnapshot_DoesNotResumeApp (tiers_test.go:205).
  • After a successful primary copy the agent runs the OS leg under the same heavy-op gate (felhom-agent/internal/localapi/server.go:904). Inferred: that is why the PBS tier was BUSY at 09:24:09 on demo-hp.

3. Options

A. One window per tier. Stop → start tier → resume at its snapshotted → let the upload finish with apps up → next tier gets its own stop.

  • Changes: the loop body; two tests are rewritten to pin the new rule.
  • Costs: two short outages instead of one long one when both tiers are due. Inferred per outage from the demo-hp parts: stop 21 s + snapshot ≤ 12 s + restart 46 s ≈ 80 s.
  • Can go wrong: a second stop is wasted if the agent refuses tier 2 as BUSY (seen today). Mitigation: run tier 2 in a later cycle, not straight after tier 1.
  • Measure first: the stop-to-snapshotted time on the PBS tier (never measured; BIGNIGHT's PBS run failed in 10 s, today's was refused).
  • Every copy stays app-consistent (each tier is taken with the apps stopped).

B. Keep one window, resume at the first tier's snapshot. The second tier then copies RUNNING apps.

  • Costs: one line of logic. Loses app-consistency on the off-site (disaster) copy. That changes risk to customer data — not CC's call.

C. Do nothing more. The page already states the measured minutes (controller v0.296.0). Every press still costs ≈ 6-8 min of no apps.

4. The pick — PROPOSAL for the operator, not a decision

Option A, with tier 2 left to the next cycle. It keeps the reason for one window (every copy app-consistent) and drops the cost (4-5 minutes of upload with apps down). It does reverse the recorded R-82 choice „never two app outages for one night", so it needs the operator's word. At night, two outages of about 80 s each are less visible than one of 6-8 min. With the operator's word, the R-82 note and §6.4 change in the same commit.

5. First slice and its proof

  • Build: quiesceAndPollTiers runs only the FIRST due tier per window and resumes at its snapshotted; the other tiers stay due and the next cycle (5 min poll) picks them up. A manual press keeps covering all tiers, but each in its own window.
  • Red test first (must FAIL on today's code): two tiers due; local reports snapshotted, snapshotted, snapshotted, done. Assert the stacks were STARTED before the second snapshotted poll answered, i.e. the apps run while local still uploads. Today it fails: start comes only after the last tier.
  • Keep green: crash marker before any stop (quiesce.go:490-494); one unquiesce per window; the max-quiesce bound.
  • Live proof on scratch 9202 (throwaway apps only): press the button; every 5 s sample (a) one throwaway app over HTTP, (b) the agent job phase. Positive observable: the app answers 200 while the phase reads snapshotted. Control from a different channel: container StartedAt from docker inspect, against the controller log. Evidence off the box before teardown.

6. Open questions for the operator

  1. Two short app stops in one night instead of one long one — acceptable? If you do nothing: every press keeps every app down for the whole local upload.
  2. On a manual press, should the off-site tier still run (second short stop), or only the local one?