d549082be9
gates / gates (push) Successful in 2m7s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
41 lines
4.8 KiB
Markdown
41 lines
4.8 KiB
Markdown
# R-518 — per-tier quiesce — design proposal (burn-down night 2026-10-05, no code)
|
|
|
|
Baselines read: felhom-controller `ef199c5`, felhom-agent `861d32a`, felhom.eu `b37902ce`. Architecture: `07-backup-architecture.md` §6.4.
|
|
|
|
## 1. The problem
|
|
„Mentés most" stops every app until the slow local copy has fully finished, although the copy's snapshot is ready within seconds. Measured 2026-10-05 on demo-hp (9 apps): the agent reported `snapshotted` at the first sample, 12 s after the local job started (09:19:29 → 09:19:41), but the apps stayed down until 09:24:09 and the last one was back at 09:24:55 — 5 min 47 s (`audits/hub-safety-2026-10-05/partE/r518-watch.log`, `r518-measure.txt`). BIGNIGHT measured ≈ 7 min 45 s on 12 apps (`evidence-bignight-2026-09-14/phase4/guest-backup-quiesce-log.txt:3-45`).
|
|
|
|
## 2. What the code does today (read in source)
|
|
- A manual press covers every tier in one window (`quiesce.go:428-431`, `allTiersForManualRun`).
|
|
- ONE stop for all due tiers, tiers run one after another (`quiesce.go:487-599`).
|
|
- Early resume at `snapshotted` happens ONLY on the last tier (`quiesce.go:630-636`). A non-last tier waits for `done` (`quiesce.go:643`), because the agent holds the guest lock until the upload ends.
|
|
- Order is primary (local) first, so the long local upload always runs with the apps down.
|
|
- This is a recorded R-82 choice: „ONE quiesce window for both due tiers (never two app outages for one night)" (`felhom.eu/CONTEXT.md:2709`). Two tests pin it: `TestBothTiersDue_ExactlyOneQuiesceWindow` (`tiers_test.go:140`), `TestNonLastTierSnapshot_DoesNotResumeApp` (`tiers_test.go:205`).
|
|
- After a successful primary copy the agent runs the OS leg under the same heavy-op gate (`felhom-agent/internal/localapi/server.go:904`). Inferred: that is why the PBS tier was BUSY at 09:24:09 on demo-hp.
|
|
|
|
## 3. Options
|
|
**A. One window per tier.** Stop → start tier → resume at its `snapshotted` → let the upload finish with apps up → next tier gets its own stop.
|
|
- Changes: the loop body; two tests are rewritten to pin the new rule.
|
|
- Costs: two short outages instead of one long one when both tiers are due. Inferred per outage from the demo-hp parts: stop 21 s + snapshot ≤ 12 s + restart 46 s ≈ 80 s.
|
|
- Can go wrong: a second stop is wasted if the agent refuses tier 2 as BUSY (seen today). Mitigation: run tier 2 in a later cycle, not straight after tier 1.
|
|
- Measure first: the stop-to-`snapshotted` time on the PBS tier (never measured; BIGNIGHT's PBS run failed in 10 s, today's was refused).
|
|
- Every copy stays app-consistent (each tier is taken with the apps stopped).
|
|
|
|
**B. Keep one window, resume at the first tier's snapshot.** The second tier then copies RUNNING apps.
|
|
- Costs: one line of logic. Loses app-consistency on the off-site (disaster) copy. That changes risk to customer data — not CC's call.
|
|
|
|
**C. Do nothing more.** The page already states the measured minutes (controller v0.296.0). Every press still costs ≈ 6-8 min of no apps.
|
|
|
|
## 4. The pick — PROPOSAL for the operator, not a decision
|
|
Option A, with tier 2 left to the next cycle. It keeps the reason for one window (every copy app-consistent) and drops the cost (4-5 minutes of upload with apps down). It does reverse the recorded R-82 choice „never two app outages for one night", so it needs the operator's word. At night, two outages of about 80 s each are less visible than one of 6-8 min. With the operator's word, the R-82 note and §6.4 change in the same commit.
|
|
|
|
## 5. First slice and its proof
|
|
- Build: `quiesceAndPollTiers` runs only the FIRST due tier per window and resumes at its `snapshotted`; the other tiers stay due and the next cycle (5 min poll) picks them up. A manual press keeps covering all tiers, but each in its own window.
|
|
- Red test first (must FAIL on today's code): two tiers due; local reports `snapshotted, snapshotted, snapshotted, done`. Assert the stacks were STARTED before the second `snapshotted` poll answered, i.e. the apps run while local still uploads. Today it fails: start comes only after the last tier.
|
|
- Keep green: crash marker before any stop (`quiesce.go:490-494`); one unquiesce per window; the max-quiesce bound.
|
|
- Live proof on scratch 9202 (throwaway apps only): press the button; every 5 s sample (a) one throwaway app over HTTP, (b) the agent job phase. Positive observable: the app answers 200 while the phase reads `snapshotted`. Control from a different channel: container `StartedAt` from `docker inspect`, against the controller log. Evidence off the box before teardown.
|
|
|
|
## 6. Open questions for the operator
|
|
1. Two short app stops in one night instead of one long one — acceptable? If you do nothing: every press keeps every app down for the whole local upload.
|
|
2. On a manual press, should the off-site tier still run (second short stop), or only the local one?
|