|
|
|
@@ -154,4 +154,70 @@ product's own path and read only when nothing else can answer.
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent
|
|
|
|
|
## ROUND 6 — the five things
|
|
|
|
|
**action:** `backup-system` (whole-guest backup; drawn app adventurelog is which row is read after)
|
|
|
|
|
**accident:** none — control round
|
|
|
|
|
|
|
|
|
|
1. **What the customer saw.** Their apps went away and came back. During the backup **4 of 26**
|
|
|
|
|
containers were up and every app answered **404** through the public route; by 22:01:13Z all 26 were
|
|
|
|
|
back and serving. No banner, no mail — and correctly so.
|
|
|
|
|
2. **What the box did by itself.** Quiesced the apps, took an LVM snapshot, ran the local tier,
|
|
|
|
|
**failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the
|
|
|
|
|
snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z, about 8½ minutes**, encrypted to
|
|
|
|
|
ep0, consuming **no local disk at all** (free stayed at 7573M throughout).
|
|
|
|
|
3. **Time to steady.** Apps down from ~21:55:30Z to **22:01:13Z ≈ 5m43s**. **This number is
|
|
|
|
|
contaminated by my own intervention** — I killed the local leg at 21:59:45, and the apps returned
|
|
|
|
|
shortly after — so it is an upper bound on the quiesce window, not a clean measurement of it.
|
|
|
|
|
4. **Alarm fired / true?** **`whole_guest_backup_failed` (error, operator-only) at 21:59:**
|
|
|
|
|
„Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)".
|
|
|
|
|
**True, and better than true: it names the TIER that failed rather than 'the backup', and it says
|
|
|
|
|
what it will do next.** The controller's own status surface agreed — `target_id:"local"`,
|
|
|
|
|
`success:false`, `size_bytes:0`, quoting the failing task's UPID. Nothing claimed success.
|
|
|
|
|
**No alarm fired for the twenty-two apps it stopped** — correct: the backup's own stack stops are
|
|
|
|
|
suppressed, so the box does not alarm about downtime it caused deliberately.
|
|
|
|
|
5. **Should have fired and did not.** **None.** The off-site tier succeeded and needed no alarm.
|
|
|
|
|
|
|
|
|
|
**Household loop:** kept sampling throughout, `ok http=301` every two minutes — but see its recorded
|
|
|
|
|
blind spot: it does not follow redirects, so it cannot see that the PUBLIC route was 404 while the
|
|
|
|
|
apps were stopped. Its silence here is not evidence the household was unaffected.
|
|
|
|
|
|
|
|
|
|
**The finding worth carrying forward:** a whole-guest backup on this box **cannot use its local tier**
|
|
|
|
|
(a ~29 GB source into a 14 GB root filesystem), and the product handles that honestly — it fails the
|
|
|
|
|
tier, says which tier, schedules a retry, and still gets the data out of the house on the off-site
|
|
|
|
|
tier. The local tier will keep retrying and keep failing on a box shaped like this one.
|
|
|
|
|
":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
|
|
|
|
|
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
|
|
|
|
|
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
|
|
|
|
|
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
|
|
|
|
|
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
|
|
|
|
|
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
|
|
|
|
|
{"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125
|
|
|
|
|
## A SAFETY GUARD added before round 7 — declared, and not a measurement
|
|
|
|
|
The failed local tier announced its own retry: „next attempt in **15m0s**" from 21:59, i.e. about
|
|
|
|
|
**22:14Z**, which falls inside round 7. That retry will fail exactly as before — a ~29 GB source into
|
|
|
|
|
a 14 GB root filesystem — but on the way it consumes `/` at ~16 MB/s, and round 7 has the box's
|
|
|
|
|
network cut for ten minutes, so nobody would be watching.
|
|
|
|
|
|
|
|
|
|
Rather than reach in a second time mid-round, a guard now runs on the box:
|
|
|
|
|
unit `diskguard`, active; floor **2500 MB**; free at start **7573 MB**
|
|
|
|
|
it kills a dump ONLY if free space falls under the floor AND a local `.tar.dat` is being written,
|
|
|
|
|
then removes the partial archive, and logs every action to /root/diskguard.log
|
|
|
|
|
|
|
|
|
|
**What it is and is not.** It is insurance against my fixture's known defect wrecking the remaining
|
|
|
|
|
rounds; it is **not** part of any round's result. If it ever fires, that is an intervention and will
|
|
|
|
|
be counted as one, with its log line quoted. If it never fires, it changed nothing. Either way the
|
|
|
|
|
product's own behaviour — fail the tier, name it, retry with backoff — is what is being measured, and
|
|
|
|
|
the guard does not touch the off-site tier at all.
|
|
|
|
|