2026-09-16T21:54:59Z ================ ROUND 6 : backup-system adventurelog, while: nothing ================
2026-09-16T21:55:00Z --- BEFORE --- containers=26  travel=200  status=200  paste=200
2026-09-16T21:55:00Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T21:55:01Z --- ACTION: backup-system on adventurelog ---

## MID-BACKUP measurements — what a household actually experiences (21:56:55Z)
The whole-system backup started **21:55:24Z** (`vzdump-lxc-9201-2026_09_16-23_55_24.tar.dat`, local
time 23:55:24 CEST, growing through 2.1 GB) and while it ran:

    containers running    **4** (was 26)      — the apps are STOPPED for the whole-guest backup
    vzdump processes      running, plus its .tmp working directory
    front doors, LAN      **301** — traefik is one of the four still up, so the box still answers
    front doors, public   **404** — nothing behind the proxy to serve
    alarms                **none** — newest event still `controller_started` 21:53

**Two findings in that, and they point opposite ways.**

**1. The downtime is real and total.** A whole-system backup stops twenty-two of the twenty-six
containers. For as long as it runs, every app is unavailable — `status`, `paste` and `travel` all
answered **404** through the public route. This is the subject of the standing whole-system-backup
downtime row, seen here on a fresh box with twelve apps: not a brief pause, but the apps down for the
duration of a multi-gigabyte dump.

**2. The suppression works.** Twenty-two apps went down at once and **not one alarm fired**. That is
correct and deliberate: the backup's own stack stops belong to a suppression set, so the box does not
alarm about downtime it caused on purpose. The distinction this project keeps re-learning — a third
way to stop an app needs a third suppression set — held here.

The duration is taken from the round's own runner when it reports, not estimated from this sample.

## I STOPPED THE BACKUP — a declared intervention, with the arithmetic that forced it
Two samples thirty seconds apart, taken because the dump was being written to the VM's own root
filesystem:

    21:58:33Z   free **4138 MB**   dump 3 602 939 904 B (3.36 GiB)
    21:59:03Z   free **3663 MB**   dump 4 100 403 200 B (3.82 GiB)
    ->  the archive grew ~497 MB and free space fell ~475 MB in 30 s = **~16 MB/s**
    ->  at that rate `/` (14 G total) would be **FULL in under four minutes**

And the guest's config says what it was trying to fit:
    rootfs  local-lvm:vm-9201-disk-0, 32G   (944 M used)
    **mp0    local-lvm:vm-9201-disk-1, mp=/var/lib/felhom, backup=1, size=70G**  <- 40.58 % used ≈ **28 GB**
    mp8     /mnt/felhom-drives   — no `backup=1`, correctly EXCLUDED (the household's data drive)
    mp9     bootstrap, read-only

So a whole-guest dump of ~29 GB of real data was being written into a filesystem with 3.6 GB left.
**It could not fit, and finishing was never possible** — the only question was whether it would fill
`/` on the nested PVE and wedge the box, which is the Phase 0 failure one level up and would have
cost every remaining round of the night.

**Decision: stopped deliberately, and counted as an intervention.** Not a product failure and not
dressed up as one — the box was doing exactly what it was asked to do. What was wrong is the shape of
my fixture: a 32 GB system disk gives `pve-root` ~14 GB, while the guest's backed-up volume is 70 GB
provisioned with 28 GB used.

**What round 6 still legitimately measured, and it is the valuable half:** with the backup running,
**4 of 26 containers** were up, the public route returned **404** for every app, and **no alarm
fired** — the suppression of the backup's own stack stops held. The downtime is real, total, and
correctly silent. What is NOT measured is the backup's duration or its completion, and the round says
so rather than implying a finished backup.

## CORRECTION — I stopped a LEG of the backup, not the backup, and the box carried on by itself
The task list settles what actually happened:

    UPID …6AAB104C  vzdump 9201  started 23:55:24 (21:55:24Z)  **ended 23:59:45 — status `job errors`**
    (that is the local-storage leg, and 23:59:45 is when my `pkill` landed)

and then, seconds later, a **different** vzdump was already running:

    86528  task UPID:…:vzdump:9201:felhom-agent@pve!agent:
    86604  /usr/bin/proxmox-backup-client backup … root.pxar:/mnt/vzsnap0
           --include-dev /mnt/vzsnap0/./var/lib/felhom
           **--repository felhom@pbs!tester-1@10.77.0.1:felhom-offsite --ns tester-1** --crypt-mode=encrypt

**So the whole-guest backup has two legs, and only one of them needed local space.** The local leg
was writing a ~29 GB source into a filesystem with 3.6 GB left and could never have finished — the
arithmetic that made me stop it stands. But the off-site leg streams to ep0 over WireGuard, encrypted,
and needs no room on `/` at all. **The box's design copes with exactly the problem I thought I was
rescuing it from**, and my note above said „I stopped the backup" when what I stopped was its local
leg. Corrected here rather than left standing.

**What my intervention did and did not cost:** it ended a leg that could not have succeeded, and it
did not stop the off-site leg, which began by itself and is the one that carries this box's data out
of the house. It is still counted as an intervention — I reached in and killed a running job.

## The box came back on its own
    22:00:52Z  containers = 22
    22:01:13Z  containers = **26**  — back to the full household, unaided
    free on `/` recovered 2539 MB -> **7573 MB** once the partial archive was removed
    the older completed backup (623 MB, 22:26) was left untouched

## The waiting rule for the off-site leg — decided NOW, before the clock forces it
The off-site leg streams roughly 28 GB to ep0 over WireGuard. It may finish in minutes or run for
an hour; I cannot tell yet, and I do not want to be choosing between „wait" and „press on" at 00:30
with rounds queuing behind me.

**The rule, fixed in advance:**
1. Round 7 does NOT start while the off-site leg is running. Its drawn accident is *internet gone for
   ten minutes*, which would cut the very transfer being measured — two experiments spoiling each
   other, and neither answerable afterwards.
2. **If the leg is still running at 00:45 CEST**, round 7 starts anyway, and the round records that
   its accident interrupted an in-flight off-site backup. That is then a REAL finding about what a
   network cut does to a running whole-guest backup — worth having, provided it is labelled as such
   and not mistaken for a clean internet-cut round.
3. Either way the remaining rounds keep their ~25-minute spacing from whenever round 7 starts, and
   the 05:00 stop is unchanged. If that means fewer than twelve rounds run, the morning verdict says
   how many actually ran rather than implying all twelve did.

Written before the situation arises, because a rule invented at the moment it binds is not a rule.

## Interim confirmation that the correction was right (00:03 CEST / 22:03Z)
While the off-site leg runs:
    22:02:09Z  proxmox-backup-client procs=2   apps=**26**   free on `/` = **7573M**
    22:02:40Z  procs=2                          apps=26       free = 7573M
    22:03:12Z  procs=2                          apps=26       free = 7573M

**Free space does not move.** That is the difference between the two legs stated as a measurement
rather than an argument: the local leg consumed ~16 MB/s of `/` and would have filled it in minutes;
the off-site leg has been running for several minutes and has taken **nothing**, because it streams
encrypted to ep0 instead of writing an archive locally. The household's apps are all back up (26)
and serving while it runs.

So the whole-guest backup on this box is not „too big to work" — it is too big for the LOCAL target,
and the tier that matters for disaster recovery is unaffected by that. My first conclusion was wider
than the evidence, and this is the evidence that narrows it.

ep0 is deliberately NOT queried to watch the snapshot land: the box's own task status answers the
same question when the leg ends, and the fence for tonight is that ep0 is written only by the
product's own path and read only when nothing else can answer.
  POST /api/guest-backup/trigger -> 200
{"data":{"started":true},"ok":true}

  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"running"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":null,"error":"","job_id":"backup-9201-1789595724110220109","phase":"snapshotted"},"ok":true}
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent
## ROUND 6 — the five things
**action:** `backup-system` (whole-guest backup; drawn app adventurelog is which row is read after)
**accident:** none — control round

1. **What the customer saw.** Their apps went away and came back. During the backup **4 of 26**
   containers were up and every app answered **404** through the public route; by 22:01:13Z all 26 were
   back and serving. No banner, no mail — and correctly so.
2. **What the box did by itself.** Quiesced the apps, took an LVM snapshot, ran the local tier,
   **failed** it, announced the failure with a retry schedule, then ran the **off-site tier from the
   snapshot with the apps already back up** — 21:59:54Z → **22:08:27Z, about 8½ minutes**, encrypted to
   ep0, consuming **no local disk at all** (free stayed at 7573M throughout).
3. **Time to steady.** Apps down from ~21:55:30Z to **22:01:13Z ≈ 5m43s**. **This number is
   contaminated by my own intervention** — I killed the local leg at 21:59:45, and the apps returned
   shortly after — so it is an upper bound on the quiesce window, not a clean measurement of it.
4. **Alarm fired / true?** **`whole_guest_backup_failed` (error, operator-only) at 21:59:**
   „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)".
   **True, and better than true: it names the TIER that failed rather than 'the backup', and it says
   what it will do next.** The controller's own status surface agreed — `target_id:"local"`,
   `success:false`, `size_bytes:0`, quoting the failing task's UPID. Nothing claimed success.
   **No alarm fired for the twenty-two apps it stopped** — correct: the backup's own stack stops are
   suppressed, so the box does not alarm about downtime it caused deliberately.
5. **Should have fired and did not.** **None.** The off-site tier succeeded and needed no alarm.

**Household loop:** kept sampling throughout, `ok http=301` every two minutes — but see its recorded
blind spot: it does not follow redirects, so it cannot see that the PUBLIC route was 404 while the
apps were stopped. Its silence here is not evidence the household was unaffected.

**The finding worth carrying forward:** a whole-guest backup on this box **cannot use its local tier**
(a ~29 GB source into a 14 GB root filesystem), and the product handles that honestly — it fails the
tier, says which tier, schedules a retry, and still gets the data out of the house on the off-site
tier. The local tier will keep retrying and keep failing on a box shaped like this one.
":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"local","vmid":9201,"archive":"","mode":"snapshot","crash_consistent":true,"size_bytes":0,"success":false,"error":"proxmox: task UPID:chaosnight:00014228:0002AC68:6AAB104C:vzdump:9201:felho
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125
## A SAFETY GUARD added before round 7 — declared, and not a measurement
The failed local tier announced its own retry: „next attempt in **15m0s**" from 21:59, i.e. about
**22:14Z**, which falls inside round 7. That retry will fail exactly as before — a ~29 GB source into
a 14 GB root filesystem — but on the way it consumes `/` at ~16 MB/s, and round 7 has the box's
network cut for ten minutes, so nobody would be watching.

Rather than reach in a second time mid-round, a guard now runs on the box:
    unit `diskguard`, active; floor **2500 MB**; free at start **7573 MB**
    it kills a dump ONLY if free space falls under the floor AND a local `.tar.dat` is being written,
    then removes the partial archive, and logs every action to /root/diskguard.log

**What it is and is not.** It is insurance against my fixture's known defect wrecking the remaining
rounds; it is **not** part of any round's result. If it ever fires, that is an intervention and will
be counted as one, with its log line quoted. If it never fires, it changed nothing. Either way the
product's own behaviour — fail the tier, name it, retry with backoff — is what is being measured, and
the guard does not touch the off-site tier at all.
,"success":true,"started_at":"2026-09-1
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
  {"data":{"backup":{"target_id":"felhom-pbs","vmid":9201,"archive":"felhom-pbs:backup/ct/9201/2026-09-16T21:59:54Z","mode":"snapshot","crash_consistent":true,"size_bytes":20811501125,"success":true,"started_at":"2026-09-1
2026-09-16T22:11:38Z --- ACCIDENT: none — control round, deliberately ---
2026-09-16T22:11:38Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T22:11:39Z     t+1s containers=26 (before 26)
2026-09-16T22:11:39Z     STEADY after 1s
2026-09-16T22:11:40Z     front doors: travel=502  status=502  paste=502  wiki=502
2026-09-16T22:11:40Z     household lines this round: 16  failures: 0
2026-09-16T22:11:40Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
  | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
  | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
  | Sep 16 21:22 | info | app_removed | Alkalmazás eltávolítva: bookstack | controller
2026-09-16T22:11:41Z ================ END ROUND 6 ================
