  round 10 armed for 2026-09-16T23:25:48Z (restore uptime-kuma + hard reset)
=== round 10 launched 2026-09-16T23:26:06Z (due 23:25:48Z) ===
2026-09-16T23:26:06Z ================ ROUND 10 : restore uptime-kuma, while: hard-reset ================
2026-09-16T23:26:08Z --- BEFORE --- containers=26  status=200  status=200  paste=200
2026-09-16T23:26:08Z     (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T23:26:08Z --- ACTION: restore on uptime-kuma ---
  POST /backup/restore (uptime-kuma) -> HTTP/1.1 302 Found
  flash: Visszaállítás elindult — az állapot itt frissül.
2026-09-16T23:26:12Z --- ACCIDENT: hard-reset (injected after the action started) ---
    2026-09-16T23:26:12Z ACCIDENT=hard-reset round=10
    2026-09-16T23:26:12Z qm reset 336 (the reset button, mid-write)
    2026-09-16T23:26:14Z reset issued
    2026-09-16T23:26:14Z accident hard-reset complete
2026-09-16T23:26:14Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T23:26:24Z     t+12s containers=0 (before 26)
2026-09-16T23:28:07Z     t+115s containers=0 (before 26)
2026-09-16T23:28:25Z     t+133s containers=25 (before 26)
2026-09-16T23:28:42Z     t+150s containers=26 (before 26)
2026-09-16T23:28:42Z     STEADY after 150s
2026-09-16T23:28:42Z     front doors, FIRST reading at 2026-09-16T23:28:42Z - TOO EARLY to trust if the accident just ended:
2026-09-16T23:28:44Z       status=200  status=200  paste=200  wiki=200
2026-09-16T23:29:44Z     front doors, SECOND reading at 2026-09-16T23:29:44Z, 60 s later - THIS is the one to trust:
2026-09-16T23:29:45Z       status=200  status=200  paste=200  wiki=200
2026-09-16T23:29:45Z     household lines this round: 0  failures: 0
2026-09-16T23:29:45Z --- alarms ---
  | Time | Severity | Type | Message | Source
  | Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
  | Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
  | Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
  | Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
  | Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
  | Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
2026-09-16T23:29:46Z ================ END ROUND 10 ================

[exited with code 0]

## WHAT HAPPENED TO THE RESTORE (asked four ways, 23:30-23:34Z)

The restore was accepted at 23:26:08Z (HTTP 302, flash "Visszaallitas elindult").
The machine was hard-reset at 23:26:12Z - FOUR SECONDS later, mid-write.

Asked afterwards, through the same doors the UI uses:
  /api/restore/status          -> 404
  /api/backup/restore/status   -> 404
  /backup/restore/status       -> 404
  /api/restore                 -> 404
  /api/backup/status           -> {"ok":true,"data":{"enabled":true,"running":false}}   (no restore field at all)
  /backups/apps (200)          -> the only "Vissza" text is a BUTTON LABEL and a JS label expression:
                                    "Visszaallitas"
                                    ": op === 'offbox-restore' ? 'Tavoli visszaallitas' : 'Visszaallitas'; }"
  /apps/uptime-kuma (200)      -> one "Vissza" (the button). "megszak" (interrupted): 0 hits.
                                  The "sikeres" hits belong to the MOVE-DATA feature, not the restore.
  NEGATIVE CONTROL "zzzznotpresent" -> 0 hits on every page, so the searches are trustworthy.

On disk, in the REAL data directory (/var/lib/docker/volumes/felhom-controller-data/_data):
  no file named *restor*, *lock* or *.pid anywhere beneath it
  NO FILE AT ALL modified in the window 23:24:00 - 23:27:30

So: there is no restore history surface of any kind. An interrupted restore and a restore that
never happened look EXACTLY the same to the customer, and to me.

### THE HONEST LIMIT OF THIS MEASUREMENT, which must not be glossed over
Only four seconds elapsed. The restore may have completed, or may never have written a byte.
I cannot tell, because:
  * the controller's log stream holds ZERO lines before 23:28:00Z - a hard reset starts it fresh,
    so the pre-reset window is unrecoverable by this route;
  * the debug ring is in memory and died with the machine.
What IS independently verifiable, and is what gets filed, is the ABSENCE OF ANY RESTORE RECORD -
that holds regardless of how far the restore got.

## FOUR CORRECTIONS TO MY OWN INSTRUMENTS, all found in this round
1. "The controller logged nothing about the restore" was UNFALSIFIABLE when I first wrote it.
   The stream cannot reach before the reset. Re-stated above with its limit.
2. My first on-disk check looked in /opt/felhom/data - A DIRECTORY THAT DOES NOT EXIST. It printed
   a tidy "no restore/lock file" that meant nothing. The real path is the docker volume above.
   Same class as the "docker: command not found -> a tidy table of absent/0" error from earlier.
3. The runner reported "household lines this round: 0  failures: 0". Both are WRONG. The log is
   never truncated (first line 21:06:30Z, continuous) and contains exactly ONE line in the round
   window - and it is a FAILURE line: "2026-09-16T23:28:09Z cloud dash UNREACHABLE" (the box was
   still booting). The runner now prints the raw before/after counts so the subtraction is checkable.
4. THE DISK GUARD WAS NOT A REAL UNIT. After the reset: "Unit diskguard.service not found" - yet it
   had reported "active" all night. It was a TRANSIENT unit, which looks identical while running and
   vanishes on reboot. It was therefore ABSENT from 23:26:12Z. It is now a file-backed, ENABLED unit
   (verified: active + enabled, MainPID cmdline "/bin/bash /root/diskguard.sh", 0 kills, 7556 MB free)
   and its script is copied off the box into this folder, which had also never been done (R-320).

## CORRECTION TO THIS ROUND'S CENTRAL CLAIM, 2026-09-17T00:24Z — I WAS WRONG
Above I wrote that four candidate status endpoints all 404 and that "there is no restore history
surface of any kind". That was built on FOUR PATHS I GUESSED, and all four were wrong. The real
route, taken from the restore page's own JavaScript rather than from my imagination, is:
    /api/backup/restore-status
It exists, it answers, and this is what it returns now:
    {"ok":true,"data":{"running":false,"started_at":"0001-01-01T00:00:00Z"}}

So the corrected finding is NARROWER and better than the one I filed:
  * a restore status surface DOES exist;
  * after the reboot it is EMPTY - `started_at` is the Go zero value 0001-01-01T00:00:00Z;
  * the payload carries no `last` field at all, yet the page's own script reads `st.last.op` and
    `st.last.message` to render "<operation> sikertelen." So there is a "last operation" branch in
    the UI with nothing to populate it after a restart.
The record is therefore IN-MEMORY ONLY and does not survive the machine stopping - which is exactly
the case a hard reset creates, and exactly when a customer would most want to know.

### The honest limit is unchanged, and now cuts both ways
Only four seconds elapsed. `started_at` may be zero because the restore never really began, not
because the reboot erased it. I cannot separate those, because the pre-reset log is unrecoverable.
What IS certain: the surface exists, it is blank now, and nothing anywhere tells the customer that a
restore they started did not finish.

### How the error happened, because that is the reusable part
I searched for the page at `/apps/uptime-kuma` and guessed API paths by pattern. The controller's
actual routes are `/backups`, `/backups/remote`, `/backups/restore` and `/stacks/<name>/backup`.
I found them in the end by asking the controller for its OWN rendered links instead of guessing -
which is what I should have done first. Guessing produced four confident 404s that I then reported
as a property of the product.
