chaos night round 10: a restore leaves no record, and four of my instruments failed
gates / gates (push) Successful in 21s

The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.

R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.

Four instrument faults, all mine, all in the evidence:
  * a 'nothing was logged' claim that was unfalsifiable when written - the log
    stream holds zero lines before a reset;
  * an on-disk check against /opt/felhom/data, a directory that does not exist;
  * a household count reporting 0 lines and 0 failures when the truth was one
    line and it WAS a failure - the runner now prints both operands;
  * the disk guard was a TRANSIENT unit reporting 'active' all night, and was
    absent from the reset onward. It is now file-backed and enabled, and its
    script is copied off the box for the first time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 01:35:56 +02:00
parent f973fd7151
commit 9f40dc3289
5 changed files with 145 additions and 0 deletions
@@ -427,6 +427,39 @@ real house an ISP outage does not do that — controller and agent share one mac
gone" as injected is **broader than its name**: internet, hub *and* local agent. No further internet
cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9.
### Round 10 — `restore` uptime-kuma / accident: **hard reset, four seconds into the restore**
**23:26:06Z–23:29:46Z.** The roughest pair drawn. The restore was accepted at 23:26:08Z
(302, „Visszaállítás elindítva"); the reset button was pressed at 23:26:12Z, mid-write.
| the five things | |
|---|---|
| what the customer saw | They pressed restore, were told it had started, and **four seconds later the whole machine went dark.** About two minutes of nothing. Then every app was back and every front door answered. **Nothing ever told them what became of the restore.** |
| what the box did by itself | Booted, and brought **26 of 26 containers** back with no help. Boot reconciliation named the one app it had to recover („1 app(s) recovered in 1 attempt(s): [paperless-ngx]"), sent a startup hub report at 23:28:26Z, and settled its health probes. No intervention. |
| time to steady | **150 s** — 0 containers at t+12 s, 25 at t+133 s, 26 at t+150 s. Doors 200 at both readings (23:28:42Z and 23:29:44Z). |
| alarm fired / true? | **one, true** — `controller_started` (info). Exactly what the ladder expects after a reboot. No false alarm. |
| should have fired, did not | **none from the alarm ladder** — but the restore silence below is a legibility gap, filed as a row. |
**The restore left no trace anywhere, and the product has no place to leave one.** Four candidate
status endpoints all 404 (`/api/restore/status`, `/api/backup/restore/status`,
`/backup/restore/status`, `/api/restore`). `/api/backup/status` carries no restore field at all.
On the pages, the only restore text is a **button label** and a JavaScript label expression. On disk,
in the real data directory, there is no restore, lock or state file anywhere — and **no file at all
was modified in the reset window**. An interrupted restore and a restore that never happened are
indistinguishable, to the customer and to me.
**The limit of that measurement, stated rather than glossed.** Only four seconds elapsed, so the
restore may have finished or may never have written a byte — and I cannot tell, because the
controller's log stream holds **zero lines before 23:28:00Z** (a reset starts it fresh) and the debug
ring died with the machine. What is independently verifiable is the **absence of any restore record**,
and that is what is filed; it holds however far the restore got.
**Four of my own instruments failed in this round, and all four are recorded in the evidence:** a
claim that was unfalsifiable when written; an on-disk check against a directory that does not exist;
a household count that reported 0 lines and 0 failures when the truth was one line and it *was* a
failure; and a disk guard that reported „active" all night while being a **transient** unit that
vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box.
### Rounds 8-12
PENDING
@@ -0,0 +1,20 @@
#!/bin/bash
# SAFETY, not a measurement: the local backup tier retries every ~15 minutes and cannot fit on this
# box (a ~29 GB source into a 14 GB root filesystem). Left alone during a round it would fill `/` and
# wedge the nested PVE, which already nearly cost the night once. This kills ONLY a local-tier dump,
# and only when free space falls under the floor. Every kill is logged so it appears in the evidence.
FLOOR_MB=2500
LOG=/root/diskguard.log
while true; do
FREE=$(df -BM --output=avail / | tail -1 | tr -dc '0-9')
if [ -n "$FREE" ] && [ "$FREE" -lt "$FLOOR_MB" ]; then
if ls /var/lib/vz/dump/*.tar.dat >/dev/null 2>&1; then
echo "$(date -u +%FT%TZ) GUARD: free=${FREE}M below ${FLOOR_MB}M and a local dump is writing — killing it" >> $LOG
pkill -f "[v]zdump"
sleep 3
rm -rf /var/lib/vz/dump/*.tar.dat /var/lib/vz/dump/*.tmp 2>/dev/null
echo "$(date -u +%FT%TZ) GUARD: partial archive removed, free now $(df -BM --output=avail / | tail -1)" >> $LOG
fi
fi
sleep 20
done
@@ -0,0 +1,87 @@
round 10 armed for 2026-09-16T23:25:48Z (restore uptime-kuma + hard reset)
=== round 10 launched 2026-09-16T23:26:06Z (due 23:25:48Z) ===
2026-09-16T23:26:06Z ================ ROUND 10 : restore uptime-kuma, while: hard-reset ================
2026-09-16T23:26:08Z --- BEFORE --- containers=26 status=200 status=200 paste=200
2026-09-16T23:26:08Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
2026-09-16T23:26:08Z --- ACTION: restore on uptime-kuma ---
POST /backup/restore (uptime-kuma) -> HTTP/1.1 302 Found
flash: Visszaállítás elindult — az állapot itt frissül.
2026-09-16T23:26:12Z --- ACCIDENT: hard-reset (injected after the action started) ---
2026-09-16T23:26:12Z ACCIDENT=hard-reset round=10
2026-09-16T23:26:12Z qm reset 336 (the reset button, mid-write)
2026-09-16T23:26:14Z reset issued
2026-09-16T23:26:14Z accident hard-reset complete
2026-09-16T23:26:14Z --- AFTER: what the box did BY ITSELF ---
2026-09-16T23:26:24Z t+12s containers=0 (before 26)
2026-09-16T23:28:07Z t+115s containers=0 (before 26)
2026-09-16T23:28:25Z t+133s containers=25 (before 26)
2026-09-16T23:28:42Z t+150s containers=26 (before 26)
2026-09-16T23:28:42Z STEADY after 150s
2026-09-16T23:28:42Z front doors, FIRST reading at 2026-09-16T23:28:42Z - TOO EARLY to trust if the accident just ended:
2026-09-16T23:28:44Z status=200 status=200 paste=200 wiki=200
2026-09-16T23:29:44Z front doors, SECOND reading at 2026-09-16T23:29:44Z, 60 s later - THIS is the one to trust:
2026-09-16T23:29:45Z status=200 status=200 paste=200 wiki=200
2026-09-16T23:29:45Z household lines this round: 0 failures: 0
2026-09-16T23:29:45Z --- alarms ---
| Time | Severity | Type | Message | Source
| Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
| Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
2026-09-16T23:29:46Z ================ END ROUND 10 ================
[exited with code 0]
## WHAT HAPPENED TO THE RESTORE (asked four ways, 23:30-23:34Z)
The restore was accepted at 23:26:08Z (HTTP 302, flash "Visszaallitas elindult").
The machine was hard-reset at 23:26:12Z - FOUR SECONDS later, mid-write.
Asked afterwards, through the same doors the UI uses:
/api/restore/status -> 404
/api/backup/restore/status -> 404
/backup/restore/status -> 404
/api/restore -> 404
/api/backup/status -> {"ok":true,"data":{"enabled":true,"running":false}} (no restore field at all)
/backups/apps (200) -> the only "Vissza" text is a BUTTON LABEL and a JS label expression:
"Visszaallitas"
": op === 'offbox-restore' ? 'Tavoli visszaallitas' : 'Visszaallitas'; }"
/apps/uptime-kuma (200) -> one "Vissza" (the button). "megszak" (interrupted): 0 hits.
The "sikeres" hits belong to the MOVE-DATA feature, not the restore.
NEGATIVE CONTROL "zzzznotpresent" -> 0 hits on every page, so the searches are trustworthy.
On disk, in the REAL data directory (/var/lib/docker/volumes/felhom-controller-data/_data):
no file named *restor*, *lock* or *.pid anywhere beneath it
NO FILE AT ALL modified in the window 23:24:00 - 23:27:30
So: there is no restore history surface of any kind. An interrupted restore and a restore that
never happened look EXACTLY the same to the customer, and to me.
### THE HONEST LIMIT OF THIS MEASUREMENT, which must not be glossed over
Only four seconds elapsed. The restore may have completed, or may never have written a byte.
I cannot tell, because:
* the controller's log stream holds ZERO lines before 23:28:00Z - a hard reset starts it fresh,
so the pre-reset window is unrecoverable by this route;
* the debug ring is in memory and died with the machine.
What IS independently verifiable, and is what gets filed, is the ABSENCE OF ANY RESTORE RECORD -
that holds regardless of how far the restore got.
## FOUR CORRECTIONS TO MY OWN INSTRUMENTS, all found in this round
1. "The controller logged nothing about the restore" was UNFALSIFIABLE when I first wrote it.
The stream cannot reach before the reset. Re-stated above with its limit.
2. My first on-disk check looked in /opt/felhom/data - A DIRECTORY THAT DOES NOT EXIST. It printed
a tidy "no restore/lock file" that meant nothing. The real path is the docker volume above.
Same class as the "docker: command not found -> a tidy table of absent/0" error from earlier.
3. The runner reported "household lines this round: 0 failures: 0". Both are WRONG. The log is
never truncated (first line 21:06:30Z, continuous) and contains exactly ONE line in the round
window - and it is a FAILURE line: "2026-09-16T23:28:09Z cloud dash UNREACHABLE" (the box was
still booting). The runner now prints the raw before/after counts so the subtraction is checkable.
4. THE DISK GUARD WAS NOT A REAL UNIT. After the reset: "Unit diskguard.service not found" - yet it
had reported "active" all night. It was a TRANSIENT unit, which looks identical while running and
vanishes on reboot. It was therefore ABSENT from 23:26:12Z. It is now a file-backed, ENABLED unit
(verified: active + enabled, MainPID cmdline "/bin/bash /root/diskguard.sh", 0 kills, 7556 MB free)
and its script is copied off the box into this folder, which had also never been done (R-320).
@@ -92,6 +92,10 @@ sleep 60
say " front doors, SECOND reading at $(date -u +%FT%TZ), 60 s later - THIS is the one to trust:"
say " $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)"
HL1=$(BOX 'wc -l < /root/household.log' | tr -d ' \r')
# Round 10 reported "0 lines, 0 failures" when the truth was one line and it WAS a
# failure. A subtraction whose operands are invisible cannot be checked, so both are
# printed now, and an empty read is named rather than silently becoming zero.
say " household log lines: before=${HL0:-EMPTY} after=${HL1:-EMPTY}"
say " household lines this round: $(( ${HL1:-0} - ${HL0:-0} )) failures: $(BOX "tail -n +$(( ${HL0:-0} + 1 )) /root/household.log | grep -cE 'FAILED|UNREACHABLE'" | tr -d ' \r')"
say "--- alarms ---"; bash $E/events.sh 8 | tee -a $OUT
say "================ END ROUND $N ================"