chaos night round 10: a restore leaves no record, and four of my instruments failed
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.
R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.
Four instrument faults, all mine, all in the evidence:
* a 'nothing was logged' claim that was unfalsifiable when written - the log
stream holds zero lines before a reset;
* an on-disk check against /opt/felhom/data, a directory that does not exist;
* a household count reporting 0 lines and 0 failures when the truth was one
line and it WAS a failure - the runner now prints both operands;
* the disk guard was a TRANSIENT unit reporting 'active' all night, and was
absent from the reset onward. It is now file-backed and enabled, and its
script is copied off the box for the first time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -427,6 +427,39 @@ real house an ISP outage does not do that — controller and agent share one mac
|
||||
gone" as injected is **broader than its name**: internet, hub *and* local agent. No further internet
|
||||
cuts are drawn, so the injector stays as it is and this caveat travels with rounds 7, 8 and 9.
|
||||
|
||||
### Round 10 — `restore` uptime-kuma / accident: **hard reset, four seconds into the restore**
|
||||
|
||||
**23:26:06Z–23:29:46Z.** The roughest pair drawn. The restore was accepted at 23:26:08Z
|
||||
(302, „Visszaállítás elindítva"); the reset button was pressed at 23:26:12Z, mid-write.
|
||||
|
||||
| the five things | |
|
||||
|---|---|
|
||||
| what the customer saw | They pressed restore, were told it had started, and **four seconds later the whole machine went dark.** About two minutes of nothing. Then every app was back and every front door answered. **Nothing ever told them what became of the restore.** |
|
||||
| what the box did by itself | Booted, and brought **26 of 26 containers** back with no help. Boot reconciliation named the one app it had to recover („1 app(s) recovered in 1 attempt(s): [paperless-ngx]"), sent a startup hub report at 23:28:26Z, and settled its health probes. No intervention. |
|
||||
| time to steady | **150 s** — 0 containers at t+12 s, 25 at t+133 s, 26 at t+150 s. Doors 200 at both readings (23:28:42Z and 23:29:44Z). |
|
||||
| alarm fired / true? | **one, true** — `controller_started` (info). Exactly what the ladder expects after a reboot. No false alarm. |
|
||||
| should have fired, did not | **none from the alarm ladder** — but the restore silence below is a legibility gap, filed as a row. |
|
||||
|
||||
**The restore left no trace anywhere, and the product has no place to leave one.** Four candidate
|
||||
status endpoints all 404 (`/api/restore/status`, `/api/backup/restore/status`,
|
||||
`/backup/restore/status`, `/api/restore`). `/api/backup/status` carries no restore field at all.
|
||||
On the pages, the only restore text is a **button label** and a JavaScript label expression. On disk,
|
||||
in the real data directory, there is no restore, lock or state file anywhere — and **no file at all
|
||||
was modified in the reset window**. An interrupted restore and a restore that never happened are
|
||||
indistinguishable, to the customer and to me.
|
||||
|
||||
**The limit of that measurement, stated rather than glossed.** Only four seconds elapsed, so the
|
||||
restore may have finished or may never have written a byte — and I cannot tell, because the
|
||||
controller's log stream holds **zero lines before 23:28:00Z** (a reset starts it fresh) and the debug
|
||||
ring died with the machine. What is independently verifiable is the **absence of any restore record**,
|
||||
and that is what is filed; it holds however far the restore got.
|
||||
|
||||
**Four of my own instruments failed in this round, and all four are recorded in the evidence:** a
|
||||
claim that was unfalsifiable when written; an on-disk check against a directory that does not exist;
|
||||
a household count that reported 0 lines and 0 failures when the truth was one line and it *was* a
|
||||
failure; and a disk guard that reported „active" all night while being a **transient** unit that
|
||||
vanished at the reset. It is now file-backed and enabled, and its script has been copied off the box.
|
||||
|
||||
### Rounds 8-12
|
||||
|
||||
PENDING
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
#!/bin/bash
|
||||
# SAFETY, not a measurement: the local backup tier retries every ~15 minutes and cannot fit on this
|
||||
# box (a ~29 GB source into a 14 GB root filesystem). Left alone during a round it would fill `/` and
|
||||
# wedge the nested PVE, which already nearly cost the night once. This kills ONLY a local-tier dump,
|
||||
# and only when free space falls under the floor. Every kill is logged so it appears in the evidence.
|
||||
FLOOR_MB=2500
|
||||
LOG=/root/diskguard.log
|
||||
while true; do
|
||||
FREE=$(df -BM --output=avail / | tail -1 | tr -dc '0-9')
|
||||
if [ -n "$FREE" ] && [ "$FREE" -lt "$FLOOR_MB" ]; then
|
||||
if ls /var/lib/vz/dump/*.tar.dat >/dev/null 2>&1; then
|
||||
echo "$(date -u +%FT%TZ) GUARD: free=${FREE}M below ${FLOOR_MB}M and a local dump is writing — killing it" >> $LOG
|
||||
pkill -f "[v]zdump"
|
||||
sleep 3
|
||||
rm -rf /var/lib/vz/dump/*.tar.dat /var/lib/vz/dump/*.tmp 2>/dev/null
|
||||
echo "$(date -u +%FT%TZ) GUARD: partial archive removed, free now $(df -BM --output=avail / | tail -1)" >> $LOG
|
||||
fi
|
||||
fi
|
||||
sleep 20
|
||||
done
|
||||
@@ -0,0 +1,87 @@
|
||||
round 10 armed for 2026-09-16T23:25:48Z (restore uptime-kuma + hard reset)
|
||||
=== round 10 launched 2026-09-16T23:26:06Z (due 23:25:48Z) ===
|
||||
2026-09-16T23:26:06Z ================ ROUND 10 : restore uptime-kuma, while: hard-reset ================
|
||||
2026-09-16T23:26:08Z --- BEFORE --- containers=26 status=200 status=200 paste=200
|
||||
2026-09-16T23:26:08Z (immich/photos is a KNOWN PRE-EXISTING failure — not caused by this round)
|
||||
2026-09-16T23:26:08Z --- ACTION: restore on uptime-kuma ---
|
||||
POST /backup/restore (uptime-kuma) -> HTTP/1.1 302 Found
|
||||
flash: Visszaállítás elindult — az állapot itt frissül.
|
||||
2026-09-16T23:26:12Z --- ACCIDENT: hard-reset (injected after the action started) ---
|
||||
2026-09-16T23:26:12Z ACCIDENT=hard-reset round=10
|
||||
2026-09-16T23:26:12Z qm reset 336 (the reset button, mid-write)
|
||||
2026-09-16T23:26:14Z reset issued
|
||||
2026-09-16T23:26:14Z accident hard-reset complete
|
||||
2026-09-16T23:26:14Z --- AFTER: what the box did BY ITSELF ---
|
||||
2026-09-16T23:26:24Z t+12s containers=0 (before 26)
|
||||
2026-09-16T23:28:07Z t+115s containers=0 (before 26)
|
||||
2026-09-16T23:28:25Z t+133s containers=25 (before 26)
|
||||
2026-09-16T23:28:42Z t+150s containers=26 (before 26)
|
||||
2026-09-16T23:28:42Z STEADY after 150s
|
||||
2026-09-16T23:28:42Z front doors, FIRST reading at 2026-09-16T23:28:42Z - TOO EARLY to trust if the accident just ended:
|
||||
2026-09-16T23:28:44Z status=200 status=200 paste=200 wiki=200
|
||||
2026-09-16T23:29:44Z front doors, SECOND reading at 2026-09-16T23:29:44Z, 60 s later - THIS is the one to trust:
|
||||
2026-09-16T23:29:45Z status=200 status=200 paste=200 wiki=200
|
||||
2026-09-16T23:29:45Z household lines this round: 0 failures: 0
|
||||
2026-09-16T23:29:45Z --- alarms ---
|
||||
| Time | Severity | Type | Message | Source
|
||||
| Sep 16 23:28 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:59 | error | whole_guest_backup_failed | Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s) | controller
|
||||
| Sep 16 21:53 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:48 | info | health_recovered | Rendszer állapot helyreállt: ok (volt: fail) | controller
|
||||
| Sep 16 21:43 | error | health_critical | Rendszer állapot kritikus (volt: ok) | controller
|
||||
| Sep 16 21:28 | info | controller_started | Controller elindult (0.245.0) | controller
|
||||
| Sep 16 21:23 | info | app_deployed | Alkalmazás telepítve: BookStack | controller
|
||||
| Sep 16 21:22 | info | app_deploy_started | Alkalmazás telepítése elindult: BookStack | controller
|
||||
2026-09-16T23:29:46Z ================ END ROUND 10 ================
|
||||
|
||||
[exited with code 0]
|
||||
|
||||
## WHAT HAPPENED TO THE RESTORE (asked four ways, 23:30-23:34Z)
|
||||
|
||||
The restore was accepted at 23:26:08Z (HTTP 302, flash "Visszaallitas elindult").
|
||||
The machine was hard-reset at 23:26:12Z - FOUR SECONDS later, mid-write.
|
||||
|
||||
Asked afterwards, through the same doors the UI uses:
|
||||
/api/restore/status -> 404
|
||||
/api/backup/restore/status -> 404
|
||||
/backup/restore/status -> 404
|
||||
/api/restore -> 404
|
||||
/api/backup/status -> {"ok":true,"data":{"enabled":true,"running":false}} (no restore field at all)
|
||||
/backups/apps (200) -> the only "Vissza" text is a BUTTON LABEL and a JS label expression:
|
||||
"Visszaallitas"
|
||||
": op === 'offbox-restore' ? 'Tavoli visszaallitas' : 'Visszaallitas'; }"
|
||||
/apps/uptime-kuma (200) -> one "Vissza" (the button). "megszak" (interrupted): 0 hits.
|
||||
The "sikeres" hits belong to the MOVE-DATA feature, not the restore.
|
||||
NEGATIVE CONTROL "zzzznotpresent" -> 0 hits on every page, so the searches are trustworthy.
|
||||
|
||||
On disk, in the REAL data directory (/var/lib/docker/volumes/felhom-controller-data/_data):
|
||||
no file named *restor*, *lock* or *.pid anywhere beneath it
|
||||
NO FILE AT ALL modified in the window 23:24:00 - 23:27:30
|
||||
|
||||
So: there is no restore history surface of any kind. An interrupted restore and a restore that
|
||||
never happened look EXACTLY the same to the customer, and to me.
|
||||
|
||||
### THE HONEST LIMIT OF THIS MEASUREMENT, which must not be glossed over
|
||||
Only four seconds elapsed. The restore may have completed, or may never have written a byte.
|
||||
I cannot tell, because:
|
||||
* the controller's log stream holds ZERO lines before 23:28:00Z - a hard reset starts it fresh,
|
||||
so the pre-reset window is unrecoverable by this route;
|
||||
* the debug ring is in memory and died with the machine.
|
||||
What IS independently verifiable, and is what gets filed, is the ABSENCE OF ANY RESTORE RECORD -
|
||||
that holds regardless of how far the restore got.
|
||||
|
||||
## FOUR CORRECTIONS TO MY OWN INSTRUMENTS, all found in this round
|
||||
1. "The controller logged nothing about the restore" was UNFALSIFIABLE when I first wrote it.
|
||||
The stream cannot reach before the reset. Re-stated above with its limit.
|
||||
2. My first on-disk check looked in /opt/felhom/data - A DIRECTORY THAT DOES NOT EXIST. It printed
|
||||
a tidy "no restore/lock file" that meant nothing. The real path is the docker volume above.
|
||||
Same class as the "docker: command not found -> a tidy table of absent/0" error from earlier.
|
||||
3. The runner reported "household lines this round: 0 failures: 0". Both are WRONG. The log is
|
||||
never truncated (first line 21:06:30Z, continuous) and contains exactly ONE line in the round
|
||||
window - and it is a FAILURE line: "2026-09-16T23:28:09Z cloud dash UNREACHABLE" (the box was
|
||||
still booting). The runner now prints the raw before/after counts so the subtraction is checkable.
|
||||
4. THE DISK GUARD WAS NOT A REAL UNIT. After the reset: "Unit diskguard.service not found" - yet it
|
||||
had reported "active" all night. It was a TRANSIENT unit, which looks identical while running and
|
||||
vanishes on reboot. It was therefore ABSENT from 23:26:12Z. It is now a file-backed, ENABLED unit
|
||||
(verified: active + enabled, MainPID cmdline "/bin/bash /root/diskguard.sh", 0 kills, 7556 MB free)
|
||||
and its script is copied off the box into this folder, which had also never been done (R-320).
|
||||
@@ -92,6 +92,10 @@ sleep 60
|
||||
say " front doors, SECOND reading at $(date -u +%FT%TZ), 60 s later - THIS is the one to trust:"
|
||||
say " $SUB=$(door $SUB) status=$(door status) paste=$(door paste) wiki=$(door wiki)"
|
||||
HL1=$(BOX 'wc -l < /root/household.log' | tr -d ' \r')
|
||||
# Round 10 reported "0 lines, 0 failures" when the truth was one line and it WAS a
|
||||
# failure. A subtraction whose operands are invisible cannot be checked, so both are
|
||||
# printed now, and an empty read is named rather than silently becoming zero.
|
||||
say " household log lines: before=${HL0:-EMPTY} after=${HL1:-EMPTY}"
|
||||
say " household lines this round: $(( ${HL1:-0} - ${HL0:-0} )) failures: $(BOX "tail -n +$(( ${HL0:-0} + 1 )) /root/household.log | grep -cE 'FAILED|UNREACHABLE'" | tr -d ' \r')"
|
||||
say "--- alarms ---"; bash $E/events.sh 8 | tee -a $OUT
|
||||
say "================ END ROUND $N ================"
|
||||
|
||||
Reference in New Issue
Block a user