Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
4.6 KiB
The accidents — five lines each (household saw · box did by itself · time to steady · alarm fired, true? · alarm owed, missing?)
All times UTC. Evidence beside this file in accidents/.
A2 — Docker's socket re-created during a whole-guest backup (demo-felhom, 20:16)
"Mentés most" (POST /api/guest-backup/trigger) at 20:16:12; the app stop began 20:16:16; systemctl restart docker.socket
in the guest at 20:16:23 (socket inode 151 → 646).
- Household saw: opengist (the one app stopped for the backup) down 20:16:16 → 20:17:50 (94 s); the dashboard briefly unable to show app states; the timeline "Controller elindult" once. No mail.
- Box did by itself: the
felhom-backupcopy finished honestly (backup: completed, 20:16:43 — vzdump does not use the socket). The unquiesce could not restart opengist (exit code 1: the controller was blind). The controller's socket watch exited it at 20:17:33 after 60 s of refusals (R-860), Docker restarted it on the new socket, boot reconciliation restarted opengist at 20:17:50, and the controller restarted traefik onto the new socket at 20:18:05. Thefelhom-pbstier answered BUSY (the agent's heavy-op gate was held) and was deferred 15 min, as designed. Then thenightOS leg ran (20:18–20:18:56: guest, host, Docker all "nothing"). - Time to steady: 102 s (20:16:23 → 20:18:05).
- Alarms fired: none. True: nothing stayed broken.
- Alarms owed and missing: none — the box healed within two minutes. (The failed restart at unquiesce left no event; acceptable because boot reconciliation repaired it 64 s later.)
A3 — the hub out of reach (Tester 1 box, 20:26:09 → 20:56:10, 30 min)
Blackhole route to the hub's address on the host and in the guest.
- Household saw: nothing (the dashboard and apps go through the tunnel, not the hub).
- Box did by itself: two host reports failed (22:40:49, 22:55:49 local —
hub: report failed; keeping current interval); the next one after the unblock arrived on time (21:10:49 UTC); the controller's report went through at 20:57:04. Missed host reports are snapshots and are not replayed — the next one replaces them. The debug OS pass could not run at all (it fetches its plan from the hub — R-866); the daemon's own leg would use its saved plan, but the 20-hour gap held it back (this box's leg ran at 19:49). - Time to steady: at the first report after the unblock, 14 min 39 s (the report interval).
- Alarms fired: none. True: 30 min is under the 45 min stale threshold.
- Alarms owed and missing: none. The 7-day OS alarms stayed quiet.
A4 — the system disk nearly full (Tester 1 box guest, 20:24)
The guest's / filled to 400 MB free; a one-package-set plan (bind9 ×3, Debian-Security) handed to the wrapper as the
agent writes it.
- Household saw: nothing (the disk was freed 3 minutes later; the apps live on the data volume).
- Box did by itself: the wrapper refused BEFORE downloading:
REFUSED: R8 free space 419430400 B is below max(500 MB, 3 x download 0 B), exit 2, nothing installed (bind9 stayed at the old version). But the download was measured as 0 B —apt-get -s --print-urisprints no URIs (R-865), so only the 500 MB floor ever applies. - Time to steady: immediate (a refusal changes nothing).
- Alarms fired: none (a refused debug plan reports nothing to the hub). True.
- Alarms owed and missing: none for a debug run; a NIGHT leg refused by R8 reports
refusedand the stale alarm fires after 7 days — not exercised.
A5 — the agent killed in the middle of a pass (demo-hp, 02:57 UTC)
Six guest packages rolled back (simulated first: 0 removals); a debug pass started; when apt-get ran, the pass and the
agent daemon were kill -9-ed (02:57:18).
- Household saw: nothing.
- Box did by itself: the root wrapper (its own process under sudo) finished all six packages;
dpkg --auditclean; systemd restarted the daemon in < 20 s (NRestarts1; the self-update rollback did nothing — no pending update). - Time to steady: < 20 s.
- Alarms fired: none. True.
- Alarms owed and missing: none — but the killed pass's report never reached the hub (R-868); no duplicate report.
A1 — a power cut in the middle of an OS update (demo-hp) — NOT RUN
The host crash (echo c > /proc/sysrq-trigger while dpkg ran) was refused by this session's permission check
("interfere with workloads"). Not attempted another way. The six rolled-back packages were brought forward by a normal
debug pass (applied, healthy); dpkg clean; demo-hp's guest package list equals the 21:47 baseline. Operator decision 2.