Files
felhom.eu/documentation/audits/night-2026-10-04/ACCIDENTS.md
T
2026-10-05 05:20:57 +02:00

4.6 KiB
Raw Blame History

The accidents — five lines each (household saw · box did by itself · time to steady · alarm fired, true? · alarm owed, missing?)

All times UTC. Evidence beside this file in accidents/.

A2 — Docker's socket re-created during a whole-guest backup (demo-felhom, 20:16)

"Mentés most" (POST /api/guest-backup/trigger) at 20:16:12; the app stop began 20:16:16; systemctl restart docker.socket in the guest at 20:16:23 (socket inode 151 → 646).

  • Household saw: opengist (the one app stopped for the backup) down 20:16:16 → 20:17:50 (94 s); the dashboard briefly unable to show app states; the timeline "Controller elindult" once. No mail.
  • Box did by itself: the felhom-backup copy finished honestly (backup: completed, 20:16:43 — vzdump does not use the socket). The unquiesce could not restart opengist (exit code 1: the controller was blind). The controller's socket watch exited it at 20:17:33 after 60 s of refusals (R-860), Docker restarted it on the new socket, boot reconciliation restarted opengist at 20:17:50, and the controller restarted traefik onto the new socket at 20:18:05. The felhom-pbs tier answered BUSY (the agent's heavy-op gate was held) and was deferred 15 min, as designed. Then the night OS leg ran (20:18–20:18:56: guest, host, Docker all "nothing").
  • Time to steady: 102 s (20:16:23 → 20:18:05).
  • Alarms fired: none. True: nothing stayed broken.
  • Alarms owed and missing: none — the box healed within two minutes. (The failed restart at unquiesce left no event; acceptable because boot reconciliation repaired it 64 s later.)

A3 — the hub out of reach (Tester 1 box, 20:26:09 → 20:56:10, 30 min)

Blackhole route to the hub's address on the host and in the guest.

  • Household saw: nothing (the dashboard and apps go through the tunnel, not the hub).
  • Box did by itself: two host reports failed (22:40:49, 22:55:49 local — hub: report failed; keeping current interval); the next one after the unblock arrived on time (21:10:49 UTC); the controller's report went through at 20:57:04. Missed host reports are snapshots and are not replayed — the next one replaces them. The debug OS pass could not run at all (it fetches its plan from the hub — R-866); the daemon's own leg would use its saved plan, but the 20-hour gap held it back (this box's leg ran at 19:49).
  • Time to steady: at the first report after the unblock, 14 min 39 s (the report interval).
  • Alarms fired: none. True: 30 min is under the 45 min stale threshold.
  • Alarms owed and missing: none. The 7-day OS alarms stayed quiet.

A4 — the system disk nearly full (Tester 1 box guest, 20:24)

The guest's / filled to 400 MB free; a one-package-set plan (bind9 ×3, Debian-Security) handed to the wrapper as the agent writes it.

  • Household saw: nothing (the disk was freed 3 minutes later; the apps live on the data volume).
  • Box did by itself: the wrapper refused BEFORE downloading: REFUSED: R8 free space 419430400 B is below max(500 MB, 3 x download 0 B), exit 2, nothing installed (bind9 stayed at the old version). But the download was measured as 0 B — apt-get -s --print-uris prints no URIs (R-865), so only the 500 MB floor ever applies.
  • Time to steady: immediate (a refusal changes nothing).
  • Alarms fired: none (a refused debug plan reports nothing to the hub). True.
  • Alarms owed and missing: none for a debug run; a NIGHT leg refused by R8 reports refused and the stale alarm fires after 7 days — not exercised.

A5 — the agent killed in the middle of a pass (demo-hp, 02:57 UTC)

Six guest packages rolled back (simulated first: 0 removals); a debug pass started; when apt-get ran, the pass and the agent daemon were kill -9-ed (02:57:18).

  • Household saw: nothing.
  • Box did by itself: the root wrapper (its own process under sudo) finished all six packages; dpkg --audit clean; systemd restarted the daemon in < 20 s (NRestarts 1; the self-update rollback did nothing — no pending update).
  • Time to steady: < 20 s.
  • Alarms fired: none. True.
  • Alarms owed and missing: none — but the killed pass's report never reached the hub (R-868); no duplicate report.

A1 — a power cut in the middle of an OS update (demo-hp) — NOT RUN

The host crash (echo c > /proc/sysrq-trigger while dpkg ran) was refused by this session's permission check ("interfere with workloads"). Not attempted another way. The six rolled-back packages were brought forward by a normal debug pass (applied, healthy); dpkg clean; demo-hp's guest package list equals the 21:47 baseline. Operator decision 2.