51ac1cb4ff
gates / gates (push) Successful in 34s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
64 lines
4.6 KiB
Markdown
64 lines
4.6 KiB
Markdown
# The accidents — five lines each (household saw · box did by itself · time to steady · alarm fired, true? · alarm owed, missing?)
|
||
|
||
All times UTC. Evidence beside this file in `accidents/`.
|
||
|
||
## A2 — Docker's socket re-created during a whole-guest backup (demo-felhom, 20:16)
|
||
|
||
"Mentés most" (`POST /api/guest-backup/trigger`) at 20:16:12; the app stop began 20:16:16; `systemctl restart docker.socket`
|
||
in the guest at 20:16:23 (socket inode 151 → 646).
|
||
- **Household saw:** opengist (the one app stopped for the backup) down 20:16:16 → 20:17:50 (94 s); the dashboard briefly
|
||
unable to show app states; the timeline "Controller elindult" once. No mail.
|
||
- **Box did by itself:** the `felhom-backup` copy finished honestly (`backup: completed`, 20:16:43 — vzdump does not use
|
||
the socket). The unquiesce could not restart opengist (`exit code 1`: the controller was blind). The controller's
|
||
socket watch exited it at 20:17:33 after 60 s of refusals (R-860), Docker restarted it on the new socket, boot
|
||
reconciliation restarted opengist at 20:17:50, and the controller restarted traefik onto the new socket at 20:18:05.
|
||
The `felhom-pbs` tier answered BUSY (the agent's heavy-op gate was held) and was deferred 15 min, as designed. Then the
|
||
`night` OS leg ran (20:18–20:18:56: guest, host, Docker all "nothing").
|
||
- **Time to steady:** 102 s (20:16:23 → 20:18:05).
|
||
- **Alarms fired:** none. True: nothing stayed broken.
|
||
- **Alarms owed and missing:** none — the box healed within two minutes. (The failed restart at unquiesce left no event;
|
||
acceptable because boot reconciliation repaired it 64 s later.)
|
||
|
||
## A3 — the hub out of reach (Tester 1 box, 20:26:09 → 20:56:10, 30 min)
|
||
|
||
Blackhole route to the hub's address on the host and in the guest.
|
||
- **Household saw:** nothing (the dashboard and apps go through the tunnel, not the hub).
|
||
- **Box did by itself:** two host reports failed (22:40:49, 22:55:49 local — `hub: report failed; keeping current
|
||
interval`); the next one after the unblock arrived on time (21:10:49 UTC); the controller's report went through at
|
||
20:57:04. Missed host reports are snapshots and are not replayed — the next one replaces them. **The debug OS pass could
|
||
not run at all** (it fetches its plan from the hub — R-866); the daemon's own leg would use its saved plan, but the
|
||
20-hour gap held it back (this box's leg ran at 19:49).
|
||
- **Time to steady:** at the first report after the unblock, 14 min 39 s (the report interval).
|
||
- **Alarms fired:** none. True: 30 min is under the 45 min stale threshold.
|
||
- **Alarms owed and missing:** none. The 7-day OS alarms stayed quiet.
|
||
|
||
## A4 — the system disk nearly full (Tester 1 box guest, 20:24)
|
||
|
||
The guest's `/` filled to 400 MB free; a one-package-set plan (bind9 ×3, Debian-Security) handed to the wrapper as the
|
||
agent writes it.
|
||
- **Household saw:** nothing (the disk was freed 3 minutes later; the apps live on the data volume).
|
||
- **Box did by itself:** the wrapper refused BEFORE downloading: `REFUSED: R8 free space 419430400 B is below max(500 MB,
|
||
3 x download 0 B)`, exit 2, nothing installed (bind9 stayed at the old version). **But the download was measured as
|
||
0 B** — `apt-get -s --print-uris` prints no URIs (R-865), so only the 500 MB floor ever applies.
|
||
- **Time to steady:** immediate (a refusal changes nothing).
|
||
- **Alarms fired:** none (a refused debug plan reports nothing to the hub). True.
|
||
- **Alarms owed and missing:** none for a debug run; a NIGHT leg refused by R8 reports `refused` and the stale alarm
|
||
fires after 7 days — not exercised.
|
||
|
||
## A5 — the agent killed in the middle of a pass (demo-hp, 02:57 UTC)
|
||
|
||
Six guest packages rolled back (simulated first: 0 removals); a debug pass started; when `apt-get` ran, the pass and the
|
||
agent daemon were `kill -9`-ed (02:57:18).
|
||
- **Household saw:** nothing.
|
||
- **Box did by itself:** the root wrapper (its own process under sudo) finished all six packages; `dpkg --audit` clean;
|
||
systemd restarted the daemon in < 20 s (`NRestarts` 1; the self-update rollback did nothing — no pending update).
|
||
- **Time to steady:** < 20 s.
|
||
- **Alarms fired:** none. True.
|
||
- **Alarms owed and missing:** none — but the killed pass's report never reached the hub (R-868); no duplicate report.
|
||
|
||
## A1 — a power cut in the middle of an OS update (demo-hp) — NOT RUN
|
||
|
||
The host crash (`echo c > /proc/sysrq-trigger` while dpkg ran) was refused by this session's permission check
|
||
("interfere with workloads"). Not attempted another way. The six rolled-back packages were brought forward by a normal
|
||
debug pass (applied, healthy); dpkg clean; demo-hp's guest package list equals the 21:47 baseline. Operator decision 2.
|