CHAOS NIGHT: household seeded, escrow done, round 1 measured
gates / gates (push) Successful in 19s

Phase 0 finished: all twelve apps deployed one at a time (the parallel burst was
my error and is recorded as such), escrow ceremony completed after the box
converged the PBS descriptor by itself (~17 min), and the R-543 bar disappeared
for good once escrowed.

Round 1 (offsite-run, control round, no accident):
- the run started on the button's own endpoint and walked every app with real
  per-volume byte counts (stop -> dump -> restart)
- three alarms fired, all TRUE and precise: app_start_failed named the one
  crash-looping app, backup_run_failures said "1 of 12 ... nextcloud", and
  offbox_repo_orphaned reported the documented rebuild behaviour rather than
  claiming a successful copy
- nextcloud's crash loop is MY damage (image layers written while the thin pool
  was 100% full -> "invalid ELF header"), not a product defect, and is recorded
  that way

The catalog bump was prepared and then REVERTED UNPUSHED: the catalog repo's own
gates returned INCONCLUSIVE (the volume-persistence prober failed its own
canary), and undetermined is never a pass. Round 7's `update` therefore becomes
`use`, decided now rather than improvised at 02:00.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 23:10:37 +02:00
parent 9fae6dfa98
commit 5f2ccec64c
3 changed files with 125 additions and 0 deletions
@@ -360,3 +360,45 @@ registered, mounted and made default, `/api/disks/candidates` STILL lists `/dev/
`initialize` — now with `data_bearing:true, mountable:true` and its durable id. The same endpoint
also lists it under `attach`. That is exactly the behaviour R-542 describes (a registered, in-use
drive still offered under „initialize"), seen again on a fresh box.
### DECISION — the catalog bump was NOT pushed, and round 7 changes because of it
The bump was prepared (`templates/nextcloud/docker-compose.yml`: `redis:7-alpine` -> `redis:7.4-alpine`,
a tag checked to exist; `catalog_since` -> 2026-09-17) and then **reverted unpushed**, because the
catalog repo's own gate runner returned:
image-resolvable INCONCLUSIVE (exit 2)
volume-persistence INCONCLUSIVE (exit 2)
UNDETERMINED (never a pass): image-resolvable, volume-persistence
with `volume-persistence` failing its OWN canary self-test —
„ERROR: the prober failed its own canary — canary-clean: expected CLEAN, got UNDETERMINED …
refusing to report a verdict: a broken detector reporting CLEAN is worse than no detector at all"
Neither inconclusive gate is caused by the one-line pin change: `catalog-since` and `engine-major`
both reported „0 compose file(s) changed" for the same range, i.e. they saw nothing of it. This is an
environment/prober problem, and it is **recorded, not worked around** — this repo's rule is that
UNDETERMINED is never a pass, and pushing past it would be exactly the habit the rule exists to stop.
Nothing was bypassed and `--no-verify` was not used; the tree was returned to clean.
**Consequence for the schedule, decided and recorded rather than improvised at 02:00:** round 7 drew
`update nextcloud`, and with no bump delivered there is no update for the guarded-update path to
apply. By the schedule's own constraint 3 — an `update` that cannot run becomes `use` — **round 7
runs as `use nextcloud` under its drawn accident (`internet gone 10min`)**, and the findings document
says so in that round's row. The guarded update is therefore NOT exercised tonight, and the morning
verdict must not claim it was.
### Reachability, settled before the rounds — and where the household loop actually runs
The customer guest is at **192.168.0.116**, on the LAN, and IS reachable from DooPlex:
ping answers · ports 80 and 443 open · `Host: wiki.enkicsifelhom.hu` -> 301 (redirect to https)
So the `use` rounds can drive the apps' own front doors directly from the drill host, which is what
the brief asked for. Note the contrast with the box itself: the nested PVE (192.168.0.115) answers,
but the CONTROLLER only listens on its container address 172.17.0.2:8080 inside the guest — that is
why every dashboard action tonight is run as a script pushed into the guest.
**The background household loop runs ON THE VM (192.168.0.115), not on DooPlex** — a deviation from
the brief's „from the drill host", declared here. It was started before reachability from DooPlex had
been established, it is a `systemd-run` transient unit (`household.service`, active, logging to
/root/household.log), and moving a working measurement mid-night to gain nothing is a worse trade
than recording where it sits. It hits the guest's traefik from one hop closer than DooPlex would, so
it does NOT exercise the LAN path between DooPlex and the box; anything that breaks only on that hop
is invisible to it. The `use` rounds, which DO run from DooPlex, cover that hop.
@@ -0,0 +1,66 @@
## Round 1 side-observations — two things that looked like defects and are not (and one that is mine)
### 1. Five front doors answered 404 — because the BACKUP was quiescing them
A front-door sweep from DooPlex during round 1 read:
wiki 404 · paste 200 · share 404 · recipes 200 · inventory 200 · status 200 · travel 200
paperless 404 · cloud 404 · photos 404 · media 200 · vault 200
The container states explained it: `bookstack Up 17 seconds`, `gokapi Up 1 second`,
`paperless-webserver Up 11 seconds`, `immich-server Created` — all seconds old. The controller log
says why, in its own words:
„[backup] Stopping nextcloud for safe volume dump" -> volume dumps -> „[backup] Restarting
nextcloud after volume dump"
So the off-site run stops each app, dumps its volumes, and starts it again — and my sweep happened to
walk the apps while the run was walking them too. **Nothing was wrong; my measurement was mistimed.**
The sweep is repeated after the run finishes, and only that later reading counts.
### 2. Nextcloud IS broken — and it is MY damage, from the disk-full episode
php: error while loading shared libraries: /lib/x86_64-linux-gnu/libxml2.so.2: **invalid ELF header**
RestartCount=6 ExitCode=127 Status=restarting
„Invalid ELF header" on a shared library means the file on disk is CORRUPT, not missing. Those layers
were written while the LVM thin pool stood at 100 % (my parallel twelve-deploy burst). The later
sequential re-deploy reported „INSTALLED" because it reused the already-pulled — and already
corrupt — image instead of pulling again. Repair: drop the image so a fresh pull happens, and
redeploy. Not done during an in-flight backup run.
**This is not filed as a product defect.** The corruption came from a full disk that I caused.
### What the PRODUCT did about it, which is the part worth measuring
„[stacks] Stack nextcloud started successfully (took 7.5s)"
„[stacks] Stack nextcloud post-start status:"
„[stacks] nextcloud nextcloud:34.0.1-apache restarting Restarting (127) 1 second ago"
This is the documented invariant working: `compose up -d` exits 0 even for a crash-loop, and **the
post-start status check is the detection**. The success line and the truth appear together, one line
apart, and the state the box records is `restarting`, not `running`.
## The alarms round 1 produced — three fired, all TRUE, and each one precise
20:51 warning **app_start_failed** „Telepített alkalmazás nem fut: Nextcloud"
21:08 error **backup_run_failures** „1 of 12 apps failed to back up in this nightly run: nextcloud"
21:09 warning **offbox_repo_orphaned** „A távoli mentési tároló elárvult: a benne lévő mentések egy
korábbi, már nem elérhető kulccsal készültek (újratelepítés).
Új mentés a tároló visszaállít…"
Scored against `08-alarm-ladder.md`:
* `app_start_failed` is exactly right for a deployed app that is not running, and it is the one
customer-switchable event that is OFF by default — it reached the operator. It fired ~5 minutes
after the crash loop started, which is the documented 5-minute crash-loop window, not a lag.
* `backup_run_failures` is operator-only by register, error severity, and **it names the app and
the count**: „1 of 12 … nextcloud". A run that partly failed said so precisely rather than
reporting a green run or a red one.
* `offbox_repo_orphaned` is the honest outcome of round 1 itself (below). Nothing pretended the
off-site copy had succeeded.
**No alarm was missed for what happened**, and no alarm fired for something that had not happened.
The only app that was broken is the only app that was named.
## What round 1 actually measured (the drawn action was `offsite-run`)
* the run STARTED on the button's own endpoint: „A távoli mentés elindult — az állapot itt frissül."
with `progress.active: true`
* it walked every app: stop -> volume dump -> restart („Stopping X for safe volume dump" …
„Restarting X after volume dump"), with real byte counts per volume
(e.g. nextcloud_db_data 148.9 MB, nextcloud_html 2.5 KB, redis_data 5.5 KB)
* one app could not be captured — the one that was crash-looping — and the run said so
* and the repository itself turned out to be **orphaned**: this box is a REBUILD for an existing
customer, so its restic password was minted fresh and the snapshots already in the remote store
were written under a key this box does not have. That is the documented rebuild behaviour
(„Guest rebuild orphans the off-site repo"), surfaced to the customer rather than hidden, with
the route out named in the message.
@@ -0,0 +1,17 @@
2026-09-16T21:07:14Z === ROUND 1: offsite-run / adventurelog / accident: nothing ===
--- BEFORE ---
apps running: 26
escrow=escrowed last_run=None last_status=None snapshots=None repo=None
--- ACTION: the off-site run, through the button's own endpoint ---
POST /backup/offbox/run -> HTTP/1.1 302 Found
flash: A távoli mentés elindult — az állapot itt frissül.
--- polling the run ---
2026-09-16T21:07:34Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
2026-09-16T21:07:54Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
2026-09-16T21:08:15Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
2026-09-16T21:08:35Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
2026-09-16T21:08:55Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap
2026-09-16T21:09:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:09:35Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:09:55Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":
2026-09-16T21:10:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":