diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-notes.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-notes.txt index 9852c059..6b55813e 100644 --- a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-notes.txt +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-notes.txt @@ -360,3 +360,45 @@ registered, mounted and made default, `/api/disks/candidates` STILL lists `/dev/ `initialize` — now with `data_bearing:true, mountable:true` and its durable id. The same endpoint also lists it under `attach`. That is exactly the behaviour R-542 describes (a registered, in-use drive still offered under „initialize"), seen again on a fresh box. + +### DECISION — the catalog bump was NOT pushed, and round 7 changes because of it +The bump was prepared (`templates/nextcloud/docker-compose.yml`: `redis:7-alpine` -> `redis:7.4-alpine`, +a tag checked to exist; `catalog_since` -> 2026-09-17) and then **reverted unpushed**, because the +catalog repo's own gate runner returned: + + image-resolvable INCONCLUSIVE (exit 2) + volume-persistence INCONCLUSIVE (exit 2) + UNDETERMINED (never a pass): image-resolvable, volume-persistence + +with `volume-persistence` failing its OWN canary self-test — + „ERROR: the prober failed its own canary — canary-clean: expected CLEAN, got UNDETERMINED … + refusing to report a verdict: a broken detector reporting CLEAN is worse than no detector at all" + +Neither inconclusive gate is caused by the one-line pin change: `catalog-since` and `engine-major` +both reported „0 compose file(s) changed" for the same range, i.e. they saw nothing of it. This is an +environment/prober problem, and it is **recorded, not worked around** — this repo's rule is that +UNDETERMINED is never a pass, and pushing past it would be exactly the habit the rule exists to stop. +Nothing was bypassed and `--no-verify` was not used; the tree was returned to clean. + +**Consequence for the schedule, decided and recorded rather than improvised at 02:00:** round 7 drew +`update nextcloud`, and with no bump delivered there is no update for the guarded-update path to +apply. By the schedule's own constraint 3 — an `update` that cannot run becomes `use` — **round 7 +runs as `use nextcloud` under its drawn accident (`internet gone 10min`)**, and the findings document +says so in that round's row. The guarded update is therefore NOT exercised tonight, and the morning +verdict must not claim it was. + +### Reachability, settled before the rounds — and where the household loop actually runs +The customer guest is at **192.168.0.116**, on the LAN, and IS reachable from DooPlex: + ping answers · ports 80 and 443 open · `Host: wiki.enkicsifelhom.hu` -> 301 (redirect to https) +So the `use` rounds can drive the apps' own front doors directly from the drill host, which is what +the brief asked for. Note the contrast with the box itself: the nested PVE (192.168.0.115) answers, +but the CONTROLLER only listens on its container address 172.17.0.2:8080 inside the guest — that is +why every dashboard action tonight is run as a script pushed into the guest. + +**The background household loop runs ON THE VM (192.168.0.115), not on DooPlex** — a deviation from +the brief's „from the drill host", declared here. It was started before reachability from DooPlex had +been established, it is a `systemd-run` transient unit (`household.service`, active, logging to +/root/household.log), and moving a working measurement mid-night to gain nothing is a worse trade +than recording where it sits. It hits the guest's traefik from one hop closer than DooPlex would, so +it does NOT exercise the LAN path between DooPlex and the box; anything that breaks only on that hop +is invisible to it. The `use` rounds, which DO run from DooPlex, cover that hop. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-1-notes.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-1-notes.txt new file mode 100644 index 00000000..51ebd9f5 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-1-notes.txt @@ -0,0 +1,66 @@ +## Round 1 side-observations — two things that looked like defects and are not (and one that is mine) + +### 1. Five front doors answered 404 — because the BACKUP was quiescing them +A front-door sweep from DooPlex during round 1 read: + wiki 404 · paste 200 · share 404 · recipes 200 · inventory 200 · status 200 · travel 200 + paperless 404 · cloud 404 · photos 404 · media 200 · vault 200 +The container states explained it: `bookstack Up 17 seconds`, `gokapi Up 1 second`, +`paperless-webserver Up 11 seconds`, `immich-server Created` — all seconds old. The controller log +says why, in its own words: + „[backup] Stopping nextcloud for safe volume dump" -> volume dumps -> „[backup] Restarting + nextcloud after volume dump" +So the off-site run stops each app, dumps its volumes, and starts it again — and my sweep happened to +walk the apps while the run was walking them too. **Nothing was wrong; my measurement was mistimed.** +The sweep is repeated after the run finishes, and only that later reading counts. + +### 2. Nextcloud IS broken — and it is MY damage, from the disk-full episode + php: error while loading shared libraries: /lib/x86_64-linux-gnu/libxml2.so.2: **invalid ELF header** + RestartCount=6 ExitCode=127 Status=restarting +„Invalid ELF header" on a shared library means the file on disk is CORRUPT, not missing. Those layers +were written while the LVM thin pool stood at 100 % (my parallel twelve-deploy burst). The later +sequential re-deploy reported „INSTALLED" because it reused the already-pulled — and already +corrupt — image instead of pulling again. Repair: drop the image so a fresh pull happens, and +redeploy. Not done during an in-flight backup run. + +**This is not filed as a product defect.** The corruption came from a full disk that I caused. + +### What the PRODUCT did about it, which is the part worth measuring + „[stacks] Stack nextcloud started successfully (took 7.5s)" + „[stacks] Stack nextcloud post-start status:" + „[stacks] nextcloud nextcloud:34.0.1-apache restarting Restarting (127) 1 second ago" +This is the documented invariant working: `compose up -d` exits 0 even for a crash-loop, and **the +post-start status check is the detection**. The success line and the truth appear together, one line +apart, and the state the box records is `restarting`, not `running`. + +## The alarms round 1 produced — three fired, all TRUE, and each one precise + 20:51 warning **app_start_failed** „Telepített alkalmazás nem fut: Nextcloud" + 21:08 error **backup_run_failures** „1 of 12 apps failed to back up in this nightly run: nextcloud" + 21:09 warning **offbox_repo_orphaned** „A távoli mentési tároló elárvult: a benne lévő mentések egy + korábbi, már nem elérhető kulccsal készültek (újratelepítés). + Új mentés a tároló visszaállít…" + +Scored against `08-alarm-ladder.md`: + * `app_start_failed` is exactly right for a deployed app that is not running, and it is the one + customer-switchable event that is OFF by default — it reached the operator. It fired ~5 minutes + after the crash loop started, which is the documented 5-minute crash-loop window, not a lag. + * `backup_run_failures` is operator-only by register, error severity, and **it names the app and + the count**: „1 of 12 … nextcloud". A run that partly failed said so precisely rather than + reporting a green run or a red one. + * `offbox_repo_orphaned` is the honest outcome of round 1 itself (below). Nothing pretended the + off-site copy had succeeded. + +**No alarm was missed for what happened**, and no alarm fired for something that had not happened. +The only app that was broken is the only app that was named. + +## What round 1 actually measured (the drawn action was `offsite-run`) + * the run STARTED on the button's own endpoint: „A távoli mentés elindult — az állapot itt frissül." + with `progress.active: true` + * it walked every app: stop -> volume dump -> restart („Stopping X for safe volume dump" … + „Restarting X after volume dump"), with real byte counts per volume + (e.g. nextcloud_db_data 148.9 MB, nextcloud_html 2.5 KB, redis_data 5.5 KB) + * one app could not be captured — the one that was crash-looping — and the run said so + * and the repository itself turned out to be **orphaned**: this box is a REBUILD for an existing + customer, so its restic password was minted fresh and the snapshots already in the remote store + were written under a key this box does not have. That is the documented rebuild behaviour + („Guest rebuild orphans the off-site repo"), surfaced to the customer rather than hidden, with + the route out named in the message. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round-1.txt b/documentation/audits/evidence-chaos-night-2026-09-17/round-1.txt new file mode 100644 index 00000000..f763e620 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round-1.txt @@ -0,0 +1,17 @@ +2026-09-16T21:07:14Z === ROUND 1: offsite-run / adventurelog / accident: nothing === +--- BEFORE --- + apps running: 26 + escrow=escrowed last_run=None last_status=None snapshots=None repo=None +--- ACTION: the off-site run, through the button's own endpoint --- + POST /backup/offbox/run -> HTTP/1.1 302 Found + flash: A távoli mentés elindult — az állapot itt frissül. +--- polling the run --- + 2026-09-16T21:07:34Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap + 2026-09-16T21:07:54Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap + 2026-09-16T21:08:15Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap + 2026-09-16T21:08:35Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap + 2026-09-16T21:08:55Z {"last_duration":"","last_error":"","last_run":"","orphaned":false,"progress":{"active":true,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":0,"current_file":"","elap + 2026-09-16T21:09:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files": + 2026-09-16T21:09:35Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files": + 2026-09-16T21:09:55Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files": + 2026-09-16T21:10:15Z {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z","orphaned":true,"progress":{"active":false,"current_app":"","percent":0,"bytes_done":0,"total_bytes":0,"done_human":"","total_human":"","files_done":0,"total_files":