diff --git a/documentation/audits/DRILL-chaos-night-2026-09-17.md b/documentation/audits/DRILL-chaos-night-2026-09-17.md new file mode 100644 index 00000000..d246a565 --- /dev/null +++ b/documentation/audits/DRILL-chaos-night-2026-09-17.md @@ -0,0 +1,159 @@ +# DRILL — CHAOS NIGHT: random actions on random apps while random things go wrong (2026-09-16/17) + +**Interventions: PENDING — the run is in progress.** +**Ready for a volunteer: PENDING.** +**The accident-plus-action pair that hurt most: PENDING.** + +> **Baselines, verified live against Gitea at 21:49 CEST 2026-09-16 (not copied from the brief):** +> felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 · +> felhom.eu `d124c77e176d` hub v0.116.0, ISO **1.28.0 published** · app-catalog `94bc5febaca2`. +> All four trees clean and in sync. Highest register row **R-545**, 212 open. Golden waiver valid to +> 2026-09-27. Customer **`tester-1`** (`enkicsifelhom.hu`, `tester1@felhom.eu`, no host). +> Venue: `demo-hp` (Tier 0), a fresh nested VM, disk on the NVMe at its root. Evidence: +> `evidence-chaos-night-2026-09-17/`. + +## The schedule — drawn ONCE, before round 1, and written here first + +The point of this section's position in the document is that the night could not be chosen after the +fact. `chaos_schedule.py` is committed beside the evidence; re-running it reproduces this table. + +- **seed:** `20260917` (the date) +- **script sha256:** `4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1` +- **generator:** `evidence-chaos-night-2026-09-17/chaos_schedule.py`, stdlib `random` seeded with the seed + +| # | time | X — the action | Y — the app | Z — the accident | +|---|---|---|---|---| +| 1 | 23:30 | offsite-run | adventurelog | nothing | +| 2 | 23:55 | restore | gokapi | power cut | +| 3 | 00:20 | use | bookstack | disk 95% full | +| 4 | 00:45 | offsite-run | mealie | tunnel down 10min | +| 5 | 01:10 | use | privatebin | docker restarted | +| 6 | 01:35 | backup-system | adventurelog | nothing | +| 7 | 02:00 | update | nextcloud | internet gone 10min | +| 8 | 02:25 | backup-app | nextcloud | internet gone 10min | +| 9 | 02:50 | use | uptime-kuma | internet gone 10min | +| 10 | 03:15 | restore | uptime-kuma | hard reset | +| 11 | 03:40 | use | paperless-ngx | drive pulled 20min | +| 12 | 04:05 | use | paperless-ngx | nothing | + +**Re-draw log** — a silent re-draw is a schedule chosen by the person running it, so every one is here: + +- r02 X=reinstall re-drawn (nothing has been removed yet) +- r06 Z=disk 95% full re-drawn (constraint 4: at most once) +- r07 Z=nothing re-drawn (constraint 6: never two in a row after r2) +- r08 X=reinstall re-drawn (nothing has been removed yet) + +**What the draw happened to give, said plainly before the night judges it:** no `remove` round was +ever drawn, so `reinstall` had nothing to reinstall and was re-drawn twice (rounds 2 and 8). Three +`internet gone` rounds land consecutively (7, 8, 9) — that is the seed's doing, and it makes rounds +7–9 a de-facto endurance test of the same accident against three different actions rather than three +independent samples. `controller killed`, `drive pulled 90s`, `memory pressure`, `agent restarted` +and `hub unreachable` were never drawn at all; **this night does not test them**, and the morning +verdict must not claim it did. + +## Phase 0 — the golden, the box, the household + +**0.1 Golden 0.245.0, baked and published.** Launched 19:52:47Z as a transient unit in the drill VM, +finished 19:58:15Z. Markers: `overlay2`=1, `including mount point`=2, `upload OK (HTTP 201)`=1, +FATAL=0, publish-skipped=0. `GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626`. +**The teardown was gated on the REGISTRY answering 200**, not on an exit code — and that mattered: +the wrapper exited **144** while every measured outcome was good. Vouched in the hub and read back +from the page (`golden currently vouched: 0.245.0`). The three bake failures of 2026-09-16 (`scp -P`, +`chmod 0700`, `GITEA_USER=admin`) were each guarded and none recurred. + +*Decision, recorded because silence reads as agreement:* the global controller floor was left at +0.244.0. The new box installs golden 0.245.0, which already carries controller 0.245.0, so no floor +was needed to deliver anything tonight; raising it would have pushed an update onto demo-felhom, a +box not in this drill. + +**0.2 The box.** VM 336 on demo-hp: 8 GiB, 4 cores, 32 G system + 100 G data disk on the NVMe at its +root, booted from the **published** ISO 1.28.0. Boot order set in its own `qm set` (combining it +silently yields `order=net0;ide2`). Install completion was judged **from the disk** — blocks used +grew 3233 → 6942 MiB then held across three checks — because „Automatically reboot" is ticked and a +finished install looks exactly like a stuck one on screen. The summary page was read before pressing +Install, and the line that made it safe was **„Disk(s): /dev/sda"** — the 32 G system disk alone. + +**The walk, as a volunteer, cost ZERO operator presses.** The box registered itself as an unclaimed +appliance and polled, visibly, until bound. The bind link came from the **waiting mail** (minted +18:17:46Z by yesterday's acknowledged host delete), the pairing code off the box's own console +(`4SY-4TX`), the „Tulajdonosi jelmondat" from the hub's customer record: + POST /bind/ -> 200, „Sikeres összekötés." +Then day-0 ran on its own and the hub recorded, without anyone pressing anything: + `appliance_bound` (customer_selfbind) · `appliance_credential_delivered` · `claim_reissued_reenroll` + · `offsite_reissued` · **`pbsdr_auto_reissue` — „Previous key destroyed (acknowledged deletion) — + credentials re-issued automatically."** + +**That last event is a first.** The brief named „the WG hook provisions by itself after an +acknowledged delete" as a claim never measured live. It ran tonight, unprompted. **Both pre-declared +presses (O1 self-bind, O2 re-issue) were therefore unnecessary.** + +The dashboard was claimed with the mailed code and **proven by logging in with the new password** — +a claim page that re-renders looks identical to success from the status code alone. The 100 GB data +drive was initialised through the wizard's own endpoint (`POST /api/storage/init`, polled to +`phase: done`), and `df` shows it mounted at `/mnt/felhom-drives/hdd_1` with 93 G free. + +**The box landed on tonight's golden with no hand upgrade:** controller **0.245.0** (healthy), agent +**0.131.0**, host `tester-1-022354` ONLINE. And the **R-543 escrow reminder bar shipped hours earlier +was live on it**, on a box nobody had touched. + +**0.3 The household — and the first real trouble.** Twelve deploys were fired; **ten were accepted, +two refused** for memory with both numbers quoted. Then nine of the ten failed: the guest's disks are +thin-provisioned over an ~11.8 GB pool carved from a 32 GB system disk, ten simultaneous image pulls +filled it, and the hub recorded `storage_fill_critical` (100 %) plus **nine `app_deploy_failed` +warnings**, one per app, each naming the failing pull. Only PrivateBin installed. + +**The product behaved; the harness did not.** R-536's failure event — shipped that same morning so an +interrupted install is not silence — fired for all nine within two minutes. The memory guard refused +rather than over-committing. The per-stack record stayed honest (`deployed: false`). The two faults +were mine: firing twelve deploys in two seconds is not household behaviour, and a 32 GB system disk +was copied from an earlier drill without checking what that drill had installed. + +**0.4 The escrow ceremony could NOT be completed — and this one is about tonight's own release.** +See „Finding: the recovery-code step cannot be done when the guide says to do it" below. + + +## Finding: the recovery-code step cannot be done when the guide says to do it (R-546) + +The box was at exactly the point of the guide this release added hours earlier — installed, bound +with no press, claimed, drive initialised, **no apps yet** — and the escrow reminder bar was on every +page telling the household to create their recovery code. It could not be done. + + POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"} + GET /api/escrow/status -> claimable:false — + detail: "exit 2: … selftest=escrow-create requires -storage (or escrow.pbs_storage…" + POST /api/escrow/claim -> **409** „A folyamat jelenlegi állapotában a kód nem kérhető le." + +**Both sides agreed on the cause.** The hub's own Backup & DR panel read „host enrolled **done** · WG +tunnel peer registered **done** · descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1) +**waiting** · ceremony possible once the descriptor is applied on the box". The box had no PBS storage +(`pvesm status`: `local`, `local-lvm` only) and no `escrow` section in `agent.json` at all. + +**It self-heals, and that was measured rather than assumed.** The box was left alone and polled: + + 20:30:12Z … 20:34:15Z pbs_storage=none escrow.pbs_storage_id=none + **20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs** + +~17 minutes after the bind. The retried ceremony passed every preflight item, the claim returned +**200** (83-character code, 129.2 bits of entropy, revealed once), and `escrow_state` became +**escrowed**. The bar then vanished from all four pages checked — the R-543 fix working through its +whole lifecycle on a box nobody had set up for the test. + +**So the defect is timing and wording, not mechanism.** For ~17 minutes a volunteer following +tonight's guide meets a stderr fragment about a `-storage` flag, while every page urges them on. +Filed **R-546** (P2). No product code was changed — this is a validation run. + +## Phase 1 — the twelve rounds + +PENDING + +## Phase 2 — the morning after + +PENDING + +## Interventions — counted, with the reason for each verdict + +PENDING + +## Teardown — three layers, stated + +PENDING diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/bake-0245.log b/documentation/audits/evidence-chaos-night-2026-09-17/bake-0245.log new file mode 100644 index 00000000..51c3436b --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/bake-0245.log @@ -0,0 +1,325 @@ +[golden] build-golden.sh v3.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.245.0 +[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) … + Logical volume "vm-9100-disk-0" created. + Logical volume pve/vm-9100-disk-0 changed. +Creating filesystem with 8388608 4k blocks and 2097152 inodes +Filesystem UUID: 7b98df7b-8d9d-43c8-ba99-0ba1075e83cc +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, + 4096000, 7962624 + Logical volume "vm-9100-disk-1" created. + Logical volume pve/vm-9100-disk-1 changed. +Creating filesystem with 6291456 4k blocks and 1572864 inodes +Filesystem UUID: e121482f-10e3-4687-b082-84db662425f4 +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, +extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst' +Total bytes read: 553512960 (528MiB, 86MiB/s) +Detected container architecture: amd64 +Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ... +done: SHA256:aX2YgyXohsbFO+/bVcsxaM32nSEJWtuQRE7krfEzqQ0 root@felhom-golden +Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ... +done: SHA256:QwivNTDkZ0thSB8KcWpCc0GTZ5He3NIFgGDmWBD5DQQ root@felhom-golden +Creating SSH host key 'ssh_host_rsa_key' - this may take some time ... +done: SHA256:HDmBrODCYm0syAs+LuleMy+QMBc/CA/GIPdmH0bppJA root@felhom-golden +[golden] starting + installing Docker (official repo, trixie channel) … +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation … +[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds … +[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) … +Unable to find image 'hello-world:latest' locally +latest: Pulling from library/hello-world +4f55086f7dd0: Pulling fs layer +4f55086f7dd0: Verifying Checksum +4f55086f7dd0: Download complete +4f55086f7dd0: Pull complete +Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8 +Status: Downloaded newer image for hello-world:latest + docker OK (overlay2; data-root /var/lib/docker) + /var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4 + /mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4 + both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576 +[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.245.0 (no registry cred at deploy) … + +WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'. +Configure a credential helper to remove this warning. See +https://docs.docker.com/go/credential-store/ + +0.245.0: Pulling from admin/felhom-controller +a8ac7f6c67ab: Pulling fs layer +bf30769d36e7: Pulling fs layer +044b66fbe46c: Pulling fs layer +b5c41a28e83f: Pulling fs layer +965b73d03024: Pulling fs layer +9dd06928817a: Pulling fs layer +b5c41a28e83f: Waiting +965b73d03024: Waiting +9dd06928817a: Waiting +044b66fbe46c: Verifying Checksum +044b66fbe46c: Download complete +b5c41a28e83f: Verifying Checksum +b5c41a28e83f: Download complete +a8ac7f6c67ab: Verifying Checksum +a8ac7f6c67ab: Download complete +965b73d03024: Verifying Checksum +965b73d03024: Download complete +9dd06928817a: Verifying Checksum +9dd06928817a: Download complete +bf30769d36e7: Verifying Checksum +bf30769d36e7: Download complete +a8ac7f6c67ab: Pull complete +bf30769d36e7: Pull complete +044b66fbe46c: Pull complete +b5c41a28e83f: Pull complete +965b73d03024: Pull complete +9dd06928817a: Pull complete +Digest: sha256:b6abd24d67be8ef1aa61b30f852f3a1f6e90ec3d3599a3824655a290c0dba875 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.245.0 +gitea.dooplex.hu/admin/felhom-controller:0.245.0 +[golden] asking the controller which infra images it manages … +[golden] baking infra images (4): traefik:v3.6.7 cloudflare/cloudflared:2026.6.0 gtstef/filebrowser:1.3.3-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 … +v3.6.7: Pulling from library/traefik +589002ba0eae: Pulling fs layer +ef63511ea6cc: Pulling fs layer +0738e5cb835e: Pulling fs layer +3e6813f70c64: Pulling fs layer +3e6813f70c64: Waiting +589002ba0eae: Verifying Checksum +589002ba0eae: Download complete +ef63511ea6cc: Verifying Checksum +ef63511ea6cc: Download complete +3e6813f70c64: Verifying Checksum +3e6813f70c64: Download complete +589002ba0eae: Pull complete +0738e5cb835e: Verifying Checksum +0738e5cb835e: Download complete +ef63511ea6cc: Pull complete +0738e5cb835e: Pull complete +3e6813f70c64: Pull complete +Digest: sha256:a9890c898f379c1905ee5b28342f6b408dc863f08db2dab20e46c267d1ff463a +Status: Downloaded newer image for traefik:v3.6.7 +docker.io/library/traefik:v3.6.7 +2026.6.0: Pulling from cloudflare/cloudflared +47de5dd0b812: Pulling fs layer +c172f21841df: Pulling fs layer +99515e7b4d35: Pulling fs layer +99ba982a9142: Pulling fs layer +d6b1b89eccac: Pulling fs layer +2780920e5dbf: Pulling fs layer +7c12895b777b: Pulling fs layer +3214acf345c0: Pulling fs layer +52630fc75a18: Pulling fs layer +dd64bf2dd177: Pulling fs layer +b839dfae01f6: Pulling fs layer +ebddc55facdc: Pulling fs layer +bdfd7f7e5bf6: Pulling fs layer +2d4d7adf6272: Pulling fs layer +40008157d8d2: Pulling fs layer +bd8962e29291: Pulling fs layer +cac2ae0193cb: Pulling fs layer +74d1dac84ecc: Pulling fs layer +dd64bf2dd177: Waiting +b839dfae01f6: Waiting +ebddc55facdc: Waiting +bdfd7f7e5bf6: Waiting +2d4d7adf6272: Waiting +40008157d8d2: Waiting +bd8962e29291: Waiting +cac2ae0193cb: Waiting +74d1dac84ecc: Waiting +2780920e5dbf: Waiting +7c12895b777b: Waiting +3214acf345c0: Waiting +52630fc75a18: Waiting +99ba982a9142: Waiting +d6b1b89eccac: Waiting +c172f21841df: Download complete +47de5dd0b812: Verifying Checksum +99515e7b4d35: Verifying Checksum +99515e7b4d35: Download complete +99ba982a9142: Verifying Checksum +99ba982a9142: Download complete +d6b1b89eccac: Verifying Checksum +d6b1b89eccac: Download complete +47de5dd0b812: Pull complete +2780920e5dbf: Verifying Checksum +2780920e5dbf: Download complete +7c12895b777b: Verifying Checksum +7c12895b777b: Download complete +3214acf345c0: Verifying Checksum +3214acf345c0: Download complete +52630fc75a18: Verifying Checksum +52630fc75a18: Download complete +dd64bf2dd177: Verifying Checksum +dd64bf2dd177: Download complete +b839dfae01f6: Verifying Checksum +b839dfae01f6: Download complete +c172f21841df: Pull complete +ebddc55facdc: Verifying Checksum +ebddc55facdc: Download complete +bdfd7f7e5bf6: Verifying Checksum +bdfd7f7e5bf6: Download complete +2d4d7adf6272: Verifying Checksum +2d4d7adf6272: Download complete +bd8962e29291: Verifying Checksum +bd8962e29291: Download complete +cac2ae0193cb: Verifying Checksum +cac2ae0193cb: Download complete +74d1dac84ecc: Verifying Checksum +74d1dac84ecc: Download complete +40008157d8d2: Verifying Checksum +40008157d8d2: Download complete +99515e7b4d35: Pull complete +99ba982a9142: Pull complete +d6b1b89eccac: Pull complete +2780920e5dbf: Pull complete +7c12895b777b: Pull complete +3214acf345c0: Pull complete +52630fc75a18: Pull complete +dd64bf2dd177: Pull complete +b839dfae01f6: Pull complete +ebddc55facdc: Pull complete +bdfd7f7e5bf6: Pull complete +2d4d7adf6272: Pull complete +40008157d8d2: Pull complete +bd8962e29291: Pull complete +cac2ae0193cb: Pull complete +74d1dac84ecc: Pull complete +Digest: sha256:ba461b8aa9c042156dbd39c38657fe7431bafa063220eab8d5330a523863da9f +Status: Downloaded newer image for cloudflare/cloudflared:2026.6.0 +docker.io/cloudflare/cloudflared:2026.6.0 +1.3.3-stable: Pulling from gtstef/filebrowser +6a0ac1617861: Pulling fs layer +ef8806083e82: Pulling fs layer +b74107c861c7: Pulling fs layer +adc935def003: Pulling fs layer +4f4fb700ef54: Pulling fs layer +18695ccc900a: Pulling fs layer +45d119d5c397: Pulling fs layer +dac52db4fc51: Pulling fs layer +6d598f86b2f2: Pulling fs layer +8aa349c8396c: Pulling fs layer +adc935def003: Waiting +4f4fb700ef54: Waiting +18695ccc900a: Waiting +45d119d5c397: Waiting +dac52db4fc51: Waiting +6d598f86b2f2: Waiting +8aa349c8396c: Waiting +6a0ac1617861: Verifying Checksum +6a0ac1617861: Download complete +b74107c861c7: Verifying Checksum +b74107c861c7: Download complete +adc935def003: Verifying Checksum +adc935def003: Download complete +4f4fb700ef54: Verifying Checksum +4f4fb700ef54: Download complete +45d119d5c397: Verifying Checksum +45d119d5c397: Download complete +ef8806083e82: Verifying Checksum +ef8806083e82: Download complete +dac52db4fc51: Verifying Checksum +dac52db4fc51: Download complete +6d598f86b2f2: Download complete +18695ccc900a: Verifying Checksum +18695ccc900a: Download complete +6a0ac1617861: Pull complete +8aa349c8396c: Download complete +ef8806083e82: Pull complete +b74107c861c7: Pull complete +adc935def003: Pull complete +4f4fb700ef54: Pull complete +18695ccc900a: Pull complete +45d119d5c397: Pull complete +dac52db4fc51: Pull complete +6d598f86b2f2: Pull complete +8aa349c8396c: Pull complete +Digest: sha256:eb3733681db8757412632c61a99ad656f0d94ed6781bb2ea114b4d70babab78c +Status: Downloaded newer image for gtstef/filebrowser:1.3.3-stable +docker.io/gtstef/filebrowser:1.3.3-stable +1.1.0: Pulling from admin/felhom-samba +897d797d2723: Pulling fs layer +3051591aa250: Pulling fs layer +ce57a3f93416: Pulling fs layer +fb94eeec2fe1: Pulling fs layer +fb94eeec2fe1: Waiting +ce57a3f93416: Verifying Checksum +ce57a3f93416: Download complete +fb94eeec2fe1: Verifying Checksum +fb94eeec2fe1: Download complete +897d797d2723: Verifying Checksum +897d797d2723: Download complete +3051591aa250: Verifying Checksum +3051591aa250: Download complete +897d797d2723: Pull complete +3051591aa250: Pull complete +ce57a3f93416: Pull complete +fb94eeec2fe1: Pull complete +Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0 +gitea.dooplex.hu/admin/felhom-samba:1.1.0 +[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'. +[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'. +[golden] baking the first-boot SSH host-key regeneration unit (F3) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'. +[golden] identity-clean + minimize … +[golden] stop + archive … +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +INFO: archive file size: 623MB +INFO: Finished Backup of VM 9100 (00:00:31) +[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_09_16-21_57_05.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive) +[golden] publishing golden (653729820 bytes, sha256 7a08aa1ad0bdd622…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.245.0/golden.tar.zst +[golden] pre-delete existing: HTTP 404 (404/204 expected) +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.245.0 +GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626 +[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.245.0 / 7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626 +[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/chaos_schedule.py b/documentation/audits/evidence-chaos-night-2026-09-17/chaos_schedule.py new file mode 100644 index 00000000..8279fbe6 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/chaos_schedule.py @@ -0,0 +1,154 @@ +#!/usr/bin/env python3 +# -*- coding: utf-8 -*- +"""chaos_schedule.py — draw the CHAOS NIGHT schedule, deterministically, ONCE. + +The point of this file is that the night cannot be chosen after the fact. The seed is the date, the +draws are stdlib `random` seeded with it, and the table this prints goes into the findings document +BEFORE round 1 runs. Anyone can re-run it and get the same night. + +A draw that breaks a constraint is RE-DRAWN and the re-draw is logged, because a silent re-draw is a +schedule chosen by the person running it. +""" +import hashlib +import random +import sys + +SEED = 20260917 + +ACTIONS = [ # (name, weight, what a household does) + ("use", 3, "ten minutes in the app: create, edit, upload, delete one thing"), + ("backup-app", 2, "„Mentés most” on the app"), + ("backup-system", 1, "whole-system „Mentés most”"), + ("offsite-run", 1, "the tier-3 leg through its own endpoint"), + ("update", 1, "the guarded Update on the app the catalog bumped"), + ("restore", 1, "restore the app through the page (wizard for off-site)"), + ("remove", 1, "remove the app with „az adataimat is töröld”"), + ("reinstall", 1, "reinstall an app removed earlier tonight, and restore it"), +] + +APPS = ["nextcloud", "immich", "bookstack", "privatebin", "gokapi", "vaultwarden", + "paperless-ngx", "jellyfin", "mealie", "uptime-kuma", "adventurelog", "homebox"] +# remove/reinstall may only touch these, so the photo and document apps survive for the morning restore +DISPOSABLE = ["privatebin", "gokapi", "homebox", "mealie"] + +ACCIDENTS = [ # (name, weight, how) + ("nothing", 3, "—"), + ("power cut", 2, "qm stop, 60 s, qm start"), + ("hard reset", 1, "qm reset"), + ("drive pulled 90s", 1, "detach the data disk, reattach after 90 s"), + ("drive pulled 20min", 1, "detach the data disk, reattach after 20 min"), + ("internet gone 10min", 2, "block the VM's outbound at the host, LAN kept"), + ("hub unreachable 15min", 1, "block only the hub's address from the VM"), + ("controller killed", 1, "docker kill felhom-controller"), + ("agent restarted", 1, "systemctl restart felhom-agent in the nested PVE"), + ("docker restarted", 1, "systemctl restart docker in the guest"), + ("disk 95% full", 1, "fill the system disk to 95 % for 10 min, then free it"), + ("tunnel down 10min", 1, "docker kill cloudflared"), + ("memory pressure", 1, "a throwaway container with a 1 GB hog for 5 min"), +] + +ROUNDS = 12 +START_MIN = 23 * 60 + 30 # 23:30 +SPACING = 25 + + +def wpick(rng, table): + names = [t[0] for t in table] + weights = [t[1] for t in table] + return rng.choices(names, weights=weights, k=1)[0] + + +def hhmm(total): + total %= 24 * 60 + return "%02d:%02d" % (total // 60, total % 60) + + +def draw(): + rng = random.Random(SEED) + rows, log = [], [] + used = {"update": 0, "controller killed": 0, "drive pulled 20min": 0, "disk 95% full": 0} + removed = [] # apps removed earlier tonight (reinstall needs one) + prev_accident = None + + for n in range(1, ROUNDS + 1): + t = hhmm(START_MIN + (n - 1) * SPACING) + + # ---- X, the action ------------------------------------------------- + for attempt in range(1, 40): + x = wpick(rng, ACTIONS) + if x == "update" and used["update"] >= 1: + log.append("r%02d X=update -> becomes 'use' (constraint 3: there is one bump)" % n) + x = "use" + if x == "reinstall" and not removed: + log.append("r%02d X=reinstall re-drawn (nothing has been removed yet)" % n) + continue + break + + # ---- Y, the app ---------------------------------------------------- + if x == "remove": + pool = [a for a in DISPOSABLE if a not in removed] + if not pool: + log.append("r%02d X=remove -> becomes 'use' (every disposable app is already removed)" % n) + x, pool = "use", APPS + y = rng.choice(pool) + elif x == "reinstall": + y = rng.choice(removed) + else: + y = rng.choice(APPS) + + # ---- Z, the accident ---------------------------------------------- + if n in (1, ROUNDS): + z = "nothing" # constraint 5: control rounds + else: + for attempt in range(1, 60): + z = wpick(rng, ACCIDENTS) + if z == "nothing" and prev_accident == "nothing" and n > 2: + log.append("r%02d Z=nothing re-drawn (constraint 6: never two in a row after r2)" % n) + continue + if z == "controller killed": + if used["controller killed"] >= 3: + log.append("r%02d Z=controller killed re-drawn (constraint 2: max 3 a night)" % n) + continue + if x == "restore": + log.append("r%02d Z=controller killed re-drawn (constraint 2: never with 'restore' — the F9/F3 class is already measured)" % n) + continue + if z == "drive pulled 20min" and used["drive pulled 20min"] >= 1: + log.append("r%02d Z=drive pulled 20min re-drawn (constraint 4: at most once)" % n) + continue + if z == "disk 95% full" and used["disk 95% full"] >= 1: + log.append("r%02d Z=disk 95%% full re-drawn (constraint 4: at most once)" % n) + continue + break + + if x in used: + used[x] += 1 + if z in used: + used[z] += 1 + if x == "remove": + removed.append(y) + if x == "reinstall" and y in removed: + removed.remove(y) + prev_accident = z + rows.append((n, t, x, y, z)) + + return rows, log + + +def main(): + rows, log = draw() + h = hashlib.sha256(open(__file__, "rb").read()).hexdigest() + print("seed: %d" % SEED) + print("script sha256: %s" % h) + print() + print("| # | time | X — the action | Y — the app | Z — the accident |") + print("|---|---|---|---|---|") + for n, t, x, y, z in rows: + print("| %d | %s | %s | %s | %s |" % (n, t, x, y, z)) + print() + print("re-draw log (%d entries):" % len(log)) + for line in log or [" (none)"]: + print(" " + line) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/events.sh b/documentation/audits/evidence-chaos-night-2026-09-17/events.sh new file mode 100755 index 00000000..7136a0c6 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/events.sh @@ -0,0 +1,33 @@ +#!/bin/bash +# events.sh — dump the hub's recorded events for the drill customer. +# +# This is the alarm-scoring surface. It reads the customer page's `events-table`, which is the +# hub's own record — NOT my notes. The hub pod has no sqlite3 and the DB is 300 MB, so the page is +# the right instrument; a copied hub.db without its -wal reads hours stale here. +# +# Usage: events.sh [n] — print the newest n rows (default 25) +set -u +N="${1:-25}" +cd /mnt/5_hdd/felhom.eu/git/felhom.eu +python3 scripts/read_credential.py HUB_PW /tmp/.ehp >/dev/null 2>&1 +HUB_PW=$(cat /tmp/.ehp); IP=$(sudo kubectl -n felhom-system get svc hub -o jsonpath='{.spec.clusterIP}') +curl -s -u ":$HUB_PW" "http://$IP:8080/customers/tester-1" -o /tmp/.ev.html +rm -f /tmp/.ehp +python3 - "$N" <<'PY' +import re,html,sys +n=int(sys.argv[1]) +t=open('/tmp/.ev.html',encoding='utf-8',errors='replace').read() +m=re.search(r'id="events-table"(.*?)', t, flags=re.S) +if not m: + print("EVENTS TABLE NOT FOUND — the instrument failed, and that is not the same as 'no events'") + raise SystemExit(1) +rows=re.findall(r']*>(.*?)', m.group(1), flags=re.S) +out=0 +for r in rows: + cells=[html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>','',c))).strip() for c in re.findall(r']*>(.*?)', r, flags=re.S)] + if not cells: continue + print(" | " + " | ".join(cells)[:200]) + out+=1 + if out>n: break +PY +rm -f /tmp/.ev.html diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/household_loop.sh b/documentation/audits/evidence-chaos-night-2026-09-17/household_loop.sh new file mode 100755 index 00000000..575aeab8 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/household_loop.sh @@ -0,0 +1,36 @@ +#!/bin/bash +# household_loop.sh — the light background household (CHAOS NIGHT Phase 0.4) +# +# Every 2 minutes, one read and one small write through a RANDOM app's front door, from the drill +# host — never from inside the box. It logs `time app op result`. It is the thing the accidents hit: +# its failures are DATA, not interventions, and the per-round summary is scored from this log. +# +# Usage: household_loop.sh +set -u +BOX="${1:?box ip}" +LOG="${2:?logfile}" +APPS="bookstack privatebin gokapi mealie homebox uptime-kuma adventurelog paperless-ngx nextcloud immich jellyfin vaultwarden" + +stamp(){ date -u +%FT%TZ; } +note(){ printf '%s %-14s %-6s %s\n' "$(stamp)" "$1" "$2" "$3" >> "$LOG"; } + +while true; do + APP=$(echo $APPS | tr ' ' '\n' | shuf -n1) + # READ: the app's own front door through the box's reverse proxy. + CODE=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 \ + -H "Host: ${APP}.enkicsifelhom.hu" "http://${BOX}/" 2>/dev/null) + case "$CODE" in + 2*|3*) note "$APP" read "ok http=$CODE" ;; + 000) note "$APP" read "UNREACHABLE (no answer within 15 s)" ;; + *) note "$APP" read "FAILED http=$CODE" ;; + esac + # WRITE: a few KB to the box's own file manager share area, which every app's drive shares. + W=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 \ + -H "Host: felhom.enkicsifelhom.hu" "http://${BOX}/api/health" 2>/dev/null) + case "$W" in + 2*) note "$APP" write "ok http=$W" ;; + 000) note "$APP" write "UNREACHABLE" ;; + *) note "$APP" write "FAILED http=$W" ;; + esac + sleep 120 +done diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh b/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh new file mode 100755 index 00000000..5be9ad99 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/inject.sh @@ -0,0 +1,76 @@ +#!/bin/bash +# inject.sh — CHAOS NIGHT accident injectors. Run from DooPlex. +# +# Only the seven accidents the seed actually drew are implemented. The other six in the brief's +# table were never drawn, and writing injectors for them would suggest this night tested them. +# +# Usage: inject.sh +# power-cut | hard-reset | drive-pulled-20min | internet-gone-10min +# tunnel-down-10min | docker-restart | disk-95-full +set -u +A="${1:?accident}"; R="${2:?round}" +VM=336 +E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17 +LOG="$E/round-${R}-accident.txt" +H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; } +say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$LOG"; } +# the customer guest lives INSIDE the nested PVE; these run one hop further in +G(){ H "qm guest exec $VM -- $*" ; } + +say "ACCIDENT=$A round=$R" +case "$A" in + power-cut) + say "qm stop $VM (the plug is pulled — no clean shutdown)" + H "qm stop $VM"; say "stopped; waiting 60 s with the box dark" + sleep 60 + H "qm start $VM"; say "power back on" + ;; + hard-reset) + say "qm reset $VM (the reset button, mid-write)" + H "qm reset $VM"; say "reset issued" + ;; + drive-pulled-20min) + F=$(H "qm config $VM | grep '^scsi1:' | sed 's/scsi1: //; s/,.*//'") + say "data disk is $F — detaching it from the RUNNING box (the cable is pulled)" + H "qm set $VM --delete scsi1"; say "detached; the disk file stays as unused0" + say "leaving it out for 20 minutes" + sleep 1200 + H "qm set $VM --scsi1 $F"; say "re-attached: $F" + ;; + internet-gone-10min) + TAP=$(H "ls /sys/class/net | grep -E \"^tap${VM}i0$\"") + [ -n "$TAP" ] || { say "NO TAP FOUND for VM $VM — accident NOT injected, and that is recorded as such"; exit 1; } + say "blocking the box's traffic off-LAN at the HOST, on $TAP; the LAN stays up" + H "sysctl -w net.bridge.bridge-nf-call-iptables=1 >/dev/null + iptables -I FORWARD 1 -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT + iptables -I FORWARD 2 -m physdev --physdev-in $TAP -j DROP" + say "blocked (LAN allowed, everything else dropped) — 10 minutes" + sleep 600 + H "iptables -D FORWARD -m physdev --physdev-in $TAP -j DROP + iptables -D FORWARD -m physdev --physdev-in $TAP -d 192.168.0.0/24 -j ACCEPT + sysctl -w net.bridge.bridge-nf-call-iptables=0 >/dev/null" + say "unblocked; host sysctl restored to 0 and both rules removed" + H "iptables -S FORWARD | head -5" | tee -a "$LOG" + ;; + tunnel-down-10min) + say "docker kill cloudflared inside the customer guest" + G "docker kill cloudflared" ; say "tunnel killed — 10 minutes" + sleep 600 + say "10 minutes up; NOT restarting it by hand — whether it returns by itself IS the measurement" + ;; + docker-restart) + say "systemctl restart docker inside the customer guest" + G "systemctl restart docker"; say "docker restarted" + ;; + disk-95-full) + say "filling the customer guest's SYSTEM disk to ~95 %" + G "bash -c 'df -h / | tail -1'" | tee -a "$LOG" + G "bash -c 'F=\$(df --output=avail -m / | tail -1); fallocate -l \$(( (F * 95 / 100) ))M /var/tmp/.chaosfill && df -h / | tail -1'" | tee -a "$LOG" + say "full — holding 10 minutes" + sleep 600 + G "bash -c 'rm -f /var/tmp/.chaosfill; df -h / | tail -1'" | tee -a "$LOG" + say "freed" + ;; + *) say "UNKNOWN ACCIDENT $A"; exit 2 ;; +esac +say "accident $A complete" diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/morning_after.sh b/documentation/audits/evidence-chaos-night-2026-09-17/morning_after.sh new file mode 100755 index 00000000..6ac8a657 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/morning_after.sh @@ -0,0 +1,34 @@ +#!/bin/bash +# morning_after.sh — CHAOS NIGHT Phase 2. Run at ~05:00. +# +# Four things, in this order, and each one asks the product rather than my notes: +# 1. every app healthy through its FRONT DOOR; every version label true; every backup page honest +# 2. one DB-backed app restored from the OFF-SITE tier onto scratch guest 9202, read back +# 3. the alarm truth table (fired / true? and should-have-fired / did it?) +# 4. the background household loop's per-round failure summary +set -u +E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17 +OUT="$E/phase2-morning-after.txt" +say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; } +H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; } + +say "=== 1. every app, through its own front door ===" +H "qm guest exec 336 -- bash -c 'pct exec 9000 -- docker ps --format \"{{.Names}}\t{{.Status}}\"'" | tee -a "$OUT" + +say "=== 1b. the backup page, per tier, read as a first-timer would ===" +say "(quoted verbatim into the findings doc; searched with ASCII fragments + controls)" + +say "=== 3. the alarm truth table inputs ===" +say "every alarm the hub RECORDED for this customer tonight, with its delivery status." +say "'suppressed' and 'never fired' are DIFFERENT and are not allowed to collapse into one." + +say "=== 4. the background household loop ===" +if [ -f "$E/household.log" ]; then + say "total household operations: $(wc -l < "$E/household.log")" + say "failures: $(grep -cE 'FAILED|UNREACHABLE' "$E/household.log")" + say "--- failures grouped by 25-minute round window ---" + awk '/FAILED|UNREACHABLE/{print substr($1,12,2)":"substr($1,15,1)"0"}' "$E/household.log" | sort | uniq -c | tee -a "$OUT" +else + say "NO household log — the loop did not run, and that is recorded as a gap, not glossed over." +fi +say "=== end of the morning-after collection ===" diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-bind.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-bind.txt new file mode 100644 index 00000000..5d6855ab --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-bind.txt @@ -0,0 +1,3 @@ +bind submitted at 2026-09-16T20:18:15Z / 22:18 CEST +POST bind -> 200 + page says: Felhom — Doboz összekötése Felhom doboz összekötése Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz. Felhom.eu diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-day0.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-day0.txt new file mode 100644 index 00000000..8e788405 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-day0.txt @@ -0,0 +1,34 @@ +2026-09-16T20:20:45Z hub-based day-0 watcher started (no SSH needed: the hub sees guests, agent and controller version) +2026-09-16T20:20:46Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:21:16Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:21:46Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:22:17Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:22:47Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:23:17Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:23:47Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:24:18Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:24:48Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:25:18Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:25:48Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:26:19Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:26:49Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:27:19Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:27:50Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:28:20Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:28:50Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:29:20Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:29:51Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:30:21Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:30:51Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:31:21Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:31:52Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:32:22Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:32:53Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:33:23Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:33:53Z tester-1-022354 Tester 1 0.131.0 ONLINE 0/0 0% 20% 31% inactive 31% local Felhom Hub 0.116.0 +2026-09-16T20:34:23Z tester-1-022354 Tester 1 0.131.0 ONLINE 1/1 16% 27% 40% inactive 100% local-lvm Felhom Hub 0.116.0 +2026-09-16T20:34:23Z GUEST CREATED +2026-09-16T20:34:23Z --- controller version the new guest landed on (asked of the customer page) --- + Controller 0.245.0 + Last report: 4 min ago +2026-09-16T20:34:24Z day-0 watcher done diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-deploy-failure.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-deploy-failure.txt new file mode 100644 index 00000000..dd9b3b76 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-deploy-failure.txt @@ -0,0 +1,38 @@ +## NINE DEPLOYS FAILED — and the product said so, correctly, within two minutes +## 2026-09-16, box tester-1-022354 / guest 9201, controller 0.245.0 + +### What actually happened (root cause, measured) +Ten deploys were admitted at 20:30:19, each passing the controller's memory check, which counts +COMMITTED memory rather than memory in use: + total=4096MB reserved=384MB usable=3712MB — committed_used climbed 2514 -> 3626MB + the 11th and 12th (mealie, adventurelog) were REFUSED with the numbers in the message +Then the image pulls began, and at **20:34 the hub recorded**: + `storage_fill_critical` (critical) — „Host tester-1-022354: storage \"local-lvm\" CRITICALLY full + at 100% (threshold 95%) — backups/writes to it will fail; free space immediately" +Between 20:32 and 20:33, **nine `app_deploy_failed` (warning) events** were pushed, one per app, each +quoting the failing pull: Gokapi, Paperless-ngx, Vaultwarden, Immich, BookStack, Nextcloud, Homebox, +Jellyfin, Uptime Kuma. Only **PrivateBin** completed (86 s, `app_deployed` at 20:31:45). + +The disk is thin-provisioned: the VM's 32 GB system disk yields an ~11.8 GB LVM thin pool, over which +the guest's 32 GB rootfs and 70 GB data volume are over-subscribed. Ten simultaneous image pulls +filled the pool. + +### What this says about the PRODUCT — it behaved, and two of tonight's own concerns are answered + * **R-536's `app_deploy_failed` works.** That event shipped this morning precisely so an + interrupted install is not silence. Nine interrupted installs produced nine warnings, each + naming the app and the failing image, within ~2 minutes of the failure. Before R-536 this was + silence, and the customer would have been left with nine cards that never resolved. + * **The fill alarm fired at CRITICAL severity** with an actionable sentence, at 95 % threshold. + * **The memory guard refused rather than over-committing**, and quoted both numbers it compared. + * The box's own per-stack record is honest: `deployed: false` for all nine, `true` only for + privatebin. Nothing claims to be installed that is not. + +### What this says about MY HARNESS — two errors, both mine + 1. **Twelve deploys fired in two seconds is not household behaviour.** A household installs an app, + waits for it, then installs another. Firing them in parallel is what drove committed memory to + the ceiling and ten image pulls onto one thin pool at once. + 2. **The drill VM's system disk (32 GB) is too small for a twelve-app household.** That size was + copied from a previous drill's VM without checking what that drill actually installed. + +Neither is a product defect and neither is recorded as one. The recovery, and the re-seed done one +app at a time, are recorded next. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-deploy-progress.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-deploy-progress.txt new file mode 100644 index 00000000..4747a583 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-deploy-progress.txt @@ -0,0 +1,602 @@ +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:31:42Z containers=4 app.yaml files=10 available_MB=3797 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:32:30Z containers=5 app.yaml files=10 available_MB=3720 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:33:18Z containers=5 app.yaml files=10 available_MB=3818 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:34:05Z containers=5 app.yaml files=10 available_MB=3873 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:34:53Z containers=5 app.yaml files=10 available_MB=3891 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:35:41Z containers=5 app.yaml files=10 available_MB=3890 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:36:29Z containers=5 app.yaml files=10 available_MB=3900 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:37:17Z containers=5 app.yaml files=10 available_MB=3900 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:38:05Z containers=5 app.yaml files=10 available_MB=3899 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:38:53Z containers=5 app.yaml files=10 available_MB=5947 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +lxc-attach: 9201: ../src/lxc/attach.c: get_attach_context: 406 Connection refused - Failed to get init pid +lxc-attach: 9201: ../src/lxc/attach.c: lxc_attach: 1474 Connection refused - Failed to get attach context +2026-09-16T20:39:41Z containers=0 app.yaml files= available_MB= +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:40:28Z containers=5 app.yaml files=10 available_MB=5933 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:41:16Z containers=5 app.yaml files=10 available_MB=5895 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:42:05Z containers=9 app.yaml files=10 available_MB=4115 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:42:53Z containers=15 app.yaml files=10 available_MB=3704 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:43:42Z containers=17 app.yaml files=10 available_MB=3714 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:44:31Z containers=21 app.yaml files=11 available_MB=3646 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:45:20Z containers=23 app.yaml files=12 available_MB=3131 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:46:09Z containers=25 app.yaml files=12 available_MB=3310 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:46:57Z containers=26 app.yaml files=12 available_MB=4344 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:47:45Z containers=26 app.yaml files=12 available_MB=4292 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:48:33Z containers=26 app.yaml files=12 available_MB=4489 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:49:21Z containers=26 app.yaml files=12 available_MB=4507 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:50:09Z containers=26 app.yaml files=12 available_MB=4503 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:50:57Z containers=26 app.yaml files=12 available_MB=4489 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:51:45Z containers=26 app.yaml files=12 available_MB=4493 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:52:33Z containers=26 app.yaml files=12 available_MB=4464 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:53:21Z containers=26 app.yaml files=12 available_MB=4492 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:54:09Z containers=26 app.yaml files=12 available_MB=4408 +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +2026-09-16T20:54:57Z containers=26 app.yaml files=12 available_MB=4479 diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-convergence.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-convergence.txt new file mode 100644 index 00000000..bf2dd590 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-convergence.txt @@ -0,0 +1,7 @@ +2026-09-16T20:30:12Z pbs_storage=none escrow.pbs_storage_id=none +2026-09-16T20:31:13Z pbs_storage=none escrow.pbs_storage_id=none +2026-09-16T20:32:14Z pbs_storage=none escrow.pbs_storage_id=none +2026-09-16T20:33:15Z pbs_storage=none escrow.pbs_storage_id=none +2026-09-16T20:34:15Z pbs_storage=none escrow.pbs_storage_id=none +2026-09-16T20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs +CONVERGED diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt new file mode 100644 index 00000000..8ffa5f0e --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt @@ -0,0 +1,60 @@ +## THE RECOVERY-CODE STEP FAILED ON A FRESH BOX — recorded verbatim, before any retry +## 2026-09-16, box `tester-1-022354` / guest 9201, controller 0.245.0, agent 0.131.0 + +Context: this is the step controller v0.245.0 added to `VOLUNTEER-first-hour.md` a few hours ago — +„A helyreállítási kód (~2 perc) — ezt ne hagyd ki", placed deliberately right after the dashboard +password and BEFORE the first app, because until it is done the off-site backup does not run. + +The box was at exactly that point of the guide: installed, bound with no press, claimed, data drive +initialised, no apps yet. The escrow reminder bar was on every page, telling the household to do +precisely this. + + GET /api/escrow/preflight -> (empty response body) + POST /api/escrow/start -> 200 {"job_id":"escrow-1789590499667361664","phase":"running"} + GET /api/escrow/status -> claimable:false, claimed:false, and: + + detail: "exit 2: exit status 2 | stderr: selftest=escrow-create requires -storage + (or escrow.pbs_storage…" + + POST /api/escrow/claim -> **409** + „A folyamat jelenlegi állapotában a kód nem kérhető le." + +So: the ceremony starts, the agent's `escrow-create` selftest refuses for a missing PBS storage id, +and the one-shot claim then correctly declines. The refusal is fail-closed and the wording is honest +— nothing pretended to succeed. What is wrong is that the household is TOLD to do this now, on every +page, and at this moment it cannot be done. + +Nothing was retried before this file was written, so the state above is the state the box was in. + +## WHY it refused — measured on both sides, not guessed + +**The hub's own Backup & DR panel says it in plain words:** + host enrolled (tester-1-022354) done + WG tunnel peer registered done + descriptor provisioned (namespace tester-1, token felhom@pbs!tester-1) **waiting** + „ceremony possible once the descriptor is applied on the box" + +**The box agrees:** + pvesm status -> only `local` (dir) and `local-lvm` (lvmthin). **No PBS storage exists yet.** + /etc/pve/storage.cfg -> no `pbs:` entry + /etc/felhom-agent/agent.json -> top-level keys are + [authz, backup, deployment_mode, hub, lan_resolver, local_api, log_level, oob, privileged, + proxmox, storage, wg_tunnel] + — there is **no `escrow` section at all**, so `escrow.pbs_storage_id` is unset, which is + precisely what the agent's selftest complained about. + +**The controller, meanwhile, already has the off-site target:** + offbox present=True, enabled=True, host=u629488-sub4.your-storagebox.de, escrow_state=**pending** + +So the chain is: the hub provisioned the DR descriptor automatically at 20:19 (no press), the +controller already knows its off-site destination, but the AGENT has not yet applied the descriptor +on the box — and the escrow ceremony depends on that. The hub documents the dependency in the very +panel that shows it as „waiting". + +## The question this does NOT yet answer, and how it is being measured +Whether the box applies the descriptor **by itself**, and how long that takes. That decides +everything about severity: a few minutes of convergence makes the new guide step slightly too early +in the journey; never converging without an operator press makes it a broken promise on every fresh +box. A watcher is now polling the box for `pvesm` gaining a PBS storage and the agent config gaining +`escrow.pbs_storage_id`, and the ceremony will be retried when it does. Nothing was pressed, and the +box is being left to do it alone. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-retry.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-retry.txt new file mode 100644 index 00000000..d6e7d03b --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-escrow-retry.txt @@ -0,0 +1,62 @@ + session=64 csrf=64 + --- preflight --- +{"data":{"agent_supported":true,"escrow_state":"pending","items":[{"id":"pbs_storage_id","ok":true,"detail":"felhom-pbs"},{"id":"dr_tier","ok":true,"detail":"DR tier applied"},{"id":"age_binary","ok":true,"detail":"/usr/bin/age"},{"id":"hub_upload","ok":true,"detail":"hub upload target configured"},{"id":"staged_secret","ok":true,"detail":"staged secret present"},{"id":"sudo_grant","ok":true,"deta + --- start (re-authenticates with the dashboard password) --- + POST /api/escrow/start -> 200 +{"data":{"job_id":"escrow-1789591140971464235","phase":"running"},"error":"","ok":true} + + --- poll --- + {"data":{"claim_expires_in_sec":598,"claimable":true,"claimed":false,"detail":"","entropy_bits":129.24070185585344,"job_id":"escrow-1789591140971464235","key_fingerprint":"6b:ca:5f:3f:ca:0f:e2:3f:fb:2 + --- claim (ONE-SHOT reveal; value goes to a file, never to stdout) --- + POST /api/escrow/claim -> 200 + recovery code claimed: 83 chars, 10 words — written to a file, NOT printed + --- escrow state after the ceremony --- + +## THE ANSWER: the gap is a TIMING gap, and the box closes it BY ITSELF +The watcher left the box alone and polled. Nothing was pressed. + + 20:30:12Z pbs_storage=none escrow.pbs_storage_id=none + 20:31:13Z pbs_storage=none escrow.pbs_storage_id=none + 20:32:14Z pbs_storage=none escrow.pbs_storage_id=none + 20:33:15Z pbs_storage=none escrow.pbs_storage_id=none + 20:34:15Z pbs_storage=none escrow.pbs_storage_id=none + **20:35:16Z pbs_storage=felhom-pbs escrow.pbs_storage_id=felhom-pbs -> CONVERGED** + +That is **~17 minutes after the bind** (20:18:15Z) and ~6 minutes after the ceremony first refused. +The retry then passed every preflight item: + pbs_storage_id ok (felhom-pbs) · dr_tier ok (DR tier applied) · age_binary ok (/usr/bin/age) + · hub_upload ok · staged_secret ok · sudo_grant ok · agent_supported true · escrow_state pending + + POST /api/escrow/start -> 200, job escrow-1789591140971464235, phase running + GET /api/escrow/status -> claimable:true, entropy_bits 129.2, key fingerprint 6b:ca:5f:3f:… + POST /api/escrow/claim -> **200** — the recovery code, 83 characters / 10 words, revealed ONCE + (written to a 0600 file; not printed, not committed — the household writes it on paper) + +**So the product is not broken here, and the fix shipped tonight is not wrong — the GUIDE's timing +is.** `VOLUNTEER-first-hour.md` §6 (written a few hours ago) places the recovery code immediately +after the dashboard password and before the first app. On a fresh box that moment is inside the +~17-minute window where the agent has not yet applied the DR descriptor, so a volunteer following the +guide literally meets „exit 2: selftest=escrow-create requires -storage " and a 409, +with nothing on the page telling them to simply wait a quarter of an hour. + +The escrow reminder bar (R-543, also shipped tonight) makes this sharper rather than softer: it is on +every page urging the household to do the very thing that cannot yet be done. + +## The state flipped: escrow_state = **escrowed** +Read from the box's own settings after the ceremony: + enabled=True host=u629488-sub4.your-storagebox.de + **escrow_state=escrowed** + last_run=None last_status=None snapshot_count=None (never run yet — honest, not "0") + +„Claimed" and „escrowed" are different facts: the claim is the household seeing the code once, the +flip to `escrowed` happens when the hub's ACK confirms `sha256(local repo_password)` matches the +stored escrow. Both happened. The off-site tier is therefore ARMED for the first time on this box, +and the escrow reminder bar shipped tonight should now be gone from every page — which the first +round will read back rather than assume. + +## An alarm caused by MY recovery, flagged so it is never scored as a round's alarm + Sep 16 20:39 **error storage_disconnected** „Meghajtó váratlanul leválasztva: Adatlemez" +That is the guest restart I performed to clear the read-only wedge. It is a TRUE alarm — the drive +really did go away for those seconds — but it belongs to my repair, not to any accident in the +schedule. Round 11's drawn accident is `drive pulled 20min`, and when that round is scored this +20:39 event must not be mistaken for it. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-events-baseline.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-events-baseline.txt new file mode 100644 index 00000000..a912fda9 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-events-baseline.txt @@ -0,0 +1,9 @@ +## events baseline, captured 2026-09-16T20:15:23Z, BEFORE round 1 +## everything above this line in later dumps belongs to yesterday's box, not tonight's + | Time | Severity | Type | Message | Source + | Sep 16 18:47 | error | node_down | No report received for 1h | hub + | Sep 16 18:17 | info | selfbind_link_sent | Self-bind link e-mailed (host delete) | hub + | Sep 16 18:17 | warning | node_stale | No report received for 30m | hub + | Sep 16 18:16 | warning | host_stale | Host tester-1-33b6a9: no report for 30m | hub + | Sep 16 17:25 | info | pbsdr_adopted | PBS DR token adopted by operator re-issue (endpoint held a token, hub had no descriptor) | hub + | Sep 16 17:16 | info | app_deployed | Alkalmazás telepítve: Nextcloud | controller diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-firstboot.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-firstboot.txt new file mode 100644 index 00000000..13136b1f --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-firstboot.txt @@ -0,0 +1,66 @@ +2026-09-16T20:11:58Z first-boot watcher started (box 192.168.0.115, VM 336) +2026-09-16T20:11:58Z PVE port 8006 answers +2026-09-16T20:14:57Z watcher restarted (the previous one killed ITSELF: pkill -f matched its own command line) +2026-09-16T20:14:58Z bootstrap=activating agent=none guests=0 +2026-09-16T20:15:20Z bootstrap=activating agent=none guests=0 +2026-09-16T20:15:42Z bootstrap=activating agent=none guests=0 +2026-09-16T20:16:04Z bootstrap=activating agent=none guests=0 +2026-09-16T20:16:25Z bootstrap=activating agent=none guests=0 +2026-09-16T20:16:47Z bootstrap=activating agent=none guests=0 +2026-09-16T20:17:09Z bootstrap=activating agent=none guests=0 +2026-09-16T20:17:30Z bootstrap=activating agent=none guests=0 +2026-09-16T20:17:52Z bootstrap=activating agent=none guests=0 +2026-09-16T20:18:14Z bootstrap=activating agent=none guests=0 +2026-09-16T20:18:36Z bootstrap=activating agent=none guests=0 +2026-09-16T20:18:58Z bootstrap=activating agent=none guests=0 +2026-09-16T20:19:27Z bootstrap= agent=none guests=0 +2026-09-16T20:19:56Z bootstrap= agent=none guests=0 +2026-09-16T20:20:25Z bootstrap= agent=none guests=0 +2026-09-16T20:20:54Z bootstrap= agent=none guests=0 +2026-09-16T20:21:23Z bootstrap= agent=none guests=0 +2026-09-16T20:21:50Z bootstrap= agent=none guests=0 +2026-09-16T20:22:19Z bootstrap= agent=none guests=0 +2026-09-16T20:22:48Z bootstrap= agent=none guests=0 +2026-09-16T20:23:17Z bootstrap= agent=none guests=0 +2026-09-16T20:23:46Z bootstrap= agent=none guests=0 +2026-09-16T20:24:15Z bootstrap= agent=none guests=0 +2026-09-16T20:24:44Z bootstrap= agent=none guests=0 +2026-09-16T20:25:12Z bootstrap= agent=none guests=0 +2026-09-16T20:25:41Z bootstrap= agent=none guests=0 +2026-09-16T20:26:10Z bootstrap= agent=none guests=0 +2026-09-16T20:26:38Z bootstrap= agent=none guests=0 +2026-09-16T20:27:07Z bootstrap= agent=none guests=0 +2026-09-16T20:27:36Z bootstrap= agent=none guests=0 +2026-09-16T20:28:05Z bootstrap= agent=none guests=0 +2026-09-16T20:28:32Z bootstrap= agent=none guests=0 +2026-09-16T20:29:01Z bootstrap= agent=none guests=0 +2026-09-16T20:29:30Z bootstrap= agent=none guests=0 +2026-09-16T20:29:59Z bootstrap= agent=none guests=0 +2026-09-16T20:30:28Z bootstrap= agent=none guests=0 +2026-09-16T20:30:58Z bootstrap= agent=none guests=0 +2026-09-16T20:31:27Z bootstrap= agent=none guests=0 +2026-09-16T20:31:56Z bootstrap= agent=none guests=0 +2026-09-16T20:32:25Z bootstrap= agent=none guests=0 +2026-09-16T20:32:54Z bootstrap= agent=none guests=0 +2026-09-16T20:33:23Z bootstrap= agent=none guests=0 +2026-09-16T20:33:52Z bootstrap= agent=none guests=0 +2026-09-16T20:34:21Z bootstrap= agent=none guests=0 +2026-09-16T20:34:50Z bootstrap= agent=none guests=0 +2026-09-16T20:35:19Z bootstrap= agent=none guests=0 +2026-09-16T20:35:48Z bootstrap= agent=none guests=0 +2026-09-16T20:36:17Z bootstrap= agent=none guests=0 +2026-09-16T20:36:46Z bootstrap= agent=none guests=0 +2026-09-16T20:37:15Z bootstrap= agent=none guests=0 +2026-09-16T20:37:44Z bootstrap= agent=none guests=0 +2026-09-16T20:38:13Z bootstrap= agent=none guests=0 +2026-09-16T20:38:42Z bootstrap= agent=none guests=0 +2026-09-16T20:39:11Z bootstrap= agent=none guests=0 +2026-09-16T20:39:40Z bootstrap= agent=none guests=0 +2026-09-16T20:40:09Z bootstrap= agent=none guests=0 +2026-09-16T20:40:38Z bootstrap= agent=none guests=0 +2026-09-16T20:41:07Z bootstrap= agent=none guests=0 +2026-09-16T20:41:36Z bootstrap= agent=none guests=0 +2026-09-16T20:42:05Z bootstrap= agent=none guests=0 +2026-09-16T20:42:25Z --- bootstrap journal, last 40 --- +2026-09-16T20:42:28Z --- what it landed on --- +2026-09-16T20:42:31Z watcher done diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-golden-bake.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-golden-bake.txt new file mode 100644 index 00000000..527bb660 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-golden-bake.txt @@ -0,0 +1,17 @@ +## 2026-09-16T19:52:02Z golden 0.245.0 — bake in the drill VM (RUNBOOK-manual-build §4.0/§4.1) + reverted to virgin (this also proves no qemu holds the qcow2) + cold boot started + ssh up: pve-manager/9.2.2/b9984c6d90a4bd80 (running kernel: 7.0.2-6-pve) + template: debian-13-standard_13.6-1_amd64.tar.zst + template downloaded + script + token landed, non-empty, executable + bake launched 2026-09-16T19:52:47Z (GITEA_USER=admin — the 'kisfenyo' namespace error cost a bake) + token leak check on the unit (needle proven non-empty, so the grep cannot match everything): 0 + unit state: inactive at 2026-09-16T19:58:15Z + log copied off the machine FIRST: 325 lines + markers: overlay2=1 mountpoints=2 upload=1 FATAL=0 skipped=0 +GOLDEN_VERSION=0.245.0 +GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626 + REGISTRY CHECK (the outcome, not the attempt): golden 0.245.0 -> http=200 + PUBLISHED; guest destroyed, qemu exited, disk reverted to virgin +## bake finished 2026-09-16T19:58:52Z diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-install-timeline.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-install-timeline.txt new file mode 100644 index 00000000..2a74b7f7 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-install-timeline.txt @@ -0,0 +1,21 @@ +install pressed at 2026-09-16T20:07:11Z +2026-09-16T20:07:33Z watcher started; completion is judged from the DISK (blocks actually used), never from the screen +2026-09-16T20:07:34Z system-disk blocks used: 3233 MiB (stable checks: 0) +2026-09-16T20:08:04Z system-disk blocks used: 5375 MiB (stable checks: 0) +2026-09-16T20:08:35Z system-disk blocks used: 6762 MiB (stable checks: 0) +2026-09-16T20:09:06Z system-disk blocks used: 6843 MiB (stable checks: 0) +2026-09-16T20:09:36Z system-disk blocks used: 6942 MiB (stable checks: 0) +2026-09-16T20:10:07Z system-disk blocks used: 6942 MiB (stable checks: 1) +2026-09-16T20:10:37Z system-disk blocks used: 6942 MiB (stable checks: 2) +2026-09-16T20:11:08Z system-disk blocks used: 6942 MiB (stable checks: 3) +2026-09-16T20:11:08Z INSTALL COMPLETE (disk stopped growing at 6942 MiB) +2026-09-16T20:11:08Z applying the post-install fix: stop, detach the CD, boot order scsi0, start from disk +disk usage before: 6942 MiB +update VM 336: -delete ide2 +update VM 336: -boot order=scsi0 +boot: order=scsi0 +scsi0: nvme-scratch:336/vm-336-disk-1.raw,size=32G +scsi1: nvme-scratch:336/vm-336-disk-0.raw,size=100G +scsihw: virtio-scsi-single +started from disk at 2026-09-16T20:11:19Z +2026-09-16T20:11:19Z post-install fix done diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-notes.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-notes.txt new file mode 100644 index 00000000..9852c059 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-notes.txt @@ -0,0 +1,362 @@ +## CHAOS NIGHT — Phase 0 notes (2026-09-16 evening, CEST) + +### Baselines, re-verified live against Gitea at 21:49 CEST (not copied from the brief) +felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) — clean, in sync +felhom-agent e98b857684f4 v0.131.0 — clean, in sync +felhom.eu d124c77e176d hub v0.116.0, ISO 1.28.0 live — clean, in sync +app-catalog 94bc5febaca2 — clean, in sync +Register: highest R-545, 212 open. Golden waiver valid to 2026-09-27. + +### A claim in the brief, CHECKED rather than inherited +The brief says the customer `tester-1` has "no host". CONFIRMED: /hosts lists exactly three hosts — +demo-felhom-8363b5, demo-hp-bb76ea, drill-r50-0a4f9a. There is no tester-1 host record. The +customers list showing „tester-1 … DOWN … 0.244.0" is the customer's LAST KNOWN state, not a live +box; reading that row as a host record would have been the mistake. + +### 0.1 — golden 0.245.0, baked and published +Launched 19:52:47Z as a transient unit in the drill VM; finished 19:58:15Z. + markers: overlay2=1 mountpoints=2 upload=1 FATAL=0 publish-skipped=0 + GOLDEN_VERSION=0.245.0 + GOLDEN_SHA256=7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626 + REGISTRY CHECK (the outcome, not the attempt): golden 0.245.0 -> http=200 + then: guest 9100 destroyed, token shredded, qemu exited, disk reverted to virgin. + + All three of 2026-09-16's bake failures were guarded against and none recurred: `scp -P` (not the + ssh `-p`), `chmod 0700` on the script, and `GITEA_USER=admin` (not the first credentials line). + The token-leak check ran with a needle PROVEN non-empty first, because `grep -F ""` matches every + line and an instrument that reports a hit on an empty needle is not a measurement. + + HONEST NOTE ON THE EXIT CODE: the wrapper script exited **144** while every measured outcome was + good. That is why teardown was gated on the REGISTRY answering 200 and not on an exit code — this + repo's own "exit codes that lie" class. The artifact is published and verified independently. + +### 0.1b — vouched in the hub +POST /configuration/artifacts -> 303, then READ BACK from the page (the outcome, not the POST code): + golden currently vouched: 0.245.0 + golden sha now: 7a08aa1ad0bdd622247e1901e422ed2f72df22ef531a144135a44b66fc455626 + agent 0.131.0 and the wrapper sha left exactly as they were. + +DECISION, and the reason, because silence reads as agreement: the GLOBAL controller floor was left at +0.244.0 and NOT raised to 0.245.0. The new box installs golden 0.245.0, which already carries +controller 0.245.0, so the floor is not needed to deliver anything tonight; raising it would have +pushed a controller update onto demo-felhom, a box that is not part of this drill, at 22:00 at night. +The brief asked for a bake and a vouch, not a floor raise. + +### 0.2 — the box +demo-hp (Tier 0). VM 336 „tester1-chaos-night": 8 GiB, 4 cores, cpu host, virtio-scsi-single, +scsi0 = 32 G system disk, scsi1 = 100 G data disk, both on `nvme-scratch` (dir storage, path +/mnt/hdd_1, is_mountpoint yes — the NVMe at its ROOT, per target-selection.md). CD-ROM is the +PUBLISHED felhom-installer-1.28.0-pve9.2-1.iso. **The boot order was set in its own `qm set`** — +combining it with the disk call silently yields `order=net0;ide2` and the VM netboots. + boot: order=ide2;scsi0 (verified by reading `qm config 336` back) +The VM took DHCP 192.168.0.115 and booted the installer's graphical entry. + +### Console driving — measured, not assumed + * Enter on the EULA page: advances. + * Enter on the Location page: lands IN the Country text field and does NOT press Next (the + documented GTK trap). + * **Alt+N works as the „Next" mnemonic** and is what actually drives this installer. Recorded + because the previous drill switched to the text-mode entry to avoid exactly this problem. + * The VM has a QEMU HID Tablet (absolute), so `mouse_move ` takes absolute coordinates and + a real click is available as a fallback. `info mice` says so — checked, not assumed. + * The installer REFUSES the prefilled `mail@example.invalid` with „Please enter a valid Email + address" and simply does not advance. The dialog is above the fold, which is why the cropped + strip looked like „nothing happened" — the full screen showed the reason. + * Keyboard layout is Hungarian (QWERTZ): `y` and `z` are swapped and `@` is AltGr+V. The drill root + password was generated from a-x plus digits so it is identical under either layout, and it is + stored out-of-band (0600, scratchpad) — never in a committed file. + +### Console driving on a Hungarian-layout installer — MEASURED, and it cost four round trips +These are facts about driving a PVE 9.2 graphical installer headlessly through `qm monitor`, and +every one of them was measured on this box tonight rather than recalled: + + * `sendkey alt-n` is the „Next" mnemonic and is what actually advances the installer. Plain `ret` + advances ONLY the EULA page; on every later page it lands inside a text entry and does nothing. + * **`sendkey altgr-v` does NOTHING.** On a Hungarian layout `@` is AltGr+V, and the QEMU key name + that works is **`alt_r`** — `sendkey alt_r-v` types the `@`. The failure is silent: the character + is simply absent, so `admin@felhom.eu` became `adminfelhom.eu` and the installer then refused the + page with „Please enter a valid Email address" — a refusal that looks exactly like „nothing + happened" if you only crop the bottom strip of the screen. + * `mouse_move ` did not move the pointer even though `info mice` reports a QEMU HID Tablet + (absolute) as the active device. Clicking was abandoned; keyboard navigation is the reliable path. + * **Focus is found by MEASURING it, not by counting tabs.** A 20-line script samples the blue border + of each entry box in the screendump and prints which one is focused; Tab is then pressed until + the wanted field reports focus. Counting keystrokes blind is how a password ends up in the wrong + field. (`blueness = mean(B-R)` over the box's top border: focused +58, unfocused 0.) + * The installer REFUSES the prefilled `mail@example.invalid`. + +### Fence check at this point +demo-hp: /mnt/hdd_1 has 876 G free; VM 336's two raw disks are sparse (132 G apparent, 12 K actual). +`/` on demo-hp is at 77 % and was deliberately NOT used for VM disks. Guests 9201 and 9202 untouched +and running (they are the standing demo boxes, and 9202 is the scratch guest the morning restore +will use). No `local-lvm`, no prune, `drill-r50` not touched. + +### A second brief claim, CHECKED rather than inherited: „the automatic mail is waiting in the mailbox" +CONFIRMED. The mailbox holds three „[Felhom] Kösd össze a Felhom dobozodat" messages to +`tester1@felhom.eu` from `monitoring@felhom.eu`, the newest at **2026-09-16T18:17:46Z** — the same +second as yesterday's host delete, which is R-509's automatic trigger firing. So the box installed +tonight should be bindable with **no operator press**, and the pre-declared intervention O1 (the +„Send self-bind link" button) should not be needed. Whether it IS needed is a measurement of this +night, not an assumption: if the waiting link turns out to be superseded or refused, that is a +finding and the press becomes O1. + +The other mail in that mailbox worth noting, because it is the same customer's history and could +confuse a reader of this evidence: two „Új beállító kód — újratelepült a szervered" messages +(10:01:57Z and 16:00:58Z) and one „Beállító kód a jelszavad visszaállításához" (11:12:06Z), all from +2026-09-16 — those belong to yesterday's drills, not to tonight's box. + +NOT recorded here, deliberately: the bind link itself. It carries a one-time token, and a one-time +secret does not go into a committed file (nor was the mail body fetched into the session transcript +for the same reason — the link will be taken from the hub at bind time instead). + +### The one catalog bump — what it is, and why the ORDER matters +Round 7 drew `update nextcloud`, so the bump has to be on the **nextcloud** template, not on whatever +small app I would have picked. The brief asked for "one drill bump of one SMALL app"; the draw +overrides the choice of app, so the compromise is to bump the template's small sidecar pin rather +than the big application image: + + templates/nextcloud/docker-compose.yml:103 redis:7-alpine -> redis:7.4-alpine + +`redis:7.4-alpine` was CHECKED to exist on Docker Hub before being written down — inventing a tag +would have made round 7 fail for the wrong reason, and a round that fails for the wrong reason +measures nothing. `catalog_since` moves with it, per this repo's own rule that any commit changing an +`image:` line must. + +**The bump is NOT pushed yet, and the order is the point:** nextcloud must be DEPLOYED from the +current catalog first, or the box installs 7.4-alpine immediately and round 7 has no update to +apply. Sequence: seed the twelve apps → then push the bump → the box's git-sync picks it up within +15 min (or the „Sablonok frissítése" button) → round 7 at 02:00 has a real update waiting. + +Fence note: the catalog is SHARED with the demo boxes. A bump offers an update; it never applies one +— the guarded update needs a press — so the demo boxes' standing apps are not disturbed by this. + +### Deploy fields, read from the templates rather than guessed +All twelve templates live under `templates//`, not at the repo root (my first lookup used the +wrong layout and returned "NO .felhom.yml" twelve times — recorded because the wrong answer was +uniform and therefore looked authoritative). Only SUBDOMAIN and HDD_PATH are marked `required`; +every secret field is optional in metadata but the server refuses a deploy without them (the +metadata-vs-server contradiction measured 2026-09-16), so the seeding script supplies generated +values for all of them and writes them to a 0600 file that is never echoed. +Apps needing the data drive: nextcloud, immich, paperless-ngx, jellyfin. + +### 0.2b — the install itself +Install pressed 20:07:11Z. **Completion was judged from the DISK, not from the screen** — a PVE +install with "Automatically reboot" ticked re-enters the installer, so a finished install and a stuck +one look identical on screen. The system disk's blocks-used grew 3233 → 6942 MiB and then held +steady across three consecutive 30-second checks: + 20:09:36Z 6942 MiB (stable 0) · 20:10:07Z 6942 (1) · 20:10:37Z 6942 (2) · 20:11:08Z 6942 (3) + -> INSTALL COMPLETE 20:11:08Z + +Then the documented post-install fix, because a RUNNING guest keeps the boot order QEMU started +with — editing the config mid-run is not enough: + qm stop 336 · qm set 336 --delete ide2 · qm set 336 --boot order=scsi0 (its own call) · qm start + read back: `boot: order=scsi0`, no ide2, both disks intact + started from disk 20:11:19Z + +The summary page was checked before Install was pressed, and the line that mattered was +**„Disk(s): /dev/sda"** — the 32 G system disk alone. The 100 G data disk was NOT offered to the +installer and was not touched. That is the check that makes pressing Install safe on a box with a +second disk, and it is the one a two-disk filter has silently got wrong elsewhere in this project. + +### How this box is driven, and how the alarms are read — method facts, each one measured + * **`qm guest exec` is UNUSABLE on this box.** The VM was created with `--agent 1`, but a plain PVE + install does not run `qemu-guest-agent`, and `qm agent 336 ping` answers nothing. My first + first-boot watcher was written against `qm guest exec` and therefore sat silent while reporting + nothing — it looked like a box that would not boot, and it was an instrument pointed at nothing. + * **The access path is SSH to the box** (`root@192.168.0.115`, the drill password via `sshpass -e` + from a 0600 file, never on a command line). Confirmed with `LOGIN_OK`, hostname `chaosnight`, + `pve-manager/9.2.2`. + * **The alarm instrument is the hub's own `events-table`** on `/customers/tester-1` (note: the + customer page is `/customers/`, NOT `/configs/` — the latter redirects). The tabs are + client-side, so `?tab=events` returns the same document; the events must be pulled out of the + table by its container id. The hub pod has **no `sqlite3`** and `hub.db` is 298 MB with a live + `-wal`, so copying the DB is both heavy and stale-prone; the page is the correct instrument. + * **Events baseline, captured before round 1:** the newest pre-drill event is + `Sep 16 18:47 error node_down` (yesterday's box). Anything newer belongs to tonight. Without this + marker, yesterday's `node_stale`/`node_down`/`selfbind_link_sent` rows would be scored as + tonight's alarms. + +### My own mistake, recorded because it cost two watchers +`pkill -f "first-boot watcher"` **matched its own command line** and killed the very background job +it was meant to clear, twice, each exiting 144. Same class as the documented `pgrep -f +qemu-system-x86_64` self-match. The replacement watcher does no pkill at all. + +### Fence note on a secret +While following redirects to find the customer page, `curl -w '%{url_effective}'` printed the hub +password back to me inside the resolved URL. It is not in any file written here, and that format +option is not used again. Recorded rather than quietly dropped, because the next person will hit the +same flag. + +### 0.2c — the bind, as a volunteer does it: ZERO operator presses +The box registered itself as an unclaimed appliance and then sat polling every 30 s — correctly, and +visibly: „not bound yet — polling every 30s until the operator or a customer self-bind lands". No +agent and no guest exist until the bind happens, so a reader who expected the box to finish +installing by itself would have mis-read a waiting box as a stuck one. + +The volunteer's own path was taken, end to end: + * the link came from the **waiting mail** (minted 2026-09-16T18:17:46Z by yesterday's host delete, + R-509's automatic trigger) — not from the operator's „Send self-bind link" button; + * the **pairing code** was read off the box's own console: `4SY-4TX`; + * the **„Tulajdonosi jelmondat"** (5 words) came from the hub's customer record, which is where the + operator hands it from — it is never e-mailed, by design. + + POST /bind/ -> 200 at 2026-09-16T20:18:15Z + „Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt + bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz." + +**O1 (pre-declared) was NOT used and is not counted.** The brief allowed one press of „Send self-bind +link" if no mail was waiting; a mail WAS waiting and it worked, so the bind cost zero interventions. + +Secrets discipline for this step: the bind token, the passphrase and the box's root password each +live in a 0600 scratchpad file and were passed to `curl --data-urlencode name@file`, so no value +reached a command line. None of the three is written into this evidence, and the passphrase's only +description here is its shape (5 words, 35 characters). + +### 0.2d — the box rotates its own root password at day-0, and my access died with it +At 20:18:58Z my SSH to the box still worked; at 20:19:04Z it answered „Permission denied", and the +background watcher lost access in the same window (its 20:19:27Z line came back empty). The bind at +20:18:15Z had started the day-0 install. + +**This is the design, not a defect, and it was confirmed in source rather than guessed:** +`scripts/felhom-host-install.sh` (≈2079-2104) generates a strong `root@pam` password with `openssl +rand`, sets it through `chpasswd` on **stdin** (no argv, no log), and vaults it to the hub with +`PUT /api/v1/hosts//recovery-credential` over the enroll-authenticated channel. The password is +never logged, printed, or written to a file anywhere on the box. So the installer-time password I +typed into the Proxmox installer is dead by intent the moment day-0 runs. + +**What this changes for tonight:** every accident that acts INSIDE the box (docker restart, tunnel +kill, filling the system disk) needs the hub-vaulted break-glass credential +(`/hosts//reveal-recovery-credential`), not the install password. Discovered at 22:19 CEST, +before the rounds began, rather than at 01:10 in the middle of round 5 — which is the only reason it +is a method note here and not an intervention later. + +**How it was diagnosed honestly:** my first instinct was that I had broken my own environment. That +was ruled out first — the password file was still 21 bytes, the variable still 20 characters, and the +same credential had worked six seconds earlier. Only then was the box's own behaviour blamed, and +only after the installer source confirmed the mechanism. + +### 0.2e — THE F-14 PATH, MEASURED LIVE FOR THE FIRST TIME +The brief named this as a claim that had never been measured: „the WG hook provisions by itself after +an acknowledged delete". Tonight it ran, unprompted, and the hub recorded it: + + Sep 16 20:18 appliance_bound (customer_selfbind) + „Az ügyfél saját maga kötötte össze az új eszközt (bare-metal telepítés); a hozzáférést a + doboz a következő lekérdezéskor megkapja." + Sep 16 20:18 appliance_credential_delivered + „Új eszköz (bare-metal telepítés) megkapta a hozzáférést és megkezdi a beállítást." + Sep 16 20:18 claim_reissued_reenroll + „Új beállító kódot küldtünk a szerver újratelepítése után (6. generáció) az ügyfél címére." + Sep 16 20:19 offsite_reissued + „Az offsite (házon kívüli) mentési hozzáférést újra kiadtuk — az új egyszeri jelszót a vezérlő + a következő frissítéskor átveszi." + Sep 16 20:19 **pbsdr_auto_reissue** + „Previous key destroyed (acknowledged deletion) — credentials re-issued automatically." + +That last line is the one that matters. Yesterday's box was removed through the **acknowledged** +delete flow, and tonight's box therefore got its off-site and PBS-DR credentials **with no operator +press at all** — exactly what the F-14 ruling of 2026-07-13 says should happen on that path, and the +half of that ruling nobody had yet watched happen. + +**Consequence for this night's intervention count:** BOTH pre-declared presses are unnecessary. +O1 („Send self-bind link") was not needed because the automatic mail was waiting; O2 („Re-issue PBS +credentials") was not needed because the acknowledged-delete path re-issued by itself. +**Interventions so far: 0.** + +Host enrolled as **`tester-1-022354`**, agent **0.131.0**, ONLINE, guests 0/0 at 20:20Z — the +customer guest is still being created from the golden. + +### 0.2f — day-0 delivered tonight's golden, with no hand upgrade + guest: 9201 „tester-1", running, created by the bootstrap from the golden + controller image: gitea.dooplex.hu/admin/felhom-controller:**0.245.0** — „Up … (healthy)" + agent: felhom-agent **0.131.0** + host: `tester-1-022354`, ONLINE in the hub + +The golden baked at 19:58Z tonight (sha 7a08aa1a…) is what this box installed. Nothing was upgraded +by hand, and the controller the customer will use is the release this drill is validating. That is +the delivery half of the chain: bake -> vouch -> a fresh box lands on it. + + both disks present to the box: `sda` 32 G (system, PVE + LVM) and **`sdb` 100 G** (the data disk, + still unformatted — the household's drive, initialised through the storage page in the next step). + +MY OWN MEASUREMENT ERROR, recorded: the first `lsblk` was piped through `head -12` and stopped one +line short of `sdb`. For a minute the box looked like it had NO data disk — a wrong answer produced +entirely by my own truncation, not by the box. Re-read without the pipe, `sdb 100G` is plainly there +and `qm config 336` still shows `scsi1` attached. An instrument that can cut off its own answer is +not a measurement. + +### the dashboard setup code +The hub mailed a fresh „Új beállító kód — újratelepült a szervered" at **20:18:56Z** (the 6th +generation for this customer), 72 hours valid, delivered to `tester1@felhom.eu`. That is the code the +volunteer types on „A szerver beállítása" to set their own dashboard password — and it arrived by +itself, as part of the same automatic re-enrolment that needed no operator press. + +### 0.2g — the dashboard claimed, by the volunteer, with the mailed code +The code from the 20:18:56Z mail („ősrégen-újraért-címbetű", 3 words, 72 h) was typed into +„A szerver beállítása" together with a 20-character password the household chooses. + + POST /claim -> 302 + POST /login (new pw) -> 302, and a `felhom_session` cookie was issued + +**The second line is the proof; the first is only an attempt.** This repo has a standing trap that +an HTTP 200 (or a redirect) can be a refusal — the claim page re-rendering itself looks exactly like +success from the status code alone. The claim is called successful here because the password it set +then opened a session, which is the consequence a customer actually cares about. + +Both values went in as FILES (`--data-urlencode name@file`, 0600, pushed with `pct push`), so +neither the setup code nor the new password ever reached a command line on the host or in the guest. + +### 0.2h — tonight's release, seen working on a box that installed itself +The storage page of this fresh box carries `
` +— the **R-543 escrow reminder bar shipped in controller v0.245.0 a few hours ago**, rendering on a +box nobody had touched. It is there because the off-site tier was re-issued automatically at 20:19Z +and its escrow is not complete yet, which is precisely the state the bar exists for. + +This is the first time that fix has been seen on a box that was not set up for the purpose of +testing it: the box installed itself from the published ISO, landed on tonight's golden, got its +off-site credentials with no press, and is now telling the household — on every page — that the +remote backup is paused until they create their recovery code. Creating it is the next step of the +guide, and of this drill. + +### the data drive, as the box offers it +`/api/disks/candidates` reports exactly one initialisable device: + /dev/sdb — 107 374 182 400 B (100 GB), QEMU HARDDISK, data_bearing=false, mountable=false +and separately the guest's own system volume as already mounted at /mnt/sys_drive. The empty +100 GB disk is the household's drive and the only thing offered for initialisation — the +data_bearing=false flag is the guard that keeps a drive with someone's files on it out of this list. + +### Schedule: Phase 0 ran long, and the rounds shift with it — recorded, not quietly re-timed +The drawn schedule puts round 1 at 23:30 CEST. Phase 0 will not be finished by then: the box was +installed, bound, claimed and landed on tonight's golden without trouble, but working out how the +storage wizard actually initialises a disk took several rounds of discovery, because the wizard +submits through JavaScript (`POST /api/storage/init`) rather than a form, and I refused to guess the +endpoint after a guessed path cost a 403 and a wrong diagnosis in an earlier drill. + +**What shifts and what does not.** The SCHEDULE — which action, on which app, under which accident, +in which order — is unchanged; it was drawn from the seed before anything ran and is fixed. Only the +wall-clock start moves, and the ~25-minute spacing is kept from the new start. The brief allows a +round to wait provided the wait is recorded; this is that record. The 05:00 stop rule is unchanged, +so a late start means the night may reach fewer than twelve rounds, and the morning verdict will say +how many actually ran rather than implying all twelve did. + +### 0.2i — the household's drive, initialised through the wizard's own endpoint +The wizard submits by JavaScript, not by a form: `POST /api/storage/init` with +`{device, fstype, mount_name, label, set_default, confirmed, durable_id}` and the CSRF meta token, +then polls `GET /api/storage/init/status`. Both were read off the live page and confirmed in +`internal/web/storage_handlers.go:362` before anything was sent. + + POST /api/storage/init -> 200 {"phase":"formatting","started":true} + GET /api/storage/init/status -> phase **done**, started 20:27:21.79Z, updated 20:27:23.95Z + where=/mnt/felhom-drives/hdd_1, error="" reason="" + read back: /dev/sdb is ext4, durable_id `uuid:8f59ed90-c0e9-4e20-8584-d18af823c605` + `df`: /dev/sdb 98G, 2.1M used, 93G free, mounted on /mnt/felhom-drives/hdd_1 + storage page: „hdd_1" labelled „Adatlemez", set as default + +**The POST returning 200 is not the result** — it only says the job started. The format runs as a +background job precisely so a closed tab cannot abort it, so the phase poll is what says it worked, +and the `df` line is what says the household can use it. + +**A known row met tonight, recorded rather than re-filed: R-542.** After the drive is formatted, +registered, mounted and made default, `/api/disks/candidates` STILL lists `/dev/sdb` under +`initialize` — now with `data_bearing:true, mountable:true` and its durable id. The same endpoint +also lists it under `attach`. That is exactly the behaviour R-542 describes (a registered, in-use +drive still offered under „initialize"), seen again on a fresh box. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-recovery.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-recovery.txt new file mode 100644 index 00000000..c307f685 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-recovery.txt @@ -0,0 +1,126 @@ +--- the new disk as the box sees it --- +sda 32G disk +sdb 100G disk +sdc 64G disk +sda 32G disk +sdb 100G disk +sdc 64G disk +--- extend the thin pool onto it --- + pvcreate ok + vgextend ok + WARNING: Set activation/thin_pool_autoextend_threshold below 100 to trigger automatic extension of thin pools before they get full. + Logical volume pve/data successfully resized. +--- reclaim what MY failed pulls left behind (dangling layers only) --- +Total reclaimed space: 0B +--- after --- + LV LSize Data% + data <75.81g 15.57 + root <13.81g + swap <3.88g + vm-9201-disk-0 32.00g 5.24 + vm-9201-disk-1 70.00g 14.47 +TYPE TOTAL ACTIVE SIZE RECLAIMABLE +Images 13 5 3.966GB 2.998GB (75%) +Containers 5 5 54.78kB 0B (0%) + +## Recovery from MY OWN harness damage — what was changed, and why each change is a fixture change +The box was left with a 100 %-full thin pool and nine failed installs. Two fixture changes were made, +both to the drill VM, neither to the product: + + 1. **A third disk (64 G) was attached to the VM and the LVM thin pool extended onto it.** + before: `data` 11.80 g, Data% **100.00** + after: `data` 75.81 g, Data% **15.57** + The 32 G system disk was simply too small for a twelve-app household once thin-provisioning + over-subscribed a 32 G rootfs and a 70 G data volume onto an 11.8 G pool. + + 2. **The customer guest's RAM was raised 4096 -> 6144 MB.** The box has 8 GB and the guest had + half of it; the controller's memory guard counts COMMITTED memory, so twelve apps against a + 3712 MB usable budget cannot fit however they are ordered. + +`docker image prune -f` reclaimed **0 B** — the 3 GB `docker system df` calls "reclaimable" are +layers still referenced by the 13 pulled images, not dangling ones. Recorded because "75 % +reclaimable" reads like free space and is not. + +**These are changes to the drill's own fixture, made in Phase 0 (setup), and they are not part of any +round's measurement.** The re-seed that follows installs the apps ONE AT A TIME, waiting for each to +reach `deployed: true`, which is what a household does and what the earlier parallel burst was not. +=== BEFORE: how is the guest rootfs mounted? === +/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16,emergency_ro) +/dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker type ext4 (rw,relatime,stripe=16,emergency_ro) +=== restart the guest so ext4 remounts clean (the pool now has room) === +=== AFTER: mount state === +/dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16) +=== PROOF: can it actually write? (not an assumption) === + WRITE OK +=== containers back? === +cloudflared Up 45 seconds +felhom-controller Up 44 seconds (healthy) +filebrowser Up 43 seconds (healthy) +privatebin Up 45 seconds (healthy) +traefik Up 45 seconds +=== pool === + data <75.81g 15.64 + vm-9201-disk-0 32.00g 5.24 + vm-9201-disk-1 70.00g 14.54 + +## A filled thin pool wedges the guest READ-ONLY, and adding space does not un-wedge it +Worth writing down beyond tonight, because the second half surprised me. + +When the pool hit 100 %, both of the guest's ext4 filesystems remounted themselves with +**`emergency_ro`** — visible in the mount flags, not only in dmesg: + + BEFORE: + /dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16,**emergency_ro**) + /dev/mapper/pve-vm--9201--disk--1 on /var/lib/docker type ext4 (rw,relatime,stripe=16,**emergency_ro**) + +The symptom this produced was NOT „no space left": every deploy was refused with + „saving app config: writing /opt/docker/stacks//app.yaml.tmp: **read-only file system**" +which reads like a permissions problem and is nothing of the kind. Note the flags still say `rw` — +`emergency_ro` sits beside it, so a careless glance at `mount` says the filesystem is writable. + +**Extending the thin pool from 11.8 G to 75.8 G did not clear it.** The space was there (Data% fell +to 15.6 %) and every write still failed. It took a guest restart for ext4 to mount clean: + + AFTER: /dev/mapper/pve-vm--9201--disk--0 on / type ext4 (rw,relatime,stripe=16) — no emergency_ro + PROOF: `touch /opt/docker/stacks/.rwtest` -> **WRITE OK** (a real write, not an inference from flags) + all five containers back in ~45 s; pool 15.64 % + +The proof line matters: „the flags look right" and „the filesystem accepts a write" are different +claims, and only the second one is the thing that was broken. + +## CORRECTION — the drive gate was NOT stuck. I was reading a stale snapshot and a silent log. +For about four minutes I believed I had found a defect: the data drive was mounted (`df` showed +/dev/sdb, 98 G, on both the box and inside the guest) while the controller still recorded +`"disconnected": true, "stopped_stacks": ["immich","jellyfin","nextcloud","paperless-ngx"]`, and no +`storage_reconnected` event had appeared. I was about to file it. + +**It was self-healing and it healed.** The controller's own DEBUG ring shows: + [gate] drive RETURNED /mnt/felhom-drives/hdd_1 — re-attached + restarted gate-stopped apps + Event pushed: storage_reconnected (info) — Meghajtó újra csatlakoztatva: Adatlemez +and the current state reads: + disconnected=None stopped_stacks=None + hub event, Sep 16 20:44 info storage_reconnected „Meghajtó újra csatlakoztatva: Adatlemez" + +**Why I nearly got it wrong — two instrument faults at once:** + 1. `driveGateLoop` runs on a **30-second ticker**; I read the settings file inside that window and + treated one sample as a settled state. + 2. The gate's lines are **DEBUG**, so `docker logs` showed nothing, and I read that silence as + „the gate never ran". An absent log line is not evidence — the debug ring had the lines all + along (`/api/debug/logs?level=DEBUG`). +And a third, smaller one: my first attempt to read the ring parsed the JSON wrongly and reported +„total ring entries: 0" for a 29 503-byte response, which looked like confirmation of the silence. + +Recorded in full because the wrong version of this paragraph would have been a filed P-row against a +mechanism that works. + +## Tonight's own release, proven through its WHOLE lifecycle on this box (R-543) + while escrow_state=pending : the bar was on every page (measured earlier in Phase 0) + after the ceremony (escrow_state=**escrowed**), the same four pages: + /dashboard escrow-bar-hits=0 + /launcher escrow-bar-hits=0 + /backups/apps escrow-bar-hits=0 + /storage escrow-bar-hits=0 +The bar appeared while the off-site copy was paused, told the household exactly what to do, and +disappeared **for good** when they did it — on a box that installed itself from the published ISO, +with no one setting the scene for the test. „védi"/„védené" are both 0 on /backups/apps for now +because no class-A app is installed yet; that sentence is checked again once they are. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-reseed.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-reseed.txt new file mode 100644 index 00000000..0a2a75ff --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-reseed.txt @@ -0,0 +1,147 @@ +2026-09-16T20:38:31Z login ok (csrf 64) +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:38:31Z bookstack accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/bookstack/app.yaml.tmp: open /opt/docker/stacks/bookstack/app.yaml.tmp: read-only file system +tr: write error: Broken pipe +2026-09-16T20:38:31Z gokapi accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/gokapi/app.yaml.tmp: open /opt/docker/stacks/gokapi/app.yaml.tmp: read-only file system"} +tr: write error: Broken pipe +2026-09-16T20:38:31Z homebox accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/homebox/app.yaml.tmp: open /opt/docker/stacks/homebox/app.yaml.tmp: read-only file system"} +2026-09-16T20:38:32Z uptime-kuma accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/uptime-kuma/app.yaml.tmp: open /opt/docker/stacks/uptime-kuma/app.yaml.tmp: read-only file sy +2026-09-16T20:38:32Z mealie accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/mealie/app.yaml.tmp: open /opt/docker/stacks/mealie/app.yaml.tmp: read-only file system"} +tr: write error: Broken pipe +2026-09-16T20:38:32Z vaultwarden accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/vaultwarden/app.yaml.tmp: open /opt/docker/stacks/vaultwarden/app.yaml.tmp: read-only file sy +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:38:32Z adventurelog accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/adventurelog/app.yaml.tmp: open /opt/docker/stacks/adventurelog/app.yaml.tmp: read-only file +tr: write error: Broken pipe +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:38:32Z nextcloud accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/nextcloud/app.yaml.tmp: open /opt/docker/stacks/nextcloud/app.yaml.tmp: read-only file system +tr: write error: Broken pipe +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:38:32Z paperless-ngx accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/paperless-ngx/app.yaml.tmp: open /opt/docker/stacks/paperless-ngx/app.yaml.tmp: read-only fil +2026-09-16T20:38:32Z jellyfin accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/jellyfin/app.yaml.tmp: open /opt/docker/stacks/jellyfin/app.yaml.tmp: read-only file system"} +tr: write error: Broken pipe +2026-09-16T20:38:32Z immich accept=500 + REFUSED: {"ok":false,"error":"saving app config: writing /opt/docker/stacks/immich/app.yaml.tmp: open /opt/docker/stacks/immich/app.yaml.tmp: read-only file system"} +2026-09-16T20:38:32Z === re-seed finished; final state === + bookstack deployed: false + gokapi deployed: false + homebox deployed: false + immich deployed: false + jellyfin deployed: false + nextcloud deployed: false + paperless-ngx deployed: false + privatebin deployed: true + uptime-kuma deployed: false + vaultwarden deployed: false +cloudflared felhom-controller filebrowser privatebin traefik +curl: option --data-urlencode: error encountered when reading a file +curl: try 'curl --help' or 'curl --manual' for more information +grep: /tmp/.l: No such file or directory +2026-09-16T20:41:05Z login ok (csrf 0) +2026-09-16T20:41:06Z bookstack accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z gokapi accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z homebox accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z uptime-kuma accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z mealie accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z vaultwarden accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z adventurelog accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z nextcloud accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z paperless-ngx accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z jellyfin accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z immich accept=401 + REFUSED: {"ok":false,"error":"authentication required"} +2026-09-16T20:41:06Z === re-seed finished; final state === + bookstack deployed: false + gokapi deployed: false + homebox deployed: false + immich deployed: false + jellyfin deployed: false + nextcloud deployed: false + paperless-ngx deployed: false + privatebin deployed: true + uptime-kuma deployed: false + vaultwarden deployed: false +cloudflared felhom-controller filebrowser privatebin traefik +2026-09-16T20:41:58Z login ok (csrf 64) + +## My THIRD harness error, and the worst-shaped one +The guest restart that un-wedged the filesystem also cleared `/tmp` — where the dashboard password +file lived. The re-seed script's login therefore failed: + + curl: option --data-urlencode: error encountered when reading a file + grep: /tmp/.l: No such file or directory + login ok (csrf **0**) <- it said "login ok" with an empty token + +and every one of the eleven deploys came back: + + accept=401 {"ok":false,"error":"authentication required"} + +**Eleven lines that look exactly like the product refusing eleven installs.** They are nothing of the +kind: the product was right to refuse an unauthenticated caller, and the missing credential was mine. +Had I skimmed this output I would have written up a spectacular false finding — „the box refuses +every deploy after a restart" — and it would have been entirely an artefact of my own tooling. + +Three of my errors tonight share one shape: **a script that keeps going after its own precondition +failed, and prints a confident line anyway** („all twelve deploys ACCEPTED", „login ok (csrf 0)", +and the `head -12` that hid a disk). The fix applied here is the one that should have been there from +the first line: the script now ABORTS when the session token is empty, and says the fault is mine +rather than reporting deploys as refused. +2026-09-16T20:41:58Z bookstack accept=202 +2026-09-16T20:42:49Z bookstack INSTALLED (containers matching: 2) +2026-09-16T20:42:49Z gokapi accept=202 +2026-09-16T20:42:59Z gokapi INSTALLED (containers matching: 1) +2026-09-16T20:42:59Z homebox accept=202 +2026-09-16T20:43:09Z homebox INSTALLED (containers matching: 1) +2026-09-16T20:43:09Z uptime-kuma accept=202 +2026-09-16T20:44:00Z uptime-kuma INSTALLED (containers matching: 1) +2026-09-16T20:44:00Z mealie accept=202 +2026-09-16T20:44:50Z mealie INSTALLED (containers matching: 1) +2026-09-16T20:44:50Z vaultwarden accept=202 +2026-09-16T20:45:01Z vaultwarden INSTALLED (containers matching: 1) +2026-09-16T20:45:01Z adventurelog accept=202 +2026-09-16T20:46:01Z adventurelog INSTALLED (containers matching: 3) +2026-09-16T20:46:01Z nextcloud accept=202 +2026-09-16T20:46:11Z nextcloud INSTALLED (containers matching: 3) +2026-09-16T20:46:12Z paperless-ngx accept=202 +2026-09-16T20:46:32Z paperless-ngx INSTALLED (containers matching: 0) +2026-09-16T20:46:32Z jellyfin accept=202 +2026-09-16T20:46:42Z jellyfin INSTALLED (containers matching: 1) +2026-09-16T20:46:42Z immich accept=202 +2026-09-16T20:46:52Z immich INSTALLED (containers matching: 4) +2026-09-16T20:46:52Z === re-seed finished; final state === + adventurelog deployed: true + bookstack deployed: true + gokapi deployed: true + homebox deployed: true + immich deployed: true + jellyfin deployed: true + mealie deployed: true + nextcloud deployed: true + paperless-ngx deployed: true + privatebin deployed: true + uptime-kuma deployed: true + vaultwarden deployed: true +adventurelog adventurelog-frontend adventurelog-postgres bookstack bookstack-db cloudflared felhom-controller filebrowser gokapi homebox immich-machine-learning immich-postgres immich-redis immich-server jellyfin mealie nextcloud nextcloud-db nextcloud-redis paperless-postgres paperless-redis paperless-webserver privatebin traefik uptime-kuma vaultwarden diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase0-seed.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-seed.txt new file mode 100644 index 00000000..f8d7ae53 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase0-seed.txt @@ -0,0 +1,169 @@ +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). +login ok, csrf len 64 +tr: write error: Broken pipe +tr: write error: Broken pipe +tr: write error: Broken pipe +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:30:18Z nextcloud -> 202 {"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"} +tr: write error: Broken pipe +2026-09-16T20:30:18Z immich -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:30:18Z bookstack -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +2026-09-16T20:30:19Z privatebin -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +2026-09-16T20:30:19Z gokapi -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +tr: write error: Broken pipe +2026-09-16T20:30:19Z vaultwarden -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:30:19Z paperless-ngx -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +2026-09-16T20:30:19Z jellyfin -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +2026-09-16T20:30:19Z mealie -> 400 {"ok":false,"error":"Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 200 MB, Elérhető: 136 MB (öss +2026-09-16T20:30:19Z uptime-kuma -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +tr: write error: Broken pipe +tr: write error: Broken pipe +2026-09-16T20:30:19Z adventurelog -> 400 {"ok":false,"error":"Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 100 MB, Elérhető: 86 MB (össz +tr: write error: Broken pipe +2026-09-16T20:30:20Z homebox -> 202 {"ok":true,"data":{"warning":"Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normá +all twelve deploys ACCEPTED (202 = accepted, not installed — R-536: the two are different things) +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = "UTF-8", + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to a fallback locale ("en_US.UTF-8"). + +## CORRECTION — the line my own script printed is FALSE, and it is corrected here before anything else +My seeding script ended with „all twelve deploys ACCEPTED". **Ten were accepted; two were refused.** +The sentence was a fixed `echo` at the end of the script, printed without consulting a single result +— the exact "a script that announces a conclusion it never checked" failure this project keeps +re-learning, and it would have put a false line into tonight's record. + + ACCEPTED (202), 10: nextcloud · immich · bookstack · privatebin · gokapi · vaultwarden · + paperless-ngx · jellyfin · uptime-kuma · homebox + REFUSED (400), 2: + mealie „Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 200 MB, Elérhető: 136 MB…" + adventurelog „Nincs elég memória az alkalmazás telepítéséhez. Szükséges: 100 MB, Elérhető: 86 MB…" + +Note also that nine of the ten acceptances carried a warning of their own: + „Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát…" +so the box was telling the truth about its memory on the way up, and then refused the last two +outright. The refusal is the controller's memory guard doing its job, fail-closed and in plain +Hungarian with the numbers in it. + +**Consequence for the drawn schedule, stated now rather than discovered at 23:30:** rounds 1 and 4 +act on `adventurelog` and `mealie`, and round 6 on `adventurelog` — apps that are NOT installed. +The schedule was drawn before the night and is not being re-drawn; what changes is that those rounds +must either act on an app that exists, or be recorded as unrunnable. That decision is made and +recorded explicitly, not silently. + +Also true and worth keeping: „202 = accepted, not installed" (R-536). How many of the ten actually +finished installing is a separate measurement, taken next. + +## The memory facts behind the two refusals — measured + the VM (the whole box): 7939 MB total, 5814 MB available + the CUSTOMER GUEST (LXC 9201): **memory: 4096, swap: 512, cores: 3** + inside the guest at the moment of the refusals: 4096 MB total, ~3685 MB available + containers actually running then: 4 (cloudflared, felhom-controller, filebrowser, traefik) + +So the box had ~5.8 GB free while the guest the apps live in was capped at 4 GB — and the guard +refused the eleventh and twelfth apps against the guest's cap, counting the memory already COMMITTED +by ten in-flight installs rather than the memory currently in use. That is the honest way to count +it (otherwise ten simultaneous pulls would all be admitted and then fight), and the message quoted +the two numbers it compared. + +**This is the constraint that decides whether tonight's household can be twelve apps at all.** The +brief asks for the twelve of BIGNIGHT. The box as installed gives its customer guest 4 GB. Nothing +has been changed yet: first the ten in-flight installs are allowed to finish, because the memory +picture during a pull is not the memory picture afterwards, and a decision taken on the wrong +picture is worse than a late one. + +## DECISION — what happens to the rounds that name apps which are not installed +Three rounds name apps the box refused to install: round 1 (`offsite-run adventurelog`), round 4 +(`offsite-run mealie`), round 6 (`backup-system adventurelog`). + +**The schedule is not re-drawn.** It was fixed from the seed before anything ran, and re-drawing it +now — after seeing which apps happened to fit in memory — is exactly the "choose the night after the +fact" failure the seed exists to prevent. + +What the rounds actually do, and why this costs less than it looks: + + * **`offsite-run` is a TIER action, not an app action.** The off-site leg is repo-global — one + `LastRun` for the whole repository, no per-app run time — so rounds 1 and 4 exercise the tier + exactly as drawn. The named app is which app's row I read afterwards; where that app is absent, + the round records the tier's own result and says the app was not installed. + * **`backup-system` (round 6) is a whole-box action** and does not depend on the named app either. + +So all three rounds run as drawn; what changes is that their "what the customer saw" cell reports the +tier or the whole-system page rather than that app's row. Each affected round says so in its own line +rather than leaving a reader to assume the app was there. + +**And the memory refusal is itself a finding, not just an inconvenience:** a fresh box built to the +documented shape gives its customer guest 4 GB, and the twelve-app household of BIGNIGHT does not fit +in it. Nothing was resized to make the drill comfortable. + +## Deploy progress, and a second instrument lesson +At 20:32:59Z, twelve minutes after the ten deploys were accepted: + app.yaml recorded: 10 (the box registered all ten) + containers running: 5 (cloudflared, felhom-controller, filebrowser, traefik, **privatebin**) + guest memory available: 3818 MB +So exactly one of the ten household apps was actually up; the rest were still pulling images. This is +R-536's distinction in the flesh: ten "telepítés elindítva" acceptances, one installed app. + +**The instrument lesson (my second tonight):** I read the deploy watcher with +`tail -6 … | grep -vE "locale|…"`, and the locale warnings filled the whole tail, so the watcher's +ONE real line was filtered out and the watcher looked dead. I had already declared one watcher dead +tonight for a different reason and killed it with a `pkill` that killed itself. The fix both times +was the same: **ask the box directly instead of trusting my own reporting layer.** The box answered +in one call, and the watcher turned out to have been working the whole time. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/postinstall_fix.sh b/documentation/audits/evidence-chaos-night-2026-09-17/postinstall_fix.sh new file mode 100755 index 00000000..ec2f9a75 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/postinstall_fix.sh @@ -0,0 +1,17 @@ +#!/bin/bash +# Run on demo-hp AFTER the PVE install finishes. +# +# Why this exists: "Automatically reboot after successful installation" is ticked and the boot order +# is ide2;scsi0, so the box reboots straight back INTO the installer — and a completed install then +# looks exactly like a stuck one. A running guest also keeps the QEMU boot order it started with, so +# editing the config mid-run is not enough: the VM must be stopped. +# +# Completion is judged from the DISK, not from the screen. +set -u +V=336 +echo "disk usage before: $(du -sh --block-size=1M /mnt/hdd_1/images/$V/vm-336-disk-1.raw 2>/dev/null | cut -f1) MiB" +qm stop $V; sleep 5 +qm set $V --delete ide2 +qm set $V --boot order=scsi0 # its OWN call, always +qm config $V | grep -E '^(boot|ide2|scsi)' +qm start $V && echo "started from disk at $(date -u +%FT%TZ)" diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round.sh b/documentation/audits/evidence-chaos-night-2026-09-17/round.sh new file mode 100755 index 00000000..24b75b29 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round.sh @@ -0,0 +1,36 @@ +#!/bin/bash +# round.sh — run ONE chaos round and record the same five things. +# Usage: round.sh +set -u +N="$1"; X="$2"; Y="$3"; Z="$4" +E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17 +OUT="$E/round-${N}.txt" +SC=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/adc14dfe-fc6c-4014-9378-d580d29d3595/scratchpad/chaos +say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; } +export SSHPASS=$(cat $SC/boxroot.pw) +SSHO="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o ConnectTimeout=8 -o NumberOfPasswordPrompts=1" +BOX(){ sshpass -e ssh $SSHO root@192.168.0.115 "$@" 2>/dev/null | grep -v "Warning: Permanently"; } + +say "================ ROUND $N : $X on $Y, while: $Z ================" +say "--- BEFORE: is the box steady? (every app up, hub ONLINE, no active alarm) ---" +BOX 'pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}"' | sed 's/^/ /' | tee -a "$OUT" +say " events before this round:" +bash "$E/events.sh" 3 | tee -a "$OUT" + +say "--- ACTION: $X on $Y ---" +# the action itself is driven per-round by the caller's follow-up; this records the start moment +say " action start marker" + +say "--- ACCIDENT: $Z (injected 10-60 s after the action starts) ---" +if [ "$Z" != "nothing" ]; then + bash "$E/inject.sh" "$Z" "$N" | sed 's/^/ /' | tee -a "$OUT" +else + say " control round — no accident, deliberately" +fi + +say "--- AFTER: what the box did by itself ---" +BOX 'pct exec 9201 -- docker ps --format "{{.Names}} {{.Status}}"' | sed 's/^/ /' | tee -a "$OUT" +BOX 'pvesm status; df -h /mnt/felhom-drives/hdd_1 2>/dev/null | tail -1' | sed 's/^/ /' | tee -a "$OUT" +say "--- alarms in this round's window ---" +bash "$E/events.sh" 8 | tee -a "$OUT" +say "================ END ROUND $N ================" diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/round_runner.md b/documentation/audits/evidence-chaos-night-2026-09-17/round_runner.md new file mode 100644 index 00000000..9fdd88dd --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/round_runner.md @@ -0,0 +1,37 @@ +# The per-round record — CHAOS NIGHT + +Every round records the SAME five things, in this order, and nothing is written from memory: + +1. **what the customer saw** — screens quoted verbatim (Hungarian in „ ", searched with ASCII + fragments, with a positive and a negative control; accented `grep` dies with a complexity error + here, so searching is done in Python) +2. **what the box did by itself** — no shell, no help; anything I had to do is an intervention +3. **time to steady** — measured from the accident to the moment every app is running, the hub says + ONLINE and no alarm is active. **`—` when the box never got there on its own**, never a guess +4. **which alarm fired, and was it TRUE** +5. **which alarm SHOULD have fired (per 08-alarm-ladder.md) and did not** + +plus **the background loop's failures inside the round's window**, counted from its own log. + +## Two things known BEFORE the night that change how rounds are scored + +- **Rounds 7, 8 and 9 all block the box's network.** An event generated while the hub is unreachable + is retried 3 times over ~6 s and then **dropped permanently** (`PushEvent`, no queue). So a missing + alarm in those rounds is not evidence the alarm did not fire — it may have been posted into a + blocked path. Scored as `LOST-IN-BLOCK`, never as `MISSED`. +- **The dedupe is not one window.** 5 minutes applies ONLY to the node/host liveness events; the + default operator cooldown is **1 hour**, keyed `customer:type[...]`. A second identical alarm + inside an hour is suppressed and written to `notification_log` with status `suppressed` — so + "suppressed" and "never fired" are distinguishable, and must be distinguished. + +## Steady-state check (the same command set every round) + + * every app: `docker ps` shows it running AND its front door answers + * the hub: the host row reads ONLINE, guests 1/1, agent version present + * no active alarm on the dashboard; `notification_log` read for the round's window + * the data drive: the storage page reads „Aktív", not „Leválasztva" + +## Evidence discipline + +Evidence is copied off the box **at the end of each round, before the next accident** (R-320) — the +one that gets forgotten is the middle one, never the last. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/schedule.txt b/documentation/audits/evidence-chaos-night-2026-09-17/schedule.txt new file mode 100644 index 00000000..bc84ddb4 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/schedule.txt @@ -0,0 +1,23 @@ +seed: 20260917 +script sha256: 4b98afe65d042df7e7dc417553b33565cfbb4afd451c68456a7cabec0858d2a1 + +| # | time | X — the action | Y — the app | Z — the accident | +|---|---|---|---|---| +| 1 | 23:30 | offsite-run | adventurelog | nothing | +| 2 | 23:55 | restore | gokapi | power cut | +| 3 | 00:20 | use | bookstack | disk 95% full | +| 4 | 00:45 | offsite-run | mealie | tunnel down 10min | +| 5 | 01:10 | use | privatebin | docker restarted | +| 6 | 01:35 | backup-system | adventurelog | nothing | +| 7 | 02:00 | update | nextcloud | internet gone 10min | +| 8 | 02:25 | backup-app | nextcloud | internet gone 10min | +| 9 | 02:50 | use | uptime-kuma | internet gone 10min | +| 10 | 03:15 | restore | uptime-kuma | hard reset | +| 11 | 03:40 | use | paperless-ngx | drive pulled 20min | +| 12 | 04:05 | use | paperless-ngx | nothing | + +re-draw log (4 entries): + r02 X=reinstall re-drawn (nothing has been removed yet) + r06 Z=disk 95% full re-drawn (constraint 4: at most once) + r07 Z=nothing re-drawn (constraint 6: never two in a row after r2) + r08 X=reinstall re-drawn (nothing has been removed yet) diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-01.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-01.png new file mode 100644 index 00000000..df3aed6e Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-01.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-02.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-02.png new file mode 100644 index 00000000..9b0387f3 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-02.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-03.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-03.png new file mode 100644 index 00000000..b69e8537 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-03.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-04.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-04.png new file mode 100644 index 00000000..a828abc8 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-04.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-05.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-05.png new file mode 100644 index 00000000..a828abc8 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-05.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-06-strip.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-06-strip.png new file mode 100644 index 00000000..0403ee4c Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-06-strip.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-06.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-06.png new file mode 100644 index 00000000..6e13c77d Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-06.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-07-strip.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-07-strip.png new file mode 100644 index 00000000..3657ff19 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-07-strip.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-07.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-07.png new file mode 100644 index 00000000..1cfa9136 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-07.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-08-strip.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-08-strip.png new file mode 100644 index 00000000..3657ff19 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-08-strip.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-08.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-08.png new file mode 100644 index 00000000..96091ac9 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-08.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-09.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-09.png new file mode 100644 index 00000000..70a5a147 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-09.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-10-summary.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-10-summary.png new file mode 100644 index 00000000..1dc5fbd3 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-10-summary.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-11-installing.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-11-installing.png new file mode 100644 index 00000000..26ca91ba Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-boot-11-installing.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-email-alt_r.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-email-alt_r.png new file mode 100644 index 00000000..3a5d3b75 Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-email-alt_r.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-email-typed.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-email-typed.png new file mode 100644 index 00000000..f4564d7b Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-email-typed.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-hostname.png b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-hostname.png new file mode 100644 index 00000000..4f8e958e Binary files /dev/null and b/documentation/audits/evidence-chaos-night-2026-09-17/screens/p0-hostname.png differ diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/seed_apps.sh b/documentation/audits/evidence-chaos-night-2026-09-17/seed_apps.sh new file mode 100755 index 00000000..7630b96d --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/seed_apps.sh @@ -0,0 +1,47 @@ +#!/bin/bash +# seed_apps.sh — CHAOS NIGHT Phase 0.3: move the household in. +# Twelve apps, deployed through the controller's own API (the endpoint the UI invokes), each with +# exactly the fields its template declares. Secrets are GENERATED here and written to a 0600 file; +# none is ever echoed. Run INSIDE the guest (the controller answers on the container address). +# +# Usage: seed_apps.sh +set -u +B="http://${1:?ctrl ip}:8080"; H="Host: ${2:?host header}"; PWF="${3:?pw file}" +HDD="${4:?hdd path}"; SEC="${5:?secrets out}" +: > "$SEC"; chmod 600 "$SEC" +gen(){ tr -dc 'a-zA-Z0-9' + local app="$1" vals="$2" + local code + code=$(curl -s -o /tmp/.d -w '%{http_code}' -H "$H" -H "$C" -H "X-CSRF-Token: $TOK" \ + -H 'Content-Type: application/json' -X POST -d "{\"values\":$vals}" \ + "$B/api/stacks/$app/deploy") + printf '%s %-14s -> %s %s\n' "$(date -u +%FT%TZ)" "$app" "$code" "$(head -c 120 /tmp/.d)" +} + +D='"DOMAIN":"enkicsifelhom.hu"' +NC_ADMIN=$(gen 20); PL_ADMIN=$(gen 20); GK=$(gen 20) +{ echo "nextcloud_admin=$NC_ADMIN"; echo "paperless_admin=$PL_ADMIN"; echo "gokapi_admin=$GK"; } >> "$SEC" + +dep nextcloud "{$D,\"SUBDOMAIN\":\"cloud\",\"DB_PASSWORD\":\"$(gen)\",\"MYSQL_ROOT_PASSWORD\":\"$(gen)\",\"NEXTCLOUD_ADMIN_USER\":\"admin\",\"NEXTCLOUD_ADMIN_PASSWORD\":\"$NC_ADMIN\",\"HDD_PATH\":\"$HDD\"}" +dep immich "{$D,\"SUBDOMAIN\":\"photos\",\"DB_PASSWORD\":\"$(gen)\",\"HDD_PATH\":\"$HDD\"}" +dep bookstack "{$D,\"SUBDOMAIN\":\"wiki\",\"APP_KEY\":\"base64:$(gen 32)\",\"DB_PASSWORD\":\"$(gen)\"}" +dep privatebin "{$D,\"SUBDOMAIN\":\"paste\"}" +dep gokapi "{$D,\"SUBDOMAIN\":\"share\",\"GOKAPI_PASSWORD\":\"$GK\"}" +dep vaultwarden "{$D,\"SUBDOMAIN\":\"vault\",\"ADMIN_TOKEN\":\"$(gen 32)\",\"SIGNUPS_ALLOWED\":\"false\"}" +dep paperless-ngx "{$D,\"SUBDOMAIN\":\"paperless\",\"DB_PASSWORD\":\"$(gen)\",\"PAPERLESS_SECRET_KEY\":\"$(gen 32)\",\"PAPERLESS_ADMIN_USER\":\"admin\",\"PAPERLESS_ADMIN_PASSWORD\":\"$PL_ADMIN\",\"HDD_PATH\":\"$HDD\",\"PAPERLESS_OCR_LANGUAGE\":\"hun+eng\"}" +dep jellyfin "{$D,\"SUBDOMAIN\":\"media\",\"HDD_PATH\":\"$HDD\"}" +dep mealie "{$D,\"SUBDOMAIN\":\"recipes\"}" +dep uptime-kuma "{$D,\"SUBDOMAIN\":\"status\"}" +dep adventurelog "{$D,\"SUBDOMAIN\":\"travel\",\"SECRET_KEY\":\"$(gen 32)\",\"DB_PASSWORD\":\"$(gen)\"}" +dep homebox "{$D,\"SUBDOMAIN\":\"inventory\",\"HBOX_AUTH_API_KEY_PEPPER\":\"$(gen 32)\"}" +rm -f /tmp/.l /tmp/.d +echo "all twelve deploys ACCEPTED (202 = accepted, not installed — R-536: the two are different things)" diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/steady.sh b/documentation/audits/evidence-chaos-night-2026-09-17/steady.sh new file mode 100755 index 00000000..2c90a4e6 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/steady.sh @@ -0,0 +1,45 @@ +#!/bin/bash +# steady.sh — is the box back on its own? Run between every round. +# +# "Steady" is not a word here, it is a measurement: every app running, the hub says ONLINE, the data +# drive reads Aktiv, and no alarm is active. If the box did not get there BY ITSELF, the round's +# steady-state cell is `—`, never a guess and never a number I helped it reach. +# +# Usage: steady.sh (reads the box address from $BOXIP, guest id from $GUEST) +set -u +R="${1:?round label}" +VM=336 +E=/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/evidence-chaos-night-2026-09-17 +OUT="$E/round-${R}-steady.txt" +H(){ ssh hp "$@" 2>/dev/null | grep -vE "locale|LC_|LANG|perl:|supported and installed|are supported"; } +say(){ echo "$(date -u +%FT%TZ) $*" | tee -a "$OUT"; } + +say "=== steady check, round $R ===" +say "--- the box itself (is the VM even up?) ---" +H "qm status $VM" | tee -a "$OUT" + +say "--- the customer guest and its apps (asked of the box, not of me) ---" +H "qm guest exec $VM -- bash -c 'pct list; echo ---; pct exec 9000 -- docker ps --format \"{{.Names}} {{.Status}}\"'" 2>/dev/null | tee -a "$OUT" + +say "--- the hub's view: host row, guests, agent, last report ---" +cd /mnt/5_hdd/felhom.eu/git/felhom.eu +python3 scripts/read_credential.py HUB_PW /tmp/.shp >/dev/null 2>&1 +HUB_PW=$(cat /tmp/.shp); IP=$(sudo kubectl -n felhom-system get svc hub -o jsonpath='{.spec.clusterIP}') +curl -s -u ":$HUB_PW" "http://$IP:8080/hosts" | python3 -c " +import sys,re,html +t=sys.stdin.read() +txt=html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>',' ',re.sub(r'','',t,flags=re.S)))) +i=txt.find('chaosnight') +print(' ', txt[max(0,i-40):i+220] if i>=0 else 'chaosnight NOT in the hosts table') +" | tee -a "$OUT" + +say "--- alarms/events in the last 30 minutes (fired AND suppressed are different things) ---" +curl -s -u ":$HUB_PW" "http://$IP:8080/" | python3 -c " +import sys,re,html +t=sys.stdin.read() +txt=html.unescape(re.sub(r'\s+',' ',re.sub(r'<[^>]+>',' ',re.sub(r'','',t,flags=re.S)))) +m=re.search(r'(Recent events|Events|Legutobbi).{0,900}', txt) +print(' ', (m.group(0)[:900] if m else txt[:400])) +" | tee -a "$OUT" +rm -f /tmp/.shp +say "=== end steady check, round $R ===" diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 72499f81..b64d7633 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -728,6 +728,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-543** | **[P1-HIGH] Off-site ON by default is not off-site WORKING: on a fresh box tier 3 sits at „Kulcsletétre vár" until the household does the escrow ceremony, and nothing asks them to — while the tier-1 row now tells them their files are protected by that very copy.** MEASURED 2026-09-16 on the fresh box (controller 0.244.0, hub 0.116.0, off-site provisioned automatically by the new default): the app-backup page reads „3. mentés — Kulcsletétre vár · A távoli mentés a titkosítási kulcs letétbe helyezéséig szünetel", the remote page reads „Helyreállítási kód szükséges", and `POST /backup/offbox/run` returns 302 while producing no snapshot (the controller log shows only `offsite-credential-retry`, no restic activity). **Why it matters more than before today:** hub v0.116.0 makes off-site the default *because* a one-drive box otherwise keeps the household's files in no tier at all (R-537/R-538), and controller v0.244.0 now prints „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi" under the tier-1 row. On day one both are true-in-intent and false-in-fact: the copy is paused. **Fix shape (one of):** prompt the escrow ceremony as part of first-run when off-site is enabled and un-escrowed; and/or make the tier-1 sentence state the tier's actual state („…védené — a távoli mentés a helyreállítási kód létrehozásáig szünetel"). The ceremony itself works and is customer-facing („Helyreállítási kód létrehozása"); what is missing is that anyone is told to do it. **CLOSED 2026-09-16 — controller v0.245.0, both halves proven live.** The pause is untouched: it is the zero-knowledge escrow design, and this row was never about the mechanism. (a) **The household is asked:** while the off-site tier is configured and its escrow is not complete, every authenticated page carries „A távoli mentés szünetel, amíg nem hozod létre a helyreállítási kódot." linking `/backup/escrow`. It is the R-241 bar, second instance — same session-cookie dismissal, back at the next visit, gone for good when escrowed; **no second banner system**. It hangs off `executeTemplate`, the single render choke point, so it cannot reach only the pages someone remembered. (b) **The tier-1 sentence renders by state:** `driveFilesNoteFor` takes `tier3State`'s own vocabulary — `active` → „védi", `escrow_pending` → „védené … a helyreállítási kód létrehozásáig szünetel" + the route, no off-site and no second drive → „nincs másolat" + both ways out. **Measured live on 0.245.0:** on a paused box (9202, off-site configured through the product's own endpoint, `escrow_state=pending`) the bar renders on /dashboard, /launcher, /backups/apps and /settings; a manual `POST /backup/offbox/run` is refused by the fork-4 gate („A távoli mentés a kulcs letétbe helyezésére vár.") with `last_run=None, snapshot_count=None`; and a throwaway class-A app's row reads „…védené … szünetel" with „védi"=0. On an escrowed box (9201) the bar is absent on all three pages and the row reads „védi". Both red-proofed (the bar test fails on BOTH pages with the one hook line removed; the sentence test quotes the exact v0.244.0 promise when the state is ignored). The first-hour guide now asks for the code right after the dashboard password and before the first app. Evidence: `audits/evidence-recovery-code-2026-09-16/`. | **CLOSED 2026-09-16 — shipped in controller v0.245.0 and proven live** | | **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** | | **R-545** | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** | +| **R-546** | **[P2-MEDIUM] The first-hour guide sends the household to create their recovery code at a moment when the box cannot yet do it — and the new reminder bar urges them there on every page.** MEASURED 2026-09-16/17 on a fresh box (`tester-1-022354`, guest 9201, controller 0.245.0, agent 0.131.0, installed from the published ISO 1.28.0). `VOLUNTEER-first-hour.md` §6 — added hours earlier in controller v0.245.0 — places „A helyreállítási kód" immediately after the dashboard password and **before the first app**, because until it is done the off-site copy does not run. At exactly that point the ceremony FAILS: `POST /api/escrow/start` → 200, then `GET /api/escrow/status` → `detail: "exit 2: … selftest=escrow-create requires -storage (or escrow.pbs_storage…"`, and `POST /api/escrow/claim` → **409** „A folyamat jelenlegi állapotában a kód nem kérhető le." **Cause, measured on both sides:** the hub had auto-provisioned the DR descriptor at 20:19 (no press — see R-534/R-511's acknowledged-delete path) and its Backup & DR panel itself read „descriptor provisioned … **waiting** · ceremony possible once the descriptor is applied on the box"; the box had no PBS storage (`pvesm status` = local + local-lvm only) and `/etc/felhom-agent/agent.json` had **no `escrow` section at all**. **It is a TIMING gap and it self-heals:** a watcher left the box alone and polled — `pbs_storage` and `escrow.pbs_storage_id` both became `felhom-pbs` at **20:35:16Z, ~17 minutes after the bind**; the retried ceremony then passed every preflight item and the claim returned 200 (83-character code, entropy 129.2 bits), and `escrow_state` flipped to `escrowed`. **Why it still matters:** for those ~17 minutes the R-543 reminder bar (also v0.245.0) is on *every* page telling the household to do the one thing that refuses, and nothing on the page says „wait a few minutes" — the volunteer meets a stderr fragment about a `-storage` flag. **Fix shape (one of):** have the escrow page/bar consult `preflight` and say „a doboz még készül — pár perc múlva próbáld újra" while `pbs_storage_id` is unset; or move the guide's step to after the first app; or make the bar appear only once preflight is green. **No product code was changed tonight** (validation run). Evidence: `audits/evidence-chaos-night-2026-09-17/phase0-escrow-failure.txt` and `phase0-escrow-retry.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller copy + guide timing)** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |