From 9b44c44f236679745bfefd3357194dd78bbb8542 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 17 Sep 2026 02:33:19 +0200 Subject: [PATCH] chaos night: teardown baseline, and the box's own logs copied off before anything stops The before-picture that cannot be retaken once the machine is gone: pvesm status, the guest list, VM 336's full config and the contents of /mnt/hdd_1. Two things it records that correct my own assumptions: * the machine has THREE disks, not the two the brief specified. The third is the 64G disk I added during Phase 0 to extend the thin pool after filling it with twelve simultaneous deploys. My damage, my remedy, and a deviation from the fixture the brief described - declared rather than quietly torn down. * the /mnt/hdd_1 claim is now earned: nvme-scratch is defined with path /mnt/hdd_1, is_mountpoint yes, and the three raw files sit in /mnt/hdd_1/images/336. The harness is stopped and disabled, its logs copied off first (R-320): household 204 lines, diskguard 0 bytes - the guard never fired all night. And a correction one minute old: I announced that the earlier log copy was twelve lines short and that re-copying rescued them. It was not short - both copies are byte-identical. I compared a line count read at 00:19 against a copy taken at 00:31. Nothing was lost; only the accuracy of the record was at risk. unproven.py: 35 of 55 not walked - NO NUMBER MOVED, which is correct for a validation night that shipped no product code. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../box-diskguard.log | 0 .../box-hloop.sh | 17 ++++ .../box-household.log | Bin 0 -> 10465 bytes .../box-units.txt | 24 ++++++ .../teardown-baseline.txt | 75 ++++++++++++++++++ 5 files changed, 116 insertions(+) create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/box-diskguard.log create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/box-hloop.sh create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/box-household.log create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/box-units.txt create mode 100644 documentation/audits/evidence-chaos-night-2026-09-17/teardown-baseline.txt diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/box-diskguard.log b/documentation/audits/evidence-chaos-night-2026-09-17/box-diskguard.log new file mode 100644 index 00000000..e69de29b diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/box-hloop.sh b/documentation/audits/evidence-chaos-night-2026-09-17/box-hloop.sh new file mode 100644 index 00000000..ac006f7f --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/box-hloop.sh @@ -0,0 +1,17 @@ +#!/bin/bash +# The light background household. Runs ON THE VM (the nested PVE), hitting the customer guest's +# traefik, because the guest is not reachable from DooPlex at all. Every 2 minutes: one read of a +# random app's front door and one read of the dashboard's health endpoint. Failures are DATA. +G="${1:?guest ip}" +LOG=/root/household.log +APPS="wiki paste share recipes inventory status travel paperless cloud photos media vault" +while true; do + A=$(echo $APPS | tr ' ' '\n' | shuf -n1) + C=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 -H "Host: ${A}.enkicsifelhom.hu" "http://${G}/" 2>/dev/null) + case "$C" in 2*|3*) R="ok http=$C";; 000) R="UNREACHABLE";; *) R="FAILED http=$C";; esac + printf '%s %-10s read %s\n' "$(date -u +%FT%TZ)" "$A" "$R" >> $LOG + H=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 -H "Host: felhom.enkicsifelhom.hu" "http://${G}/api/health" 2>/dev/null) + case "$H" in 2*|3*) R="ok http=$H";; 000) R="UNREACHABLE";; *) R="FAILED http=$H";; esac + printf '%s %-10s dash %s\n' "$(date -u +%FT%TZ)" "$A" "$R" >> $LOG + sleep 120 +done diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/box-household.log b/documentation/audits/evidence-chaos-night-2026-09-17/box-household.log new file mode 100644 index 0000000000000000000000000000000000000000..797764724fc7a16ae9befdbc8c6ba8282f80bb1d GIT binary patch literal 10465 zcmchd&1xe@6os?SQxtfS1Y%JC$#%mmoQX^X#svRB@UCt3pk`#X=x)iGT^=G&m?xQA zb#18Dv2?07(bD+S*Hx$P)$NwjNjg29T%0D;`!p#h(=wYpny!9sHg#|0)z#I)nD(n# z4#VbCHc5U=39+d9r7>5R^Xu8~2dQ)Cd2PNnRw6Iz5h8gs9ueY6CXQ;>yLvse-8b$* z;$zRr$%*-VeR+2`znaf(&DH!bS?zdhlB7(F$I{fhTFk3vU7PBu-PLATHobXnx9g!^ zn99he%JLVD?D}A;UjAa0Pi@sL+%`Vk&dh)R{A*smnY!!R?pL#J2YHj)JhgxS)b9nN zR`VinScPpZ61wEWp(mMLndygW*t%0f<3oZv3g_?wXGbIff;!HeW9X_~y$U*p2@uqA z;T)T)AL_t7fM5=(T${&c*$!=39!OA!W*$*WDG=zeW%TFl_HKT6KfAd%4>$9B^QCKF z_DU%|D~mHJC41;sZM%^Qcs2dD+ci6BQ(gVkwnJ}zZU(bxq^4e*Vd?(2lOkK3re@Q= z)}49Y4!`VGi@venR`irT+}zGCKmT$0+x3hRx~e0@g}rzaaRO2xaOcLoCQD)+NKi-i z=GfHTs_x~Wv5L%qppL>hB2HNb1a-K>gqa`icNq}Wp$SHG;S31o$lo&$&I1YRcuz36 zd*wh-hbH*EYPT|K*cUrjIS|yL2}X3t90=;r1S3ANav+$a(99zSG$g1)6O1^!1rXGs z3HHmX3;a|pfS?XdFk(t6fS?XdFk*rzfH>|*%ZX+lypE3u>Ugi(z+NH_AgDtV4DRjX z0D?L+!59Y+)S(Fm_bXWk5X>PndDP61mPk;CBzQi$r8?es#Q_9$NP^-Bt`BhlK^>By zI5t)A*_A1NkDv}oP#nSiO4b1cbI43?9l^alEs>xOO)$8(ivtMifCSmON`asbO)w%4 zJL!)Tqz+9mX0L!?4m*X*Rqd+nY6zKqQhA=SPf&*@7(Dxk0|@HS1S5K)o!e#Q9zh+N zU_^b`xqVzw)S(H+bGtY^f;u2U zx>rC@hb9;?GuXLZ*5MJ%VW)5*B4#)ws6!KsIJd*vZ)&T@{XoA7dOHvt? z+&VmhIyAwExx~)xvR58)+>yDtT?p^0j-TC+2Q0orm7nW^QgD6VxFIUPSZ~aR5OblAt)e&cOD{&h2A@IwZj{K^;I)ha@PD z;ClnvD+wVgUqoOm!XM=c)jLI?kOVq6P~fjyTTlCsO2^xp$X( zG406`sZ2?oug>HvZ|G{La>N*zEjhx{bR)?ipKaSkM? zLlX=;yXpXfIyAwEvzq`x9hzW79tjZCp$UfdMBOVOm?Mqp?aqM&b!dX#QykVyQXr^9 s6AbI^x(*recRtLKHvj+t literal 0 HcmV?d00001 diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/box-units.txt b/documentation/audits/evidence-chaos-night-2026-09-17/box-units.txt new file mode 100644 index 00000000..3c573df6 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/box-units.txt @@ -0,0 +1,24 @@ +[Unit] +Description=CHAOS NIGHT background household (survives power cuts and resets) +After=network-online.target pve-guests.service + +[Service] +Type=simple +ExecStart=/root/hloop.sh 192.168.0.116 +Restart=always +RestartSec=10 + +[Install] +WantedBy=multi-user.target +[Unit] +Description=chaos-night disk guard (harness safety net, not product) +After=network.target + +[Service] +Type=simple +ExecStart=/bin/bash /root/diskguard.sh +Restart=always +RestartSec=5 + +[Install] +WantedBy=multi-user.target diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/teardown-baseline.txt b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-baseline.txt new file mode 100644 index 00000000..257f438b --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/teardown-baseline.txt @@ -0,0 +1,75 @@ +# TEARDOWN BASELINE - taken 2026-09-17T00:30:34Z, BEFORE anything is destroyed +# The brief asks for a three-layer teardown with a before and an after. This is the before. +# It cannot be taken again once the machine is gone. + +## LAYER 2 - the host (demo-hp) +pvesm status, BEFORE: + Name Type Status Total (KiB) Used (KiB) Available (KiB) % + felhom-pbs pbs active 0 0 0 0.00% + local dir active 40453376 29325896 9040364 72.49% + local-lvm lvmthin active 56487936 25278351 31209584 44.75% + nvme-scratch dir active 983379700 66624884 866728204 6.78% + +guests, BEFORE: + VM 336 tester1-chaos-night running 8192 MB bootdisk 32.00 G pid 3515585 + CT 9201 demo-hp running <- a STANDING demo box, must survive + CT 9202 demo-hp-scratch running <- the scratch guest, must survive + +/mnt/hdd_1 contents, BEFORE: + dump 4.8G · e2d-images 8.0K · felhom-data 994M · images 59G · lost+found 16K + private 4.0K · scratch-drives 1.9M + df: /dev/nvme0n1 938G total, 64G used, 827G available, 8% + +## LAYER 1 - the machine (VM 336 "tester1-chaos-night") + agent: 1 cores: 4 memory: 8192 cpu: host + boot: order=scsi0 net0: virtio=BC:24:11:BB:C2:8F,bridge=vmbr0 + scsi0: nvme-scratch:336/vm-336-disk-1.raw 32G <- system disk + scsi1: nvme-scratch:336/vm-336-disk-0.raw 100G <- data disk (the one pulled in round 11) + scsi2: nvme-scratch:336/vm-336-disk-2.raw 64G <- SEE THE DEVIATION BELOW + +## A DEVIATION FROM THE BRIEF, DECLARED RATHER THAN QUIETLY TORN DOWN +The brief specified "system disk + one data disk". This machine has THREE disks. +The third (scsi2, 64 G) was added by me during Phase 0, to extend the LVM thin pool after I filled +it to 100% by firing twelve app deploys at once. The damage was mine, the remedy was mine, and it +changed the fixture from what the brief described. It appears in the nested guest as `sdc`, feeding +`pve-data_tdata`. +Recorded here because a teardown that silently removes an undeclared disk would erase the only +evidence that the fixture was not what the brief asked for. + +## THE CLAIM, NOW EARNED (storage definition read 2026-09-17T00:31:06Z) + dir: nvme-scratch + path /mnt/hdd_1 + content images,rootdir + is_mountpoint yes +and the disks resolve to real files at the storage root: + /mnt/hdd_1/images/336/vm-336-disk-0.raw 107374182400 bytes (100G, the data disk) + /mnt/hdd_1/images/336/vm-336-disk-1.raw 34359738368 bytes ( 32G, the system disk) + /mnt/hdd_1/images/336/vm-336-disk-2.raw 68719476736 bytes ( 64G, the disk I added) +So the brief's requirement - the disk on /mnt/hdd_1 at its root - was met, and that is now a +measurement rather than an inference. The teardown must leave /mnt/hdd_1/images/336 gone. + +## END-OF-SESSION TOOL READING (required whatever else happens) +`python3 scripts/unproven.py --summary`, 2026-09-17T00:30:40Z: + where felhom stands - 55 claims, verified_on 2026-08-22 + walked 20 · partial 17 (11 cite evidence, 6 prose only) + built 14 (1 cite evidence, 13 prose only) · missing 4 (0 cite evidence, 4 prose only) + NOT WALKED: 35 of 55 +NO NUMBER MOVED. 35 of 55 is exactly the figure the repo already documents, so tonight's work did +not change any claim's walked status - which is correct: this was a validation night, and it shipped +no product code. + +## THE HARNESS IS STOPPED (2026-09-17T00:32:07Z), and its logs are off the box + household inactive / disabled + diskguard inactive / disabled + household.log frozen at 204 lines, 10465 bytes; 7 lines flagged as failures, of which only 2 are + real events (see household-summary.txt - three were my own classifier bug, and the log says so). + diskguard.log 0 bytes: the guard never fired once, all night. + +### AND A CORRECTION TO MY OWN ACCOUNT, ONE MINUTE OLD +I announced that my earlier copy of the household log was "twelve lines short" and that re-copying +had rescued the missing lines. IT WAS NOT SHORT. Both copies are byte-identical: 204 lines, 10465 +bytes. The "192" I compared against was a line count read at 00:19, twelve minutes BEFORE the copy +was taken at 00:31. I compared a stale number with a fresh one and reported a rescue that never +happened. Nothing was lost and nothing was recovered; the only thing at risk was the accuracy of +this record, which is why it is written down. Same class as the night's other slips: a number is +only meaningful next to the time it was taken.