chaos night: teardown baseline, and the box's own logs copied off before anything stops
gates / gates (push) Successful in 21s

The before-picture that cannot be retaken once the machine is gone: pvesm
status, the guest list, VM 336's full config and the contents of /mnt/hdd_1.

Two things it records that correct my own assumptions:
  * the machine has THREE disks, not the two the brief specified. The third is
    the 64G disk I added during Phase 0 to extend the thin pool after filling
    it with twelve simultaneous deploys. My damage, my remedy, and a deviation
    from the fixture the brief described - declared rather than quietly torn
    down.
  * the /mnt/hdd_1 claim is now earned: nvme-scratch is defined with
    path /mnt/hdd_1, is_mountpoint yes, and the three raw files sit in
    /mnt/hdd_1/images/336.

The harness is stopped and disabled, its logs copied off first (R-320):
household 204 lines, diskguard 0 bytes - the guard never fired all night.

And a correction one minute old: I announced that the earlier log copy was
twelve lines short and that re-copying rescued them. It was not short - both
copies are byte-identical. I compared a line count read at 00:19 against a copy
taken at 00:31. Nothing was lost; only the accuracy of the record was at risk.

unproven.py: 35 of 55 not walked - NO NUMBER MOVED, which is correct for a
validation night that shipped no product code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 02:33:19 +02:00
parent 62f6b7b0b0
commit 9b44c44f23
5 changed files with 116 additions and 0 deletions
@@ -0,0 +1,17 @@
#!/bin/bash
# The light background household. Runs ON THE VM (the nested PVE), hitting the customer guest's
# traefik, because the guest is not reachable from DooPlex at all. Every 2 minutes: one read of a
# random app's front door and one read of the dashboard's health endpoint. Failures are DATA.
G="${1:?guest ip}"
LOG=/root/household.log
APPS="wiki paste share recipes inventory status travel paperless cloud photos media vault"
while true; do
A=$(echo $APPS | tr ' ' '\n' | shuf -n1)
C=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 -H "Host: ${A}.enkicsifelhom.hu" "http://${G}/" 2>/dev/null)
case "$C" in 2*|3*) R="ok http=$C";; 000) R="UNREACHABLE";; *) R="FAILED http=$C";; esac
printf '%s %-10s read %s\n' "$(date -u +%FT%TZ)" "$A" "$R" >> $LOG
H=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 -H "Host: felhom.enkicsifelhom.hu" "http://${G}/api/health" 2>/dev/null)
case "$H" in 2*|3*) R="ok http=$H";; 000) R="UNREACHABLE";; *) R="FAILED http=$H";; esac
printf '%s %-10s dash %s\n' "$(date -u +%FT%TZ)" "$A" "$R" >> $LOG
sleep 120
done
@@ -0,0 +1,24 @@
[Unit]
Description=CHAOS NIGHT background household (survives power cuts and resets)
After=network-online.target pve-guests.service
[Service]
Type=simple
ExecStart=/root/hloop.sh 192.168.0.116
Restart=always
RestartSec=10
[Install]
WantedBy=multi-user.target
[Unit]
Description=chaos-night disk guard (harness safety net, not product)
After=network.target
[Service]
Type=simple
ExecStart=/bin/bash /root/diskguard.sh
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.target
@@ -0,0 +1,75 @@
# TEARDOWN BASELINE - taken 2026-09-17T00:30:34Z, BEFORE anything is destroyed
# The brief asks for a three-layer teardown with a before and an after. This is the before.
# It cannot be taken again once the machine is gone.
## LAYER 2 - the host (demo-hp)
pvesm status, BEFORE:
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40453376 29325896 9040364 72.49%
local-lvm lvmthin active 56487936 25278351 31209584 44.75%
nvme-scratch dir active 983379700 66624884 866728204 6.78%
guests, BEFORE:
VM 336 tester1-chaos-night running 8192 MB bootdisk 32.00 G pid 3515585
CT 9201 demo-hp running <- a STANDING demo box, must survive
CT 9202 demo-hp-scratch running <- the scratch guest, must survive
/mnt/hdd_1 contents, BEFORE:
dump 4.8G · e2d-images 8.0K · felhom-data 994M · images 59G · lost+found 16K
private 4.0K · scratch-drives 1.9M
df: /dev/nvme0n1 938G total, 64G used, 827G available, 8%
## LAYER 1 - the machine (VM 336 "tester1-chaos-night")
agent: 1 cores: 4 memory: 8192 cpu: host
boot: order=scsi0 net0: virtio=BC:24:11:BB:C2:8F,bridge=vmbr0
scsi0: nvme-scratch:336/vm-336-disk-1.raw 32G <- system disk
scsi1: nvme-scratch:336/vm-336-disk-0.raw 100G <- data disk (the one pulled in round 11)
scsi2: nvme-scratch:336/vm-336-disk-2.raw 64G <- SEE THE DEVIATION BELOW
## A DEVIATION FROM THE BRIEF, DECLARED RATHER THAN QUIETLY TORN DOWN
The brief specified "system disk + one data disk". This machine has THREE disks.
The third (scsi2, 64 G) was added by me during Phase 0, to extend the LVM thin pool after I filled
it to 100% by firing twelve app deploys at once. The damage was mine, the remedy was mine, and it
changed the fixture from what the brief described. It appears in the nested guest as `sdc`, feeding
`pve-data_tdata`.
Recorded here because a teardown that silently removes an undeclared disk would erase the only
evidence that the fixture was not what the brief asked for.
## THE CLAIM, NOW EARNED (storage definition read 2026-09-17T00:31:06Z)
dir: nvme-scratch
path /mnt/hdd_1
content images,rootdir
is_mountpoint yes
and the disks resolve to real files at the storage root:
/mnt/hdd_1/images/336/vm-336-disk-0.raw 107374182400 bytes (100G, the data disk)
/mnt/hdd_1/images/336/vm-336-disk-1.raw 34359738368 bytes ( 32G, the system disk)
/mnt/hdd_1/images/336/vm-336-disk-2.raw 68719476736 bytes ( 64G, the disk I added)
So the brief's requirement - the disk on /mnt/hdd_1 at its root - was met, and that is now a
measurement rather than an inference. The teardown must leave /mnt/hdd_1/images/336 gone.
## END-OF-SESSION TOOL READING (required whatever else happens)
`python3 scripts/unproven.py --summary`, 2026-09-17T00:30:40Z:
where felhom stands - 55 claims, verified_on 2026-08-22
walked 20 · partial 17 (11 cite evidence, 6 prose only)
built 14 (1 cite evidence, 13 prose only) · missing 4 (0 cite evidence, 4 prose only)
NOT WALKED: 35 of 55
NO NUMBER MOVED. 35 of 55 is exactly the figure the repo already documents, so tonight's work did
not change any claim's walked status - which is correct: this was a validation night, and it shipped
no product code.
## THE HARNESS IS STOPPED (2026-09-17T00:32:07Z), and its logs are off the box
household inactive / disabled
diskguard inactive / disabled
household.log frozen at 204 lines, 10465 bytes; 7 lines flagged as failures, of which only 2 are
real events (see household-summary.txt - three were my own classifier bug, and the log says so).
diskguard.log 0 bytes: the guard never fired once, all night.
### AND A CORRECTION TO MY OWN ACCOUNT, ONE MINUTE OLD
I announced that my earlier copy of the household log was "twelve lines short" and that re-copying
had rescued the missing lines. IT WAS NOT SHORT. Both copies are byte-identical: 204 lines, 10465
bytes. The "192" I compared against was a line count read at 00:19, twelve minutes BEFORE the copy
was taken at 00:31. I compared a stale number with a fresh one and reported a rescue that never
happened. Nothing was lost and nothing was recovered; the only thing at risk was the accuracy of
this record, which is why it is written down. Same class as the night's other slips: a number is
only meaningful next to the time it was taken.