chaos night: teardown baseline, and the box's own logs copied off before anything stops
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
The before-picture that cannot be retaken once the machine is gone: pvesm
status, the guest list, VM 336's full config and the contents of /mnt/hdd_1.
Two things it records that correct my own assumptions:
* the machine has THREE disks, not the two the brief specified. The third is
the 64G disk I added during Phase 0 to extend the thin pool after filling
it with twelve simultaneous deploys. My damage, my remedy, and a deviation
from the fixture the brief described - declared rather than quietly torn
down.
* the /mnt/hdd_1 claim is now earned: nvme-scratch is defined with
path /mnt/hdd_1, is_mountpoint yes, and the three raw files sit in
/mnt/hdd_1/images/336.
The harness is stopped and disabled, its logs copied off first (R-320):
household 204 lines, diskguard 0 bytes - the guard never fired all night.
And a correction one minute old: I announced that the earlier log copy was
twelve lines short and that re-copying rescued them. It was not short - both
copies are byte-identical. I compared a line count read at 00:19 against a copy
taken at 00:31. Nothing was lost; only the accuracy of the record was at risk.
unproven.py: 35 of 55 not walked - NO NUMBER MOVED, which is correct for a
validation night that shipped no product code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,17 @@
|
||||
#!/bin/bash
|
||||
# The light background household. Runs ON THE VM (the nested PVE), hitting the customer guest's
|
||||
# traefik, because the guest is not reachable from DooPlex at all. Every 2 minutes: one read of a
|
||||
# random app's front door and one read of the dashboard's health endpoint. Failures are DATA.
|
||||
G="${1:?guest ip}"
|
||||
LOG=/root/household.log
|
||||
APPS="wiki paste share recipes inventory status travel paperless cloud photos media vault"
|
||||
while true; do
|
||||
A=$(echo $APPS | tr ' ' '\n' | shuf -n1)
|
||||
C=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 -H "Host: ${A}.enkicsifelhom.hu" "http://${G}/" 2>/dev/null)
|
||||
case "$C" in 2*|3*) R="ok http=$C";; 000) R="UNREACHABLE";; *) R="FAILED http=$C";; esac
|
||||
printf '%s %-10s read %s\n' "$(date -u +%FT%TZ)" "$A" "$R" >> $LOG
|
||||
H=$(curl -s -o /dev/null -w '%{http_code}' --max-time 15 -H "Host: felhom.enkicsifelhom.hu" "http://${G}/api/health" 2>/dev/null)
|
||||
case "$H" in 2*|3*) R="ok http=$H";; 000) R="UNREACHABLE";; *) R="FAILED http=$H";; esac
|
||||
printf '%s %-10s dash %s\n' "$(date -u +%FT%TZ)" "$A" "$R" >> $LOG
|
||||
sleep 120
|
||||
done
|
||||
Binary file not shown.
@@ -0,0 +1,24 @@
|
||||
[Unit]
|
||||
Description=CHAOS NIGHT background household (survives power cuts and resets)
|
||||
After=network-online.target pve-guests.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
ExecStart=/root/hloop.sh 192.168.0.116
|
||||
Restart=always
|
||||
RestartSec=10
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
[Unit]
|
||||
Description=chaos-night disk guard (harness safety net, not product)
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
ExecStart=/bin/bash /root/diskguard.sh
|
||||
Restart=always
|
||||
RestartSec=5
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -0,0 +1,75 @@
|
||||
# TEARDOWN BASELINE - taken 2026-09-17T00:30:34Z, BEFORE anything is destroyed
|
||||
# The brief asks for a three-layer teardown with a before and an after. This is the before.
|
||||
# It cannot be taken again once the machine is gone.
|
||||
|
||||
## LAYER 2 - the host (demo-hp)
|
||||
pvesm status, BEFORE:
|
||||
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
|
||||
felhom-pbs pbs active 0 0 0 0.00%
|
||||
local dir active 40453376 29325896 9040364 72.49%
|
||||
local-lvm lvmthin active 56487936 25278351 31209584 44.75%
|
||||
nvme-scratch dir active 983379700 66624884 866728204 6.78%
|
||||
|
||||
guests, BEFORE:
|
||||
VM 336 tester1-chaos-night running 8192 MB bootdisk 32.00 G pid 3515585
|
||||
CT 9201 demo-hp running <- a STANDING demo box, must survive
|
||||
CT 9202 demo-hp-scratch running <- the scratch guest, must survive
|
||||
|
||||
/mnt/hdd_1 contents, BEFORE:
|
||||
dump 4.8G · e2d-images 8.0K · felhom-data 994M · images 59G · lost+found 16K
|
||||
private 4.0K · scratch-drives 1.9M
|
||||
df: /dev/nvme0n1 938G total, 64G used, 827G available, 8%
|
||||
|
||||
## LAYER 1 - the machine (VM 336 "tester1-chaos-night")
|
||||
agent: 1 cores: 4 memory: 8192 cpu: host
|
||||
boot: order=scsi0 net0: virtio=BC:24:11:BB:C2:8F,bridge=vmbr0
|
||||
scsi0: nvme-scratch:336/vm-336-disk-1.raw 32G <- system disk
|
||||
scsi1: nvme-scratch:336/vm-336-disk-0.raw 100G <- data disk (the one pulled in round 11)
|
||||
scsi2: nvme-scratch:336/vm-336-disk-2.raw 64G <- SEE THE DEVIATION BELOW
|
||||
|
||||
## A DEVIATION FROM THE BRIEF, DECLARED RATHER THAN QUIETLY TORN DOWN
|
||||
The brief specified "system disk + one data disk". This machine has THREE disks.
|
||||
The third (scsi2, 64 G) was added by me during Phase 0, to extend the LVM thin pool after I filled
|
||||
it to 100% by firing twelve app deploys at once. The damage was mine, the remedy was mine, and it
|
||||
changed the fixture from what the brief described. It appears in the nested guest as `sdc`, feeding
|
||||
`pve-data_tdata`.
|
||||
Recorded here because a teardown that silently removes an undeclared disk would erase the only
|
||||
evidence that the fixture was not what the brief asked for.
|
||||
|
||||
## THE CLAIM, NOW EARNED (storage definition read 2026-09-17T00:31:06Z)
|
||||
dir: nvme-scratch
|
||||
path /mnt/hdd_1
|
||||
content images,rootdir
|
||||
is_mountpoint yes
|
||||
and the disks resolve to real files at the storage root:
|
||||
/mnt/hdd_1/images/336/vm-336-disk-0.raw 107374182400 bytes (100G, the data disk)
|
||||
/mnt/hdd_1/images/336/vm-336-disk-1.raw 34359738368 bytes ( 32G, the system disk)
|
||||
/mnt/hdd_1/images/336/vm-336-disk-2.raw 68719476736 bytes ( 64G, the disk I added)
|
||||
So the brief's requirement - the disk on /mnt/hdd_1 at its root - was met, and that is now a
|
||||
measurement rather than an inference. The teardown must leave /mnt/hdd_1/images/336 gone.
|
||||
|
||||
## END-OF-SESSION TOOL READING (required whatever else happens)
|
||||
`python3 scripts/unproven.py --summary`, 2026-09-17T00:30:40Z:
|
||||
where felhom stands - 55 claims, verified_on 2026-08-22
|
||||
walked 20 · partial 17 (11 cite evidence, 6 prose only)
|
||||
built 14 (1 cite evidence, 13 prose only) · missing 4 (0 cite evidence, 4 prose only)
|
||||
NOT WALKED: 35 of 55
|
||||
NO NUMBER MOVED. 35 of 55 is exactly the figure the repo already documents, so tonight's work did
|
||||
not change any claim's walked status - which is correct: this was a validation night, and it shipped
|
||||
no product code.
|
||||
|
||||
## THE HARNESS IS STOPPED (2026-09-17T00:32:07Z), and its logs are off the box
|
||||
household inactive / disabled
|
||||
diskguard inactive / disabled
|
||||
household.log frozen at 204 lines, 10465 bytes; 7 lines flagged as failures, of which only 2 are
|
||||
real events (see household-summary.txt - three were my own classifier bug, and the log says so).
|
||||
diskguard.log 0 bytes: the guard never fired once, all night.
|
||||
|
||||
### AND A CORRECTION TO MY OWN ACCOUNT, ONE MINUTE OLD
|
||||
I announced that my earlier copy of the household log was "twelve lines short" and that re-copying
|
||||
had rescued the missing lines. IT WAS NOT SHORT. Both copies are byte-identical: 204 lines, 10465
|
||||
bytes. The "192" I compared against was a line count read at 00:19, twelve minutes BEFORE the copy
|
||||
was taken at 00:31. I compared a stale number with a fresh one and reported a rescue that never
|
||||
happened. Nothing was lost and nothing was recovered; the only thing at risk was the accuracy of
|
||||
this record, which is why it is written down. Same class as the night's other slips: a number is
|
||||
only meaningful next to the time it was taken.
|
||||
Reference in New Issue
Block a user