The acknowledged delete went through at 07:25:13Z (confirm_host_id +
delete_escrow=1 -> 303). Every line of the after-state written down before the
act matched: the host 404s; drill-r50, both demo hosts and the tester-1
customer still 200; the customer lists zero hosts; ep0 identical across three
readings - six snapshots, 16G, nothing removed.
The automatic connect mail arrived two seconds later and is quoted with its
token redacted. It is provably tonight's: the mailbox held no such mail newer
than 18:17:46Z when checked at 00:38Z.
Why the hub layer finished six hours late is mine: the retry guard refused to
post while the host page contained the word ONLINE, and that word sits in a
JavaScript string that is always on the page. The hub's structured status said
'down' from about 00:54Z. It gave up at 01:20Z and nothing ran again until
07:24Z. Eleventh instrument fault of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The two the brief flagged were both TRUE and both checked tonight rather than
assumed: the automatic self-bind mail really was waiting (18:17:46Z, zero
operator presses), and the WG hook really does re-issue by itself after an
acknowledged delete - pbsdr_auto_reissue at 20:19Z, the F-14 path measured live
for the first time. Neither pre-declared press was needed.
The ones that were wrong: the off-site app restore was impossible on a rebuild
box whose repository is orphaned by design; the fixture ended with three disks
rather than two, which is my deviation and not the brief's; round 7's drawn
'update' could not run because the catalog's own canary failed, so 'use' ran
instead and was logged; 'an internet cut tests hub unreachability' was false
here because the hub resolves to a LAN address, which is my error against my
own recorded warning; and the schedule table's clock column was nominal - the
night's twelve rounds finished at 00:17Z, about four hours earlier than the
table suggests, with the drawn order, apps and accidents never changed.
Plus one the brief did not make and the night could not answer: the
dropped-event path remains unmeasured, because no event coincided with any of
the three hub outages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 2: all twelve front doors 200 and all 26 containers healthy (traefik and
cloudflared carry no healthcheck, so a bare Up is correct). The off-site
restore could not be done and two independent instruments agree why - restic
itself and the product's own status surface both say the repository is orphaned
with zero readable snapshots, because this box is a rebuild whose restic
password was minted fresh. The product surfaced that honestly within seconds.
What DID leave the house: the whole-guest copy on ep0, two intact snapshots
including tonight's 21:59:54Z one. Stated as a limit: that is a listing, not a
verification, and a PBS verify writes state so it was not run.
Household loop: 204 probes, 7 flagged, only 2 real events - both single-sample
outages during the two abrupt stops. Three were my own classifier counting a
301 as a failure, corrected in the log's own words. Ten of twelve rounds left
no mark, which is the 2-minute sampling rate and not proof of nothing.
Catalog bump verified reverted. Mail delivery proof recorded: raised and
delivered are two different claims and only one had evidence before tonight.
Interventions: 1 of 4. Both pre-declared presses unused - both prompt claims
they insured against turned out true, and the F-14 path was measured live for
the first time. Phase 0's seeding repairs listed separately because that damage
was mine; harness acts excluded with the reason stated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Headline, three lines. Interventions: 1 - round 6's local backup leg, whose
off-site leg then succeeded unaided; both pre-declared presses went unused.
Ready for a volunteer: still yes - nothing cost a byte of customer data, the
box healed itself every time with no human, and all 17 alarms were true, none
missing, every one delivered. The pair that hurt most: restore + hard reset,
not because the box suffered (26/26 containers back in 150 s) but because it is
the only pair where the household is left not knowing what happened.
Capability map: a new PROVEN-LIVE row for a random night of household actions
under accidents, carrying what it does NOT claim - per-app off-site restore
untested (orphaned repo by design), the dropped-event path still unmeasured
because no event coincided with any hub outage, twelve rounds is a sample not
coverage, and the household loop's 2-minute sampling means ten rounds left no
mark in it.
The unaided-recovery-journey row gets a second scope note rather than a change:
tonight did not walk it and could not have, so its PROVEN-LIVE still stands on
0.206.0 only - neither re-proven nor contradicted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The row first claimed a restore leaves no record anywhere, citing four status
endpoints that 404'd. All four were paths I guessed. The real route, read out
of the restore page's own JavaScript, is /api/backup/restore-status and it
exists.
The corrected finding is narrower and better: the endpoint answers with the Go
zero value (started_at 0001-01-01T00:00:00Z) and carries no 'last' field at
all, while the page's own script renders '<operation> sikertelen.' from
st.last.message. The restore record is in-memory only and does not survive the
machine stopping - exactly the case a hard reset creates.
The original wording is left visible in the audit with the correction beside
it; the register row is corrected in place because a register must be accurate.
The reusable lesson: I found the real routes by asking the controller for its
own rendered links. Guessing produced four confident 404s that I then reported
as a property of the product.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Steady at t+1s, every door 200 at both readings, household 2 lines 0 failures,
and no new alarm - the newest feed entry is still round 11's recovery. Nothing
was raised about a box nothing was done to.
Checked rather than assumed: inject.sh's default branch exits 2 on an unknown
accident, so the control rounds never reach it - the runner handles the
no-accident case itself. A broken injector produces exactly the same result as
a control round, and only the code path distinguishes them.
Alarm truth table, all twelve rounds: 17 alarms fired, 17 true, 0 missing.
Three design gaps filed (R-547, R-549, R-550).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The richest alarm round of the night, and every alarm was true and correctly
paired: storage_disconnected naming the drive by the household's own label,
four app_start_failed naming exactly the four apps whose data lives on that
drive, health_degraded, then storage_reconnected and health_recovered.
The other eleven apps kept serving throughout. The box recovered unaided in
67 s after the drive was plugged back in, with the front door following at
128 s. The drive came back clean: 98 G, 2% used, mountpoint config unchanged.
The round's own snapshot showed health_degraded with no recovery, which would
have been the first missing alarm of the night. The recovery had fired seconds
after the snapshot. A re-read taken after the precondition found it. Not a
missing alarm - a premature reading, caught by the discipline the earlier
mistimed readings forced.
Truth table now: 17 alarms fired, 17 true, 0 missing, across eleven rounds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.
R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.
Four instrument faults, all mine, all in the evidence:
* a 'nothing was logged' claim that was unfalsifiable when written - the log
stream holds zero lines before a reset;
* an on-disk check against /opt/felhom/data, a directory that does not exist;
* a household count reporting 0 lines and 0 failures when the truth was one
line and it WAS a failure - the runner now prints both operands;
* the disk guard was a TRANSIENT unit reporting 'active' all night, and was
absent from the reset onward. It is now file-backed and enabled, and its
script is copied off the box for the first time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 9's recovery confirmed end to end: 23:23:43Z 'Hub report pushed
successfully (15354 bytes)', exactly 15 minutes after the cycle that failed,
with nothing done to the box. The reading was deliberately taken after the
report was due so it could not be premature.
Also recorded: round 1's poll loop reported 'completed, exit 0' two hours after
it stopped working. It went silent during round 2's power cut and its SSH hung
until TCP reset it. Three familiar classes in one: silence is not completion,
the exit code described the local shell not the remote work, and a very late
completion notice can be mistaken for a fresh result. It touched nothing after
21:24:56Z, so no round is contaminated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
With the injector corrected, the hub really was unreachable. The controller
built its 23:08:42Z report, retried the push three times over 1m40.8s and
gave up at 23:10:23Z - 31 seconds before the link returned. Nothing queued,
which is correct: a report is a snapshot, not a fact.
The box passed. 26 containers throughout, every front door serving, both the
hub link and the host-agent link repaired unaided the moment the block lifted,
no alarm fired and none should have.
R-549 filed (P2): the staleness threshold (30 min) is exactly twice the report
cadence (15 min), so ONE failed push spends the entire budget. The measured gap
was 29m59s - one second inside the alarm. A healthy, self-repaired box came that
close to paging the operator.
Also recorded: the injected cut is broader than its name - it severed the
controller from its own host agent too, which a real ISP outage would not do.
The caveat travels with rounds 7, 8 and 9. The event-drop path remains
unmeasured, because no event was raised during any cut.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 8 (backup-app nextcloud + internet cut) passed on the product side:
26 containers throughout, LAN doors served the whole ten minutes, public
path restored unaided in <=43 s, no alarm fired and none should have.
The finding is against my own instrument. The hub report due at 22:38:43Z
fell inside the cut and SUCCEEDED, because hub.felhom.eu resolves to a LAN
address (192.168.0.192, measured from guest and host) and the injector
allowed the whole LAN. So rounds 7 and 8 never tested hub unreachability,
and the dropped-event behaviour is still unmeasured.
Two fixes, both to the harness, neither to the product:
* inject.sh now blocks the hub address from the VM's side (the host tap
rule). The hub itself is untouched - the fence is kept. It refuses to
inject at all if it cannot resolve the hub.
* run_round.sh no longer reads the front doors 3 s after an unblock. Every
door reading now states its own timestamp and a second reading is taken
60 s later. This was the sixth mistimed reading of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 7 (use nextcloud, internet cut ten minutes). What the customer saw depends
entirely on where they stood: at home nothing at all - traefik answered 301
throughout and every app kept serving; away from home, ten minutes of 502, then
200 again about a minute after the block lifted. The box kept all 26 containers
running, needed no repair, and re-established the way in unaided in ~64 seconds.
No alarm fired, and none should have: the ladder puts node_stale at 30 minutes
and this was ten.
The question the round was meant to answer is recorded as NOT EXERCISED rather
than passed. Are alarms raised while the hub is unreachable retried and then
silently dropped? The box reports every 15m0s (measured: 21:53:49Z, 22:08:43Z)
and the cut fell entirely between two reports, so nothing was attempted and
nothing could be lost. Rounds 8 and 9 are also ten-minute cuts and one should
contain a scheduled report naturally - they are not re-timed to force it.
The fence held, checked against a baseline taken BEFORE the round: demo-hp is
back to -P FORWARD ACCEPT, zero physdev rules, sysctl 0. That host also carries
guests 9201 and 9202, so an abandoned rule would have been a fence breach rather
than an untidy drill.
Also labelled honestly: the runner's own 530 reading was taken three seconds
after unblocking and measures nothing about recovery.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 6 (whole-guest backup, control round). The local tier failed - a ~29GB
source into a 14GB root filesystem - and the product handled it well:
whole_guest_backup_failed (error): "Whole-guest backup FAILED on the LOCAL
TIER - retrying with backoff (next attempt in 15m0s)"
It names the tier rather than "the backup", says what it will do next, and its
status surface agrees (target_id local, success false, size_bytes 0). Then the
OFF-SITE tier ran from the same snapshot with the apps already back up,
finishing in ~8.5 minutes, encrypted to ep0, consuming no local disk at all.
During the backup 4 of 26 containers were up and every app answered 404
publicly - and NO alarm fired for those stops, which is correct: the backup's
own stack stops are suppressed, so the box does not alarm about downtime it
caused deliberately.
The downtime number (~5m43s) is recorded as CONTAMINATED by my own intervention
rather than presented as clean: I killed the local leg partway through.
A safety guard now runs on the box before round 7, declared and not a
measurement: the failed local tier retries every ~15 minutes and would consume /
at ~16MB/s while round 7 has the network cut. The guard kills only a local-tier
dump, only below a 2500M floor, and logs every action. If it fires it is an
intervention and will be counted as one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 4 (tunnel killed 10 min): cloudflared came back in ~97 seconds and the
CONTROLLER did it, not Docker - RestartCount=0 proves the unless-stopped policy
never acted, and the log shows the protected-infra recovery redeploying it.
health_critical fired at 21:43 and health_recovered closed it at 21:48, so the
alarm was not a dead end.
Round 4's drawn ACTION never ran, and that is recorded rather than re-run: the
round-2 power cut rebooted the guest, /tmp is cleared on boot, and the dashboard
password lived there. The off-site run was never triggered. Re-running a round
after watching it fail is how a drill starts choosing its own results. The file
now lives in /root, so round 10's hard reset cannot disarm rounds 6, 8 and 10.
Round 5 (docker restarted): all 26 containers back in 16 seconds, doors serving
again in ~40, controller_started true and correct, nothing missed. The household
loop is recorded as NOT SAMPLED - 0 lines because the round was shorter than its
2-minute sampling interval, which is not the same as 0 failures.
Two traps avoided and written down: round 5's own alarm snapshot was taken six
seconds after the controller started, from which controller_started looked
missing (it fired); and my check for the password file put its redirect on the
wrong host, reporting "not there" for a file that was present all along. Tenth
and eleventh of the same family tonight - a check whose own precondition was
wrong, answering confidently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 3 in the findings document: the disk sat at 96% for ten minutes, twelve
apps kept serving, the household loop logged 12 operations with zero failures,
and NOTHING was ever raised. The silence is the finding, and it was predicted
from the ladder before the round: the fill-watch is a daily sweep plus one check
~90s after a controller start, and that single check ran about twenty seconds
before the disk filled.
Recorded with it: I twice labelled a mid-window reading "end of window",
estimating the clock instead of reading it. The readings were unchanged but the
label was wrong, and "nothing yet" is not "nothing ever".
And a limit of my own instrument, stated before its numbers get quoted: the
household loop does not follow redirects, so it measures "is the app serving on
the box" and never "can the household reach it from outside". It logged zero
failures straight through round 4's tunnel outage while the public route was
returning 530. So "0 household failures in round 4" must not be read as "the
household was unaffected" - someone away from home would have met 530 for about
ninety seconds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 2 (restore gokapi + power cut, 20s into the restore): the box came back
BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being
restored when the plug came out - returned healthy. The only alarm was
controller_started, which is what the ladder expects for a 60-second outage:
no node_stale (30 min threshold), no app_start_failed (90s boot grace). No
false alarm, none missed.
Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the
loop died with the box and zero lines is not zero failures.
The round also handed over immich's whole diagnosis. app_oom fired - "immich
(immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the
app and the exact container. That is why immich saw CONNECTION_CLOSED and
crash-looped twelve times. It is added as tonight's line on the EXISTING OOM
row rather than filed as a new one, because this project's standing finding is
that those signals are invisible inside LXC guests and on this box the scan
caught one. The diagnosis I spent twenty minutes reaching from logs was sitting
in the alarm feed, correctly labelled, the whole time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The five apps my disk-full burst broke are removed with their data and
redeployed cleanly, one at a time. Two distinct faults of mine, with different
cures, and separating them is what made either fixable:
- image layers written while the pool was full -> "invalid ELF header",
exit 127; cured by dropping the image so compose re-pulls
- my re-seed's FRESH database passwords over volumes initialised with the
first set -> Postgres auth_failed / MariaDB "Access denied"; cured only by
removing the app with its data and deploying once
gokapi proves they are different: a new image left it Restarting(1), a new
database made it healthy.
Headroom measured so the disk-full failure cannot quietly repeat: pool 39% of
75.8G, docker filesystem 29%, data drive 1%, 4.3G guest RAM free.
Round 2's script hardened before it runs unattended: its "steady" test now
compares against the count it measured itself in the same round, instead of a
hardcoded 24 that the rebuilt household might never reach.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 1 (offsite-run, control round) is written up with its five things: the
run worked end to end in 1m45s, and on this REBUILT box the remote repository
is orphaned - the documented behaviour, surfaced honestly instead of reported
as a successful copy. Three alarms fired, all true and precise.
Corrections to my own earlier claims, each recorded where the wrong version
was written:
- "the 404s were my mistimed sweep" - wrong for four of five. Proven with a
negative control (a no-such-host request returns the identical 404, 19 bytes)
that traefik simply has no route to an unhealthy container.
- "nextcloud is repaired" - wrong. The re-pull fixed the corrupt library, but
the app still cannot reach its database, and the container reports HEALTHY
the whole time. A health signal is not a data signal.
- "all the broken apps are corrupt layers" - wrong. bookstack logged a clean
startup, gokapi logged nothing, and immich shows a Postgres auth_failed.
The real cause of most of it is mine: my re-seed generated FRESH database
passwords over volumes whose databases were initialised with the first set.
The affected apps are being removed with their data and redeployed cleanly.
Also recorded: an HTTP 000 is "no answer", not "it failed" - the action still
took effect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The schedule was drawn from seed 20260917 and written into the findings document
BEFORE round 1, with its re-draw log.
Phase 0 measured:
- golden 0.245.0 baked, published (registry 200, not an exit code) and vouched;
the box installed itself from the published ISO 1.28.0 and landed on it with
no hand upgrade (controller 0.245.0, agent 0.131.0).
- ZERO operator presses: the waiting self-bind mail worked, and the acknowledged
-delete path re-issued off-site AND PBS-DR credentials by itself
(pbsdr_auto_reissue) - the F-14 half nobody had watched happen live.
- R-546 filed (P2): tonight's own guide sends the household to create the
recovery code ~17 minutes before the box can do it. It self-heals; the bar
urges them there the whole time. Measured on both sides, not inferred.
- R-543 proven through its whole lifecycle on a fresh box: bar present while
paused, gone for good once escrowed.
- Known rows met and recorded, not re-filed: R-542, R-536's failure events.
Also recorded honestly: three harness errors of mine (a script that announced
"all twelve deploys ACCEPTED" without checking, a "login ok (csrf 0)" that
turned eleven of my own 401s into what looked like product refusals, and a
head -12 that hid a disk), and a near-miss where I almost filed a defect
against a drive gate that was working and logging at DEBUG.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS