Architecture: 08 records ruling A (45 m, the round-9 arithmetic, the cost) and
the two new event types with their audiences; 03 records the slow counter as
built; 07 records the restore-record persistence as a REVERSED design for the
restore record only; CONTEXT.md carries the day's rulings.
Guide: the recovery code moves after the first apps and waits for the yellow bar.
Register: R-549 and R-550 closed PROVEN-LIVE; R-546 closed on red-proofed tests
with its live walk owed by R-551 (no Tier-0 box is paused AND agent-connected).
R-552 filed: an interrupted-restore notice for a removed app never clears -
found in my own v0.246.0 after the release was built.
Evidence: Part A (hub prints 45m/1h30m), Part C delivery on HP and N100, B.4(a)
live proof and its teardown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-549 (operator ruling A): controllerStatus hardcoded 30m/1h while both
staleness checkers and hostStatus read alerting.stale_threshold. Moving the
threshold to 45m would have painted a customer amber 15 minutes before the
alarm could fire - the second definition rollup.go's header forbids. It now
reads the same value, down at 2x. Both 'checker initialized' log lines print
the threshold, which no line did before.
R-539 (ruling 3 of 2026-09-16): controller_slow_crashloop (warning,
operator-only), minted when the agent's slow_crashloop_since moves, with the
fast sibling's first-sight rule.
R-550: restore_interrupted (warning, for the household) allowlisted with a
Hungarian customer message.
Red-proofs, each seen failing then passing: the status test with the old
hardcoded numbers; the checker test with the movement branch removed; the
operator-only test with the registration removed; the household-message test
with the Hungarian entry removed (asserted on the SUBJECT - the body
legitimately repeats the raw message, which my first version of the test
mistook for a fallback).
go build/vet/test ./... green, 18 packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Eight per-round accident logs (rounds 3, 4, 5, 7, 8, 9, 10, 11) sat untracked
in the evidence directory, plus round 1's poll log's final three lines - the
TCP reset that ended the ghost task. Found by the clean-tree check at the start
of the next task. Evidence, no content change to any finding.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The acknowledged delete went through at 07:25:13Z (confirm_host_id +
delete_escrow=1 -> 303). Every line of the after-state written down before the
act matched: the host 404s; drill-r50, both demo hosts and the tester-1
customer still 200; the customer lists zero hosts; ep0 identical across three
readings - six snapshots, 16G, nothing removed.
The automatic connect mail arrived two seconds later and is quoted with its
token redacted. It is provably tonight's: the mailbox held no such mail newer
than 18:17:46Z when checked at 00:38Z.
Why the hub layer finished six hours late is mine: the retry guard refused to
post while the host page contained the word ONLINE, and that word sits in a
JavaScript string that is always on the page. The hub's structured status said
'down' from about 00:54Z. It gave up at 01:20Z and nothing ran again until
07:24Z. Eleventh instrument fault of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
drill-r50 and the tester-1 CUSTOMER record both captured at 200 while the
delete is still pending, with the four host records listed and the customer
page showing exactly one host - tonight's box.
The expected after-state is written down BEFORE the act, so it cannot be
adjusted to fit what happens: three host records left, the other three still
200, the customer record still 200 with zero hosts, and ep0 untouched. A fence
is only proven by a comparison, which is why ep0 was listed twice too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The two the brief flagged were both TRUE and both checked tonight rather than
assumed: the automatic self-bind mail really was waiting (18:17:46Z, zero
operator presses), and the WG hook really does re-issue by itself after an
acknowledged delete - pbsdr_auto_reissue at 20:19Z, the F-14 path measured live
for the first time. Neither pre-declared press was needed.
The ones that were wrong: the off-site app restore was impossible on a rebuild
box whose repository is orphaned by design; the fixture ended with three disks
rather than two, which is my deviation and not the brief's; round 7's drawn
'update' could not run because the catalog's own canary failed, so 'use' ran
instead and was logged; 'an internet cut tests hub unreachability' was false
here because the hub resolves to a LAN address, which is my error against my
own recorded warning; and the schedule table's clock column was nominal - the
night's twelve rounds finished at 00:17Z, about four hours earlier than the
table suggests, with the drawn order, apps and accidents never changed.
Plus one the brief did not make and the night could not answer: the
dropped-event path remains unmeasured, because no event coincided with any of
the three hub outages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Phase 2: all twelve front doors 200 and all 26 containers healthy (traefik and
cloudflared carry no healthcheck, so a bare Up is correct). The off-site
restore could not be done and two independent instruments agree why - restic
itself and the product's own status surface both say the repository is orphaned
with zero readable snapshots, because this box is a rebuild whose restic
password was minted fresh. The product surfaced that honestly within seconds.
What DID leave the house: the whole-guest copy on ep0, two intact snapshots
including tonight's 21:59:54Z one. Stated as a limit: that is a listing, not a
verification, and a PBS verify writes state so it was not run.
Household loop: 204 probes, 7 flagged, only 2 real events - both single-sample
outages during the two abrupt stops. Three were my own classifier counting a
301 as a failure, corrected in the log's own words. Ten of twelve rounds left
no mark, which is the 2-minute sampling rate and not proof of nothing.
Catalog bump verified reverted. Mail delivery proof recorded: raised and
delivered are two different claims and only one had evidence before tonight.
Interventions: 1 of 4. Both pre-declared presses unused - both prompt claims
they insured against turned out true, and the F-14 path was measured live for
the first time. Phase 0's seeding repairs listed separately because that damage
was mine; harness acts excluded with the reason stated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Headline, three lines. Interventions: 1 - round 6's local backup leg, whose
off-site leg then succeeded unaided; both pre-declared presses went unused.
Ready for a volunteer: still yes - nothing cost a byte of customer data, the
box healed itself every time with no human, and all 17 alarms were true, none
missing, every one delivered. The pair that hurt most: restore + hard reset,
not because the box suffered (26/26 containers back in 150 s) but because it is
the only pair where the household is left not knowing what happened.
Capability map: a new PROVEN-LIVE row for a random night of household actions
under accidents, carrying what it does NOT claim - per-app off-site restore
untested (orphaned repo by design), the dropped-event path still unmeasured
because no event coincided with any hub outage, twelve rounds is a sample not
coverage, and the household loop's 2-minute sampling means ten rounds left no
mark in it.
The unaided-recovery-journey row gets a second scope note rather than a change:
tonight did not walk it and could not have, so its PROVEN-LIVE still stands on
0.206.0 only - neither re-proven nor contradicted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The mailbox closes a gap the truth table could not: 17 alarms fired and all 17
were true, but that was the FEED. The mailbox shows they reached a person.
Round 11's full set arrived - storage_disconnected naming the drive, four
app_start_failed naming exactly the four apps whose data was on it, then
health_degraded. Round 6's whole_guest_backup_failed names the TIER. Round 1's
offbox_repo_orphaned was mailed within seconds of the run.
What is absent matters too: NO node_stale mail for tonight's box after round
9's outage, exactly as R-549 predicts - the gap was 29m59s against a 30-minute
threshold. One second the other way and this would be a page-out for a healthy,
self-repaired box.
Teardown layer 3 began with a DELIBERATE un-acknowledged delete, to see the
gate refuse: HTTP 409, 'Host is ONLINE - deletion is refused'. Gate one fired,
not the escrow gate - the box died inside the hub's 30-minute liveness window.
The record survived, verified by its own URL returning 200 rather than by
counting substrings on a list page (grep -c counts lines, not occurrences - my
'3 then 2' was my error, not a deletion).
The wait is the product's and not mine to shortcut: the acknowledged delete is
armed for 00:55Z behind a guard that will not post while the host reads ONLINE.
The fence says the hub is never changed, so a liveness gate is waited for.
Also recorded: the mailbox baseline proving no connect mail exists from tonight,
so the one quoted after the delete is provably new - the exact check the brief
said had been skipped before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Machine layer: VM 336 destroyed with all three disks purged, gated on its NAME
rather than its number because 9201 and 9202 share the host. qm list now shows
no VMs; /mnt/hdd_1/images/336 is gone; both standing guests still run.
Host layer, before and after: nvme-scratch 6.78% -> 1.61% (~48.5 GB released),
/mnt/hdd_1/images 59G -> 9.4G with only the scratch guest's own 9202 directory
left, free space 827G -> 875G. local-lvm UNCHANGED at 44.75% - the fence that
said 'local-lvm never' held. Firewall back at baseline with 0 physdev rules, so
none of the three network accidents left a rule on a host carrying two standing
guests. Both harness units stopped and disabled before the box died; the disk
guard's log was 0 bytes - it never fired once.
9202: nothing to remove, shown rather than asserted - three infrastructure
containers, 55 catalog TEMPLATES none of which was touched since 21:00, no
app.yaml marked deployed, no offbox config. A false label in my own transcript
is corrected there: I printed '(nothing listed above = no app stacks)' directly
beneath 55 names.
ep0: read again immediately before the delete and identical to the baseline.
The single-host delete handler shows no ep0 cascade, but a grep returning
nothing is the weakest evidence there is, so the store gets a before and an
after rather than an inference.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The before-picture that cannot be retaken once the machine is gone: pvesm
status, the guest list, VM 336's full config and the contents of /mnt/hdd_1.
Two things it records that correct my own assumptions:
* the machine has THREE disks, not the two the brief specified. The third is
the 64G disk I added during Phase 0 to extend the thin pool after filling
it with twelve simultaneous deploys. My damage, my remedy, and a deviation
from the fixture the brief described - declared rather than quietly torn
down.
* the /mnt/hdd_1 claim is now earned: nvme-scratch is defined with
path /mnt/hdd_1, is_mountpoint yes, and the three raw files sit in
/mnt/hdd_1/images/336.
The harness is stopped and disabled, its logs copied off first (R-320):
household 204 lines, diskguard 0 bytes - the guard never fired all night.
And a correction one minute old: I announced that the earlier log copy was
twelve lines short and that re-copying rescued them. It was not short - both
copies are byte-identical. I compared a line count read at 00:19 against a copy
taken at 00:31. Nothing was lost; only the accuracy of the record was at risk.
unproven.py: 35 of 55 not walked - NO NUMBER MOVED, which is correct for a
validation night that shipped no product code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The defect I almost filed: I suspected the household is never told the off-site
repository is orphaned, because 'elarvult' appears zero times in 60 KB of HTML
and my search was sound (negative control 0, two positive controls finding real
Hungarian text). Wrong measurement. The page's script fetches
/backup/offbox/status and the page carries offbox-orphan-card, orphan-reveal,
orphan-confirm and a triangle-alert icon. The household IS told, in a card
rendered client-side. Nothing filed - caught BEFORE the row existed, unlike
R-550.
The worthless probe: my attempt to read a verify state out of the ep0 manifest
returned nothing, and so did its negative control. With a compressed blob,
'no match' and 'unreadable' are indistinguishable, and I had no positive
control. So the verify state is UNKNOWN, not absent, and the only claim that
stands is that the copies are present and well-formed.
The catalog: verified reverted rather than remembered - clean tree, level with
origin, original redis pin and catalog_since intact, newest commit 2026-09-15.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
restic with the box's own credentials: 'Fatal: wrong password or no key found',
exit status 1. The product's own surface: orphaned true, snapshots 0, status
error. The repository is orphaned because this box is a REBUILD - its restic
password was minted fresh, so the existing snapshots cannot be opened. The
product surfaced that honestly as a true alarm in round 1.
So Phase 2's 'restore one DB-backed app from off-site' has nothing to restore
from. That is a fact about the fixture, not a product failure.
Recorded alongside: ep0 holds two intact whole-guest snapshots for this box,
including tonight's 21:59:54Z copy - round 6's off-site leg, the one that ran
by itself after I killed the local leg. Its file index is four times the size
of the afternoon copy. So the data did leave the house.
Stated as a limit, not glossed: that is a LISTING, not a verification. A PBS
verify would prove restorability and writes state, so it was not run - ep0 is
read-only for evidence tonight.
Tenth instrument slip recorded: my first restic probe printed 'exit status 0'
beneath a fatal error, because the zero belonged to the head at the end of the
pipe. restic 0.14 has sftp.command, not sftp.args - established by asking
'restic options' rather than assuming.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The row first claimed a restore leaves no record anywhere, citing four status
endpoints that 404'd. All four were paths I guessed. The real route, read out
of the restore page's own JavaScript, is /api/backup/restore-status and it
exists.
The corrected finding is narrower and better: the endpoint answers with the Go
zero value (started_at 0001-01-01T00:00:00Z) and carries no 'last' field at
all, while the page's own script renders '<operation> sikertelen.' from
st.last.message. The restore record is in-memory only and does not survive the
machine stopping - exactly the case a hard reset creates.
The original wording is left visible in the audit with the correction beside
it; the register row is corrected in place because a register must be accurate.
The reusable lesson: I found the real routes by asking the controller for its
own rendered links. Guessing produced four confident 404s that I then reported
as a property of the product.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
192 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of
which only TWO are real events - and both are outage pairs caused by an
injected accident: wiki during round 2's power cut, cloud during round 10's
hard reset. Each lasted less than one probe interval and healed by itself.
Three of the seven were my own classifier counting a 301 redirect as a
dashboard failure. The log carries the correction in its own words at
21:11:25Z, and the wrong lines were left in place so the correction is visible.
Stated rather than glossed: ten of the twelve rounds left no mark in this log
at all, including the twenty minutes with the drive pulled. The loop samples
each name every two minutes, so that silence is a limit of the instrument, not
proof the household saw nothing.
Ninth instrument slip recorded: the first listing printed nothing and said
'binary file matches' while the COUNT had already printed, so the summary looked
complete while the detail was dropped. Not corruption - zero null bytes, one
line of padding spaces. Re-read with grep -a.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Steady at t+1s, every door 200 at both readings, household 2 lines 0 failures,
and no new alarm - the newest feed entry is still round 11's recovery. Nothing
was raised about a box nothing was done to.
Checked rather than assumed: inject.sh's default branch exits 2 on an unknown
accident, so the control rounds never reach it - the runner handles the
no-accident case itself. A broken injector produces exactly the same result as
a control round, and only the code path distinguishes them.
Alarm truth table, all twelve rounds: 17 alarms fired, 17 true, 0 missing.
Three design gaps filed (R-547, R-549, R-550).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
ep0 baseline (read-only, nothing removed, no prune): three namespaces, each
with ct/9201 holding 2 snapshots, six in total, 16G used of 98G. The teardown
repeats this listing so 'the backups stayed' is a comparison, not an assertion.
Phase 2 readiness: scratch guest 9202 is running with a healthy controller and
restic 0.14.0 inside the controller container, but has NO off-site target
configured - so the restore will need the box's own repository address and
password handed to it.
Recorded decision: I did NOT read those from the box while round 12 ran. Round
12 is the closing control round and its whole value is that nothing was done to
the box during it. The read costs nothing to defer; the round cannot be re-run.
Also recorded: the eighth instrument slip - perl locale warnings plus a head
that cut the output before the marker returned an empty block that looked like
'9202 has no containers'.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The richest alarm round of the night, and every alarm was true and correctly
paired: storage_disconnected naming the drive by the household's own label,
four app_start_failed naming exactly the four apps whose data lives on that
drive, health_degraded, then storage_reconnected and health_recovered.
The other eleven apps kept serving throughout. The box recovered unaided in
67 s after the drive was plugged back in, with the front door following at
128 s. The drive came back clean: 98 G, 2% used, mountpoint config unchanged.
The round's own snapshot showed health_degraded with no recovery, which would
have been the first missing alarm of the night. The recovery had fired seconds
after the snapshot. A re-read taken after the precondition found it. Not a
missing alarm - a premature reading, caught by the discipline the earlier
mistimed readings forced.
Truth table now: 17 alarms fired, 17 true, 0 missing, across eleven rounds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The box is steady and the fences are clean: 26 containers, both safety units
active and enabled, 0 guard kills, 7556 MB free, all real front doors serving,
demo-hp back at -P ACCEPT with 0 physdev rules.
The slip: my check probed 'docs', a hostname that does not exist on this box.
Traefik's 404 for an unknown Host is correct behaviour, not damage. The real
names came from the household loop's own list. A probe with the wrong target
produces a confident number that means nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
One intervention in ten measured rounds: round 6's killed local backup leg.
Both pre-declared presses went unused - the automatic self-bind mail was
waiting, and the acknowledged-delete path re-issued PBS credentials by itself.
Phase 0's seeding repairs are listed separately and in full: that damage was
mine, the product behaved correctly throughout, and every repair went through
the product's own endpoints. Acts on my own instruments are excluded, with the
reason stated, so the count cannot be gamed in either direction.
Standing against the stop rule: 1 of 4.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The box passed the roughest pair drawn. A hard reset four seconds into a
restore: 26/26 containers back in 150 s, boot reconciliation naming the app it
recovered, every front door serving, one true controller_started alarm, no
false one, no intervention.
R-550 filed (P2): there is no restore record anywhere. Four candidate status
endpoints 404, no restore field in the status JSON, only a button label on the
pages, and no file at all modified in the reset window. An interrupted restore
and one that never happened look identical to the customer. Honest limit
recorded: only four seconds elapsed and the pre-reset log is unrecoverable, so
the absence of a record is what is filed, not a claim about how far it got.
Four instrument faults, all mine, all in the evidence:
* a 'nothing was logged' claim that was unfalsifiable when written - the log
stream holds zero lines before a reset;
* an on-disk check against /opt/felhom/data, a directory that does not exist;
* a household count reporting 0 lines and 0 failures when the truth was one
line and it WAS a failure - the runner now prints both operands;
* the disk guard was a TRANSIENT unit reporting 'active' all night, and was
absent from the reset onward. It is now file-backed and enabled, and its
script is copied off the box for the first time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 9's recovery confirmed end to end: 23:23:43Z 'Hub report pushed
successfully (15354 bytes)', exactly 15 minutes after the cycle that failed,
with nothing done to the box. The reading was deliberately taken after the
report was due so it could not be premature.
Also recorded: round 1's poll loop reported 'completed, exit 0' two hours after
it stopped working. It went silent during round 2's power cut and its SSH hung
until TCP reset it. Three familiar classes in one: silence is not completion,
the exit code described the local shell not the remote work, and a very late
completion notice can be mistaken for a fresh result. It touched nothing after
21:24:56Z, so no round is contaminated.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Nine rounds, 8 alarms fired, 8 true, 0 missing. Two design gaps (R-547, R-549).
Also records what this night will probably NOT answer: the event-drop path has
never been exercised, because no event was raised during any of the three
internet cuts, and rounds 10-12 draw no further cut.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
With the injector corrected, the hub really was unreachable. The controller
built its 23:08:42Z report, retried the push three times over 1m40.8s and
gave up at 23:10:23Z - 31 seconds before the link returned. Nothing queued,
which is correct: a report is a snapshot, not a fact.
The box passed. 26 containers throughout, every front door serving, both the
hub link and the host-agent link repaired unaided the moment the block lifted,
no alarm fired and none should have.
R-549 filed (P2): the staleness threshold (30 min) is exactly twice the report
cadence (15 min), so ONE failed push spends the entire budget. The measured gap
was 29m59s - one second inside the alarm. A healthy, self-repaired box came that
close to paging the operator.
Also recorded: the injected cut is broader than its name - it severed the
controller from its own host agent too, which a real ISP outage would not do.
The caveat travels with rounds 7, 8 and 9. The event-drop path remains
unmeasured, because no event was raised during any cut.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 8 (backup-app nextcloud + internet cut) passed on the product side:
26 containers throughout, LAN doors served the whole ten minutes, public
path restored unaided in <=43 s, no alarm fired and none should have.
The finding is against my own instrument. The hub report due at 22:38:43Z
fell inside the cut and SUCCEEDED, because hub.felhom.eu resolves to a LAN
address (192.168.0.192, measured from guest and host) and the injector
allowed the whole LAN. So rounds 7 and 8 never tested hub unreachability,
and the dropped-event behaviour is still unmeasured.
Two fixes, both to the harness, neither to the product:
* inject.sh now blocks the hub address from the VM's side (the host tap
rule). The hub itself is untouched - the fence is kept. It refuses to
inject at all if it cannot resolve the hub.
* run_round.sh no longer reads the front doors 3 s after an unblock. Every
door reading now states its own timestamp and a second reading is taken
60 s later. This was the sixth mistimed reading of the night.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Built from the verdicts recorded in each round's write-up. 8 alarms fired,
8 true, 0 missing. The one design gap (a transient full disk is never
mentioned to anyone) is already filed as R-547.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The post-block hub report fired at 22:23:43Z - "Hub report pushed successfully
(10538 bytes)" - and the cadence held to the second across the whole accident:
21:53:49Z, 22:08:43Z, 22:23:43Z, fifteen minutes apart, with a ten-minute
network cut sitting between the second and third.
So round 7's complete answer: the box lost its way out for ten minutes, kept
every app serving at home, restored the public path unaided in ~64 seconds,
raised no alarm (correctly - staleness is 30 minutes), attempted no hub contact
during the outage because none was due, and then reported on schedule. Nothing
was dropped because nothing was sent.
Good, and deliberately narrow: this did NOT test what happens to an alarm raised
WHILE the hub is unreachable. Rounds 8 and 9 are also ten-minute cuts and one
should contain a scheduled report naturally.
Also recorded: the fifth mistimed reading of the night, caught this time by the
measurement labelling itself - the command printed its own timestamp next to the
due time, so a reading taken fifteen seconds early announced itself instead of
becoming "the box never resumed reporting". Same principle as the marker blocks
that caught the password-file check and the command lines that caught the vzdump
self-match: make the instrument say what it actually did.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 7 (use nextcloud, internet cut ten minutes). What the customer saw depends
entirely on where they stood: at home nothing at all - traefik answered 301
throughout and every app kept serving; away from home, ten minutes of 502, then
200 again about a minute after the block lifted. The box kept all 26 containers
running, needed no repair, and re-established the way in unaided in ~64 seconds.
No alarm fired, and none should have: the ladder puts node_stale at 30 minutes
and this was ten.
The question the round was meant to answer is recorded as NOT EXERCISED rather
than passed. Are alarms raised while the hub is unreachable retried and then
silently dropped? The box reports every 15m0s (measured: 21:53:49Z, 22:08:43Z)
and the cut fell entirely between two reports, so nothing was attempted and
nothing could be lost. Rounds 8 and 9 are also ten-minute cuts and one should
contain a scheduled report naturally - they are not re-timed to force it.
The fence held, checked against a baseline taken BEFORE the round: demo-hp is
back to -P FORWARD ACCEPT, zero physdev rules, sysctl 0. That host also carries
guests 9201 and 9202, so an abandoned rule would have been a fence breach rather
than an untidy drill.
Also labelled honestly: the runner's own 530 reading was taken three seconds
after unblocking and measures nothing about recovery.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
R-547 (P3): a disk that fills and empties between sweeps is never mentioned to
anyone. The guest's root filesystem sat at 96% for ten minutes and no alarm of
any kind fired - checked twice, once by the round's runner and once
independently after the fill was released. disk_critical is defined at >=95%
used, but the fill-watch is a DAILY sweep plus one check ~90s after a controller
start, so a ten-minute window contains no check. The timing was almost comic:
the controller restarted at 21:28 after the previous round's power cut, so its
single opportunistic check ran about twenty seconds before the disk filled.
This is the ladder working as designed, not a missed alarm - it is filed because
the honest answer to "would the household be told?" is no, and that is written
down nowhere.
R-548 (P3): the whole-guest backup's LOCAL tier cannot fit on a
small-system-disk box and retries on that tier for ever. A ~29GB source into a
14GB pve-root, measured falling at ~16MB/s - under four minutes to a full / on
the nested PVE. The product's behaviour is correct throughout: it failed the
tier, named it, scheduled a retry, its status surface agreed, and the off-site
tier then succeeded from the same snapshot in ~8.5 minutes taking no local disk.
What is filed is the loop: on a box this shape the local tier can never succeed.
Honest caveat recorded in the row - the 32GB system disk is this drill's own
fixture choice - but nothing checks whether the local target could hold the
source before starting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Measured from the box's own log rather than recalled from documentation:
21:53:41Z Registered periodic job: hub-report (every 15m0s)
21:53:49Z Hub report pushed successfully (30535 bytes)
22:08:43Z Hub report pushed successfully (15934 bytes) - exactly 15m later
The block runs ~22:10Z to ~22:20Z and the next report is due ~22:23:43Z, after
it lifts. So the box never attempts a push while cut off: nothing was tried,
nothing failed, nothing was lost.
The honest verdict for the question I wanted this round to answer - are alarms
raised while the hub is unreachable retried and then silently dropped? - is NOT
EXERCISED, not "passed". Recorded that way.
It is still a finding of its own: a ten-minute internet outage is invisible to
the fleet view because the box had nothing due to say, and the hub's staleness
threshold (30 minutes) is set well beyond it. The two mechanisms agree.
Rounds 8 and 9 are also ten-minute cuts at ~25-minute spacing against a
15-minute cycle, so one will very likely contain a scheduled report and exercise
the drop behaviour properly. They are NOT re-timed to make that happen -
re-timing a round to get a better result is choosing the night after the fact.
Also captured while blocked: internet unreachable from the box, LAN reachable,
26 containers up, free space unmoved, disk guard silent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Mid-cut measurements, asked of the box while its network was blocked:
from the box: internet blocked (ping 1.1.1.1 fails), LAN reachable
26 containers still running - the apps do not care the internet is gone
free / unchanged at 7573M; diskguard active with zero log lines
from DooPlex: LAN 301, public 502
The two paths separate cleanly, and the public failure code differs from round
4's on purpose: 530 when cloudflared was dead (Cloudflare had no tunnel at all),
502 now (the tunnel lives but can reach nothing). Two different failures of the
same journey, reported differently without being asked to.
And the twelfth self-inflicted reading of the night, corrected: every
"vzdump procs: 2" was MY OWN COMMAND. The [v]zdump bracket trick stops the
pattern matching itself, but the label I echoed - "vzdump procs:" - contains the
word, so ps listed my own shell and the grep counted it. No second backup was
ever running; the box has been idle since the off-site leg finished at 22:08:27Z.
Same family as the pkill -f that killed my own watcher earlier: a pattern that
matches the hand holding it. The cure that worked: ask for command lines, not a
count.
The disk guard stays regardless - the local tier really did announce a retry with
backoff, that retry is still due, and the guard has cost nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 6 (whole-guest backup, control round). The local tier failed - a ~29GB
source into a 14GB root filesystem - and the product handled it well:
whole_guest_backup_failed (error): "Whole-guest backup FAILED on the LOCAL
TIER - retrying with backoff (next attempt in 15m0s)"
It names the tier rather than "the backup", says what it will do next, and its
status surface agrees (target_id local, success false, size_bytes 0). Then the
OFF-SITE tier ran from the same snapshot with the apps already back up,
finishing in ~8.5 minutes, encrypted to ep0, consuming no local disk at all.
During the backup 4 of 26 containers were up and every app answered 404
publicly - and NO alarm fired for those stops, which is correct: the backup's
own stack stops are suppressed, so the box does not alarm about downtime it
caused deliberately.
The downtime number (~5m43s) is recorded as CONTAMINATED by my own intervention
rather than presented as clean: I killed the local leg partway through.
A safety guard now runs on the box before round 7, declared and not a
measurement: the failed local tier retries every ~15 minutes and would consume /
at ~16MB/s while round 7 has the network cut. The guard kills only a local-tier
dump, only below a 2500M floor, and logs every action. If it fires it is an
intervention and will be counted as one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
While the off-site leg streams to ep0, free space on / does not move:
22:02:09Z procs=2 apps=26 free=7573M
22:02:40Z procs=2 apps=26 free=7573M
22:03:12Z procs=2 apps=26 free=7573M
That is the difference between the two legs of the whole-guest backup, stated as
a measurement. The local leg consumed ~16MB/s of / and would have filled it in
under four minutes; the off-site leg has run for several minutes and taken
nothing, because it streams encrypted to ep0 rather than writing an archive
locally. All 26 apps are back up and serving while it runs.
So the whole-guest backup is not "too big to work" - it is too big for the LOCAL
target, and the tier that matters for disaster recovery is unaffected. My first
conclusion was wider than the evidence; this narrows it.
ep0 is deliberately not queried to watch the snapshot land: the box's own task
status answers the same question when the leg ends, and tonight's fence is that
ep0 is written only by the product's own path and read only when nothing else
can answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Correction to my own account, recorded where the wrong version stood. I wrote
"I stopped the backup". What I stopped was its LOCAL leg: the task list shows
that job ending 23:59:45 with status "job errors" - my pkill - while a second
vzdump was already running, streaming encrypted to ep0
(--repository felhom@pbs!tester-1@...:felhom-offsite --ns tester-1).
The arithmetic that forced the intervention still stands: the local leg was
writing a ~29GB source into a filesystem with 3.6GB free, falling at ~16MB/s,
which gave under four minutes before / filled and the nested PVE wedged - the
Phase 0 failure one level up. But the off-site leg needs no local space at all,
so the box's design copes with exactly the problem I thought I was rescuing it
from. It is still counted as an intervention: I reached in and killed a job.
The box then recovered unaided: 22 containers at 22:00:52Z, 26 at 22:01:13Z,
and / went from 2539MB free back to 7573MB once the partial archive was removed.
The older completed backup was left untouched.
Round 6's valuable half stands: during the backup 4 of 26 containers were up,
every app returned 404 through the public route, and NO alarm fired - the
suppression of the backup's own stack stops held.
And the waiting rule for the off-site leg is fixed in advance: round 7 does not
start while it runs, unless it is still running at 00:45, in which case round 7
proceeds and records that it cut an in-flight backup - labelled as that, not as
a clean internet-cut round.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Measured while the vzdump was in flight (started 21:55:24Z, 2.1GB and growing):
- containers running: 4 of 26. The whole-guest backup stops twenty-two apps.
- front doors: LAN 301 (traefik is one of the four still up) but public 404 -
nothing behind the proxy to serve. Every app unavailable for the duration.
- alarms: NONE. Twenty-two apps went down at once and not one alarm fired.
Those two findings point opposite ways and both matter. The downtime is real
and total, not a brief pause - this is the standing whole-system-backup downtime
row, seen on a fresh box with twelve apps. And the suppression is correct: the
backup's own stack stops belong to a suppression set, so the box does not alarm
about downtime it caused deliberately.
The duration itself is taken from the round's runner when it reports, not
estimated from a single sample.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 4 (tunnel killed 10 min): cloudflared came back in ~97 seconds and the
CONTROLLER did it, not Docker - RestartCount=0 proves the unless-stopped policy
never acted, and the log shows the protected-infra recovery redeploying it.
health_critical fired at 21:43 and health_recovered closed it at 21:48, so the
alarm was not a dead end.
Round 4's drawn ACTION never ran, and that is recorded rather than re-run: the
round-2 power cut rebooted the guest, /tmp is cleared on boot, and the dashboard
password lived there. The off-site run was never triggered. Re-running a round
after watching it fail is how a drill starts choosing its own results. The file
now lives in /root, so round 10's hard reset cannot disarm rounds 6, 8 and 10.
Round 5 (docker restarted): all 26 containers back in 16 seconds, doors serving
again in ~40, controller_started true and correct, nothing missed. The household
loop is recorded as NOT SAMPLED - 0 lines because the round was shorter than its
2-minute sampling interval, which is not the same as 0 failures.
Two traps avoided and written down: round 5's own alarm snapshot was taken six
seconds after the controller started, from which controller_started looked
missing (it fired); and my check for the password file put its redirect on the
wrong host, reporting "not there" for a file that was present all along. Tenth
and eleventh of the same family tonight - a check whose own precondition was
wrong, answering confidently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 3 in the findings document: the disk sat at 96% for ten minutes, twelve
apps kept serving, the household loop logged 12 operations with zero failures,
and NOTHING was ever raised. The silence is the finding, and it was predicted
from the ladder before the round: the fill-watch is a daily sweep plus one check
~90s after a controller start, and that single check ran about twenty seconds
before the disk filled.
Recorded with it: I twice labelled a mid-window reading "end of window",
estimating the clock instead of reading it. The readings were unchanged but the
label was wrong, and "nothing yet" is not "nothing ever".
And a limit of my own instrument, stated before its numbers get quoted: the
household loop does not follow redirects, so it measures "is the app serving on
the box" and never "can the household reach it from outside". It logged zero
failures straight through round 4's tunnel outage while the public route was
returning 530. So "0 household failures in round 4" must not be read as "the
household was unaffected" - someone away from home would have met 530 for about
ninety seconds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Killed cloudflared and deliberately never restarted it, because whether it
returns by itself IS the measurement. It returned in ~97 seconds (killed
~21:41:30Z, running again 21:43:07.478Z), and the box did it, not Docker:
RestartCount=0 proves the unless-stopped policy never acted, and the controller
log shows "[infra] deploying cloudflared -> /opt/docker/stacks/cloudflared".
That is the product's protected-infra recovery repairing one of its own
infrastructure stacks unasked.
health_critical (error) fired at 21:43 - "Rendszer allapot kritikus (volt: ok)"
- which is exactly what the ladder predicts for a missing protected container,
and it is true. Whether health_recovered closes the pair is checked at the end
of the round, not guessed at now.
A METHOD CORRECTION that retro-labels every front-door reading tonight: the 530
during the outage is a Cloudflare status, which exposed that `curl -sL` was
following traefik's 301 out to the public hostname and back down the tunnel. So
every "front door" reading so far measured the PUBLIC path, not the LAN. It does
not invalidate the readings - a 200 by that route proves more, not less - but it
invalidates the label, and with it any claim of the form "the app is fine, only
the tunnel is down". The two paths are now measured separately: during this
outage the public route gave 530 while traefik answered 301 locally throughout.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 3 (use bookstack, system disk held at 96% for ten minutes):
- the household saw nothing wrong: wiki, status and paste all answered 200
before, during and after, and the background loop logged 12 operations with
ZERO failures on a 96%-full disk
- the box kept all 26 containers running and released the space cleanly
(29G used -> 944M used) with the thin pool untouched at 39.69% throughout
- NO alarm fired at any point, checked twice independently after the fill was
released
That silence is the finding, and it was predicted from the ladder before the
round rather than discovered after: disk_critical is defined at >=95% used, but
the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller
start. The controller happened to restart at 21:28, so its single opportunistic
check ran about twenty seconds BEFORE the disk filled. A disk that fills and
empties between sweeps is invisible - by design, but the honest answer to
"would the household be told?" is no.
Also fixed and explained: my injector printed "unexpected EOF" while the
accident itself completed. bash -n passes, so it was not local syntax - G()
flattens its argument through `pct exec`, so a nested bash -c '...' has its
quoting re-parsed remotely. Both instances were in the disk branch only; the
accidents still to come use plain commands. The experiment was verified on the
box, not from the script's own account.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Fence check mid-night: demo-hp carries only this drill's VM 336; guests 9201 and
9202 are running with their full standing sets intact, bentopdf included.
Nothing of theirs was stopped, removed or redeployed tonight.
Firewall baseline recorded BEFORE the rounds that need it (7-9 block the box's
internet by flipping a host sysctl and inserting two physdev rules):
iptables -S FORWARD -> "-P FORWARD ACCEPT" and nothing else
physdev rules -> 0
net.bridge.bridge-nf-call-iptables = 0
A control taken before the experiment, so that "it looks clean afterwards" can
be a measurement rather than an assertion - on a host that also carries the two
standing demo guests.
And a third mistimed reading of my own, recorded: I labelled a 21:36:00Z check
"end of window" when the fill runs to ~21:39:49Z. The readings are unchanged
(nothing fired), but the label is the point - "nothing yet, five minutes in" and
"nothing in the whole window" are different findings. From here the end-of-window
check is taken when the round's runner reports completion, because the runner
knows when it released the fill and I was guessing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Five minutes into the 96%-full disk: the apps keep serving (26 containers), the
pool is untouched at 39.69%, the filesystem is writable, and nothing has been
raised - the newest event is still controller_started from 21:28.
I nearly filed this as the "end of window" check. The fill began 21:29:49Z, so
the ten-minute hold runs to ~21:39:49Z and this reading was taken at 21:34:23Z,
halfway through. It is recorded as INTERIM and the end-of-window check stays
owed, because an alarm arriving late is a different finding from one that never
arrives - and a mid-window reading standing in for the final one would have
quietly turned "not yet" into "never".
What it already establishes: twelve apps keep running and serving with the
system disk at 96% full, and after five minutes nobody has been told anything.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Reviewed before it runs unattended at ~01:11, not after. The accident flips a
host-wide sysctl on demo-hp and inserts two FORWARD rules, and its cleanup ran
only on the happy path: if the script were killed during its ten-minute sleep,
or the SSH dropped, the rules and the sysctl would have stayed. demo-hp is a
Tier-0 host that guests 9201 and 9202 also live on, so an abandoned FORWARD
rule is a fence breach rather than a measurement.
It now traps EXIT, INT and TERM, removes both rules and restores the sysctl
whatever happens, and clears the trap on the normal path so the cleanup does
not run twice. The rules still match --physdev-in on this VM's own tap, resolved
at run time, and the default FORWARD policy is never touched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Measured while the guest's root filesystem was held at 96%:
- the app's view is a genuinely full disk (29G used, 1.5G free, 28G fill file)
- the shared thin pool stayed at 39.69% - fallocate reserves blocks without
writing them, so this round is NOT a repeat of the pool exhaustion that
wedged the box in Phase 0. Recorded explicitly, because "disk 95% full"
invites exactly that wrong reading.
- the filesystem stayed writable (a real touch, not the mount flags)
No disk alarm fired, and that was PREDICTED from the ladder before the round:
the fill-watch is a daily sweep at 03:30 plus one check ~90s after a controller
start. The timing is sharper still - the controller restarted at 21:28 after
round 2's power cut, so its single opportunistic check ran about twenty seconds
BEFORE the disk filled. A disk that fills and empties between checks is
invisible; that is by design, but it is the honest answer to "would the
household be told?" - no.
Household lines are attributed to the right round: the two UNREACHABLE entries
at 21:27:57Z are round 2's recovery tail, not round 3's accident.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 2 (restore gokapi + power cut, 20s into the restore): the box came back
BY ITSELF in 148 seconds, 0 -> 25 -> 26 containers, and gokapi - the app being
restored when the plug came out - returned healthy. The only alarm was
controller_started, which is what the ladder expects for a 60-second outage:
no node_stale (30 min threshold), no app_start_failed (90s boot grace). No
false alarm, none missed.
Round 2's household measure is recorded as NOT COLLECTED, not as a pass: the
loop died with the box and zero lines is not zero failures.
The round also handed over immich's whole diagnosis. app_oom fired - "immich
(immich-postgres) - egy folyamatat a memoriakorlat leallitotta" - naming the
app and the exact container. That is why immich saw CONNECTION_CLOSED and
crash-looped twelve times. It is added as tonight's line on the EXISTING OOM
row rather than filed as a new one, because this project's standing finding is
that those signals are invisible inside LXC guests and on this box the scan
caught one. The diagnosis I spent twenty minutes reaching from logs was sitting
in the alarm feed, correctly labelled, the whole time.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 2 (restore gokapi + power cut) is running: the restore started, the plug
came out 20s later, and the box is recovering on its own.
Found at round 2 and fixed for the rest of the night: the background household
loop ran as a TRANSIENT unit on the VM, so the first accident that could have
produced household failures - a power cut - instead killed the loop and produced
no lines at all. Zero lines is not zero failures, and round 2's household
measure is recorded as NOT COLLECTED rather than as a pass. It is now a real
systemd unit with Restart=always, enabled at boot, so it returns with the box
after the power cuts, hard reset and docker restart still to come.
Machinery for the remaining rounds written in advance rather than mid-round:
one generic runner covering every action and accident the seed actually drew,
so no round is measured a different way from another. It records the same five
things each time, counts the household loop's lines and failures for its own
window, and names immich as a known pre-existing failure so nothing later is
misattributed to an accident.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The household the accidents actually hit: 11 of 12 front doors serve 200,
verified with a no-such-host control so a 200 means a real route to a real app.
bookstack's 500 was my seventh error and the first I confirmed before acting:
the repair script DECODED the stored APP_KEY and printed "prefix ok: False,
body length: 96, NOT valid base64" - Laravel could never have used it. A clean
redeploy with base64:$(openssl rand -base64 32) had it healthy in 45 seconds.
immich is diagnosed (CONNECTION_CLOSED to its postgres during reverse-geocoding
init, RestartCount=12) and deliberately LEFT BROKEN: no round in the drawn
schedule acts on it, and chasing the one app the schedule never touches would
cost rounds that were drawn before the night began. It is named as a known
pre-existing condition so no later failure is misattributed to an accident.
Machinery fixed before its round arrives, not during it:
- inject.sh used `qm guest exec`, which this box cannot do (no guest agent,
measured in Phase 0). Three drawn accidents depend on in-guest work, so it now
goes over SSH + pct exec.
- the disk-95%-full accident gained a POOL GUARD: the guest's disks are thin
provisioned over the pool that hit 100% and remounted the box read-only
earlier tonight, so the fill is capped and any cap is declared in the round's
own evidence rather than silently applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Verified after the repairs: 26 containers running, all twelve apps deployed,
ZERO auth-failure lines on the rebuilt DB-backed apps, and 10 of 12 front doors
answering 200 - including cloud (nextcloud) and share (gokapi), both of which
were broken an hour ago. The cures are confirmed at the front door, not by a
health badge.
bookstack diagnosed properly rather than guessed at: its healthcheck exits 22
(curl's "server returned an HTTP error"), a direct request to the container
returns 500, its migrations completed cleanly and it has no auth failures. So
neither the database nor the image is at fault - the app itself errors. The
likely cause is mine: I passed APP_KEY=base64: plus 32 random alphanumerics,
which is not a base64-encoded 32-byte key. The repair decodes the stored key
and prints its true length BEFORE redeploying, so the hypothesis is confirmed
or refuted in the evidence.
Also recorded: my sixth slip, running docker over SSH on the VM instead of
inside the guest, which printed a tidy table of "absent" and "0" that read like
"nothing is wrong" and was produced by a shell with no docker at all. Six of my
errors tonight share one shape - a command whose precondition failed, still
printing a confident answer - and the same discipline caught every one: ask the
box directly, with a control.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
The five apps my disk-full burst broke are removed with their data and
redeployed cleanly, one at a time. Two distinct faults of mine, with different
cures, and separating them is what made either fixable:
- image layers written while the pool was full -> "invalid ELF header",
exit 127; cured by dropping the image so compose re-pulls
- my re-seed's FRESH database passwords over volumes initialised with the
first set -> Postgres auth_failed / MariaDB "Access denied"; cured only by
removing the app with its data and deploying once
gokapi proves they are different: a new image left it Restarting(1), a new
database made it healthy.
Headroom measured so the disk-full failure cannot quietly repeat: pool 39% of
75.8G, docker filesystem 29%, data drive 1%, 4.3G guest RAM free.
Round 2's script hardened before it runs unattended: its "steady" test now
compares against the count it measured itself in the same round, instead of a
hardcoded 24 that the rebuilt household might never reach.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
Round 1 (offsite-run, control round) is written up with its five things: the
run worked end to end in 1m45s, and on this REBUILT box the remote repository
is orphaned - the documented behaviour, surfaced honestly instead of reported
as a successful copy. Three alarms fired, all true and precise.
Corrections to my own earlier claims, each recorded where the wrong version
was written:
- "the 404s were my mistimed sweep" - wrong for four of five. Proven with a
negative control (a no-such-host request returns the identical 404, 19 bytes)
that traefik simply has no route to an unhealthy container.
- "nextcloud is repaired" - wrong. The re-pull fixed the corrupt library, but
the app still cannot reach its database, and the container reports HEALTHY
the whole time. A health signal is not a data signal.
- "all the broken apps are corrupt layers" - wrong. bookstack logged a clean
startup, gokapi logged nothing, and immich shows a Postgres auth_failed.
The real cause of most of it is mine: my re-seed generated FRESH database
passwords over volumes whose databases were initialised with the first set.
The affected apps are being removed with their data and redeployed cleanly.
Also recorded: an HTTP 000 is "no answer", not "it failed" - the action still
took effect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS