R-539 closed: five real controller kills on demo-hp 9201 with the production 24 h window raised controller_slow_crashloop, and exactly one operator mail arrived (09:29:40Z). The fast brake never armed. Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed). unproven.py: 35 of 55 not walked, no number moved. The report names the brief's wrong claims first and one recommendation not followed: the controller floor was not raised - validated on one guest, a gap of my own found during validation, and a floor above the golden reaches Peti's box too. The operator's call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
6.7 KiB
REPORT — the chaos-night fixes: the quiet alarm, the restore record, the recovery-code timing, the slow crash loop
2026-09-17. R-549, R-550, R-546, R-539. Per-repo detail: felhom-controller/REPORT.md (v0.246.0),
felhom-agent/REPORT.md (v0.132.0). This file covers felhom.eu (hub v0.117.0, the setting, the documents)
and the whole task. Written as REPORT-<topic>.md because REPORT.md holds the earlier session's work.
Claims in the brief that turned out wrong — named first
- „Readiness =
escrow.pbs_storage_idset." The agent's preflightokcovers five blocking items;pbs_storage_idis the one R-546's box showed. The controller reads the combinedok. - „The backup tiers persist their records atomically." Checked at
felhom-controller/controller/internal/settings/settings.go:727-752— TRUE (tmp + rename,.bakrecovery). - „Ruling A is a config change plus a document line, not code." WRONG.
hub/internal/web/rollup.gocontrollerStatushardcoded 30 m / 1 h while both checkers andhostStatusread the config — the dashboard would have turned amber 15 minutes before the alarm could fire. Hub v0.117.0 fixes it. - „Verify from the log line that the checker runs with 45 m." No log line printed the threshold. Hub v0.117.0 adds it to both „checker initialized" lines.
- „The failure shows a raw error." Not in a browser — the page hid its start form behind the
checklist. The raw
-storagestderr came from the chaos-night harness calling the API directly. - „Emit the event from the agent" / „make the interval configurable for the test." The agent has no event channel (the hub mints from heartbeat timestamps); and no test interval was needed — five real kills 8 minutes apart reached the production threshold in the session.
1. Confirmed baselines (re-verified at start)
felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) · felhom-agent e98b857684f4 v0.131.0 ·
felhom.eu ea25c6c5b1f1 hub v0.116.0 · highest row R-550 · golden waiver to 2026-09-27.
2. Files (felhom.eu)
hub/internal/web/{rollup.go,configs.go,server.go,rollup_test.go}, hub/internal/monitor/{controller_supervisor.go,controller_supervisor_test.go,staleness.go,host_staleness.go},
hub/internal/notify/{dispatcher.go,templates.go}, hub/internal/api/{handler.go,chaosnight_events_test.go},
hub/CHANGELOG.md, manifests/hub.yaml, .claude/rules/hub.md, CONTEXT.md, STATUS.md,
documentation/architecture/{03-host-agent.md,05-hub-architecture.md,07-backup-architecture.md,08-alarm-ladder.md},
documentation/runbooks/VOLUNTEER-first-hour.md, documentation/backlog/{OPEN-ITEMS.md,CLOSED-ITEMS.md},
documentation/audits/evidence-chaos-fixes-2026-09-17/ (6 files), this report.
3. Commits
felhom.eu: 469bfa5 (chaos-night leftover evidence), 37ae31f (hub v0.117.0), 06334e1 (manifest: 0.117.0 + 45m), 3c1882a (docs, register, evidence), and the commit carrying this report.
felhom-agent: 18d03bd (v0.132.0), release-record CHANGELOG commit, 77cd70f (REPORT). Tag v0.132.0.
felhom-controller: 0fe315b (v0.246.0), 29e2acb (REPORT).
4–5. Tests
Hub: go build/vet/test ./... green, 18 packages. Agent: 30 packages. Controller: 28 packages.
Red-proofs, each seen failing then passing — hub: status follows threshold; slow-crashloop movement; operator-only registration; Hungarian subject. Agent: no counter; once-per-24h guard; persistence. Controller: restore record across restart; startup helper; main() wiring; restore page card; bar held back; waiting card; start refusal. Two of my own test mistakes corrected and recorded (a body-based fallback check; a -storage needle matching a menu id).
6. Deployed versions
Hub 0.117.0 (ArgoCD Synced/Healthy, revision == manifest HEAD, pod ready). Agent 0.132.0 on demo-hp and the N100 (signed jobs, committed). Controller 0.246.0 on demo-hp guest 9201. Proof lines in the evidence directory.
7. NOT live-validated
- R-546's readiness branches (bar held back, waiting card, start refusal) — no Tier-0 box is paused AND agent-connected. R-551.
- The escrow ceremony passing once ready — deliberately not run: the hub keeps one escrow per host, and a ceremony on demo-hp would supersede the standing box's escrow.
- Controller 0.246.0 is on 9201 only; the fleet floor was not raised (below).
8. Evidence copied off before each teardown
Yes. B.4(a) log written by the run itself to DooPlex before teardown; teardown logged separately; Part C logs off the box as they ran.
9. Teardown — three layers
- Machine: throwaway
homeboxon 9201 removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up; password and scripts shredded in the guest. - Host: nothing provisioned; demo-hp agent config untouched (the
pbs_storage_idremoval was considered and not done). - Hub: no host or customer record created or deleted; the test events stay as history. Stated effect: 9201's slow crash-loop stays raised until 2026-09-18 09:22Z and cannot mail again before then.
An error of mine, caught by a gate before it was pushed
Closing R-539, I wrote open(CLOSED-ITEMS.md,'w').write(open(CLOSED-ITEMS.md).read() + row). Python opens
— and empties — the file for writing before it reads it, so the closed register fell from 215 rows to
1. The pre-push instructions gate refused the push: citations of closed rows (R-549, and R-320 in
unprompted-work.md) suddenly pointed at nothing. Nothing damaged was pushed. The file was restored
from pushed commit 3c1882a and R-539 appended with read-then-write; the diff against 3c1882a is exactly
one added line, and the open register exactly one removed line. The earlier closures (R-546/R-549/R-550)
used read-then-write and were intact. Lesson kept: never open a file for writing in the same expression
that reads it.
A recommendation not followed, with its reason
The brief's rules say „floor raised to deliver it". The controller floor was not raised to 0.246.0. The release was validated on one guest; raising a floor above golden 0.245.0 would roll it to Peti's box as well, and R-552 (my own gap) was found during validation. One line of evidence is not a fleet decision made unattended — the operator's to take, and cheap to take.
Observations
- An interrupted-restore notice for a removed app never clears. FILED: R-552
- No Tier-0 box can exercise the paused + agent-connected escrow state. FILED: R-551
- Controller
handler.goand agentcontrollersupervisor_test.gowere notgofmt-clean before this task. NOT-A-FINDING: pre-existing formatting only; reformatting unrelated lines was left out to keep the diffs reviewable.