# REPORT — the chaos-night fixes: the quiet alarm, the restore record, the recovery-code timing, the slow crash loop **2026-09-17.** R-549, R-550, R-546, R-539. Per-repo detail: `felhom-controller/REPORT.md` (v0.246.0), `felhom-agent/REPORT.md` (v0.132.0). This file covers felhom.eu (hub v0.117.0, the setting, the documents) and the whole task. Written as `REPORT-.md` because `REPORT.md` holds the earlier session's work. ## Claims in the brief that turned out wrong — named first 1. **„Readiness = `escrow.pbs_storage_id` set."** The agent's preflight `ok` covers **five** blocking items; `pbs_storage_id` is the one R-546's box showed. The controller reads the combined `ok`. 2. **„The backup tiers persist their records atomically."** Checked at `felhom-controller/controller/internal/settings/settings.go:727-752` — **TRUE** (tmp + rename, `.bak` recovery). 3. **„Ruling A is a config change plus a document line, not code."** WRONG. `hub/internal/web/rollup.go` `controllerStatus` hardcoded 30 m / 1 h while both checkers and `hostStatus` read the config — the dashboard would have turned amber 15 minutes before the alarm could fire. Hub v0.117.0 fixes it. 4. **„Verify from the log line that the checker runs with 45 m."** No log line printed the threshold. Hub v0.117.0 adds it to both „checker initialized" lines. 5. **„The failure shows a raw error."** Not in a browser — the page hid its start form behind the checklist. The raw `-storage` stderr came from the chaos-night harness calling the API directly. 6. **„Emit the event from the agent" / „make the interval configurable for the test."** The agent has no event channel (the hub mints from heartbeat timestamps); and no test interval was needed — five real kills 8 minutes apart reached the production threshold in the session. ## 1. Confirmed baselines (re-verified at start) felhom-controller `714d5bce0920` v0.245.0 (MinAgent 0.131.0) · felhom-agent `e98b857684f4` v0.131.0 · felhom.eu `ea25c6c5b1f1` hub v0.116.0 · highest row R-550 · golden waiver to 2026-09-27. ## 2. Files (felhom.eu) `hub/internal/web/{rollup.go,configs.go,server.go,rollup_test.go}`, `hub/internal/monitor/{controller_supervisor.go,controller_supervisor_test.go,staleness.go,host_staleness.go}`, `hub/internal/notify/{dispatcher.go,templates.go}`, `hub/internal/api/{handler.go,chaosnight_events_test.go}`, `hub/CHANGELOG.md`, `manifests/hub.yaml`, `.claude/rules/hub.md`, `CONTEXT.md`, `STATUS.md`, `documentation/architecture/{03-host-agent.md,05-hub-architecture.md,07-backup-architecture.md,08-alarm-ladder.md}`, `documentation/runbooks/VOLUNTEER-first-hour.md`, `documentation/backlog/{OPEN-ITEMS.md,CLOSED-ITEMS.md}`, `documentation/audits/evidence-chaos-fixes-2026-09-17/` (6 files), this report. ## 3. Commits felhom.eu: `469bfa5` (chaos-night leftover evidence), `37ae31f` (hub v0.117.0), `06334e1` (manifest: 0.117.0 + 45m), `3c1882a` (docs, register, evidence), and the commit carrying this report. felhom-agent: `18d03bd` (v0.132.0), release-record CHANGELOG commit, `77cd70f` (REPORT). Tag `v0.132.0`. felhom-controller: `0fe315b` (v0.246.0), `29e2acb` (REPORT). ## 4–5. Tests Hub: `go build/vet/test ./...` green, 18 packages. Agent: 30 packages. Controller: 28 packages. **Red-proofs, each seen failing then passing** — hub: status follows threshold; slow-crashloop movement; operator-only registration; Hungarian subject. Agent: no counter; once-per-24h guard; persistence. Controller: restore record across restart; startup helper; `main()` wiring; restore page card; bar held back; waiting card; start refusal. Two of my own test mistakes corrected and recorded (a body-based fallback check; a `-storage` needle matching a menu id). ## 6. Deployed versions Hub **0.117.0** (ArgoCD Synced/Healthy, revision == manifest HEAD, pod ready). Agent **0.132.0** on demo-hp and the N100 (signed jobs, committed). Controller **0.246.0** on demo-hp guest 9201. Proof lines in the evidence directory. ## 7. NOT live-validated - R-546's readiness branches (bar held back, waiting card, start refusal) — no Tier-0 box is paused AND agent-connected. **R-551.** - The escrow ceremony passing once ready — deliberately **not run**: the hub keeps one escrow per host, and a ceremony on demo-hp would supersede the standing box's escrow. - Controller 0.246.0 is on 9201 only; **the fleet floor was not raised** (below). ## 8. Evidence copied off before each teardown Yes. B.4(a) log written by the run itself to DooPlex before teardown; teardown logged separately; Part C logs off the box as they ran. ## 9. Teardown — three layers - **Machine:** throwaway `homebox` on 9201 removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up; password and scripts shredded in the guest. - **Host:** nothing provisioned; demo-hp agent config untouched (the `pbs_storage_id` removal was considered and not done). - **Hub:** no host or customer record created or deleted; the test events stay as history. Stated effect: 9201's slow crash-loop stays raised until 2026-09-18 09:22Z and cannot mail again before then. ## An error of mine, caught by a gate before it was pushed Closing R-539, I wrote `open(CLOSED-ITEMS.md,'w').write(open(CLOSED-ITEMS.md).read() + row)`. Python opens — and empties — the file for writing **before** it reads it, so the closed register fell from **215 rows to 1**. The pre-push `instructions` gate refused the push: citations of closed rows (R-549, and R-320 in `unprompted-work.md`) suddenly pointed at nothing. **Nothing damaged was pushed.** The file was restored from pushed commit `3c1882a` and R-539 appended with read-then-write; the diff against `3c1882a` is exactly one added line, and the open register exactly one removed line. The earlier closures (R-546/R-549/R-550) used read-then-write and were intact. Lesson kept: never open a file for writing in the same expression that reads it. ## A recommendation not followed, with its reason The brief's rules say „floor raised to deliver it". **The controller floor was not raised to 0.246.0.** The release was validated on one guest; raising a floor above golden 0.245.0 would roll it to Peti's box as well, and R-552 (my own gap) was found during validation. One line of evidence is not a fleet decision made unattended — the operator's to take, and cheap to take. ## Observations 1. An interrupted-restore notice for a removed app never clears. **FILED: R-552** 2. No Tier-0 box can exercise the paused + agent-connected escrow state. **FILED: R-551** 3. Controller `handler.go` and agent `controllersupervisor_test.go` were not `gofmt`-clean before this task. **NOT-A-FINDING: pre-existing formatting only; reformatting unrelated lines was left out to keep the diffs reviewable.** ## Addendum — the operator's two „yes" answers, carried out (2026-09-17, afternoon) **1. Controller floor raised to 0.246.0** (declared MinAgent 0.131.0). Hub: `Global controller-version floor set to "0.246.0"` and `managed floor SERVED for demo-felhom … from declared`; the N100 ran 0.246.0 within seconds (`at/above floor 0.246.0 (we are 0.246.0)`). demo-hp runs 0.246.0 (it has a per-customer override at 0.243.0). **Peti's box is DOWN on the hub** (last controller 0.115.0) and receives it when it reports. Evidence: `audits/evidence-chaos-fixes-2026-09-17/decision1-floor-0246.txt`. The „recommendation not followed" above is therefore superseded by the operator's decision. **2. Agent 0.132.0 vouched — which required a golden.** The first vouch was refused by the hub's R-120 gate (`golden 0.245.0 is older than the newest controller the fleet reports (0.246.0)`); nothing was stored. Asked, the operator chose to bake. **Golden 0.246.0** baked by RUNBOOK-manual-build §4.1, sha `05b7559d…`, amd64, all markers, token leak 0 (control 1), registry 200 before teardown; then vouched together with agent 0.132.0: `Artifact manifest set: agent=0.132.0 golden=0.246.0 min_agent="0.131.0" wrapper_sha=true`, read back exactly. `golden_currency_gate.py` moved from WAIVED to **OK**. Evidence: `tests/golden-0.246.0-2026-09-17/`. **Mistakes of mine on the way, none reaching the registry:** the first bake attempt used the **arm64** template (my version sort), aborted before anything was built; stopping it, a self-matching `pkill` killed my own shell; the second attempt's watcher was killed for low memory and its post-bake steps were done by hand. **Observation.** 4. `golden_currency_gate.py` reported the newest bake as 0.242.0 although 0.243.0–0.245.0 were baked and published, because those bakes filed their logs under `audits/` rather than `documentation/tests/golden--/`. **NOT-A-FINDING: a recording slip in earlier bakes (including 0.245.0, mine), not a gate defect — the gate reads the place the runbook names; 0.246.0 is filed there and the gate reads it.**