Decision 1 delivered: floor 0.246.0 served, the N100 on 0.246.0 within seconds; Peti's box is DOWN on the hub and receives it when it reports. Decision 2: the agent vouch was refused by R-120 until a newer golden existed. On the operator's choice, golden 0.246.0 was baked (sha 05b7559d, amd64, all markers, token leak 0 with control 1, registry 200 before teardown) and vouched together with agent 0.132.0. golden_currency_gate: WAIVED -> OK. Bake evidence filed where the gate and runbook read it: tests/golden-0.246.0-2026-09-17/. Recorded, none reaching the registry: a first attempt on the arm64 template (my version sort), a self-matching pkill, and an OOM-killed watcher whose post-bake steps were done by hand. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
8.7 KiB
REPORT — the chaos-night fixes: the quiet alarm, the restore record, the recovery-code timing, the slow crash loop
2026-09-17. R-549, R-550, R-546, R-539. Per-repo detail: felhom-controller/REPORT.md (v0.246.0),
felhom-agent/REPORT.md (v0.132.0). This file covers felhom.eu (hub v0.117.0, the setting, the documents)
and the whole task. Written as REPORT-<topic>.md because REPORT.md holds the earlier session's work.
Claims in the brief that turned out wrong — named first
- „Readiness =
escrow.pbs_storage_idset." The agent's preflightokcovers five blocking items;pbs_storage_idis the one R-546's box showed. The controller reads the combinedok. - „The backup tiers persist their records atomically." Checked at
felhom-controller/controller/internal/settings/settings.go:727-752— TRUE (tmp + rename,.bakrecovery). - „Ruling A is a config change plus a document line, not code." WRONG.
hub/internal/web/rollup.gocontrollerStatushardcoded 30 m / 1 h while both checkers andhostStatusread the config — the dashboard would have turned amber 15 minutes before the alarm could fire. Hub v0.117.0 fixes it. - „Verify from the log line that the checker runs with 45 m." No log line printed the threshold. Hub v0.117.0 adds it to both „checker initialized" lines.
- „The failure shows a raw error." Not in a browser — the page hid its start form behind the
checklist. The raw
-storagestderr came from the chaos-night harness calling the API directly. - „Emit the event from the agent" / „make the interval configurable for the test." The agent has no event channel (the hub mints from heartbeat timestamps); and no test interval was needed — five real kills 8 minutes apart reached the production threshold in the session.
1. Confirmed baselines (re-verified at start)
felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) · felhom-agent e98b857684f4 v0.131.0 ·
felhom.eu ea25c6c5b1f1 hub v0.116.0 · highest row R-550 · golden waiver to 2026-09-27.
2. Files (felhom.eu)
hub/internal/web/{rollup.go,configs.go,server.go,rollup_test.go}, hub/internal/monitor/{controller_supervisor.go,controller_supervisor_test.go,staleness.go,host_staleness.go},
hub/internal/notify/{dispatcher.go,templates.go}, hub/internal/api/{handler.go,chaosnight_events_test.go},
hub/CHANGELOG.md, manifests/hub.yaml, .claude/rules/hub.md, CONTEXT.md, STATUS.md,
documentation/architecture/{03-host-agent.md,05-hub-architecture.md,07-backup-architecture.md,08-alarm-ladder.md},
documentation/runbooks/VOLUNTEER-first-hour.md, documentation/backlog/{OPEN-ITEMS.md,CLOSED-ITEMS.md},
documentation/audits/evidence-chaos-fixes-2026-09-17/ (6 files), this report.
3. Commits
felhom.eu: 469bfa5 (chaos-night leftover evidence), 37ae31f (hub v0.117.0), 06334e1 (manifest: 0.117.0 + 45m), 3c1882a (docs, register, evidence), and the commit carrying this report.
felhom-agent: 18d03bd (v0.132.0), release-record CHANGELOG commit, 77cd70f (REPORT). Tag v0.132.0.
felhom-controller: 0fe315b (v0.246.0), 29e2acb (REPORT).
4–5. Tests
Hub: go build/vet/test ./... green, 18 packages. Agent: 30 packages. Controller: 28 packages.
Red-proofs, each seen failing then passing — hub: status follows threshold; slow-crashloop movement; operator-only registration; Hungarian subject. Agent: no counter; once-per-24h guard; persistence. Controller: restore record across restart; startup helper; main() wiring; restore page card; bar held back; waiting card; start refusal. Two of my own test mistakes corrected and recorded (a body-based fallback check; a -storage needle matching a menu id).
6. Deployed versions
Hub 0.117.0 (ArgoCD Synced/Healthy, revision == manifest HEAD, pod ready). Agent 0.132.0 on demo-hp and the N100 (signed jobs, committed). Controller 0.246.0 on demo-hp guest 9201. Proof lines in the evidence directory.
7. NOT live-validated
- R-546's readiness branches (bar held back, waiting card, start refusal) — no Tier-0 box is paused AND agent-connected. R-551.
- The escrow ceremony passing once ready — deliberately not run: the hub keeps one escrow per host, and a ceremony on demo-hp would supersede the standing box's escrow.
- Controller 0.246.0 is on 9201 only; the fleet floor was not raised (below).
8. Evidence copied off before each teardown
Yes. B.4(a) log written by the run itself to DooPlex before teardown; teardown logged separately; Part C logs off the box as they ran.
9. Teardown — three layers
- Machine: throwaway
homeboxon 9201 removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up; password and scripts shredded in the guest. - Host: nothing provisioned; demo-hp agent config untouched (the
pbs_storage_idremoval was considered and not done). - Hub: no host or customer record created or deleted; the test events stay as history. Stated effect: 9201's slow crash-loop stays raised until 2026-09-18 09:22Z and cannot mail again before then.
An error of mine, caught by a gate before it was pushed
Closing R-539, I wrote open(CLOSED-ITEMS.md,'w').write(open(CLOSED-ITEMS.md).read() + row). Python opens
— and empties — the file for writing before it reads it, so the closed register fell from 215 rows to
1. The pre-push instructions gate refused the push: citations of closed rows (R-549, and R-320 in
unprompted-work.md) suddenly pointed at nothing. Nothing damaged was pushed. The file was restored
from pushed commit 3c1882a and R-539 appended with read-then-write; the diff against 3c1882a is exactly
one added line, and the open register exactly one removed line. The earlier closures (R-546/R-549/R-550)
used read-then-write and were intact. Lesson kept: never open a file for writing in the same expression
that reads it.
A recommendation not followed, with its reason
The brief's rules say „floor raised to deliver it". The controller floor was not raised to 0.246.0. The release was validated on one guest; raising a floor above golden 0.245.0 would roll it to Peti's box as well, and R-552 (my own gap) was found during validation. One line of evidence is not a fleet decision made unattended — the operator's to take, and cheap to take.
Observations
- An interrupted-restore notice for a removed app never clears. FILED: R-552
- No Tier-0 box can exercise the paused + agent-connected escrow state. FILED: R-551
- Controller
handler.goand agentcontrollersupervisor_test.gowere notgofmt-clean before this task. NOT-A-FINDING: pre-existing formatting only; reformatting unrelated lines was left out to keep the diffs reviewable.
Addendum — the operator's two „yes" answers, carried out (2026-09-17, afternoon)
1. Controller floor raised to 0.246.0 (declared MinAgent 0.131.0). Hub: Global controller-version floor set to "0.246.0" and managed floor SERVED for demo-felhom … from declared; the N100 ran 0.246.0 within
seconds (at/above floor 0.246.0 (we are 0.246.0)). demo-hp runs 0.246.0 (it has a per-customer override at
0.243.0). Peti's box is DOWN on the hub (last controller 0.115.0) and receives it when it reports.
Evidence: audits/evidence-chaos-fixes-2026-09-17/decision1-floor-0246.txt. The „recommendation not followed"
above is therefore superseded by the operator's decision.
2. Agent 0.132.0 vouched — which required a golden. The first vouch was refused by the hub's R-120 gate
(golden 0.245.0 is older than the newest controller the fleet reports (0.246.0)); nothing was stored. Asked,
the operator chose to bake. Golden 0.246.0 baked by RUNBOOK-manual-build §4.1, sha 05b7559d…, amd64,
all markers, token leak 0 (control 1), registry 200 before teardown; then vouched together with agent 0.132.0:
Artifact manifest set: agent=0.132.0 golden=0.246.0 min_agent="0.131.0" wrapper_sha=true, read back
exactly. golden_currency_gate.py moved from WAIVED to OK. Evidence: tests/golden-0.246.0-2026-09-17/.
Mistakes of mine on the way, none reaching the registry: the first bake attempt used the arm64
template (my version sort), aborted before anything was built; stopping it, a self-matching pkill killed my
own shell; the second attempt's watcher was killed for low memory and its post-bake steps were done by hand.
Observation. 4. golden_currency_gate.py reported the newest bake as 0.242.0 although 0.243.0–0.245.0
were baked and published, because those bakes filed their logs under audits/ rather than
documentation/tests/golden-<ver>-<date>/. NOT-A-FINDING: a recording slip in earlier bakes (including
0.245.0, mine), not a gate defect — the gate reads the place the runbook names; 0.246.0 is filed there and the
gate reads it.