Files
felhom.eu/REPORT-chaos-fixes-2026-09-17.md
T
admin 2c96a84f12
gates / gates (push) Successful in 21s
chaos-night fixes: R-539 closed PROVEN-LIVE, the morning note, the report
R-539 closed: five real controller kills on demo-hp 9201 with the production
24 h window raised controller_slow_crashloop, and exactly one operator mail
arrived (09:29:40Z). The fast brake never armed.

Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed).
unproven.py: 35 of 55 not walked, no number moved.

The report names the brief's wrong claims first and one recommendation not
followed: the controller floor was not raised - validated on one guest, a gap
of my own found during validation, and a floor above the golden reaches Peti's
box too. The operator's call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 11:33:11 +02:00

6.7 KiB
Raw Blame History

REPORT — the chaos-night fixes: the quiet alarm, the restore record, the recovery-code timing, the slow crash loop

2026-09-17. R-549, R-550, R-546, R-539. Per-repo detail: felhom-controller/REPORT.md (v0.246.0), felhom-agent/REPORT.md (v0.132.0). This file covers felhom.eu (hub v0.117.0, the setting, the documents) and the whole task. Written as REPORT-<topic>.md because REPORT.md holds the earlier session's work.

Claims in the brief that turned out wrong — named first

  1. „Readiness = escrow.pbs_storage_id set." The agent's preflight ok covers five blocking items; pbs_storage_id is the one R-546's box showed. The controller reads the combined ok.
  2. „The backup tiers persist their records atomically." Checked at felhom-controller/controller/internal/settings/settings.go:727-752 — TRUE (tmp + rename, .bak recovery).
  3. „Ruling A is a config change plus a document line, not code." WRONG. hub/internal/web/rollup.go controllerStatus hardcoded 30 m / 1 h while both checkers and hostStatus read the config — the dashboard would have turned amber 15 minutes before the alarm could fire. Hub v0.117.0 fixes it.
  4. „Verify from the log line that the checker runs with 45 m." No log line printed the threshold. Hub v0.117.0 adds it to both „checker initialized" lines.
  5. „The failure shows a raw error." Not in a browser — the page hid its start form behind the checklist. The raw -storage stderr came from the chaos-night harness calling the API directly.
  6. „Emit the event from the agent" / „make the interval configurable for the test." The agent has no event channel (the hub mints from heartbeat timestamps); and no test interval was needed — five real kills 8 minutes apart reached the production threshold in the session.

1. Confirmed baselines (re-verified at start)

felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) · felhom-agent e98b857684f4 v0.131.0 · felhom.eu ea25c6c5b1f1 hub v0.116.0 · highest row R-550 · golden waiver to 2026-09-27.

2. Files (felhom.eu)

hub/internal/web/{rollup.go,configs.go,server.go,rollup_test.go}, hub/internal/monitor/{controller_supervisor.go,controller_supervisor_test.go,staleness.go,host_staleness.go}, hub/internal/notify/{dispatcher.go,templates.go}, hub/internal/api/{handler.go,chaosnight_events_test.go}, hub/CHANGELOG.md, manifests/hub.yaml, .claude/rules/hub.md, CONTEXT.md, STATUS.md, documentation/architecture/{03-host-agent.md,05-hub-architecture.md,07-backup-architecture.md,08-alarm-ladder.md}, documentation/runbooks/VOLUNTEER-first-hour.md, documentation/backlog/{OPEN-ITEMS.md,CLOSED-ITEMS.md}, documentation/audits/evidence-chaos-fixes-2026-09-17/ (6 files), this report.

3. Commits

felhom.eu: 469bfa5 (chaos-night leftover evidence), 37ae31f (hub v0.117.0), 06334e1 (manifest: 0.117.0 + 45m), 3c1882a (docs, register, evidence), and the commit carrying this report. felhom-agent: 18d03bd (v0.132.0), release-record CHANGELOG commit, 77cd70f (REPORT). Tag v0.132.0. felhom-controller: 0fe315b (v0.246.0), 29e2acb (REPORT).

4–5. Tests

Hub: go build/vet/test ./... green, 18 packages. Agent: 30 packages. Controller: 28 packages. Red-proofs, each seen failing then passing — hub: status follows threshold; slow-crashloop movement; operator-only registration; Hungarian subject. Agent: no counter; once-per-24h guard; persistence. Controller: restore record across restart; startup helper; main() wiring; restore page card; bar held back; waiting card; start refusal. Two of my own test mistakes corrected and recorded (a body-based fallback check; a -storage needle matching a menu id).

6. Deployed versions

Hub 0.117.0 (ArgoCD Synced/Healthy, revision == manifest HEAD, pod ready). Agent 0.132.0 on demo-hp and the N100 (signed jobs, committed). Controller 0.246.0 on demo-hp guest 9201. Proof lines in the evidence directory.

7. NOT live-validated

  • R-546's readiness branches (bar held back, waiting card, start refusal) — no Tier-0 box is paused AND agent-connected. R-551.
  • The escrow ceremony passing once ready — deliberately not run: the hub keeps one escrow per host, and a ceremony on demo-hp would supersede the standing box's escrow.
  • Controller 0.246.0 is on 9201 only; the fleet floor was not raised (below).

8. Evidence copied off before each teardown

Yes. B.4(a) log written by the run itself to DooPlex before teardown; teardown logged separately; Part C logs off the box as they ran.

9. Teardown — three layers

  • Machine: throwaway homebox on 9201 removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up; password and scripts shredded in the guest.
  • Host: nothing provisioned; demo-hp agent config untouched (the pbs_storage_id removal was considered and not done).
  • Hub: no host or customer record created or deleted; the test events stay as history. Stated effect: 9201's slow crash-loop stays raised until 2026-09-18 09:22Z and cannot mail again before then.

An error of mine, caught by a gate before it was pushed

Closing R-539, I wrote open(CLOSED-ITEMS.md,'w').write(open(CLOSED-ITEMS.md).read() + row). Python opens — and empties — the file for writing before it reads it, so the closed register fell from 215 rows to 1. The pre-push instructions gate refused the push: citations of closed rows (R-549, and R-320 in unprompted-work.md) suddenly pointed at nothing. Nothing damaged was pushed. The file was restored from pushed commit 3c1882a and R-539 appended with read-then-write; the diff against 3c1882a is exactly one added line, and the open register exactly one removed line. The earlier closures (R-546/R-549/R-550) used read-then-write and were intact. Lesson kept: never open a file for writing in the same expression that reads it.

A recommendation not followed, with its reason

The brief's rules say „floor raised to deliver it". The controller floor was not raised to 0.246.0. The release was validated on one guest; raising a floor above golden 0.245.0 would roll it to Peti's box as well, and R-552 (my own gap) was found during validation. One line of evidence is not a fleet decision made unattended — the operator's to take, and cheap to take.

Observations

  1. An interrupted-restore notice for a removed app never clears. FILED: R-552
  2. No Tier-0 box can exercise the paused + agent-connected escrow state. FILED: R-551
  3. Controller handler.go and agent controllersupervisor_test.go were not gofmt-clean before this task. NOT-A-FINDING: pre-existing formatting only; reformatting unrelated lines was left out to keep the diffs reviewable.