Files
felhom.eu/REPORT-chaos-fixes-2026-09-17.md
admin 851198af7f
gates / gates (push) Successful in 22s
golden 0.246.0 baked and vouched with agent 0.132.0; controller floor 0.246.0 (operator decisions)
Decision 1 delivered: floor 0.246.0 served, the N100 on 0.246.0 within seconds;
Peti's box is DOWN on the hub and receives it when it reports.

Decision 2: the agent vouch was refused by R-120 until a newer golden existed.
On the operator's choice, golden 0.246.0 was baked (sha 05b7559d, amd64, all
markers, token leak 0 with control 1, registry 200 before teardown) and vouched
together with agent 0.132.0. golden_currency_gate: WAIVED -> OK. Bake evidence
filed where the gate and runbook read it: tests/golden-0.246.0-2026-09-17/.

Recorded, none reaching the registry: a first attempt on the arm64 template
(my version sort), a self-matching pkill, and an OOM-killed watcher whose
post-bake steps were done by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-17 12:25:52 +02:00

8.7 KiB
Raw Permalink Blame History

REPORT — the chaos-night fixes: the quiet alarm, the restore record, the recovery-code timing, the slow crash loop

2026-09-17. R-549, R-550, R-546, R-539. Per-repo detail: felhom-controller/REPORT.md (v0.246.0), felhom-agent/REPORT.md (v0.132.0). This file covers felhom.eu (hub v0.117.0, the setting, the documents) and the whole task. Written as REPORT-<topic>.md because REPORT.md holds the earlier session's work.

Claims in the brief that turned out wrong — named first

  1. „Readiness = escrow.pbs_storage_id set." The agent's preflight ok covers five blocking items; pbs_storage_id is the one R-546's box showed. The controller reads the combined ok.
  2. „The backup tiers persist their records atomically." Checked at felhom-controller/controller/internal/settings/settings.go:727-752 — TRUE (tmp + rename, .bak recovery).
  3. „Ruling A is a config change plus a document line, not code." WRONG. hub/internal/web/rollup.go controllerStatus hardcoded 30 m / 1 h while both checkers and hostStatus read the config — the dashboard would have turned amber 15 minutes before the alarm could fire. Hub v0.117.0 fixes it.
  4. „Verify from the log line that the checker runs with 45 m." No log line printed the threshold. Hub v0.117.0 adds it to both „checker initialized" lines.
  5. „The failure shows a raw error." Not in a browser — the page hid its start form behind the checklist. The raw -storage stderr came from the chaos-night harness calling the API directly.
  6. „Emit the event from the agent" / „make the interval configurable for the test." The agent has no event channel (the hub mints from heartbeat timestamps); and no test interval was needed — five real kills 8 minutes apart reached the production threshold in the session.

1. Confirmed baselines (re-verified at start)

felhom-controller 714d5bce0920 v0.245.0 (MinAgent 0.131.0) · felhom-agent e98b857684f4 v0.131.0 · felhom.eu ea25c6c5b1f1 hub v0.116.0 · highest row R-550 · golden waiver to 2026-09-27.

2. Files (felhom.eu)

hub/internal/web/{rollup.go,configs.go,server.go,rollup_test.go}, hub/internal/monitor/{controller_supervisor.go,controller_supervisor_test.go,staleness.go,host_staleness.go}, hub/internal/notify/{dispatcher.go,templates.go}, hub/internal/api/{handler.go,chaosnight_events_test.go}, hub/CHANGELOG.md, manifests/hub.yaml, .claude/rules/hub.md, CONTEXT.md, STATUS.md, documentation/architecture/{03-host-agent.md,05-hub-architecture.md,07-backup-architecture.md,08-alarm-ladder.md}, documentation/runbooks/VOLUNTEER-first-hour.md, documentation/backlog/{OPEN-ITEMS.md,CLOSED-ITEMS.md}, documentation/audits/evidence-chaos-fixes-2026-09-17/ (6 files), this report.

3. Commits

felhom.eu: 469bfa5 (chaos-night leftover evidence), 37ae31f (hub v0.117.0), 06334e1 (manifest: 0.117.0 + 45m), 3c1882a (docs, register, evidence), and the commit carrying this report. felhom-agent: 18d03bd (v0.132.0), release-record CHANGELOG commit, 77cd70f (REPORT). Tag v0.132.0. felhom-controller: 0fe315b (v0.246.0), 29e2acb (REPORT).

4–5. Tests

Hub: go build/vet/test ./... green, 18 packages. Agent: 30 packages. Controller: 28 packages. Red-proofs, each seen failing then passing — hub: status follows threshold; slow-crashloop movement; operator-only registration; Hungarian subject. Agent: no counter; once-per-24h guard; persistence. Controller: restore record across restart; startup helper; main() wiring; restore page card; bar held back; waiting card; start refusal. Two of my own test mistakes corrected and recorded (a body-based fallback check; a -storage needle matching a menu id).

6. Deployed versions

Hub 0.117.0 (ArgoCD Synced/Healthy, revision == manifest HEAD, pod ready). Agent 0.132.0 on demo-hp and the N100 (signed jobs, committed). Controller 0.246.0 on demo-hp guest 9201. Proof lines in the evidence directory.

7. NOT live-validated

  • R-546's readiness branches (bar held back, waiting card, start refusal) — no Tier-0 box is paused AND agent-connected. R-551.
  • The escrow ceremony passing once ready — deliberately not run: the hub keeps one escrow per host, and a ceremony on demo-hp would supersede the standing box's escrow.
  • Controller 0.246.0 is on 9201 only; the fleet floor was not raised (below).

8. Evidence copied off before each teardown

Yes. B.4(a) log written by the run itself to DooPlex before teardown; teardown logged separately; Part C logs off the box as they ran.

9. Teardown — three layers

  • Machine: throwaway homebox on 9201 removed with data and backups (0 containers, 0 volumes); 10 of 10 standing apps up; password and scripts shredded in the guest.
  • Host: nothing provisioned; demo-hp agent config untouched (the pbs_storage_id removal was considered and not done).
  • Hub: no host or customer record created or deleted; the test events stay as history. Stated effect: 9201's slow crash-loop stays raised until 2026-09-18 09:22Z and cannot mail again before then.

An error of mine, caught by a gate before it was pushed

Closing R-539, I wrote open(CLOSED-ITEMS.md,'w').write(open(CLOSED-ITEMS.md).read() + row). Python opens — and empties — the file for writing before it reads it, so the closed register fell from 215 rows to 1. The pre-push instructions gate refused the push: citations of closed rows (R-549, and R-320 in unprompted-work.md) suddenly pointed at nothing. Nothing damaged was pushed. The file was restored from pushed commit 3c1882a and R-539 appended with read-then-write; the diff against 3c1882a is exactly one added line, and the open register exactly one removed line. The earlier closures (R-546/R-549/R-550) used read-then-write and were intact. Lesson kept: never open a file for writing in the same expression that reads it.

A recommendation not followed, with its reason

The brief's rules say „floor raised to deliver it". The controller floor was not raised to 0.246.0. The release was validated on one guest; raising a floor above golden 0.245.0 would roll it to Peti's box as well, and R-552 (my own gap) was found during validation. One line of evidence is not a fleet decision made unattended — the operator's to take, and cheap to take.

Observations

  1. An interrupted-restore notice for a removed app never clears. FILED: R-552
  2. No Tier-0 box can exercise the paused + agent-connected escrow state. FILED: R-551
  3. Controller handler.go and agent controllersupervisor_test.go were not gofmt-clean before this task. NOT-A-FINDING: pre-existing formatting only; reformatting unrelated lines was left out to keep the diffs reviewable.

Addendum — the operator's two „yes" answers, carried out (2026-09-17, afternoon)

1. Controller floor raised to 0.246.0 (declared MinAgent 0.131.0). Hub: Global controller-version floor set to "0.246.0" and managed floor SERVED for demo-felhom … from declared; the N100 ran 0.246.0 within seconds (at/above floor 0.246.0 (we are 0.246.0)). demo-hp runs 0.246.0 (it has a per-customer override at 0.243.0). Peti's box is DOWN on the hub (last controller 0.115.0) and receives it when it reports. Evidence: audits/evidence-chaos-fixes-2026-09-17/decision1-floor-0246.txt. The „recommendation not followed" above is therefore superseded by the operator's decision.

2. Agent 0.132.0 vouched — which required a golden. The first vouch was refused by the hub's R-120 gate (golden 0.245.0 is older than the newest controller the fleet reports (0.246.0)); nothing was stored. Asked, the operator chose to bake. Golden 0.246.0 baked by RUNBOOK-manual-build §4.1, sha 05b7559d…, amd64, all markers, token leak 0 (control 1), registry 200 before teardown; then vouched together with agent 0.132.0: Artifact manifest set: agent=0.132.0 golden=0.246.0 min_agent="0.131.0" wrapper_sha=true, read back exactly. golden_currency_gate.py moved from WAIVED to OK. Evidence: tests/golden-0.246.0-2026-09-17/.

Mistakes of mine on the way, none reaching the registry: the first bake attempt used the arm64 template (my version sort), aborted before anything was built; stopping it, a self-matching pkill killed my own shell; the second attempt's watcher was killed for low memory and its post-bake steps were done by hand.

Observation. 4. golden_currency_gate.py reported the newest bake as 0.242.0 although 0.243.0–0.245.0 were baked and published, because those bakes filed their logs under audits/ rather than documentation/tests/golden-<ver>-<date>/. NOT-A-FINDING: a recording slip in earlier bakes (including 0.245.0, mine), not a gate defect — the gate reads the place the runbook names; 0.246.0 is filed there and the gate reads it.