chaos-night fixes: R-539 closed PROVEN-LIVE, the morning note, the report
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
R-539 closed: five real controller kills on demo-hp 9201 with the production 24 h window raised controller_slow_crashloop, and exactly one operator mail arrived (09:29:40Z). The fast brake never armed. Register 215 -> 213 open (R-551, R-552 filed; R-539, R-546, R-549, R-550 closed). unproven.py: 35 of 55 not walked, no number moved. The report names the brief's wrong claims first and one recommendation not followed: the controller floor was not raised - validated on one guest, a gap of my own found during validation, and a floor above the golden reaches Peti's box too. The operator's call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,33 @@
|
||||
# Part C LIVE - the slow crash loop with the PRODUCTION 24 h window (no test-only interval)
|
||||
# 9201 on demo-hp, agent 0.132.0. Restart #1 already persisted from proof A (08:52:29Z).
|
||||
# Four more kills, 8 min apart: any 15-minute window holds at most two restarts, so the fast brake never arms.
|
||||
2026-09-17T08:57:38Z persisted before: {"restarts":["2026-09-17T08:52:29.969961515Z"],"slow_crashloop_since":"0001-01-01T00:00:00Z"}
|
||||
2026-09-17T08:57:39Z KILL for restart #2: felhom-controller
|
||||
2026-09-17T08:58:36Z controller back: running after 56 s (not by hand)
|
||||
Sep 17 10:58:32 demo-hp felhom-agent[1195783]: time=2026-09-17T10:58:32.336+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=2
|
||||
2026-09-17T09:05:40Z KILL for restart #3: felhom-controller
|
||||
2026-09-17T09:06:37Z controller back: running after 55 s (not by hand)
|
||||
Sep 17 11:06:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:06:32.247+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=3
|
||||
2026-09-17T09:13:42Z KILL for restart #4: felhom-controller
|
||||
2026-09-17T09:14:33Z controller back: running after 50 s (not by hand)
|
||||
Sep 17 11:14:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:14:32.222+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=4
|
||||
2026-09-17T09:21:43Z KILL for restart #5: felhom-controller
|
||||
2026-09-17T09:22:35Z controller back: running after 50 s (not by hand)
|
||||
Sep 17 11:22:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:22:32.184+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" restarts_24h=5
|
||||
Sep 17 11:22:32 demo-hp felhom-agent[1195783]: time=2026-09-17T11:22:32.184+02:00 level=WARN msg="controller-supervisor: SLOW CRASH-LOOP — the controller keeps dying; still restarting it, raising controller_slow_crashloop" vmid=9201 restarts_24h=5 window=24h0m0s threshold=5
|
||||
2026-09-17T09:22:35Z persisted after: {"restarts":["2026-09-17T08:52:29.969961515Z","2026-09-17T08:58:30.062356117Z","2026-09-17T09:06:29.97157901Z","2026-09-17T09:14:29.970146017Z","2026-09-17T09:22:29.963142311Z"],"slow_crashloop_since":"2026-09-17T09:22:29.963142311Z"}
|
||||
2026-09-17T09:22:35Z waiting for the hub to mint controller_slow_crashloop (the agent's heartbeat carries it) ...
|
||||
2026-09-17T09:29:50Z hub customer page mentions of controller_slow_crashloop: 3
|
||||
warning controller_slow_crashloop Host demo-hp-bb76ea guest 9201: the controller keeps dying — the agent restarted it 5 times in 24 hours (each one too far apart for the 15-minute brake). It is still being restarted; this is the warning that it will not stay up. Last reason: controller container exited on 2 consecutive sweeps hub
|
||||
hub log: 2026/09/17 11:29:40 [INFO] Controller supervisor: controller_slow_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps)
|
||||
hub log: 2026/09/17 11:29:40 [INFO] Operator email sent for demo-hp/controller_slow_crashloop
|
||||
|
||||
## DELIVERED, not just "sent" - the mailbox itself (read 2026-09-17T09:31Z)
|
||||
exactly ONE mail for the five restarts:
|
||||
2026-09-17T09:29:40Z monitoring@felhom.eu -> admin@felhom.eu
|
||||
"[Felhom] ⚠️ demo-hp: controller_slow_crashloop"
|
||||
"Customer: demo-hp Event: controller_slow_crashloop Severity: warning ... Host demo-hp-bb76ea guest 9201:
|
||||
the controller keeps dying — the agent restarted it 5 times in 24 hours ..."
|
||||
## After: 9201 controller 0.246.0 Up (healthy), 24 containers running.
|
||||
## Standing effect on the demo box, stated: slow_crashloop stays raised on 9201 until 09:22Z tomorrow and
|
||||
## the counter holds 5 restarts; it cannot re-mail inside that window (moves at most once per 24 h).
|
||||
@@ -321,3 +321,4 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
| **R-546** | **The first-hour guide and the reminder bar sent the household to create their recovery code before the box could (P2).** Closed in controller **v0.246.0** (`0fe315b`): the bar consults the agent's OWN preflight `ok` (all five blocking items, not a copy of `pbs_storage_id`), cached 60 s, probed only while paused, held back while not ready; `/backup/escrow` shows a waiting card that polls and reloads; `POST /api/escrow/start` refuses 409 before staging (the direct path that produced the raw `-storage` stderr); unknown readiness keeps the bar. The guide moves the step after the first apps: „amikor a sárga sáv megjelenik”. Red-proofs: bar held back, waiting card, start refusal; controls ready and unknown. **Proven by tests through ServeHTTP, NOT live** — no Tier-0 box is paused and agent-connected (**R-551**); chaos night measured live the ~17-minute red window and its self-heal. | **CLOSED 2026-09-17 — PROVEN (tests); live walk owed by R-551** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-549** | **The staleness alarm's budget was two report cycles, so one failed push spent all of it (P2).** Closed by operator ruling A (2026-09-17): `alerting.stale_threshold` 30 m → **45 m**, `node_down`/`host_down` at 90 m (`manifests/hub.yaml`, commit `06334e1`). The dashboard's customer status hardcoded 30 m / 1 h and would have disagreed with the alarms, so hub **v0.117.0** (`37ae31f`) makes `controllerStatus` read the same value — red-proof `TestControllerStatus_FollowsConfiguredThreshold` (report 40m old: status warn, want ok). **Proven live:** the running hub printed `node_stale after 45m0s, node_down after 1h30m0s` and `host_stale after 45m0s, host_down after 1h30m0s` at 08:22Z. **Reasoning kept:** the threshold is configuration and every reader — both checkers, host status, customer status — reads the one value; a dead box now pages 15 minutes later, a cost the ruling accepts. `audits/evidence-chaos-fixes-2026-09-17/partA-hub-45m.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-550** | **The restore record was in-memory only: after the machine stopped, nothing told the household their restore did not finish (P2).** Closed in controller **v0.246.0** (`0fe315b`) by operator ruling „fix” — a reversal of the in-memory design for the restore record only (cooldowns stay in memory): `restore-status.json` in DataDir, atomic at both ends of an op; a record still running at startup becomes a failed, interrupted result per app, shown on `/backups/restore` until that app's next restore, raised once as `restore_interrupted` (hub v0.117.0, household). Red-proofs: record across restart (`StartedAt:0001-01-01`), `main()` wiring (AST), startup helper, page card. **Proven live on demo-hp 9201:** a throwaway homebox restore killed 2 s in; after the supervisor's restart the status read `ok:false … megszakadt … interrupted:true`, the card showed, the event reached the hub (HTTP 200, stored under demo-hp); a second restore cleared the card. **Known gap filed: R-552** (a removed app keeps its notice). `audits/evidence-chaos-fixes-2026-09-17/partB4a-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 06334e1:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-539** | **The restart brake caught a FAST crash loop and was blind to a SLOW one (P3).** Closed by operator ruling 3 of 2026-09-16 in agent **v0.132.0** (`18d03bd`, tag `v0.132.0`, sha256 `4afe8157…`) + hub **v0.117.0** (`37ae31f`): beside the unchanged 3-in-15 brake, restarts in the last 24 h, persisted per guest; at the fifth `slow_crashloop_since` moves (at most once per 24 h) and the hub mints `controller_slow_crashloop` (warning, operator-only). Deliberate kills count. Red-proofs: no counter; once-per-24h guard removed; save removed; negative control 7 h apart. Delivered by operator-signed `agent_update` to demo-hp and the N100 (ruling 1 of 2026-09-16), both logging `slow_crashloop_max=5 slow_crashloop_window=24h0m0s`. **Proven live with the PRODUCTION window (no test-only interval):** five real controller kills on demo-hp 9201, ~8 min apart, each restarted by the agent; the fast brake never armed; at #5 `SLOW CRASH-LOOP … restarts_24h=5`; the hub minted the event and exactly ONE operator mail arrived (09:29:40Z). **Reasoning kept:** the slow record is persisted and the fast one is not, because persisting a give-up could outlive the fix while a counter that only warns cannot. `audits/evidence-chaos-fixes-2026-09-17/partC-*.txt` | **CLOSED 2026-09-17 — PROVEN-LIVE** | full text: `git show 3c1882a:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -721,7 +721,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. **THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled:** `DatastorePowerUser` carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries `Datastore.Modify` is `DatastoreAdmin`. Applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` ONLY; the per-customer `DatastoreBackup` entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt`. **CLOSED 2026-09-16 — the grant works, proven END TO END on a fresh box.** After the narrow grant (DatastoreAdmin for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only — `DatastorePowerUser` was measured to carry Backup+Prune and PBS has no custom roles), a newly installed box for the same rebuilt customer hit the very refusal this row describes, by itself: „pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit Re-issue PBS credentials action" (19:01 CEST). Pressing that action then SUCCEEDED: „tenantsync: reissue ok … token_id=felhom@pbs!tester-1", „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, gen 2; fresh consume-once secret stored)", „pbs token secret consumed by host … (single-use)" — no permission error. This morning the identical action returned „missing Datastore.Modify … status 255" → 502. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt` and `phaseC-reissue.txt`. | **CLOSED 2026-09-16 — grant given and proven end to end** |
|
||||
| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom.<domain> címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. **CLOSED 2026-09-16 — shipped in ISO 1.28.0, published the same day on the operator's explicit yes.** `felhom-bootstrap.sh` prints `print_bound_banner` the moment the bind delivery lands: „a doboz össze van kötve", „a beállítás magától folytatódik", „ezen a gépen nincs több teendőd" — replacing the pairing code on the console. The payload in the published image is byte-identical to repo HEAD and the string is present in it; the boot menu and the install were walked on that exact file. **What it deliberately does NOT do, recorded rather than implied away:** it does not name the dashboard URL (the one-shot bind delivery carries the customer id, passphrase and mode — not the domain), and it does not reflect the later CLAIM, because this unit has exited by then. **And the new banner was never SEEN on a screen** — the box bound itself while the walk was driving it headlessly, so the proof is the shipped payload plus the gate, not a photograph. | **CLOSED 2026-09-16 — shipped in ISO 1.28.0 (published); on-screen effect not photographed** |
|
||||
| **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks/<app>/app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0.** `app_deploy_started` is emitted beside the 202; `app_deployed` now fires from the async path's own end, and `app_deploy_failed` (warning) replaces the silence an interrupted install used to get. Both new types are registered in `allowedEventTypes` AND `customerMessages`. The accept-time `app.yaml` is deliberately NOT deleted on failure — it is the crash-safe record with `Deployed:false` and it holds the settings the customer typed; the state every surface reads is `not_deployed`. Red-proofs: the accept-time call back → `TestDeployAcceptance_DoesNotClaimTheAppIsInstalled` fails; the success hook removed → `TestDeployDoneHook_...` fails at „the deploy ended and nothing was told about it". | **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0** |
|
||||
| **R-539** | **[P3-LOW] The restart brake catches a FAST crash loop and is blind to a SLOW one — add a second, slower counter (operator ruling 2026-09-16).** MEASURED 2026-09-16 (R-531): four controller restarts 20 minutes apart, none accumulating, because the budget window is 15 minutes; the only trace is an `info` `controller_restarted_by_agent` event, which mails nobody. A box whose controller dies every 20 minutes is restarted forever and nothing tells the operator. **The ruling:** a second counter — N restarts in 24 h → `controller_slow_crashloop` (warning) — kept beside the existing 3-in-15-minutes brake, which is unchanged. **Not built in the 2026-09-16 task** (its brief said the budget is a design and this task measures it); it is the nightly's to build. Needs: the agent-side counter, the new event type in `allowedEventTypes` + `customerMessages`, and a red-proof that a 20-minute cycle raises it while a healthy box never does. | **READY — rank P3-LOW; owner: CC (agent + hub)** |
|
||||
| **R-540** | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** |
|
||||
| **R-541** | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** |
|
||||
| **R-542** | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** |
|
||||
|
||||
Reference in New Issue
Block a user