R-534 and R-511 CLOSED — the ep0 grant proven end to end on a fresh box
gates / gates (push) Successful in 20s

The rebuilt-customer case reproduced by itself: the WG-registration hook refused
exactly as R-511 describes and named the Re-issue action. Pressing it then worked —
reissue ok, pbsdr ADOPTED (gen 2), and the box consumed the single-use secret two
seconds later. No permission error.

This morning the identical action returned „missing Datastore.Modify … status 255"
and a 502. The only change in between is the narrow grant on ep0, and the narrowest
role was measured rather than recalled: DatastorePowerUser carries Backup+Prune only,
and PBS has no custom roles.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 19:26:54 +02:00
parent 63e2de9b9e
commit a73abf04db
3 changed files with 73 additions and 2 deletions
+2 -2
View File
@@ -710,7 +710,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-507** | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** |
| **R-508** | **[P2-MEDIUM] Customer `tester-1` has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer.** MEASURED 2026-09-14: the edit form's `email` value is empty; on bind the hub logged `[ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one`. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. **What it needs:** the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning)** |
| **R-509** | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). **SHIPPED hub v0.114.0 (2026-09-15), NOT YET PROVEN BY A REAL MAIL:** triggers added — e-mail set/changed on a customer with no box, and host delete — each re-checking no bound host; every send recorded as `selfbind_link_sent` and shown on the Setup tab. Unit-proven with a red-proof (`TestSelfBind_EmailSetOnWaitingCustomerSendsLink`). The live check with the Gmail-read mailbox was NOT run: it needs a throwaway customer, and deleting one runs the RESET cascade (ep0 `deprovision` + Cloudflare), which is fenced without the operator's word. **Closes on:** one real mail from either trigger. | **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** |
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. **SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes **CONSEQUENCE MEASURED 2026-09-16:** the code fix shipped for this row is sound and currently INERT - the adopt path runs and the ENDPOINT refuses it for a missing `Datastore.Modify` grant (**R-534**). On the drill box the result is that a fresh install has NO off-site tier at all, and read-only listing of ep0 shows `ns/tester-1/ct` empty before and after the whole drill. This row stays open until R-534's grant is given and a re-issue is seen to succeed. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. **SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes **CONSEQUENCE MEASURED 2026-09-16:** the code fix shipped for this row is sound and currently INERT - the adopt path runs and the ENDPOINT refuses it for a missing `Datastore.Modify` grant (**R-534**). On the drill box the result is that a fresh install has NO off-site tier at all, and read-only listing of ep0 shows `ns/tester-1/ct` empty before and after the whole drill. This row stays open until R-534's grant is given and a re-issue is seen to succeed. **CLOSED 2026-09-16 — the adopt path is PROVEN LIVE.** The rebuilt-customer case reproduced by itself on a fresh box (the WG-registration hook refused exactly as this row describes and named the Re-issue action), and the adopt then completed: reissue ok → pbsdr ADOPTED (gen 2) → the box consumed the single-use secret two seconds later. The code shipped 2026-09-15 was sound; what made it inert was the ep0 grant (R-534), now given. | **CLOSED 2026-09-16 — proven live on a fresh box** |
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. | **READY — rank P3-LOW; owner: CC (controller + catalog)** |
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
@@ -718,7 +718,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. **THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled:** `DatastorePowerUser` carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries `Datastore.Modify` is `DatastoreAdmin`. Applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` ONLY; the per-customer `DatastoreBackup` entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt`. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (ep0 grant) · CC (verify after)** |
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. **THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled:** `DatastorePowerUser` carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries `Datastore.Modify` is `DatastoreAdmin`. Applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` ONLY; the per-customer `DatastoreBackup` entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt`. **CLOSED 2026-09-16 — the grant works, proven END TO END on a fresh box.** After the narrow grant (DatastoreAdmin for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only — `DatastorePowerUser` was measured to carry Backup+Prune and PBS has no custom roles), a newly installed box for the same rebuilt customer hit the very refusal this row describes, by itself: „pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit Re-issue PBS credentials action" (19:01 CEST). Pressing that action then SUCCEEDED: „tenantsync: reissue ok … token_id=felhom@pbs!tester-1", „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, gen 2; fresh consume-once secret stored)", „pbs token secret consumed by host … (single-use)" — no permission error. This morning the identical action returned „missing Datastore.Modify … status 255" → 502. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt` and `phaseC-reissue.txt`. | **CLOSED 2026-09-16 — grant given and proven end to end** |
| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom.<domain> címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. | **READY — rank P2-MEDIUM; owner: CC (ISO payload `felhom-bootstrap.sh` — ships with the next ISO)** |
| **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks/<app>/app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0.** `app_deploy_started` is emitted beside the 202; `app_deployed` now fires from the async path's own end, and `app_deploy_failed` (warning) replaces the silence an interrupted install used to get. Both new types are registered in `allowedEventTypes` AND `customerMessages`. The accept-time `app.yaml` is deliberately NOT deleted on failure — it is the crash-safe record with `Deployed:false` and it holds the settings the customer typed; the state every surface reads is `not_deployed`. Red-proofs: the accept-time call back → `TestDeployAcceptance_DoesNotClaimTheAppIsInstalled` fails; the success hook removed → `TestDeployDoneHook_...` fails at „the deploy ended and nothing was told about it". | **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0** |
| **R-539** | **[P3-LOW] The restart brake catches a FAST crash loop and is blind to a SLOW one — add a second, slower counter (operator ruling 2026-09-16).** MEASURED 2026-09-16 (R-531): four controller restarts 20 minutes apart, none accumulating, because the budget window is 15 minutes; the only trace is an `info` `controller_restarted_by_agent` event, which mails nobody. A box whose controller dies every 20 minutes is restarted forever and nothing tells the operator. **The ruling:** a second counter — N restarts in 24 h → `controller_slow_crashloop` (warning) — kept beside the existing 3-in-15-minutes brake, which is unchanged. **Not built in the 2026-09-16 task** (its brief said the budget is a design and this task measures it); it is the nightly's to build. Needs: the agent-side counter, the new event type in `allowedEventTypes` + `customerMessages`, and a red-proof that a 20-minute cycle raises it while a healthy box never does. | **READY — rank P3-LOW; owner: CC (agent + hub)** |