R-534 and R-511 CLOSED — the ep0 grant proven end to end on a fresh box
gates / gates (push) Successful in 20s
gates / gates (push) Successful in 20s
The rebuilt-customer case reproduced by itself: the WG-registration hook refused exactly as R-511 describes and named the Re-issue action. Pressing it then worked — reissue ok, pbsdr ADOPTED (gen 2), and the box consumed the single-use secret two seconds later. No permission error. This morning the identical action returned „missing Datastore.Modify … status 255" and a 502. The only change in between is the narrow grant on ep0, and the narrowest role was measured rather than recalled: DatastorePowerUser carries Backup+Prune only, and PBS has no custom roles. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -10,3 +10,33 @@
|
||||
## correctly rather than a blocked step. The grant is proven end-to-end by the FRESH BOX in Part E:
|
||||
## its WG registration must provision `felhom-pbs` with no hand — which is exactly the path that
|
||||
## failed on 2026-09-16 with „missing Datastore.Modify". R-534 and R-511 stay open until that run.
|
||||
## 2026-09-16T17:25:34Z PART C.2 — pressing „Re-issue PBS credentials" on the FRESH box (the hub itself asked for it)
|
||||
submitting the customer form to the re-issue action, fields: _csrf, cf_api_token, cf_tunnel_token, customer_id, customer_name, domain, dr_tier, email, git_token, git_username, offsite_box_type, offsite_enabled, offsite_quota_gb, offsite_type, pbsdr_storage_id
|
||||
POST pbsdr-reissue -> http=200
|
||||
hub log right after:
|
||||
2026/09/16 19:25:36 [INFO] tenantsync: reissue ok for tester-1 (ns=tester-1, token_id=felhom@pbs!tester-1; secret withheld from logs)
|
||||
2026/09/16 19:25:36 [INFO] pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, ns tester-1, token_id felhom@pbs!tester-1, gen 2; fresh consume-once secret stored, withheld from logs)
|
||||
2026/09/16 19:25:38 [INFO] host-report from tester-1-33b6a9 (1 guests, 2 storage targets, 1 backups, 0 restore-tests, 0 pbs-snapshots, 11186 bytes)
|
||||
2026/09/16 19:25:38 [INFO] DR-recipe host-half stored for customer tester-1 (host tester-1-33b6a9, v1)
|
||||
2026/09/16 19:25:38 [INFO] pbs token secret consumed by host tester-1-33b6a9 (single-use; value withheld from logs)
|
||||
## R-511 REPRODUCED ON THE FRESH BOX, from the hub's own log (2026-09-16 18:01:19 CEST):
|
||||
## „[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a
|
||||
## PBS token for tester-1 but the hub has no descriptor — use the explicit „Re-issue PBS
|
||||
## credentials" action — save the customer config to retry"
|
||||
## That is R-511's exact shape: a customer whose box was rebuilt keeps the ep0 token, the automatic
|
||||
## hook refuses, and the product names the one action that helps. It happened by itself on a box
|
||||
## installed an hour after the ep0 grant, so the walk handed back the very test that could not be
|
||||
## run this morning (the button renders only once a descriptor exists, and there was no host then).
|
||||
## AND THE ADOPT SUCCEEDED — the ep0 grant proven END TO END (2026-09-16 19:25:36 CEST):
|
||||
## „tenantsync: reissue ok for tester-1 (ns=tester-1, token_id=felhom@pbs!tester-1; secret withheld
|
||||
## from logs)"
|
||||
## „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, ns tester-1, token_id felhom@pbs!tester-1,
|
||||
## gen 2; fresh consume-once secret stored, withheld from logs)"
|
||||
## „pbs token secret consumed by host tester-1-33b6a9 (single-use; value withheld from logs)"
|
||||
## Two seconds later the box reported back and the hub stored the DR-recipe host half.
|
||||
## THE COMPARISON THAT MATTERS: this morning the same action ended with
|
||||
## „Process exited with status 255 (stderr: Error: permission check failed — missing
|
||||
## Datastore.Modify on /datastore/felhom-offsite)" -> HTTP 502, nothing written.
|
||||
## The only thing that changed in between is the narrow grant on ep0 (DatastoreAdmin for the hub's
|
||||
## felhom@pbs at the datastore root). R-534 and R-511 are therefore both closed by a live run, not
|
||||
## by an argument — and the rebuilt-customer case they describe reproduced BY ITSELF on a fresh box.
|
||||
|
||||
@@ -53,3 +53,44 @@
|
||||
## nothing asks the household to perform. `POST /backup/offbox/run` -> 302 and produced no snapshot;
|
||||
## the controller log shows only `offsite-credential-retry`, no restic activity. So on day one the
|
||||
## sentence above promises protection by a copy that does not exist yet.
|
||||
## A TRAP THIS PROJECT ALREADY KNOWS, hit again and recorded: my `grep -oE` with an accented pattern
|
||||
## („Helyre…") died with „exceeds complexity limits" and told me nothing, while two guessed paths
|
||||
## (/backups/escrow, /escrow) 404'd. Parsing the page in Python with the ASCII fragment „Helyre"
|
||||
## found the link at once: /backup/escrow („Helyreállítási kód létrehozása").
|
||||
## The standing rule is „never let an accented pattern gate a conclusion" — the failure mode here
|
||||
## was not a false 0 but a dead tool, and the cost was the same: two wrong guesses.
|
||||
## THE ESCROW CEREMONY PAGE (/backup/escrow), in the household's own words:
|
||||
## „1. Előfeltételek ellenőrzése — Ellenőrzés folyamatban…"
|
||||
## „2. Fontos tudnivalók — A helyreállítási kód a mentései utolsó kulcsa. Pontosan egyszer jelenik
|
||||
## meg — a rendszer sehol nem tárolja, és a Felhom sem ismeri. Ha a szerver megsemmisül, a távoli
|
||||
## mentések CSAK ezzel a kóddal állíthatók vissza."
|
||||
## So the ceremony is customer-facing, honest about what the code is, and shows it exactly once.
|
||||
## That is also why this walk captures it into an out-of-band 0600 file at the moment it appears:
|
||||
## the off-site restore later in this phase cannot happen without it, and nothing can re-issue it.
|
||||
## 2026-09-16T17:2xZ THE ESCROW CEREMONY CANNOT START YET — preflight, verbatim:
|
||||
## agent_supported: true · escrow_state: „pending"
|
||||
## pbs_storage_id NOT OK — „escrow.pbs_storage_id not configured"
|
||||
## dr_tier NOT OK — „DR tier not applied on this host"
|
||||
## age_binary ok — /usr/bin/age
|
||||
## hub_upload ok — hub upload target configured
|
||||
## staged_secret ok — staged secret present
|
||||
## sudo_grant ok — sudo grant listed (list-mode)
|
||||
## So four of six prerequisites are already in place on a box that installed itself an hour ago;
|
||||
## the two red ones are the PBS-DR cascade (host → WG peer → apply), which is precisely what the
|
||||
## ep0 grant of this morning (R-534) exists to unblock. The hub had said at 17:05 that the
|
||||
## descriptor „applies once the cascade is ready (host → WG peer → apply)" — a host now exists.
|
||||
## Checked next on the hub side; whatever it says is the end-to-end verdict on that grant.
|
||||
## THE CASCADE, in the hub's own words (customer page, DR tier widget):
|
||||
## done „host enrolled (tester-1-33b6a9)"
|
||||
## done „WG tunnel peer registered"
|
||||
## waiting „provisions automatically when the WG peer registers"
|
||||
## waiting „ceremony possible once the descriptor is applied on the box"
|
||||
## and the agent's capability list on the SAME page:
|
||||
## „escrow-ceremony critical inactive — customer recovery-code ceremony (controller-driven)
|
||||
## — disabled by configuration"
|
||||
## „pbsdr-create / pbsdr-grant / pbsdr-read inactive — disabled by configuration"
|
||||
## The WG peer really is registered: ep0 shows FIVE wg peers with fresh handshakes, and the hub
|
||||
## pushes them every five minutes („wgsync: pushed 5 peers").
|
||||
## So the chain stops at the APPLY step, because the agent's PBS-DR capabilities are switched off by
|
||||
## configuration — not because ep0 refused anything. This morning's grant (R-534) is therefore still
|
||||
## unproven end-to-end: nothing has yet asked ep0 to do the thing the grant allows.
|
||||
|
||||
@@ -710,7 +710,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-507** | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** |
|
||||
| **R-508** | **[P2-MEDIUM] Customer `tester-1` has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer.** MEASURED 2026-09-14: the edit form's `email` value is empty; on bind the hub logged `[ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one`. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. **What it needs:** the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning)** |
|
||||
| **R-509** | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). **SHIPPED hub v0.114.0 (2026-09-15), NOT YET PROVEN BY A REAL MAIL:** triggers added — e-mail set/changed on a customer with no box, and host delete — each re-checking no bound host; every send recorded as `selfbind_link_sent` and shown on the Setup tab. Unit-proven with a red-proof (`TestSelfBind_EmailSetOnWaitingCustomerSendsLink`). The live check with the Gmail-read mailbox was NOT run: it needs a throwaway customer, and deleting one runs the RESET cascade (ep0 `deprovision` + Cloudflare), which is fenced without the operator's word. **Closes on:** one real mail from either trigger. | **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** |
|
||||
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. **SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes **CONSEQUENCE MEASURED 2026-09-16:** the code fix shipped for this row is sound and currently INERT - the adopt path runs and the ENDPOINT refuses it for a missing `Datastore.Modify` grant (**R-534**). On the drill box the result is that a fresh install has NO off-site tier at all, and read-only listing of ep0 shows `ns/tester-1/ct` empty before and after the whole drill. This row stays open until R-534's grant is given and a re-issue is seen to succeed. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
|
||||
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. **SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes **CONSEQUENCE MEASURED 2026-09-16:** the code fix shipped for this row is sound and currently INERT - the adopt path runs and the ENDPOINT refuses it for a missing `Datastore.Modify` grant (**R-534**). On the drill box the result is that a fresh install has NO off-site tier at all, and read-only listing of ep0 shows `ns/tester-1/ct` empty before and after the whole drill. This row stays open until R-534's grant is given and a re-issue is seen to succeed. **CLOSED 2026-09-16 — the adopt path is PROVEN LIVE.** The rebuilt-customer case reproduced by itself on a fresh box (the WG-registration hook refused exactly as this row describes and named the Re-issue action), and the adopt then completed: reissue ok → pbsdr ADOPTED (gen 2) → the box consumed the single-use secret two seconds later. The code shipped 2026-09-15 was sound; what made it inert was the ep0 grant (R-534), now given. | **CLOSED 2026-09-16 — proven live on a fresh box** |
|
||||
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. | **READY — rank P3-LOW; owner: CC (controller + catalog)** |
|
||||
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
@@ -718,7 +718,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
|
||||
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
|
||||
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
|
||||
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. **THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled:** `DatastorePowerUser` carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries `Datastore.Modify` is `DatastoreAdmin`. Applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` ONLY; the per-customer `DatastoreBackup` entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt`. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (ep0 grant) · CC (verify after)** |
|
||||
| **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. **THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled:** `DatastorePowerUser` carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries `Datastore.Modify` is `DatastoreAdmin`. Applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` ONLY; the per-customer `DatastoreBackup` entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt`. **CLOSED 2026-09-16 — the grant works, proven END TO END on a fresh box.** After the narrow grant (DatastoreAdmin for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only — `DatastorePowerUser` was measured to carry Backup+Prune and PBS has no custom roles), a newly installed box for the same rebuilt customer hit the very refusal this row describes, by itself: „pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit Re-issue PBS credentials action" (19:01 CEST). Pressing that action then SUCCEEDED: „tenantsync: reissue ok … token_id=felhom@pbs!tester-1", „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, gen 2; fresh consume-once secret stored)", „pbs token secret consumed by host … (single-use)" — no permission error. This morning the identical action returned „missing Datastore.Modify … status 255" → 502. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt` and `phaseC-reissue.txt`. | **CLOSED 2026-09-16 — grant given and proven end to end** |
|
||||
| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom.<domain> címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. | **READY — rank P2-MEDIUM; owner: CC (ISO payload `felhom-bootstrap.sh` — ships with the next ISO)** |
|
||||
| **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks/<app>/app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0.** `app_deploy_started` is emitted beside the 202; `app_deployed` now fires from the async path's own end, and `app_deploy_failed` (warning) replaces the silence an interrupted install used to get. Both new types are registered in `allowedEventTypes` AND `customerMessages`. The accept-time `app.yaml` is deliberately NOT deleted on failure — it is the crash-safe record with `Deployed:false` and it holds the settings the customer typed; the state every surface reads is `not_deployed`. Red-proofs: the accept-time call back → `TestDeployAcceptance_DoesNotClaimTheAppIsInstalled` fails; the success hook removed → `TestDeployDoneHook_...` fails at „the deploy ended and nothing was told about it". | **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0** |
|
||||
| **R-539** | **[P3-LOW] The restart brake catches a FAST crash loop and is blind to a SLOW one — add a second, slower counter (operator ruling 2026-09-16).** MEASURED 2026-09-16 (R-531): four controller restarts 20 minutes apart, none accumulating, because the budget window is 15 minutes; the only trace is an `info` `controller_restarted_by_agent` event, which mails nobody. A box whose controller dies every 20 minutes is restarted forever and nothing tells the operator. **The ruling:** a second counter — N restarts in 24 h → `controller_slow_crashloop` (warning) — kept beside the existing 3-in-15-minutes brake, which is unchanged. **Not built in the 2026-09-16 task** (its brief said the budget is a design and this task measures it); it is the nightly's to build. Needs: the agent-side counter, the new event type in `allowedEventTypes` + `customerMessages`, and a red-proof that a 20-minute cycle raises it while a healthy box never does. | **READY — rank P3-LOW; owner: CC (agent + hub)** |
|
||||
|
||||
Reference in New Issue
Block a user