From 78126575cb6ca843b98be1ec6d4048753f104486 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 29 Sep 2026 21:56:48 +0200 Subject: [PATCH] Register: R-719..R-725 from the new-household drill; R-505 closed (box-side hop proven); R-718 measured again Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- documentation/backlog/OPEN-ITEMS.md | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index a3287e52..84fb6430 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -687,7 +687,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-502** | **[P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner.** MEASURED 2026-09-14: `grep -rn bootstrap-modes scripts/*.py .gitea/workflows` → nothing; the harness is run by hand in `felhom-iso-assistant:trixie`. While adding the R-496 checks, its fake hub's `register` reply turned out to carry no `pairing_code`, so `print_pairing_banner` returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. **Fix shape:** register the harness in `repo_gates.py` behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. | **READY — rank P3-LOW; owner: CC** | | **R-503** | **[P3-LOW] SPIKE (not built): an install-time disk rule for the public ISO — "exactly one internal disk → install; otherwise stop in Hungarian".** Offered to the operator 2026-09-14 and **not chosen**: the 2026-07-31 ruling (a person chooses the disk) stands. Recorded so a reversal starts from measurements, not from the offer. **What must be measured first:** (1) whether the Proxmox auto-installer's HTTP answer mode can serve a per-machine answer from posted system info without network being a precondition a volunteer can miss; (2) whether USB transport is reliably visible in sysfs (`/sys/block/*/device` path, `removable`) where udev properties were measured blind (SPIKE-universal-iso-1 §3.2); (3) whether any refusal can be shown in Hungarian without modifying the Proxmox installer squashfs. **Reverses two rulings if built — needs an operator word.** | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (ruling), CC (spike)** | | **R-504** | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** | -| **R-505** | **[P1-HIGH] A fresh box on customer `tester-1` connects its Cloudflare tunnel but receives NO routes, so the dashboard answers 503 — the doorstep walk's intervention I1, again.** MEASURED 2026-09-14 on VM 331 (ISO 1.27.0, bound 15:47:26Z): `cloudflared` in the guest registered four connections (bud01, vie06, vie05, bud01) and then logged `No ingress rules were defined in provided config (if any) nor from the cli, cloudflared will return 503 for all incoming HTTP requests`; it ran `tunnel run` with the token, the same start mode as demo-hp's guest, and received **0** `Updated to new configuration` events. From DooPlex through public DNS, 12 requests 16:00:49–16:01:46Z → **12 × 503**, and the box logged **exactly 12** "No ingress" warnings in the same window — every request reached this connector. The guest's own front door answers the name (302). **The operator reports the dashboard loads on a phone over mobile data**; that is not explained by these measurements (no Cloudflare access to list the tunnel's connectors). **Leading hypothesis, not established:** the tunnel's routes live somewhere this connector never receives — a locally-managed config on another connector, or routes defined for a different tunnel. **Why P1:** from a network that reaches this connector, a volunteer cannot open their dashboard. **What it needs:** the operator checks the `tester-1` tunnel in Cloudflare Zero Trust (connectors, and public hostnames), or rules which tunnel this record should carry. **CAUSE CONFIRMED AND FIXED BY THE OPERATOR 2026-09-14 (evening), measured with no box:** a throwaway `cloudflared` on DooPlex with the `tester-1` token connected (4 connections) and at 17:12Z received NO configuration — `No ingress rules` per request, public link 503 — so the tunnel had no public hostnames; the box was not at fault. The operator then added one published application route, copied from a demo tunnel. Re-test 17:17Z: `Updated to new configuration … {"hostname":"*.enkicsifelhom.hu","service":"https://traefik"}, {"service":"http_status:404"}`; public requests to `felhom.` and `wiki.` hit ingress rule 0 and failed only on `lookup traefik … no such host` (502), as expected with no box. The earlier report that the phone loaded the dashboard is not reproduced by any measurement. **Remaining:** the box-side hop (cloudflared → traefik in the guest) is proven by the next drill. | **FIXED (operator) — box-side proof in the next drill; owner: CC** | +| **R-505** | **[P1-HIGH] A fresh box on customer `tester-1` connects its Cloudflare tunnel but receives NO routes, so the dashboard answers 503 — the doorstep walk's intervention I1, again.** MEASURED 2026-09-14 on VM 331 (ISO 1.27.0, bound 15:47:26Z): `cloudflared` in the guest registered four connections (bud01, vie06, vie05, bud01) and then logged `No ingress rules were defined in provided config (if any) nor from the cli, cloudflared will return 503 for all incoming HTTP requests`; it ran `tunnel run` with the token, the same start mode as demo-hp's guest, and received **0** `Updated to new configuration` events. From DooPlex through public DNS, 12 requests 16:00:49–16:01:46Z → **12 × 503**, and the box logged **exactly 12** "No ingress" warnings in the same window — every request reached this connector. The guest's own front door answers the name (302). **The operator reports the dashboard loads on a phone over mobile data**; that is not explained by these measurements (no Cloudflare access to list the tunnel's connectors). **Leading hypothesis, not established:** the tunnel's routes live somewhere this connector never receives — a locally-managed config on another connector, or routes defined for a different tunnel. **Why P1:** from a network that reaches this connector, a volunteer cannot open their dashboard. **What it needs:** the operator checks the `tester-1` tunnel in Cloudflare Zero Trust (connectors, and public hostnames), or rules which tunnel this record should carry. **CAUSE CONFIRMED AND FIXED BY THE OPERATOR 2026-09-14 (evening), measured with no box:** a throwaway `cloudflared` on DooPlex with the `tester-1` token connected (4 connections) and at 17:12Z received NO configuration — `No ingress rules` per request, public link 503 — so the tunnel had no public hostnames; the box was not at fault. The operator then added one published application route, copied from a demo tunnel. Re-test 17:17Z: `Updated to new configuration … {"hostname":"*.enkicsifelhom.hu","service":"https://traefik"}, {"service":"http_status:404"}`; public requests to `felhom.` and `wiki.` hit ingress rule 0 and failed only on `lookup traefik … no such host` (502), as expected with no box. The earlier report that the phone loaded the dashboard is not reproduced by any measurement. **Remaining:** the box-side hop (cloudflared → traefik in the guest) is proven by the next drill. **BOX-SIDE HOP PROVEN 2026-09-29 (new-household drill):** a fresh box on `tester-1` answered `https://felhom.enkicsifelhom.hu` through Cloudflare (302 → /claim, 200) 1 min after its controller started; the claim, every app front door and the gate all worked through the tunnel. | **CLOSED 2026-09-29 — proven end to end on a fresh box** | | **R-506** | **[P3-LOW] `day0-install.md` A.1 says "the controller manages per-app hostnames itself via the tunnel" — it does not.** MEASURED 2026-09-14: no code under `felhom-controller/controller/internal` creates tunnel ingress, DNS records or tunnel configurations (`grep -i 'ingress\|cfd_tunnel\|/configurations\|dns_records'` → only comments saying cloudflared is deployed when a token exists; positive control: the geo-restriction CF API use IS found in `cmd/controller/main.go`). The controller's only Cloudflare act is geo-restriction; the tunnel runs `tunnel run` with the token, so its routes come from Cloudflare's remote config set by the operator. A reader following A.1 skips the one step that makes the dashboard reachable (R-505). **Fix shape:** A.1 names the public-hostname step and its service settings, copied from a working tunnel. | **READY — rank P3-LOW; owner: CC (doc), operator (the settings to copy)** | | **R-507** | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** | | **R-508** | **[P2-MEDIUM] Customer `tester-1` has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer.** MEASURED 2026-09-14: the edit form's `email` value is empty; on bind the hub logged `[ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one`. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. **What it needs:** the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning)** | @@ -829,7 +829,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-715** | **[P3-LOW] The setup gate's probe reads only an HTTP-200 JSON object, so three apps with a real status get the button.** MEASURED 2026-09-29 on 9202: ghost (`{"setup":[{"status":…}]}` — a list), home-assistant (`/api/onboarding` — a top-level list), gramps-web (405 after the setup). And a probe that never flips BLOCKS the household's press (fail closed — measured on gramps-web while its check was still in the catalog): a wrong probe in a template would keep an app closed to everyone but the household until the catalog is fixed. **Fix direction:** list indexes in `field`, an optional `status:` to match, and a catalog gate that refuses a probe without a before/after measurement in its comment. **Fixed in controller v0.282.0** (RP29, RP30): list indexes in `field`, `done_status:`. ghost, home-assistant, gramps-web measured before/after on fresh installs; each gate opened by itself; a press before the setup refused on all three (`audits/signup-lock-2026-09-29/`E). New catalog gate `check-probe-measured.py` (5 decoys). | **CLOSED — 2026-09-29** | | **R-716** | **[P3-LOW] Apps installed before controller 0.281.0 keep their open sign-up — decision 47 closes it only on apps whose gate the box opened.** READ 2026-09-29 on the demo boxes after catalog `6faf432` synced: demo-hp's adventurelog and opengist, demo-felhom's opengist carry `signup_block:` in their synced template and no gate record, so no block (`audits/gate-rollout-2026-09-29/0/P0-3-demo-boxes-after-push.txt`). This is Part 0's rule working as designed (a catalog change never touches an installed app). **Needs an operator word** before anything changes on an installed app: a one-time "close sign-up now" press on the app page for an installed app, or leave them. Only the demo boxes have such installs today. **Operator ruled A (decision 49); built in controller v0.282.0 and pressed** on demo-hp's adventurelog and opengist and demo-felhom's opengist: before, sign-up served; after, refused; adventurelog's own switch on (`audits/signup-lock-2026-09-29/`C). | **CLOSED — 2026-09-29** | | **R-717** | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | -| **R-718** | **[P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so.** MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. **Fix direction:** the close card and the gate-open moment say "the app restarts once" where `after_setup.env` exists. | **OPEN — P3; owner: CC** | +| **R-718** | **[P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so.** MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. **Fix direction:** the close card and the gate-open moment say "the app restarts once" where `after_setup.env` exists. **ALSO MEASURED 2026-09-29 (new-household drill):** the gate-open press on a fresh vikunja restarted it for its own switch — the front door answered 404 for ~2 s and nothing said so. | **OPEN — P3; owner: CC** | +| **R-719** | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. | **READY — rank P2-MEDIUM; owner: operator (which shape) / CC** | +| **R-720** | **[P2-MEDIUM] A new household's apps are not in the off-site copy: every app starts with „3. mentés Kikapcsolva", and nothing the household is told says to switch it on.** MEASURED 2026-09-29 on a fresh box (customer with off-site ON, the default since 2026-09-16): `/backups/remote` read „Aktív — nincs kijelölt alkalmazás"; `/backups/apps` read „3. mentés Kikapcsolva — Ez az alkalmazás nincs kijelölve távoli mentésre" for all three apps. From source, an app is off-site only after `POST /backup/offbox/toggle` (`settings.SetAppOffbox`); no deploy or claim path sets it. The volunteer guide §7 says that after the recovery code „a távoli mentés magától elindul" — it starts, and copies nothing. So a household following the guide has **no off-site copy of its apps on night one**, on a one-drive box where the whole-guest tiers do not carry the data drive (`07` §6). The drill pressed „Bekapcsolás" for BookStack only, as a household reading the page might, to measure both paths on the night. R-240 (the empty run's wording) is the same gap seen from the other end. **Needs an operator decision** (it changes what the product promises): apps default to off-site ON when the customer has off-site, or the guide adds the step. | **READY — rank P2-MEDIUM; owner: operator (decide) / CC** | +| **R-721** | **[P2-MEDIUM] The household presses Stop during a whole-guest backup, and the backup starts the app again.** MEASURED 2026-09-29 19:37 UTC on a fresh box: the first off-site whole-guest backup quiesced three apps at 19:37:01; the household pressed „Leállítás" on actualbudget at 19:37:16 and the controller recorded `desired state for actualbudget recorded as "stopped"`; the same second the backup's early resume ran `unquiescing … restarting 3 stack(s)` and `Starting stack: actualbudget`. The app page then read „Fut" while `app.yaml` kept `desired_state: stopped` (so the next reboot would stop it). The household had to press Stop again, and the removal was refused „Az alkalmazáson mentés vagy visszaállítás fut" for ~4 min (honest). Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase1/step10-stop-undone-by-quiesce.log`. **Fix direction:** unquiesce restarts only stacks whose desired state is still running; a test pins it (Stop during a quiesce → still stopped after the resume). | **READY — rank P2-MEDIUM; owner: CC** | +| **R-722** | **[P2-MEDIUM] The volunteer guide is stale in five places a volunteer reads literally.** MEASURED 2026-09-29 walking `runbooks/VOLUNTEER-first-hour.md` on golden 0.282.0: (1) §8 says BookStack's login is `admin@admin.com / password` — since the random first password (controller v0.280.0) that login is REFUSED; the app page says the right thing („a Beállítások oldalon látható első jelszó"). (2) §7 says the recovery-code bar comes „néhány perccel" after setup — measured ~15 min after enrolment (the off-site tier arrives with the agent's next 15-minute host report). (3) Nothing says to switch each app's off-site copy on (R-720). (4) §2 says a 2 GB USB stick and Rufus; the download page says at least 4 GB and Balena Etcher. (5) The operator part says no button is needed for the link (R-719), and §5 names the mail „Elindult a Felhom szervered" while a customer with an earlier box gets „Új beállító kód — újratelepült a szervered … A korábbi jelszavad már nem érvényes". Also, from the screens: the installer pre-selects `/dev/sda` (the guide says it never chooses), and its Summary screen rests on „Previous", not „Install". **Fix:** a guide edit; the operator approves the text. | **READY — rank P2-MEDIUM; owner: CC (text) / operator (approve)** | +| **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** MEASURED 2026-09-29: `Operator email sent for tester-1/node_recovered` 2 s after the new box's first controller report (the customer's previous box had been silent 12 days — the new box is not a recovery); and `backup_tier_skipped (warning) — Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed)` → operator mail at 19:27 UTC, 7 min after enrolment, because the first whole-guest run fired before the off-site tier's descriptor arrived (applied ~19:35 UTC, backed up fine at 19:37). Operator-only, so no household is alarmed, but an operator learns to ignore both. **Fix direction:** `node_recovered` not for a host enrolled < N min ago; the skipped tier inside the first hour after enrolment is `info`, not a mailed warning. | **READY — rank P3-LOW; owner: CC** | +| **R-724** | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). | **READY — rank P3-LOW; owner: CC** | +| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). | **READY — rank P3-LOW; owner: CC** |