{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
killed at 2026-09-15T09:05:00Z (5 s into the deploy)
health 200 at 2026-09-15T09:09:08Z — 248 s after the kill
Sep 15 11:09:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:09:22.719+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
Sep 15 11:09:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:09:52.697+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
Sep 15 11:10:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:10:22.739+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s
## teardown: remove http=502 at 2026-09-15T09:10:40Z
## CORRECTION 2026-09-15T09:11:57Z: the 'health 200 at 09:09:08Z — 248 s after the kill' line is WRONG. docker inspect: felhom-controller exited 09:05:02Z (exit 137, the kill) and was NOT restarted. What answered 200 once is unexplained. What happened instead, and it is the guard working as designed: the agent had restarted this controller at 08:54:24Z, 08:56:53Z and 08:58:24Z (my idle, unpark and failed-swap tests; the swap's own rollback at 09:01:57Z was correctly NOT counted), so on the second not-running sweep after this kill it logged 'CRASH-LOOP … restarts_in_window=3' at 09:05:52Z and paused restarts for 30 min. Mid-deploy restart timing is therefore NOT measured here; the resume after the pause is captured separately.
## 2026-09-15T08:56:29Z after 100 s parked: health=502 container=exited
Sep 15 10:55:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:55:22.724+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
Sep 15 10:55:52 demo-hp felhom-agent[1526161]: time=2026-09-15T10:55:52.720+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
Sep 15 10:56:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:22.751+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
unparked at 2026-09-15T08:56:31Z
health 200 at 2026-09-15T08:56:56Z — 25 s after unpark
Sep 15 10:56:22 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:22.751+02:00 level=INFO msg="controller-supervisor: controller is not running and the guest is PARKED — leaving it" vmid=9201 status=exited marker=/var/lib/felhom-agent/guests/9201/controller-parked
Sep 15 10:56:52 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:52.716+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exited unit=felhom-controller-bootstrap.service
Sep 15 10:56:53 demo-hp felhom-agent[1526161]: time=2026-09-15T10:56:53.962+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps"
## CORRECTION 2026-09-15T09:00:18Z: the block above is NOT a swap test — the endpoint has no scheme, curl spoke HTTP to the HTTPS local API, the agent answered 400 and no swap started. It is one more IDLE kill (killed 08:57:30Z, restarted 08:58:24Z). Re-run with https:// follows.
## 2026-09-15T09:03:00Z swap status: {"ok":true,"data":{"current":"gitea.dooplex.hu/admin/felhom-controller:0.243.0","error":"new controller did not become healthy within timeout","in_flight":false,"previous":"gitea.dooplex.hu/admin/felhom-controller:0.243.0","state":"failed","target":"gitea.dooplex.hu/admin/felhom-controller:0.243.0"}}
Sep 15 11:00:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:00:52.731+02:00 level=INFO msg="controller-supervisor: controller is not running during a controller SWAP — the swap owns it" vmid=9201 status=exited
Sep 15 11:01:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:22.734+02:00 level=INFO msg="controller-supervisor: controller is not running during a controller SWAP — the swap owns it" vmid=9201 status=exited
Sep 15 11:01:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:52.712+02:00 level=INFO msg="controller-supervisor: controller is not running during a controller SWAP — the swap owns it" vmid=9201 status=exited
Sep 15 11:01:55 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:55.784+02:00 level=WARN msg="controller-swap: health verdict negative — rolling back" vmid=9201 target=gitea.dooplex.hu/admin/felhom-controller:0.243.0
Sep 15 11:01:55 demo-hp felhom-agent[1526161]: time=2026-09-15T11:01:55.784+02:00 level=WARN msg="controller-swap: rolling back" vmid=9201 previous=gitea.dooplex.hu/admin/felhom-controller:0.243.0 reason="new controller did not become healthy within timeout"
Sep 15 11:02:06 demo-hp felhom-agent[1526161]: time=2026-09-15T11:02:06.759+02:00 level=INFO msg="controller-swap: rolled back to previous, controller healthy" vmid=9201 previous=gitea.dooplex.hu/admin/felhom-controller:0.243.0
2026/09/15 08:47:09 infra.go:74: [WARN] [infra] filebrowser admin probe: Post "http://filebrowser:80/api/auth/login?username=admin": dial tcp: lookup filebrowser on 192.168.0.250:53: no such host (will retry)
2026/09/15 08:52:10 filebrowser_password.go:90: [INFO] [infra] filebrowser: admin/admin is refused (HTTP 401) — the password was set by someone; leaving it and recording "operator"
## CORRECTION 2026-09-15T08:40:31Z: the 08:40:19Z block above is INVALID — the dashboard password loaded as an empty string (read_credential.py call wrong), so there was no login, no session and no reveal; 'revealed=401' was an EMPTY password. Only admin/admin=401 and wrong=401 stand from that block. Re-run follows.
## RE-RUN 2026-09-15T08:40:48Z via LAN https://192.168.0.114, the endpoints the UI calls (login → /apps/filebrowser → POST reveal → FileBrowser login):
Rendszermentés (teljes mentés) A teljes szerver — alkalmazások, beállítások és adatbázisok együtt — időszakos mentése, amelyből az egész készülék visszaállítható. Ezt a host-ügynök készíti és kezeli. ✓ Utolsó teljes mentés 2026-09-15 06:26 (4 órája) 0 B Helyi tároló (local) Naprakész Következő mentés 4 órája — a mentési ablakon belül – Visszaállítás ellenőrizve Még nem futott Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-15 06:26 (4 órája) Naprakész Biztonsági szerver – külön hardver (PBS) ✓ Utolsó sikeres mentés: 2026-09-15 06:11 (4 órája) Naprakész Mentés most A mentés alatt az alkalmazások leállnak — általában néhány perc, nagyobb adatnál több.
## remote tile: lass="stat-label">Adatmentés aktív ✓ Távoli rendszermentés külön hardveren (PBS) <div class=
{"ok":true,"message":"Telepítés elindítva – az állapot a kártyán követhető"}
2026-09-15T09:10:53Z vaultwarden health: healthy
container env SIGNUPS_ALLOWED: false
stranger register via traefik (vault.enkisfelhom.hu): 422 body: <!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="color-scheme" content="light dark">
<title>422 Unprocessable Entity</titl
app page: invite fragment 'Invite User'=1 'meghiv'=1; negative control 'Hozd l'=0
## NOTE 2026-09-15T09:11:25Z: the 'stranger register … 422' line above is NOT a proof of refusal — 422 means the request body did not parse. Re-run with the spike's exact body follows.
## RE-RUN 2026-09-15T09:12:20Z stranger send-verification-email via traefik (vault.enkisfelhom.hu, catalog-deployed, SIGNUPS_ALLOWED=false): HTTP 400 message: Registration not allowed or user already exists
## 2026-09-15 — the big night's P1 fixes (agent v0.131.0, controller v0.243.0, hub v0.114.0, catalog templates, ISO 1.27.1 published)
Nine rows closed. **Full original text: `git show <this commit>^ -- documentation/backlog/OPEN-ITEMS.md`.** Evidence folder: `documentation/audits/evidence-p1fixes-2026-09-15/`.
| ID | Title | Shipped | Evidence |
|---|---|---|---|
| **R-493** | [P1-HIGH] There are NO customer-facing install instructions, so a volunteer cannot begin — this blocks inviting anyone. | ISO 1.27.1 published + `felhom.eu/letoltes` live | 2026-09-15: round trip over `https://iso.felhom.eu` sha256 `25637007…c053`, 1 705 322 496 B; `felhom.eu/letoltes` 200 naming 1.27.1 (`documentation/tests/iso-release-1.27.1-2026-09-14/README.md`) |
| **R-495** | [P2-MEDIUM] The public installer asks a stranger four questions nothing answers, and REFUSES its own default on one of them. | answered by the guide; ISO 1.27.1 published 2026-09-15 | same as R-493 |
| **R-496** | [P2-MEDIUM] The box's console tells a stranger, in English and FIRST, to open the Proxmox admin page — and calls the owner passphrase „a jelszavadat”. | ISO 1.27.1 (console Felhom-only) published 2026-09-15 | same as R-493; G15 live proof on VM 332 |
| **R-512** | [P2-MEDIUM] Vaultwarden is installed with open registration, and the one control the page tells the customer to use to close it is read-only. | catalog template (no version change), 2026-09-15 | spike: stranger 400 / invite 200 / invited 200 (E1-vaultwarden-spike.txt); live on 9202 from the catalog: `SIGNUPS_ALLOWED=false`, stranger 400 „Registration not allowed" (E1-vaultwarden-9202-live.txt) |
| **R-513** | [P1-HIGH — SECURITY] Every box's file manager (FileBrowser, `files.<domain>`, a launcher tile) accepts the login `admin` / `admin`, and on demo-hp that login page is on the public internet. | controller v0.243.0 | 9202 (default login): generated, stored encrypted, reveal 200 (16 chars), revealed=200 admin=401 wrong=401; 9201 (hand-set): recorded `operator` 08:52:10Z, untouched; public `files.enkisfelhom.hu` admin → 401 (B1/B4 files) |
| **R-514** | [P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut". | catalog template: 1 worker × 1 thread, 1280M | live on 9202: 20 PDFs at once → 20/20 SUCCESS, memory.peak 772 370 432 B, no OOM (E2-paperless-9202-live.txt). The OOM-visibility half is NOT proven live → R-528 |
| **R-515** | [P3-LOW] The Paperless-ngx app page tells the customer to log in with `admin / admin`, and that login does not exist: the deploy form generates the admin password. | catalog template, 2026-09-15 | `default_creds` removed; first steps point at „Automatikusan generált értékek" (app-catalog commit e6aa443) |
| **R-517** | [P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true. | controller v0.243.0 + agent v0.131.0 | live on 9201: per-tier rows „Helyi tároló (local) ✓ … Naprakész", „PBS ✓ … Naprakész", remote tick on a real PBS success (C4-backup-page-9201.txt). Found live and fixed after the release: an unknown size printed „0 B" → „–" (controller main d3eacbb, unreleased) |
| **R-523** | [P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule. | agent v0.131.0 + hub v0.114.0 (+ golden script `--restart always`, no bake) | measured first: kill leaves both `unless-stopped` and `always` exited (A1). Live on 9201: idle kill → dashboard 200 in 59 s; parked → stayed dead 100 s with the PARKED line, unpark → 200 in 25 s; kill during a swap → supervisor deferred ×3, swap rolled back itself; the crash-loop guard tripped for real after 3 test restarts and the hub mailed `controller_crashloop` (A4-* files). NOT measured: restart timing during a deploy (the budget was spent) → R-531 |
| **R-442** | **`remove_hdd_data: true` was INERT — the customer's data stayed on the drive while the API reported success (HTTP 200, `hdd_paths_removed: null`, 128 MB of Nextcloud left; demo-hp 2026-09-01).** Shipped in controller **v0.236.0** (2026-09-13). Removal resolved the drive from the GLOBAL `cfg.Paths.HDDPath` (no default, set on NO box — demo-hp AND demo-felhom both measured 0 `hdd_path` / 0 `FELHOM_PATHS_*`, so the fleet shares the shape); it now reads the app's OWN `app.yaml``HDD_PATH` — the `07-backup-architecture.md` ~L437 rule that deploy, the start gate and the backup destination already implemented — and a data removal it cannot resolve, or whose drive is absent, is REFUSED (409, exact Hungarian sentence, typed `stacks.RemoveRefusedError`) BEFORE `compose down`, app kept. SSD app → `hdd_paths_removed: []`, never `null`, plus `hdd_note`. Missing folders stated in `hdd_paths_missing`. The backup-half refusal reaches the response (`backup_paths_refused`) and its base follows the same rule. **Reasoning kept:***"declares no drive" ≠ "could not resolve the drive" — the first is a fact, the second a refusal*; *an app gone with its data left behind is unrecoverable from the UI — the customer cannot even re-run the removal*; *no fallback to the global — that silent fallback is the exact path this closes*. **Observation carried:** 8 of the 13 `needs_hdd` catalog apps bind ONLY `${USERDATA_PATH}` (the shared library) and no `${HDD_PATH}` folder, so for them "delete my data" correctly removes nothing on the drive and the modal shows no checkbox. | **CLOSED 2026-09-13 — shipped controller v0.236.0, proven live on demo-hp** | `audits/R442-2026-09-13/` — A: 63 MB written by Nextcloud ITSELF, gone after removal and listed with its size; C: 409 + sentence, all 69 files untouched, app still deployed, `[ERROR] … refused` logged; D: gokapi `[]` + note; ASCII controls (`llap`: C=2 A=0 D=0). 15 tests + two red-proofs in `felhom-controller/REPORT.md`. Full original text: `git show d6837d98ee24:documentation/backlog/OPEN-ITEMS.md`. |
| **R-449** | **UPDATE ARC SLICE 5 — an upgrade test that runs again. BUILT AND RUN 2026-09-06:**`app-catalog-felhom.eu/scripts/upgrade-test.py` + `upgrade_fixtures.py`, 7 edges across 3 apps, evidence in `audits/upgrade-spike-2026-09-06/`. **Reasoning kept — success is an APPLICATION-LEVEL READBACK, never file identity:**`survive2.py`'s sha256+inode rule is right for a redeploy and WRONG for an upgrade, because a migration is supposed to rewrite files and that rule would fail every correct upgrade. **Reasoning kept — nothing is ever seeded into a volume by hand (R-156);** an app with no non-browser route is recorded `inconclusive`, which is a result and not a licence to plant a file. **Reasoning kept — run the negative control FIRST:** C3's TO image exits immediately and came back `failed`; a harness that cannot fail a known-broken upgrade proves nothing with its greens. **WHAT IT MEASURED:** all five real catalog upgrades kept the customer's data; and **whether an upgrade can be undone is a property of the individual APP, not of upgrades** — docmost REFUSES (*"corrupted migrations: previously executed migration 20260213T085259-notifications is missing"*), privatebin does not, which independently reproduces the Nextcloud finding on a second app by a DIFFERENT mechanism and puts two measurements behind §4's ruling that "rollback" is the wrong word. **It also found a defect in our own catalog (R-459).****What stays open, as its own rows rather than inside this one:** R-459 (the skipped MariaDB datadir upgrade), R-460 (bookstack's file half is unprovable headlessly), R-462 (the widening, costed). Full original text: `git show 417df06f3529:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — harness built, run, and proven by a red negative control** | `audits/SPIKE-upgrade-test-2026-09-06.md`; `audits/upgrade-spike-2026-09-06/evidence/`; catalog `0474ce387e6f` |
| **R-438** | **The catalog sync rewrote a DEPLOYED app's `docker-compose.yml` and no architecture document recorded that it did. BOTH HALVES NOW DISCHARGED — the document was written 2026-09-02, the behaviour was changed in controller v0.235.0 (2026-09-06).**`Syncer.copyTemplates` copied into every stack folder on a 15-minute cycle with **no deployed check**, so a deployed app's file and its running containers disagreed from that moment, and the next `compose up -d` from any of thirteen call sites resolved the disagreement by upgrading — measured live: the sync rewrote the file at 17:45:17Z while the container went on running the old image, a restart then upgraded it in 18.3 s **with a pull**, and a boot reconciliation upgraded it **with nobody pressing anything**. **Reasoning kept — the distinction this row existed to protect:**`RestartStack`'s use of `up -d` to pick up template changes was **CHOSEN and written down in its own comment**, so reversing it was an operator DECISION, not a bug fix; that is why the row stayed open through v0.233.0 and v0.234.0 while only the documentation half was done. **Reasoning kept — one fear was measured SMALLER than stated:** a plain power cut does NOT upgrade anything, because Docker's `restart: unless-stopped` restores the containers on the old image and the reconciler logs `no boot-orphaned apps`; the unattended upgrade needs the narrower precondition *"and the app did not come back"*. Full original text: `git show bc47dd4ef997:documentation/backlog/OPEN-ITEMS.md`. | **CLOSED 2026-09-06 — documented 2026-09-02, behaviour changed in controller v0.235.0** | `audits/SPIKE-app-update-2026-09-01.md` §2, §3, §8; `architecture/09-update-architecture.md`; `tests/VALIDATION-update-slice3-2026-09-06.md` |
@@ -696,10 +696,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-488** | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** |
| **R-489** | **[P3-LOW] `POST /api/stacks/{name}/remove` reports `volumes_removed: null` over named volumes it DID remove.** MEASURED 2026-09-13 on demo-hp five times (gokapi, actualbudget, adventurelog ×2, glance): `docker compose down --volumes` removed the app's named volumes (`docker volume ls` count 2 → 0) and the response carried `"volumes_removed":null`. The customer's confirmation dialog therefore cannot say what it deleted. Split out of R-474 (closed in v0.240.0 for the backups half). **Fix shape:** list the volumes before `down --volumes`, diff after, and report the difference (`[]` when none, never `null`). **PARTLY SHIPPED in v0.242.0 (`d698ce3`), measured live on 9202 the same night:** the difference is computed and a fresh compose-created volume IS reported (`["opengist_opengist_data"]`), but the listing filters on the compose project LABEL and a volume recreated by a unit restore (`docker volume create <name>`, `restore.go:154`) carries no labels — compose still removes it and the response says `[]` (`audits/v0242-2026-09-14/19-R489-cause.txt`). **Remaining fix:** list by the `<project>_` name prefix as well (union), or label the recreated volume as compose would. | **READY — rank P3-LOW; owner: CC (residual)** |
| **R-492** | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** |
| **R-493** | **[P1-HIGH] There are NO customer-facing install instructions, so a volunteer cannot begin — this blocks inviting anyone.** MEASURED 2026-09-14 (drill `audits/DRILL-fresh-install-0242-2026-09-14.md`, Phase 0.2), each source with what was searched: the website's 9 pages carry no mention of the installer, of `iso.felhom.eu`, or of any install step (`iso`, `letolt`, `telepit`); `https://iso.felhom.eu/` and `/index.html` both return **404** — only the exact object name `felhom-installer-1.26.1-pve9.2-1.iso` answers, so the file cannot be found without being told its name; `RUNBOOK-onboarding-draft-v4.md` is an operator-attended script; the R-11 tester one-pager was ruled 2026-07-21 and never written; no `email.md` exists in the workspace; the two hub mails a new customer receives (`Kösd össze a Felhom dobozodat`, `Elindult a Felhom szervered — beállító kód`) assume the box is already installed. **Stopgap written by the drill:**`documentation/runbooks/VOLUNTEER-first-hour.md` (Hungarian, the steps the product really needs, each addition over today listed at its top). **What it needs:** the operator to choose the channel and approve the text. **2026-09-14:** guide approved by the operator and aligned to ISO 1.27.1; download page planned at `felhom.eu/letoltes` (R-504). Stays open until the page and ISO are live. | **WAITING-ON-OPERATOR — publish yes; rank P1-HIGH** |
| **R-494** | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).***Original finding, kept:***[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.****What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** |
| **R-495** | **[P2-MEDIUM] The public installer asks a stranger four questions nothing answers, and REFUSES its own default on one of them.** MEASURED 2026-09-14 on `felhom-installer-1.26.1` (drill screens `s04`–`s22`): every screen after the GRUB menu is **English** Proxmox (EULA, disk, locale, password, network, summary); **three disks are offered with no guidance which is the system disk** (the default happened to be right); the administrator e-mail is prefilled `mail@example.invalid`; the hostname is prefilled `pve.example.invalid` and pressing Next on it returns **„Invalid values: hostname does not look valid”** — a volunteer who accepts the defaults cannot continue; at the end, with the stick still in, the default-ticked auto-reboot boots back into the installer. None of this is wrong for Proxmox; all of it is unanswered for a Felhom household. **Fix shape (cheap first):** the volunteer instructions answer each question (done in `runbooks/VOLUNTEER-first-hour.md`); later, prefill hostname and e-mail from the ISO build. **ANSWERED 2026-09-14 by operator ruling + guide, not by code:** the interactive installer stays (the 2026-07-31 ruling re-affirmed); `VOLUNTEER-first-hour.md` answers every screen (disk, keyboard, password, e-mail, host name, stick) and release gate G14 pins the disk rule. The English screens remain by ruling. Closes when the guide and ISO are published. | **ANSWERED — awaiting publish; owner: operator (publish)** |
| **R-496** | **[P2-MEDIUM] The box's console tells a stranger, in English and FIRST, to open the Proxmox admin page — and calls the owner passphrase „a jelszavadat”.** MEASURED 2026-09-14 (drill screen `s29`): above the Hungarian pairing banner the console prints Proxmox's own `Welcome to the Proxmox Virtual Environment. Please use your web browser to configure this server - connect to: https://192.168.0.134:8006/` — the operator admin UI, which a household must never be sent to; the Felhom banner below says „add meg ezt a kódot és **a jelszavadat**”, while the self-bind mail and page name the same secret **„Tulajdonosi jelmondat”** (R-323 renamed the mail; `scripts/iso/felhom-bootstrap.sh:71-72` was not renamed). **Fix shape:** replace the Proxmox `/etc/issue` block on appliance installs; rename the console word. Both live in `felhom-bootstrap.sh`, a frozen ISO payload (G9), so it ships with the next ISO. **FIXED IN ISO 1.27.1, NOT YET PUBLISHED (2026-09-14).** 1.27.0 masked `pvebanner` at first boot and still showed the Proxmox block on that boot (VM 331 screen s20); 1.27.1 masks by symlink in the postinst and writes a no-ő/ű `/etc/issue`. Live on VM 332: first-boot console Felhom-only; after a proven reboot `pvebanner` masked, `/etc/issue` 0 × 8006. Harness 55/55, red first. Closes when 1.27.1 is published. | **FIXED — awaiting publish (operator yes); owner: CC** |
| **R-497** | **[P2-MEDIUM] No product channel ever gives the customer the „Tulajdonosi jelmondat”, yet the self-bind mail says they received it at setup.** MEASURED 2026-09-14: the 5-word phrase is minted at customer creation (`hub/internal/web/configs.go`, `RandomPassphrase(5)`) and shown **only on the operator's customer page**; the self-bind mail (`FormatSelfBindEmail`) says „amelyet a beállításkor kaptál” and the console asks for it; nothing sends it. Day-0 runbook A.2 says the operator must dictate or hand it over — a step that exists only in an operator document. A volunteer whose operator forgets cannot bind the box and is told they already have the phrase. **Fix shape:** the operator's customer-create confirmation says, in one line, to hand this phrase to the customer now; the volunteer instructions say where it comes from (done in `VOLUNTEER-first-hour.md`). **CLOSED 2026-09-14 — hub v0.113.0 deployed (ArgoCD Synced, image 0.113.0). Tests red first (`passphrase_handover_test.go`); live: `tester-1`'s page renders the hand-over sentence 1× (2× with `?flash=created`), negative control 0. The mail change is pinned by test; no mail could be read live (no mailbox).** Evidence: `audits/DOORSTEP-walk-1270-2026-09-14.md` §5. | **CLOSED — hub v0.113.0** |
| **R-498** | **[P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal `wiki.DOMAIN` — the placeholder is never filled in.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a **wiki.DOMAIN** címet a böngészőben", PrivateBin's „Nyisd meg a **paste.DOMAIN** címet". `grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml` → **52 of 53** templates carry it in `first_steps`; the controller renders the string as text (`internal/stacks/metadata.go`). A stranger reading their first instruction meets a word that is not an address. **Fix shape:** substitute the stack's real `SUBDOMAIN.DOMAIN` at render time (one place in the controller), with a render test per template that fails on a literal `DOMAIN`. | **READY — rank P3-LOW; owner: CC** |
| **R-499** | **[P2-MEDIUM] Every app without a data drive is told its data is „already in the full system backup (PBS)" and there is „nothing to do" — on a box with no PBS, whose only whole-box copy sits on the same disk.** MEASURED 2026-09-14 on a fresh 0.242.0 box with the DR tier off (drill step 7): `GET /stacks/bookstack/backup` renders „Ennek az alkalmazásnak az adatai a belső rendszerlemezen vannak, amelyek **már szerepelnek a teljes rendszermentésben (PBS)** … ehhez az alkalmazáshoz nincs külön teendő." The sentence sits under `{{if not .IsHDDApp}}` in `controller/internal/web/templates/tier2_config.html:20-26` and consults nothing about where the whole-guest backup goes. On the same box `/backups` says, correctly, „Helyi tároló (local)" and „A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem." **Two pages of one product contradict each other, and the reassuring one is the false one.****Fix shape:** branch the sentence on the box's actual whole-guest target (PBS vs local, same-disk flag the overview already computes); a render test per branch. | **READY — rank P2-MEDIUM; owner: CC** |
@@ -712,22 +709,25 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-506** | **[P3-LOW] `day0-install.md` A.1 says "the controller manages per-app hostnames itself via the tunnel" — it does not.** MEASURED 2026-09-14: no code under `felhom-controller/controller/internal` creates tunnel ingress, DNS records or tunnel configurations (`grep -i 'ingress\|cfd_tunnel\|/configurations\|dns_records'` → only comments saying cloudflared is deployed when a token exists; positive control: the geo-restriction CF API use IS found in `cmd/controller/main.go`). The controller's only Cloudflare act is geo-restriction; the tunnel runs `tunnel run` with the token, so its routes come from Cloudflare's remote config set by the operator. A reader following A.1 skips the one step that makes the dashboard reachable (R-505). **Fix shape:** A.1 names the public-hostname step and its service settings, copied from a working tunnel. | **READY — rank P3-LOW; owner: CC (doc), operator (the settings to copy)** |
| **R-507** | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** |
| **R-508** | **[P2-MEDIUM] Customer `tester-1` has no registered e-mail, so neither the self-bind link nor the setup code can reach a volunteer.** MEASURED 2026-09-14: the edit form's `email` value is empty; on bind the hub logged `[ERROR] [claim] claim code generated (gen 1) but customer tester-1 has NO registered email — deliver via resend after setting one`. A volunteer onboarded on this record would sit at „A szerver beállítása" with no code. **What it needs:** the operator sets the volunteer's address on the record before sending the guide (day-0 A.2). The hub's customer page could warn when a record with an unclaimed box has no e-mail — the log line exists, the page says nothing. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (record), CC (page warning)** |
| **R-509** | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2). | **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** |
| **R-510** | **[P1-HIGH] `tester-1`'s tunnel now has its route, and still gives a fresh box 502: the route sends traffic to `https://traefik` WITH certificate checking, and traefik answers the name `traefik` with its default certificate.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0), after the operator's R-505 fix: from DooPlex `https://felhom.enkicsifelhom.hu` → **502 ×3** (18:07:48Z, `server: cloudflare`). The box's `cloudflared` logs `Request failed … tls: failed to verify certificate: x509: certificate is valid for 544346c4….traefik.default, not traefik … ingressRule=0 originService=https://traefik`. Traefik itself holds a valid Let's Encrypt `CN=*.enkicsifelhom.hu` (openssl on 127.0.0.1:443 with SNI). **Control:** demo-hp's working tunnel config reads `{"hostname":"*.enkisfelhom.hu", "originRequest":{"noTLSVerify":true}, "service":"https://traefik"}` — the same route WITH `noTLSVerify`. So the Cloudflare-side public hostname for `*.enkicsifelhom.hu` lacks „No TLS Verify" (or an origin server name). The Cloudflare side is not visible to the session; the inference rests on the log line and the control. Intervention **I2** of the big night: the claim and every dashboard request go to the guest's LAN address with the name forced. **Fix:** operator ticks „No TLS Verify" on that public hostname, then day-0 A.1 names the setting beside the route. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (Cloudflare route), CC (day-0 A.1 wording after)** |
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
| **R-512** | **[P2-MEDIUM] Vaultwarden is installed with open registration, and the one control the page tells the customer to use to close it is read-only.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, controller 0.242.0, catalog vaultwarden 1.36.0-alpine): the deploy form sends `SIGNUPS_ALLOWED=true` (catalog default); after install, „Vaultwarden — Beállítások" says „Ez az alkalmazás már telepítve van. **Az alábbi beállítások csak olvashatók.**" and, below it, „Regisztráció engedélyezése — Igen / Nem — Új fiókok regisztrálásának engedélyezése. **Az első fiók létrehozása után állítsd 'Nem'-re.**" At 18:27:20Z a stranger with no invite and no login registered `idegen.probe@example.com` → **200** (`phase3/vaultwarden-stranger-signup.txt`). Once the dashboard is reachable through the tunnel, anyone who guesses `vault.<domain>` can open an account on the household's password server. A route does exist (the Vaultwarden admin panel's own setting, token under „Megjelenítés"), and no screen names it. **Fix shape:** default `SIGNUPS_ALLOWED=false` with an invite-first first step, or make that one field editable after install; the page must not instruct an act it forbids. | **READY — rank P2-MEDIUM; owner: CC (catalog + controller)** |
| **R-513** | **[P1-HIGH — SECURITY] Every box's file manager (FileBrowser, `files.<domain>`, a launcher tile) accepts the login `admin` / `admin`, and on demo-hp that login page is on the public internet.** MEASURED 2026-09-14 (BIGNIGHT): `POST /api/auth/login?username=admin` with `X-Password: admin` → **200 + a session token** on VM 333 (fresh ISO 1.27.1 install, controller 0.242.0) and on demo-hp guests **9201 and 9202** (loopback, `Host: files.enkisfelhom.hu`); negative control `admin` / wrong → **401** on all three. On VM 333 that token lists both sources — „Adatlemez" (the data drive's `userdata`: documents, media, photos…) and „Beolvasás" (with `paperless`) — `GET /api/users?id=self` 200. **Public exposure, measured by GET only:**`https://files.enkisfelhom.hu/` from DooPlex through Cloudflare → 200, FileBrowser Quantum, `passwordAvailable:true, noAuth:false` (`phase3/filebrowser-public-reachability-demo-hp.txt`); no login was attempted over the internet. The generated `config.yaml` sets no admin credential (FileBrowser's own default applies); no screen shows the customer any FileBrowser login. Geo-restriction narrows who can reach it, it does not authenticate. **Not changed tonight** (9201's standing state is fenced; no product code). **Fix shape:** the controller sets a generated admin password (or proxy auth behind the dashboard session) at stack creation and on every existing box, and shows it where the customer finds app credentials. | **READY — rank P1-HIGH; owner: CC (controller) · operator (rotate on live boxes first)** |
| **R-514** | **[P2-MEDIUM] Paperless-ngx dies silently when a family uploads 20 documents at once: its worker is OOM-killed inside the catalog's 768 MB cap, 11 uploads fail, 8 wait forever, and the app still reads „Fut".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, catalog paperless-ngx 2.20.15, `paperless-webserver` limit 805 306 368 B): 20 small 3-page PDFs posted through `/api/documents/post_document/` at 18:29:57Z. At 18:31:22Z the VM kernel logged `Memory cgroup out of memory: Killed process … (gs)`×2, `([celeryd: celer)`, `([celery beat] -)` — constraint MEMCG of that container; `docker inspect` → `oomkilled=true restarts=0`. Paperless's own task list: **11 FAILURE (`WorkerLostError`), 1 STARTED, 8 PENDING**, unchanged 13 minutes later; **0 documents**. The controller shows the app running and healthy; no event, no alarm (`phase3/paperless-tasks.txt`, `paperless-oom-check.txt`). A household scanning a drawer of bills sees nothing arrive and no reason. **Fix shape:** raise the cap or set `PAPERLESS_TASK_WORKERS=1` / `PAPERLESS_THREADS_PER_WORKER=1` in the template, and let the dead-app/health check see an OOM-killed worker. | **READY — rank P2-MEDIUM; owner: CC (catalog)** |
| **R-515** | **[P3-LOW] The Paperless-ngx app page tells the customer to log in with `admin / admin`, and that login does not exist: the deploy form generates the admin password.** MEASURED 2026-09-14 (BIGNIGHT, VM 333): `/apps/paperless-ngx` „Első lépések — Jelentkezz be: admin / admin" and „Alapértelmezett belépés admin / admin"; the catalog template generates `PAPERLESS_ADMIN_PASSWORD` (`password:16`) and the token call with the generated value succeeded (`phase3/seed-paperless.txt`). A household following the page is refused at its first login. Same card also sends documents to „FileBrowser … import/paperless" — the FileBrowser login is R-513. **Fix:** the card points at „Beállítások → Automatikusan generált értékek" as gokapi's does. | **READY — rank P3-LOW; owner: CC (catalog)** |
| **R-509** | **[P1-HIGH] A box installed for an EXISTING customer never gets the self-bind e-mail the console tells the volunteer to open.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1): customer `tester-1` now has `tester1@felhom.eu` registered; the box registered as appliance 28 at 17:44:53Z and its console says „Nyisd meg az e-mailben kapott linket"; **ten minutes later the mailbox (read through the Gmail connector) held 0 messages to that address.** Cause, from source: the hub auto-sends the link only at customer creation (`hub/internal/web/configs.go:725`) and at RESET completion (`customer_reset.go:162`); a customer whose e-mail was added later, or whose previous box was destroyed, never receives one unless the operator presses „Send self-bind link". The volunteer guide's operator prerequisites do not list that press. Intervention **I1** of the big night (the operator's button pressed). **Fix shape (for the operator to choose):** send the link when an unclaimed appliance registers and a customer with no host is waiting, or add the press to the guide's operator prerequisites (day-0 A.2).**SHIPPED hub v0.114.0 (2026-09-15), NOT YET PROVEN BY A REAL MAIL:** triggers added — e-mail set/changed on a customer with no box, and host delete — each re-checking no bound host; every send recorded as `selfbind_link_sent` and shown on the Setup tab. Unit-proven with a red-proof (`TestSelfBind_EmailSetOnWaitingCustomerSendsLink`). The live check with the Gmail-read mailbox was NOT run: it needs a throwaway customer, and deleting one runs the RESET cascade (ep0 `deprovision` + Cloudflare), which is fenced without the operator's word. **Closes on:** one real mail from either trigger. | **READY — rank P1-HIGH; owner: CC (hub fix) · operator (which fix shape)** |
| **R-510** | **[P1-HIGH] `tester-1`'s tunnel now has its route, and still gives a fresh box 502: the route sends traffic to `https://traefik` WITH certificate checking, and traefik answers the name `traefik` with its default certificate.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0), after the operator's R-505 fix: from DooPlex `https://felhom.enkicsifelhom.hu` → **502 ×3** (18:07:48Z, `server: cloudflare`). The box's `cloudflared` logs `Request failed … tls: failed to verify certificate: x509: certificate is valid for 544346c4….traefik.default, not traefik … ingressRule=0 originService=https://traefik`. Traefik itself holds a valid Let's Encrypt `CN=*.enkicsifelhom.hu` (openssl on 127.0.0.1:443 with SNI). **Control:** demo-hp's working tunnel config reads `{"hostname":"*.enkisfelhom.hu", "originRequest":{"noTLSVerify":true}, "service":"https://traefik"}` — the same route WITH `noTLSVerify`. So the Cloudflare-side public hostname for `*.enkicsifelhom.hu` lacks „No TLS Verify" (or an origin server name). The Cloudflare side is not visible to the session; the inference rests on the log line and the control. Intervention **I2** of the big night: the claim and every dashboard request go to the guest's LAN address with the name forced. **Fix:** operator ticks „No TLS Verify" on that public hostname, then day-0 A.1 names the setting beside the route.**2026-09-15, after the operator's tick:** three GETs from DooPlex 08:09:07–08:09:14Z → **530 ×3** (`server: cloudflare`) — no tunnel is connected, because tester-1 has no box since the BIGNIGHT teardown. The fix cannot be observed until a box exists. `day0-install.md` A.1 now names „No TLS Verify" beside the route, marked unproven. **Closes on:** 200 or the claim page through the tunnel on the next tester-1 box. | **WAITING-ON-OPERATOR — rank P1-HIGH; owner: operator (Cloudflare route), CC (day-0 A.1 wording after)** |
| **R-511** | **[P2-MEDIUM] A customer whose box is rebuilt keeps its ep0 PBS token, and then the DR tier can be neither provisioned nor re-issued: the hub's error advises the one action that refuses.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, `tester-1`, DR tier ticked): on the new box's WireGuard registration the hub logged `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action — save the customer config to retry`. The operator's `POST /configs/tester-1/pbsdr-reissue` → **400 `No provisioned PBS DR tier for this customer`** (`hub/internal/web/pbsdr.go` ~411). The token was left by the doorstep walk's host delete (a host delete does not deprovision tenancy; only RESET does, which also removes the tunnel). So a box rebuilt for an existing customer — the reinstall journey — has no whole-guest off-site tier and no button that restores it. **Fix shape:** let re-issue adopt an existing endpoint token when the descriptor is absent (the message already assumes it does), or have host delete offer to drop the PBS token.**SHIPPED hub v0.114.0 (2026-09-15):** re-issue ADOPTS the endpoint token when the descriptor is absent and the DR flag is on (re-key + descriptor rebuilt from the endpoint + `pbsdr_adopted` audit row); DR flag off still refuses, endpoint untouched (both pinned, red-proofed). **NOT proven on tester-1's real state:** re-issue needs an enrolled host and tester-1 has none. **ep0 cleanup DONE on the operator's yes (2026-09-15 08:33Z):** the one doorstep snapshot `ns tester-1 / ct/9201/2026-09-14T16:04:53Z` (logical 1 928 820 672 B) forgotten via the local PBS API; namespace lists empty; token `felhom@pbs!tester-1` kept; nothing outside tester-1 read or touched (`D3-*` in `audits/evidence-p1fixes-2026-09-15/`). The token-only release on host delete was NOT built → R-526. **Closes on:** adopt ending in a descriptor the next tester-1 box consumes. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
| **R-516** | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. | **READY — rank P3-LOW; owner: CC (controller + catalog)** |
| **R-517** | **[P1-HIGH] After a failed off-site whole-system backup, „Biztonsági mentés" tells the customer the full backup is current and that a remote copy on separate hardware exists — neither is true.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, controller 0.242.0, customer `tester-1` with the DR tier ticked but never provisioned — R-511): „Mentés most" at 19:03:23Z; the local tier succeeded (8 877 619 753 B, 362 s); the `felhom-pbs` tier then failed — agent: `could not activate storage 'felhom-pbs': storage 'felhom-pbs' does not exist`; `pvesm status` lists only `local` and `local-lvm`. At 19:15:36Z `/backups` read: „✗ · **Utolsó teljes mentés 2026-09-14 21:09 (5 perce) · 0 B · Biztonsági szerver – külön hardver (PBS) · Naprakész**" and „✓ **Távoli rendszermentés — külön hardveren (PBS)**" (`phase4/backup-pages-after.txt`). The successful 8.9 GB local backup is no longer shown; a 0-byte failed attempt is labelled up to date; the remote tier is ticked as present. The CLAUDE.md rule „presence is not success" in page form: an attempt's timestamp stands in for a result. The hub did raise a true `whole_guest_backup_failed (error)` to the operator; the customer's page says the opposite. **Fix shape:** the tile shows the newest SUCCESSFUL backup per tier, a failed tier as failed, and a tier whose storage does not exist as „nincs beállítva". **Also measured after F2's reboot (19:45:41Z):** the whole-system tile read „– · Utolsó teljes mentés – · Méret / cél · **Naprakész**" — no backup listed at all, still labelled up to date (the 21:03 CEST local success no longer shown). | **READY — rank P1-HIGH; owner: CC (controller)** |
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-518** | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-519** | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql`**19:40:07**, `volume-dumps/*.tar`**19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-520** | **[P3-LOW] A power cut during a guarded Update leaves no record that an update was running, so whether the journal resumes or aborts honestly is unmeasured for a real version change.** MEASURED 2026-09-14 (BIGNIGHT F3, VM 333): the only Update the catalog allowed after the Phase 4 revert was a same-version one on nextcloud; power was cut 2 s after `update nextcloud: phase pulling` (safety dump written, pin advanced to the unchanged definition). After boot: no log line resumes, aborts or names the interrupted update; `app.yaml``pinned_images` = `installed_images`, no hold, no verdict; page „Fut · Naprakész". Nothing wrong was produced — and nothing could have been, with identical images. **What it needs:** the same cut on a real bump (a throwaway app with a one-step catalog move on a scratch branch or the scratch guest), asserting the page's message and the pin after boot. | **READY — rank P3-LOW; owner: CC (drill)** |
| **R-521** | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** |
| **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-523** | **[P1-HIGH] If the controller container is killed, nothing restarts it: the household's dashboard is gone and no screen can bring it back — the big night's stop rule.** MEASURED 2026-09-14 (BIGNIGHT F9, VM 333, controller 0.242.0): `docker kill felhom-controller` at 21:34:39Z, 4 s into a deploy. The container stays `Exited (137)` with `restart=unless-stopped` (Docker does not restart a container stopped by kill); the in-guest `felhom-controller-bootstrap` unit is a one-shot (`active (exited)` since boot) and does not watch it; the host agent does not either. The dashboard answered **502** for 33 min until the harness power-cycled the box. The app being deployed came up by itself (`homebox … (healthy)`). The hub raised `node_stale` at 22:04:43Z (30 min after the last report, which reached it 3 s before the kill) and **suppressed the operator mail by cooldown** (F8's `node_stale` 39 min earlier) — so for over half an hour neither the household nor the operator was told. Recovery by power-cycle (`qm reset` 22:08:17Z): the bootstrap started the controller at boot (22:10:23Z), dashboard 200 at +131 s, all 12 apps running at +229 s, the deployed homebox `running · deploying false · deployed true` — „Fut · Naprakész", **not stuck**. Hub `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z with its mail **also suppressed by cooldown**. `docker kill` is the brief's injection; the same state follows any stop that Docker records as deliberate (an operator's `docker stop`, a failed self-update that stops the old container). Memory note „controller DOES auto-recover — test with kill -9/OOM, never docker kill" describes the mechanism, not the consequence: nothing watches for a controller that is simply not running. **Fix shape:** a systemd watchdog (or the agent) that starts `felhom-controller` whenever it is not running and the operator has not parked it; restart policy `always`. | **READY — rank P1-HIGH; owner: CC (controller bootstrap / agent)** |
| **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.****Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** |
| **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** |
| **R-526** | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** |
| **R-527** | **[P3-LOW] The catalog flag `locked_after_deploy` is read by no controller code — every setting is read-only after install whatever the catalog says.** FOUND 2026-09-15: `stacks/metadata.go` parses it; `grep -rn LockedAfterDeploy` finds no reader; `deploy.html` renders „Az alábbi beállítások csak olvashatók" for every field. Recorded as the design in `02-controller-module-map.md`; the flag is a seam never wired. **Fix shape:** remove the flag from the catalog, or wire an editable-after-install allow-list (a bigger change). | **READY — rank P3-LOW; owner: CC** |
| **R-528** | **[P2-MEDIUM] Docker does not report an OOM kill inside a Felhom LXC guest: `OOMKilled` stays false and no `oom` event fires, so the v0.243.0 OOM line is not proven live.** MEASURED 2026-09-15 on scratch 9202 (Docker 29.8.0): Paperless capped at 128M restarted 11 times with `OOMKilled=false` and zero `docker events --filter event=oom`; a memory hog inside the running container was killed (rc 137) with the same silence (`E2-oom-signal-measure-9202.txt`). BIGNIGHT VM 333 did read `oomkilled=true`, so the shape differs by case. **Fix shape:** the agent reads the guest container cgroups' `memory.events oom_kill` counters (host-side, reliable), or the controller alarms on a restart-count trend. | **READY — rank P2-MEDIUM; owner: CC** |
| **R-529** | **[P3-LOW] The agent-plane `host_stale` / `host_down` / `host_recovered` mails still wait out the one-hour quiet rule.** The 2026-09-15 ruling (decision A) named `node_*` only, and the task fenced „a design is not a defect". The same F9 silence can happen on the host plane. **What it needs:** the operator's word whether the ruling extends to `host_*`. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** |
| **R-530** | **[P2-MEDIUM] A floor does not deliver an agent: agents update only by an operator-signed `agent_update` job per box, and nothing records which boxes still run 0.130.0.** MEASURED 2026-09-15: the hub HOLDS a floor whose declared MinAgent is above the box's agent (`api/handler.go` ResolveManagedFloor); the agent's only update path is `signedjobs` + `selfupdate.Executor`. demo-hp reached 0.131.0 by `felhom-opsign -op agent_update` (key `felhom-op-1`) at 08:44:16Z and its controller floor was then SERVED in 3 s. **demo-felhom (N100) and Peti's box still run 0.130.0** — not touched (Peti fenced; N100 not asked). **What it needs:** the operator signs per box, or rules a fleet rollout step. | **WAITING-ON-OPERATOR — rank P2-MEDIUM; owner: operator (signing)** |
| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** |
| **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** |
| **R-533** | **[P3-LOW] The operator signing keys were placed on DooPlex world-readable (mode 664) and sit outside any documented location.** OBSERVED 2026-09-15: `/mnt/5_hdd/felhom.eu/felhom-op-operational`, `felhom_op_ed25519`, `felhom-rec-recovery` arrived 664; CC set them to 600 (no other change). Their fingerprints match the signers demo-hp pins. **What it needs:** the operator decides where the keys live between sessions (hardware key, or a documented 0600 path) and records it in `operations/nodes.md`. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
## 2026-09-15 — New page: /letoltes (the installer download)
The volunteer guide sends people to `felhom.eu/letoltes`; the page did not exist. It now names installer
**1.27.1**, its size and SHA-256, how to check the sum on Linux/Mac and Windows, and — first — the disk
warning: the installer erases the disk you choose, never picks one, unplug the backup drive, call the
operator if unsure. `noindex` (it is for people holding the guide). Same skeleton as the other pages, no
nav entry; added to `sitemap.xml` and to `scripts/site_gates.py` PAGES; site gate OK. Published together
with the ISO (round trip verified first), so the link never pointed at a missing file.
## 2026-07-19 — Restored: the index grid background
The subtle grid behind the hero was **not** removed on purpose. It lived as a fixed `body::before` in
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.