# Journal — BIGNIGHT 2026-09-14 (a household's first month in one night) Times UTC unless marked. Hub logs print CEST (UTC+2). Brief: `drills/BIGNIGHT-2026-09-14.md`. ## Phase 1 — baselines (read 17:29–17:40Z from live source) | repo | main == origin/main | version | |---|---|---| | felhom-controller | `406755fa8fba` | v0.242.0 (CHANGELOG top) | | felhom-agent | `4586f0f7f6d1` | v0.130.0 (tree: one untracked `scripts/__pycache__/`, no tracked change) | | felhom.eu | `a4d684412b49` | hub v0.113.0 (CHANGELOG top; live image `felhom-hub:0.113.0`) | | app-catalog-felhom.eu | `6d6eec307939` | — | Register: highest id R-508; 221 table rows (grep `^| R-n`). ISO 1.27.1: `/mnt/5_hdd/felhom.eu/felhom-iso/out/felhom-installer-1.27.1-pve9.2-1.iso`, 1 705 322 496 B, sha256 `25637007d5a7120ff9faa6b5b7ead3e33c0a361ac2d67e9fd4e0ee77c034c053` = `.sha256` = manifest `output-sha256`; manifest `repo-commit 27e8ec86…`, mode release, answer-file NONE. **Found, not rebuilt.** DooPlex headroom: `/mnt/5_hdd` 37 %, `/` 52 %. demo-hp (17:29Z): `qm list` empty; 9201 + 9202 running (each `memory: 25898`, `cores: 7`); host RAM 29 994 MB total, 24 652 MB available; 8 threads (Ryzen V1756B); storages `local` 56.96 %, `local-lvm` 44.17 %, `nvme-scratch` (dir `/mnt/hdd_1`, is_mountpoint yes) 1.03 %. Gmail connector: authenticates as **`felhom.eu@gmail.com`** — the catch-all for `@felhom.eu` (threads to `admin@felhom.eu` and `drill0242@felhom.eu` land there). Read only; nothing sent. ## Phase 1 — the Tester 1 record (hub `GET /customers/tester-1`, Basic auth, 17:30Z) Customer ID `tester-1` · Name „Tester 1" · Domain `enkicsifelhom.hu` · **Email `tester1@felhom.eu`** · Config MANAGED. Status tile still reads `down`, „Last report 1h ago · Controller 0.242.0" — the stale report of the destroyed VM 331 (host record deleted 16:46:52Z). Claim: „Claimed 1h ago · generation 1". Edit tab: `dr_tier` **ON** (all four DR steps „waiting — no host enrolled yet"); **`offsite_enabled` OFF** („Enable offsite (provisions a Hetzner Storage Box on save)"). Controller floor v0.242.0. Mailbox: `to:tester1@felhom.eu` → 0 threads before tonight. **Brief vs record — off-site.** The brief says "Off-site (Tier 3) is ON … its own namespace on ep0". On the record, the ep0 namespace is the **DR tier (PBS whole-guest)**; the **Tier-3 restic off-site is OFF**, and turning it on provisions a Hetzner Storage Box (money, and a new external resource). **Not followed:** I do not tick it — money is fenced by the rules file §1. "Off-site" tonight = the DR tier to ep0 as configured. Consequence, stated up front: the Phase 4 off-site integrity check and the Phase 6 app restore "from off-site" have no restic tier to act on; each is recorded as such when reached. **Harness slip:** the first read of the customer page printed the page's text into this session's tool output, which includes the customer's hub **API key** (not the retrieval passphrase — that is masked). It stayed on DooPlex; not written to any file. ## Tunnel before any box (17:31:19Z, from DooPlex) — `tunnel-before-no-box.txt` DNS resolves to Cloudflare (both resolvers); `https://felhom.enkicsifelhom.hu` → **530, `error code: 1033`** (no connector). Expected with no box. ## Venue — VM 333 `bignight-household` on demo-hp (created 17:31:52Z) q35 / OVMF (`pre-enrolled-keys=0`), **4 cores**, `cpu=host`, **16 384 MB**, virtio NIC `BC:24:11:F2:E2:95` on vmbr0, `scsi0` **200 G** qcow2 on `nvme-scratch` (dir at `/mnt/hdd_1`, its root), ISO 1.27.1 on `ide2` (sha verified on the HP = `25637007…`). **One disk at install**; the data disk is added after. **Why these numbers.** System disk = doorstep VM 331's 200 G. Memory: the HP has 29 994 MB; 9201 and 9202 together used ≈ 5.3 GB at 17:29Z with 24 652 MB available. Both LXCs carry a 25 898 MB *limit*, so no VM size leaves both limits whole; 16 GB leaves ≈ 8 GB free plus 8 GB swap for 9201's real load. Cores: the 0242 drill's 4 of 8 threads. Power-on **17:32:13Z**. GRUB (`s01`, `s02`): the two Felhom entries; **text mode** chosen (H4). ## Phase 2.1 — install from 1.27.1 (text mode), one disk | UTC | screen | what it said / what was done | |---|---|---| | 17:34:5x | `s03` | English Proxmox EULA → „I agree" | | 17:35:11 | `s04` | „Target harddisk: /dev/sda (QEMU HARDDISK) (200.00 GiB)" — **one disk, no choice to make** → Next | | 17:35:24 | `s05` | Country Hungary · Timezone Europe/Budapest · Keyboard layout **Hungarian** (H2: changed to U.S. English, `s06`–`s09`; one dropped keystroke landed on „Turkish" first, corrected) | | 17:37:00 | `s11` | „Root password [at least 8 characters]" · Confirm · „Administrator email" **prefilled `mail@example.invalid`** → 20-char generated password (0600 file on DooPlex), `tester1@felhom.eu` (the guide: „a saját e-mail címedet") | | 17:39:01 | `s13` | nic0 `bc:24:11:f2:e2:95, virtio_net` · Hostname **`pve.example.invalid`** · IP `192.168.0.136`/24 (the DHCP lease, frozen static) · GW `192.168.0.1` · DNS `192.168.0.250` · „[X] Pin network interface names" → hostname set `felhom.enkicsifelhom.hu` per the guide (`s15`) | | 17:40:36 | `s16` | Summary: ext4 · /dev/sda · Europe/Budapest · U.S. English · tester1@felhom.eu · nic0 · felhom.enkicsifelhom.hu · 192.168.0.136/24 · 192.168.0.1 · 192.168.0.250 · „[X] Automatically reboot after successful installation" (H3: unticked, `s17`) | | 17:41:02 | — | Install pressed | | ≤17:43:47 | `s18` | „Success — Installation finished - reboot now?" (**≤ 2 m 45 s** copying; qcow2 4 058 MB, settled) | | 17:44:xx | first boot | H3: `qm stop`, `--delete ide2`, `--boot order=scsi0`, `qm start` | Seeds prepared on DooPlex (scratchpad, tmpfs): 200 JPEG 1600×1200 noise + caption (264 MB), 20 three-page PDFs (2.3 MB), a 50 MB random file, a 150 s 1280×720 H.264 video (98.9 MB). ## Phase 2.2 — first screen, and waiting for the mail **17:44:53Z — hub lists Unclaimed appliance 28** (38 s after power-on): smbios uuid `1ecd1c40…`, pairing code `***-***` (redacted), MAC `bc:24:11:f2:e2:95`, „Standard PC (Q35 + ICH9, 2009)", 15.6 GB, three SSH host keys. Bind list offers „Tester 1 (0 hosts)". **The console, 17:45:31Z (`s19`, code redacted), verbatim:** ``` Felhom otthoni szerver Ezen a gépen most nincs dolgod, és bejelentkezni sem kell. A beállításhoz kövesd a Felhomtól kapott útmutatót. felhom login: ============================================== Felhom — a doboz készen áll, és a párosításra vár. Párosító kód: ***-*** Nyisd meg az e-mailben kapott linket, és add meg ezt a kódot és a Tulajdonosi jelmondatodat (az 5 szót a Felhom üzemeltetőjétől kaptad). Ez a képernyő magától frissül — nincs teendő a doboznál, és nyugodtan itt hagyhatod bekapcsolva. ============================================== ``` No Proxmox `:8006` line on the first boot (the 1.27.1 fix holds on a fresh install). Text says „otthoni szerver"; the guide §3 quotes „Ezen a képernyőn nincs teendőd" — **the guide's quote does not match the screen word for word** (meaning is the same). **Mailbox, 17:46Z:** `to:tester1@felhom.eu` → 0 threads. The console asks for „az e-mailben kapott linket". Hub source: the self-bind link is auto-sent only at **customer creation** (`configs.go:725`) and at **RESET completion** (`customer_reset.go:162`); `tester-1` was created 32 days ago with no e-mail, so no link was ever sent to `tester1@felhom.eu`. Waiting the brief's 10 minutes (to 17:55Z) before acting. **17:55:11Z — ten minutes, 0 messages to `tester1@felhom.eu`.** Filed **R-509 (P1)** at 17:56Z, before acting. **Intervention I1:** the operator's „Send self-bind link" button on the customer's Setup tab is pressed (a volunteer cannot press it). **17:55:42Z** `POST /customers/tester-1/selfbind-link` → 303 `flash=selfbind-sent`; hub: `self-bind link (hash effe179d…, valid 7 days) emailed to the registered address of tester-1`. **The mail, as the volunteer reads it** (Gmail connector, message `1a0a10f8cd07d8f6`, Date 17:55:43Z — **1 s** after the press; Gmail threads it under the older drill0242 mail of the same subject, a mailbox quirk): > **[Felhom] Kösd össze a Felhom dobozodat** — from `monitoring@felhom.eu` to `tester1@felhom.eu` > > Kedves Ügyfél! > Elkészült a Felhom dobozod, és készen áll az összekötésre. Az alábbi hivatkozáson tudod te magad > összekötni a fiókoddal — nincs szükség bejelentkezésre: > `https://hub.felhom.eu/bind/<64 hex, redacted>` > A hivatkozás megnyitása után két adatot kell megadnod: > 1. A párosító kódot, amely a doboz képernyőjén (a monitoron) látható. > 2. A tulajdonosi jelmondatodat (az 5 szóból álló kifejezést), amelyet a Felhom üzemeltetőjétől kaptál — > személyesen vagy telefonon, e-mailben soha. Ez igazolja, hogy a fiók a tiéd. > A hivatkozás 7 napig érvényes. Biztonsági okból 5 sikertelen próbálkozás után zárolódik — ilyenkor vedd > fel a kapcsolatot az ügyfélszolgálattal. > Ha nem te kérted ezt, hagyd figyelmen kívül ezt az e-mailt. > Üdvözlettel, Felhom.eu The Tulajdonosi jelmondat is taken from the hub customer page into a 0600 file (the operator's hand-over, per the guide's prerequisite 3 — not a volunteer act, not counted). ## Phase 2.2 (cont.) — following the link (the self-bind page, first time exercised by any walk) **17:56:11Z `GET https://hub.felhom.eu/bind/` → 200** (`screen-bind-page.txt`), verbatim: „Felhom — Doboz összekötése" · „Kösd össze a most telepített Felhom dobozodat a fiókoddal. Add meg a doboz képernyőjén látható párosító kódot és a tulajdonosi jelmondatodat." · **Párosító kód** („A doboz monitorán jelenik meg, a telepítés után.", placeholder `ABC-234`) · **Tulajdonosi jelmondat** („Az öt szóból álló kifejezés, **amelyet a beállításkor kaptál**.", placeholder „öt szó, kötőjellel vagy szóközzel") · **Összekötés** · „Biztonsági okból 5 sikertelen próbálkozás után a hivatkozás zárolódik. …" Observation: the page still says the phrase was received „a beállításkor" (at setup); the mail (hub v0.113.0) says „a Felhom üzemeltetőjétől kaptál — személyesen vagy telefonon". Two wordings for one hand-over. **17:56:29Z `POST /bind/` (pairing code + Tulajdonosi jelmondat) → 200** (`screen-bind-result.txt`): „**Sikeres összekötés.** A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt bezárhatod — a beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz." **The self-bind link works end to end — first time any walk exercised it.** One try, no error. Hub (CEST): 19:56:30 `self-bind SUCCESS: appliance 28 bound to customer tester-1 by customer self-service` · 19:56:47 credentials DELIVERED once · 19:57:23 `host enrolled: tester-1-a61396` · **19:57:24 `[claim] reenroll code (gen 2) emailed … reset code re-issued (gen 2) for tester-1 on box re-enrollment (clean-slate reinstall)`** · 19:57:25 manifest agent 0.130.0 golden 0.242.0 · 19:57:42 `wg registered … ip=10.77.0.5/32` · **19:57:42 `[ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action`** · 19:59:06 `controller_started (info) — Controller elindult (0.242.0)` · 19:59:43 `tester-1 down → ok`. Console after bind (`s21`, codes redacted): the pairing banner is printed **three times** with the same code; it never changes to "bound" (R-214 class). **Which claim path — the brief asked "fresh claim or reset flow".** **Neither a fresh claim nor the RESET:** the hub treated the box as a **re-enrolment of a claimed customer** and sent the *reinstall* mail (gen 2). Quoted (Gmail `1a0a111205fa7fe8`, 17:57:24Z, 54 s after the bind): > **[Felhom] Új beállító kód — újratelepült a szervered** > Kedves Ügyfél! > A Felhom szervered újratelepült, ezért a vezérlőpultod belépését újra be kell állítani. **A korábbi jelszavad > már nem érvényes.** > Beállító kód: `<3 words, redacted>` > A kód 72 óráig érvényes, és egyszer használható fel. Nyisd meg a vezérlőpultot — "A szerver beállítása" oldal > fogad —, add meg a kódot, majd válassz új jelszót: > https://felhom.enkicsifelhom.hu > Ha nem te telepítetted újra a szervered, vedd fel a kapcsolatot az üzemeltetővel. For a volunteer on their **first** box this mail is wrong in two places: the guide §5 tells them to look for „Elindult a Felhom szervered — beállító kód", and this one says their server was *reinstalled* and their *previous password* no longer works — a password they never had. A careful stranger reads „Ha nem te telepítetted újra … vedd fel a kapcsolatot" and stops. (Cause: the record was claimed once today by VM 331.) **Tunnel after bind, before claim — 18:07:48Z from DooPlex (`tunnel-after-bind.txt`): 502 ×3** (`server: cloudflare`, body `error code: 502`). Changed from 530/1033 (no connector) to 502 (connector present, origin failing). ## Phase 2.4 — the gate: the dashboard through the tunnel — **FAIL (P1)** `box-logs-phase2/cloudflared-before-claim.txt`, `tunnel-origin-diagnosis.txt`. Guest 9201 on the box: `192.168.0.150`; containers `felhom-controller:0.242.0 (healthy)`, `traefik:v3.6.7`, `cloudflared:2026.6.0`, `filebrowser 1.3.3-stable`. cloudflared: 4 connections registered (17:59:02Z, vie06 …), pre-checks all PASS. Then every request: `Request failed … tls: failed to verify certificate: x509: certificate is valid for 544346c4….traefik.default, not traefik … ingressRule=0 originService=https://traefik`. Traefik on 127.0.0.1:443 with SNI `felhom.enkicsifelhom.hu` serves **`CN=*.enkicsifelhom.hu`, Let's Encrypt YR1, valid 2026-09-14 17:01 → 12-13** — the right certificate exists. **Control (demo-hp 9201, read only):** its route `{"hostname":"*.enkisfelhom.hu","originRequest":{"noTLSVerify":true},"service":"https://traefik"}`. **Conclusion:** the operator's route reaches the box now (530 → 502) but lacks „No TLS Verify". Filed **R-510 (P1)** at 18:1xZ **before acting**. **Intervention I2:** from here the claim and all dashboard traffic go to `192.168.0.150:443` with the name forced (`--resolve`), from DooPlex on the same household LAN as the box. ## Phase 2.2 (cont.) — the claim, with the mailed code (I2 route) `screen-claim-page.txt`. **18:10:06Z `GET /`** → 302 `/claim`; the page: „**A szerver beállítása** — Tester 1 — Add meg az e-mailben kapott beállító kódot, majd válassz saját jelszót a vezérlőpult védelméhez." · Beállító kód · Új jelszó (min. 12 karakter) · Új jelszó megerősítése · **Beállítás és belépés** · „Nem kaptad meg a kódot? Új kód kérése". (The page does not mention the reinstall the mail talked about — the two agree on the code, not the story.) **Harness slip:** the first `POST /claim` (18:10:06Z) sent both of the page's `_csrf` values joined (two forms, two tokens) → 200 „Érvénytelen űrlap — töltsd újra az oldalt." (form refused, not a code attempt). Re-fetched, one token. **18:10:28Z `POST /claim`** (the **mailed** code, a generated 24-char password twice) → **302 `/`** + `felhom_session`. **The mailed setup code works — first walk to use the real mail instead of a box-printed code.** `/launcher` 200: „Indítópult — Felhom.eu", tile **Filebrowser**, sharing off; menu Vezérlőpult · Alkalmazások · Tárhely (Meghajtók, Hálózati tárhely) · Biztonsági mentés (Áttekintés, Távoli mentés, Alkalmazások, Visszaállítás) · Megosztás (Hálózati megosztás) · Rendszermonitor · **Debug** · Beállítások (Rendszer, Értesítések, Biztonság és hozzáférés) · `0.242.0`. **Power-on → claimed dashboard: 38 m 15 s** (17:32:13 → 18:10:28), of which 10 min were the brief's mail wait. ## Phase 2.3 — version Box: `felhom-controller:0.242.0 (healthy)`, `felhom-agent 0.130.0`. Hub: „Controller 0.242.0 · Registry latest v0.242.0 — up to date · Effective floor v0.242.0 — at/above floor". **PASS — landed on golden 0.242.0, reports current, no self-update needed** (golden == floor). Bind → `controller_started` 2 m 36 s. ## Phase 2.4 — the gate after the claim — **FAIL, still R-510** `tunnel-after-claim.txt`: 18:10:41Z from DooPlex `https://felhom.enkicsifelhom.hu/login` → **502 ×6**; box cloudflared at 18:10:42Z: the same `x509 … traefik.default, not traefik` for each. Continued on the LAN address (I2). ## Phase 2.1 (cont.) — the data disk (100 G `scsi1`, hot-attached 17:46:42Z, before the bind) - Box `lsblk`: `sdb 100G disk`, no partitions. - **Dashboard** (`screen-storage-dashboard-after-claim.txt`): „Lemezek állapota — QEMU QEMU HARDDISK · Nincs adat · 0 °C" (one line, which disk is not said); **nothing mentions a new disk.** - **Tárhely → Meghajtók:** „**Nincs regisztrált adattároló.** Adjon hozzá egyet az alábbi űrlappal." · „Új meghajtó inicializálása" · „Meglévő meghajtó csatolása" · „Nem regisztrált meghajtók — az ügynök által észlelt, még nem regisztrált adatmeghajtók" · „Rendszermeghajtók — védett …" · „Már csatlakoztatott tárhely hozzáadása kézzel — Elérési út — Pl. /mnt/hdd_1 …". Formal „Adjon" (the rest of the product says „te"). - `GET /api/disks/candidates`: `initialize: [/dev/sdb 107374182400 B, QEMU HARDDISK, data_bearing false]`; `attach: [/dev/mapper/pve-vm--9201--disk--1, mount_source /mnt/sys_drive, already_mounted true]` (the guest's own system volume offered for attach — observed in the 0242 drill too). - **Would a household know to enrol it? No.** Nothing on the dashboard, the launcher or any mail says a disk appeared; the volunteer guide has no step for it; the page that lists it is two menus deep and speaks of „inicializálás" and „ügynök". A volunteer who plugs in a second disk has to go looking. ## The DR tier on this box — stuck (operator act O1, not a customer step) Hub 19:57:42 CEST: `pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action`. **18:11:57Z operator `POST /configs/tester-1/ pbsdr-reissue` → 400 `No provisioned PBS DR tier for this customer`.** Neither path provisions it. Filed **R-511 (P2)**. Not forced further: removing the old ep0 token is a destructive act on ep0 tenancy, and RESET would also remove the tunnel (fenced by the brief: the customer is kept). **Consequence for tonight, stated once:** this box has **no off-site tier of either kind** — restic Tier-3 is off (money) and the PBS DR tier cannot provision. Nothing is written to ep0 tonight. Phase 4's off-site integrity check and Phase 6's restore „from off-site" are recorded as not possible when reached, with the local tiers walked instead. No „A vezérlőpultod mostantól jelszóval védett" mail arrived for the claim (the 0242 drill's box received one at 13:31:23Z, 12 s after its claim); at 18:12Z the mailbox holds only the bind mail and the reinstall mail. ## Phase 2.1 (cont.) — enrolling the data drive, as a household could from the screens 18:12:57Z `POST /api/storage/init` (the „Új meghajtó inicializálása" wizard's call) `{device /dev/sdb, ext4, mount_name hdd_1, label Adatlemez, set_default true}` → `formatting` → **`done` at 18:12:59Z (2 s)**, where `/mnt/felhom-drives/hdd_1` (`drive-init.txt`). The wizard's „Csatlakoztatási név" field has **no value, only the placeholder `hdd_1`** — a person must type a name. Tárhely now: „Adatlemez · /mnt/felhom-drives/hdd_1 · **Alapértelmezett** · Aktív · 0.0 GB / 97.9 GB · ext4 · /dev/sdb[/felhom-data] · QEMU HARDDISK · Nincs alkalmazás ezen a tárolón". `/api/disks`: role `user-data`, durable id `uuid:c4b530fd…`. **Phase 2 evidence pulled off the box 18:14Z:** `box-logs-phase2/` (bootstrap + agent journals, controller log, box state, cloudflared). Secret patterns searched (`passphrase=|retrieval_password|password:`): 0 hits. ### Phase 2 interventions: **2** - **I1** (R-509): the self-bind mail never came for an existing customer; the operator's send button was pressed. - **I2** (R-510): the tunnel answers 502; claim and dashboard over the LAN address with the name forced. Operator acts that are prerequisites, not counted: passphrase hand-over; PBS re-issue attempt (O1, refused). ## Phase 3 — a family moves in (`phase3/`) Deploy pages captured before any deploy (`phase3/deploy-pages-before.txt`); the memory line each showed at 18:14Z with 3 infra containers running: bookstack 1525 / 11444 MB · docmost 1568 · privatebin 1404 · gokapi 1408 · nextcloud 1633 · **immich 3422** · vaultwarden 1427 · paperless-ngx 1880 · jellyfin 1892 · mealie 1578 · uptime-kuma 1431 · adventurelog 1476 (guest limit 11 444 MB of the VM's 16 GB). Every deploy goes through the page's own call, `POST /api/stacks//deploy {"values":…}` with the form's server-generated values; a required empty password (gokapi, nextcloud admin) is typed as a customer would — a generated value in a 0600 file. | app | deploy | seed through the front door | household use | |---|---|---|---| | **bookstack** | 18:14:35 → running **80 s** | default login → admin e-mail + password changed (old default refused); book „Családi tudástár BIGNIGHT"; 5 Hungarian pages (Töltött káposzta, Wifi jelszó helye, Kutya oltási naptár, Nyaralás Balaton 2026, Autó szerviz); 2 attachments 256 KiB + 512 KiB, **sha equal ×2** | 2nd user „Kovács Péter" (Editor) created, logs in, edits Balaton page (edit visible); **Péter's attachment upload returned an HTML page, not JSON — inconclusive (harness: page id hard-coded)**; Anna deletes „Kutya oltási naptár" → 404 → recycle bin → restore → 200 with its text | | **docmost** | 18:16:10 → running ≈ 90 s | `/api/auth/setup` workspace „Kovács család"; space „Háztartás"; 5 docs via Markdown import (titles taken from each file's `#` heading: Rezsi, Iskola, Lista, Orvos, Kert); read back „ELMŰ", „árvíztűrő" present | renamed „Bevásárlólista — szombat"; deleted „Kert" → trash → restored → listed; invited peter@ (200) | | **privatebin** | 18:18:2x → running **15 s** | 5 encrypted pastes (browser v2 format): 1week, never, 1month, never, **10min (the expiring one)**; every one decrypts equal; TTL meta 604800 / [] / 2592000 / [] / 600. Harness: PrivateBin's 10-second post limit refused one post (status 1), retried | — (file-only; pastes are the use) | | **gokapi** | 18:18:43 → running **20 s** | deploy page asks „Admin jelszó *" (customer types it); login admin; 3 uploads via the dropzone's chunk calls: **50 MB (2 chunks, 1.6 s)**, a PDF, a JPEG; admin lists 3 | a stranger downloads all three, **sha equal ×3**; the PDF deleted through the UI's API → its link no longer serves it (78-byte page) | App pages: „Fut · **Naprakész**" (title „Ez az alkalmazás a legfrissebb elérhető változatot futtatja."). Gokapi's password: app page „Első lépések" says „a jelszó a Beállítások oldalon található"; **Beállítások** shows „Admin jelszó 👁 — Telepítéskor beállított kezdeti jelszó — ha az alkalmazásban megváltoztattad, az itt nem frissül." Findable. BookStack page „Első lépések": „Nyisd meg a **wiki.DOMAIN** címet" (R-498, still). **nextcloud** 18:20:14 deploy (HDD_PATH offered: „Adatlemez — 92.9 GB szabad (alapértelmezett)") → +50 s `starting` → **+121 s `unhealthy`** (the harness stopped polling there — its own mistake) → at 18:23:43Z `running`, all three containers `(healthy)`. So the card reads „unhealthy" for a while on a first start; recorded as observed, not a defect. **nextcloud** seed (WebDAV, 4 parallel, as the admin the deploy page asked for): folders Fotók/Balaton 2026, Számlák, Közös mappa; **200 JPEG + 20 PDF, 220 × 201, 86 s** (264 MB of photos); 2nd user „peter" (OCS 100), „Közös mappa" shared with him (OCS 200); Péter uploads into it and Anna reads it back (sha equal); foto-123 read back sha equal. Household use: a child deletes foto-007 (204 → 404); trashbin lists it; restore (201); photo back, **sha equal**. Admin used 343 161 297 B. Nextcloud 34.0.1. **immich** 18:22:57 deploy → +116 s `degraded` → +126 s `starting` … (continues below). **immich** running after 146 s (degraded 116 s on the way). Seed through its API: admin sign-up (201), **200 JPEG uploads, 200 × 201 in 16 s**; stats photos 200, 275 682 444 B; ML queues draining (faceDetection / smartSearch / ocr). **vaultwarden** 18:25:37 → running **15 s**. Seed with the Bitwarden client crypto (PBKDF2 600 000, AES-CBC + HMAC): register (200), token (200), **10 login entries** (200 × 10), `/api/sync` → 10 items, **all names decrypt to the originals**. Attachment: the v2 slot's returned `url` is `/ciphers//attachment/` — posting it to the web root 404s, under `/api` 200 (a client detail, the harness's first try left one empty attachment slot on „Ügyfélkapu+"). The retried attachment downloads 200 with its recorded size; byte-equality not proven (harness lost the key). **Finding R-512 (P2):** registration stays open and the page's instruction to close it points at a read-only field; a stranger registered with no invite (200). ### SECURITY — R-513 (P1), found 18:30Z while looking for a way to put a video on the drive The launcher's **Filebrowser** tile opens `files.enkicsifelhom.hu`: password login, no signup. **No screen gives the customer a FileBrowser login.** A customer's first guess, `admin` / `admin`, **works** (200 + token); wrong password 401. With it, both sources list (Adatlemez → `documents`, …; Beolvasás → `paperless`). The same login works on demo-hp's **9201** and **9202**, and 9201's login page answers **200 through Cloudflare from the internet** (GET only, no login attempted remotely). Filed R-513 before anything else; nothing changed on any box. **immich** (cont.) ML queues empty at 18:32:09Z (≈ 6 min after upload); asset 42 original read back **sha equal**. **paperless-ngx** 18:27:xx deploy → running. Token with the deploy-generated admin password (the app card's „admin / admin" is not what was set — the form generates a 16-char password). Tags Számla / Garancia / Adó (201 ×3); **20 PDFs posted 18:29:57Z → 0 documents**. Kernel: the container's cgroup OOM-killed `gs` and the celery worker at 18:31:22Z; tasks 11 FAILURE (WorkerLostError) · 1 STARTED · 8 PENDING, unchanged at 18:44Z. App still „Fut". **R-514 (P2).** **jellyfin** 18:29:34 → running **50 s**. Startup wizard through its REST calls (Hungarian UI culture, user „anna"), login 200; libraries none; the container sees `/media/{audiobooks,books,comics,movies,music,photos,podcasts,tv}` (the drive's `userdata/media`, read-only). **mealie** → running **85 s**. The card's default login `changeme@example.com / MyPassword` works (200); password changed (old now 401); 5 Hungarian recipes with ingredients + steps (201/200 ×5); meal plan 15–19 Sept (201 ×5) reads back all five; Gulyásleves reads back its accents. **uptime-kuma** → running **50 s**. Its socket.io endpoint answers the polling transport with the SPA page (websocket only); DooPlex has no websocket library — a raw client is written next. **uptime-kuma** (cont.) The first screen of a fresh Uptime Kuma 2.4 is an **English „Which database would you like to use?"** page (SQLite / Embedded MariaDB / …, „Next") — the controller's card says nothing about it; the socket API is not up until it is answered. SQLite chosen through its own call (`POST /setup-database` → `{"ok":true}`), then over a websocket: setup „anna" (successAdded), login ok, **3 monitors** (the box's own dashboard, the family wiki, the router by ping). Heartbeats after 90 s: **router UP (0.5 ms)**; **dashboard and wiki DOWN — „Request failed with status code 502"**: from inside the box, its own public names go out to Cloudflare and hit R-510. A household that adds its own apps as monitors gets red for every one. **jellyfin media via the product's own route — „Hálózati megosztás" (SMB):** enable (303 „Beállítás mentve…"), password (first try used wrong field names → „legalább 8 karakter" flash, harness; second 303 „Megosztási jelszó beállítva"), share „Filmek" = existing folder `userdata/media/movies` (303 „A megosztás létrehozva."). State „fut"; list shows `beolvasas` (rendszer) and `Filmek` (Felhőmentés **bekapcsolva**). From demo-hp on the LAN (a household laptop): `smbclient //192.168.0.150/Filmek -c put` **98 898 905 B in 2.6 s**. The drive root itself is refused for sharing („Ez a mappa nem osztható meg."). **jellyfin** (cont.) the video, copied through \\FELHOM\Filmek, is visible inside the container at `/media/movies/Csaladi-nyaralas-2026.mp4`; library „Családi videók" (movies, `/media/movies`) added (204); the scan finds „Csaladi-nyaralas" (RunTimeTicks 1 500 000 000 = 150 s) **in 5 s**; `static` stream of the first MiB → 206, 1 048 576 B. **adventurelog** → running **161 s**. Account creation took five harness attempts (evidence kept in order): the headless `/auth/browser/v1/auth/signup` through the frontend refuses without a usable CSRF token (the frontend clears `csrftoken` on every response); the backend's own `/accounts/signup/` form works (302) — a household would use the frontend's sign-up page, which this harness did not drive. Then the frontend login form (`POST /login`) → `sessionid` (Domain `.enkicsifelhom.hu`) and through it: collection „Balaton 2026 nyár" 201; Tihany, Badacsony, Keszthely 201 with a visit each 201; Badacsony edited (rating 4) 200; read back 3 locations with their visits. **Photos: 502 then 500** — see below. (An anonymous POST to `/api/collections/` makes the backend raise `ValueError … AnonymousUser` → 500: the app's own behaviour, noted, not filed.) **Version labels** (`phase3/version-labels.txt`, 18:56:27Z): all 12 app pages „Fut · **Naprakész**" (ASCII fragment `Naprak` = 1 each, control 0). **The memory guard never refused.** All twelve deployed on the default box; the last deploy page before immich read 3 422 / 11 444 MB and before vaultwarden 3 734 / 11 444 MB; `free` in the guest at 18:45Z: 3 520 MB used of 11 828. **How many apps fit on a default box: at least these twelve, with room left**; the one app that ran out of memory did so inside its own container cap (R-514), which the guard does not see. **adventurelog photos** — the frontend proxy logs `RequestContentLengthMismatchError: Request body length does not match content-length header` for the multipart POST: the **R-483** class („photo upload fails from any non-browser client"; the operator confirmed the browser path on 2026-09-13). Not re-filed; the harness cannot prove the browser path. **Hub during Phase 3:** 12 × `app_deployed (info)` for tester-1, 20:14:35 → 20:45:42 CEST; nothing else — no alarm for the Paperless OOM (R-514). ### Phase 3 interventions: **0** (every seed through the app's own front door or the product's own SMB share; harness retries are recorded, none needed an act a customer could not make). Rows: R-512, R-513, R-514, R-515, R-516. Evidence pulled off the box 19:00Z: `box-logs-phase3/`. ## Phase 4 — a month of routines (`phase4/`) Backup pages read before (`phase4/backup-pages-before.txt`). The nightly's legs, each through its own endpoint, in the nightly's order (window: db-dump W, Tier-2 W+60 m, off-site W+105 m, whole-system after): | leg | endpoint | start → end | result | |---|---|---|---| | Tier 1 — DB dumps + recovery units („Mentés most" on Alkalmazások) | `POST /api/backup/run` | 19:00:26 → 19:02:38 | `db_dump count 6, 2m8s, success true` | | Tier 2 — off-drive copies | `POST /api/backup/tier2` (no page calls it; the nightly does) | 19:02:5x → (below) | | | Tier 3 — off-site | none: not configured on this record (restic off; PBS DR stuck, R-511) | — | not run | | whole-system (Plane-2 local) | `POST /api/guest-backup/trigger` („Mentés most" on Áttekintés) | (below) | | Catalog: **no app had a newer version** (12 × template == catalog, all `catalog_since 2026-07-18`). So a real one-step bump: **privatebin `2.0.5 → 2.0.6`** (upstream Docker Hub 2026-08-08), catalog commit `d5d91e0` pushed **18:59:36Z** with `catalog_since 2026-09-14`, gates OK. At 19:02:41Z the box still reads 2.0.5 (the 15-minute sync). R-499 still present on `/stacks/bookstack/backup` („már szerepelnek a teljes rendszermentésben (PBS)"). Tier 2 result (controller log, `phase4/controller-log-t1-t2.txt`): `Tier 2 run complete: 12 app(s) processed` at 19:03:06Z, **17 s**. The eight system-disk apps copied to the data drive (`hdd_1/backups/secondary/…`, e.g. bookstack 154.9 MB); the four drive apps copied to the internal SSD marked **`[SSD: state-only]`** (immich 1.7 GB, nextcloud 1.3 GB, paperless 68.8 MB, jellyfin 0.6 MB) — definitions and databases, not the photos, as the Nextcloud page said („csak az adatbázis és a konfiguráció másolódik"). The status API's `running` flag covers only the DB leg, so it read false during Tier 2 (the harness loop exited on it; the log is the observable). **Whole-system „Mentés most"** (`POST /api/guest-backup/trigger` 19:03:23Z; `phase4/guest-backup-run.txt`, `guest-backup-quiesce-log.txt`): the controller quiesced **all 12 apps** („backup due on 2 tier(s)"); tier `local` 19:03:49 → 19:09:59Z, **8 877 619 753 B, 362 s, success**; then tier `felhom-pbs` started with the apps still stopped, failed (`storage 'felhom-pbs' does not exist` — the DR tier that never provisioned, R-511), unquiesce 19:10:09Z, last app started 19:11:12Z → **apps down ≈ 7 m 45 s** under a page line promising „csak néhány másodpercre" → **R-518 (P2)**. Backoff logged: next PBS attempt in 15 min „so the apps are not stopped again for a backup that cannot succeed". Harness: the status API kept the previous job's `done` until the new job id appeared; the first poll stopped on it. **Backup pages, read as a first-timer after the run** (`phase4/backup-pages-after.txt`, 19:15:36Z): - Áttekintés: „✗ · Utolsó teljes mentés 2026-09-14 21:09 (5 perce) · **0 B** · Biztonsági szerver – külön hardver (PBS) · **Naprakész**" and „✓ **Távoli rendszermentés — külön hardveren (PBS)**" — **false on both counts**; the real 8.9 GB local backup is not shown → **R-517 (P1)**. Dates are local time (21:09 CEST) ✓. „DB mentések 6 fájl" ✓. „Következő mentés — 0 órája" (reads backwards, seen in the 0242 drill). Both honest warnings about one disk remain although a data drive is now enrolled — one of them offers „Adatlemez · Kijelölöm" as the whole-system target ✓. - Alkalmazások: every app „1. mentés Auto helyi Utolsó: …"; drive apps' second copy state-only as said. - `/stacks/bookstack/backup` still claims PBS coverage (**R-499**). - **Off-site integrity check: not possible on this record** — „Távoli mentés": „Még nincs beállítva távoli mentési cél"; `/backup/offbox/status` → `snapshots 0, status ""`. Verdict and duration: **none — no tier to check** (recorded, not run). **Hub after the routines** (`phase4/hub-check.txt`, 19:15:31Z): „Tester 1 · **ok** · Controller 0.242.0 · Last report 2 min ago · Containers **24/24**" (all 12 apps' containers listed), storage „SSD 33 % · **Adatlemez 4 %**", „Last DB dump 13 min ago", „No off-site data reported … this page cannot say". 13 × `crossdrive_completed (info)`; **21:11:12 `whole_guest_backup_failed (error)` + operator e-mail — a TRUE alarm** (the tier really cannot succeed), with no word of the local success. The hub page shows no whole-system backup line at all. **Catalog travel:** the box synced at 18:59:00Z (36 s before the push) and picked the bump up at **19:14:01Z** („Sablonok frissítve — frissítve: privatebin") — 14 m 25 s after the push. App page: „Frissítés elérhető — **ma**" (title „Újabb változat érhető el …"). PrivateBin before the update: 2 pastes decrypt equal; the 10-minute paste answers „Document does not exist, has expired or has been deleted." — honest expiry. **The guarded Update, privatebin 2.0.5 → 2.0.6** (`phase4/privatebin-update.txt`): `POST /api/stacks/privatebin/update` 19:16:33Z → „Frissítés elindult – az állapot a kártyán követhető". Phases: checking → **precondition met — Tier 2 (second drive) copy from 19:03:05Z (13 m old, limit 24 h)** → safety-dump (no database, no-op) → **pinning** („pin advanced to the catalog's current definition") → pulling → starting → verifying → „healthy after 5s" → installed-images recorded `privatebin/pdo:2.0.6 (sha256:4c141b23…)` → **DONE in 11 s**. Container `privatebin/pdo:2.0.6 (healthy)`. App page „Fut · **Naprakész**". **Data read back after: both surviving pastes decrypt equal**; the expired one still expired. PASS. (Observation: `template_images` in the API still read 2.0.5 for ≈ 15 s after `updating` went false.) **Catalog revert** `a161ccb` pushed 19:17:55Z (image back to 2.0.5; `catalog_since` stays 2026-09-14 by the gate). Phase 4 evidence pulled off the box 19:18Z (`box-logs-phase4/`). ## Phase 5 — the accidents (`phase5/`) Method, the same for every fault: a **family loop** (`family.py`) keeps three apps busy (BookStack page loads each second, a Nextcloud note written every 2 s, Mealie API each second) and logs ok/fail per 5 s; `watch_recovery.py` records from the fault's T0 when the dashboard's `/api/health` answers and when each of the 12 apps is `running` again with the **same images** as the Phase 4 baseline; `alarms.sh` reads the hub's events, staleness changes and operator mails for tester-1; the Gmail connector is read for what reached the customer's and the operator's mailbox. For each fault: what the customer saw · what the box did alone · time to steady state · which alarm fired and was it true · which alarm should have fired and did not. Evidence off the box before the next fault. **Before F1 — does the failed PBS tier stop the apps again?** The controller deferred the tier 15 min at 19:11:12Z („so the apps are not stopped again for a backup that cannot succeed"). Read at 19:27:40Z (below). **Quiesce retry check (19:27:40Z):** no quiesce or stack stop between 19:18 and 19:27Z — the deferred PBS tier did not stop the apps again. The agent's own attempt still failed at 21:27:43 CEST (hub `WARN host … backup FAILED: target=felhom-pbs … storage 'felhom-pbs' does not exist`), no operator mail for it. ### F1 — power cut while the family uses three apps (`phase5/F1-power-cut.txt`, `phase5/F1/`, `family-activity-F1.txt`) | | | |---|---| | how | `qm stop 333` 19:28:08Z (hard) · 61 s off · `qm start` 19:29:09Z | | what the customer saw | from 19:28:12Z every request to wiki / cloud / recipes fails (connection refused/timeouts) for ≈ 3 m 30 s; BookStack answers again 19:31:37Z, Nextcloud 19:31:42Z, Mealie after 19:32:02Z; the dashboard asks for login again | | what the box did alone | dashboard `/api/health` 200 at **+133 s**; privatebin, vaultwarden, jellyfin +134–135 s … immich, adventurelog +191 s; **paperless-ngx left „boot-orphaned"** and started by the boot reconciler at 19:32:26Z („1 app(s) recovered in 1 attempt(s)"); **all 12 running at +243 s, every one on the same images** as before | | time to steady state | **4 m 03 s** from power-on | | alarms fired | hub: `controller_started (info) — Controller elindult (0.242.0)` only — **true**. Re-read at +10 min (past the 90 s dead-app grace): **no `app_start_failed`, no false alarm** | | should have fired, did not | none — a 61 s outage is under the hub's 30-minute staleness threshold, by design | Verdict **PASS**. Jellyfin's health probe answered 503 / refused for ≈ 45 s after its start (WARN lines), then healthy. ### F2 — power cut during the nightly backup (`phase5/F2-power-cut-during-backup.txt`, `phase5/F2/`) | | | |---|---| | how | `POST /api/backup/run` 19:40:03Z; the harness waited for the first „Stopping … for safe volume dump" (adventurelog, 19:40:07Z) and cut power at **19:40:08Z**; 33 s off; `qm start` 19:40:41Z | | what the customer saw | every app down ≈ 3½ min, as F1. Afterwards `/backups/apps`: „Utolsó adatbázis mentés **2026-09-14 21:40 (5 perce)**", six databases „21:40 · OK", every app „Utolsó: 5 perce"; `/backups` and `/dashboard`: **no word of an interruption**; dashboard tile „Utolsó mentés: 2026-09-14 **19:40**" (UTC beside local times — R-500); whole-system tile „– · Naprakész" with no backup listed (added to R-517) | | what the box did alone | controller 19:42:48Z: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted and left 1 app(s) stopped — restarting them: [adventurelog]` → restarted 19:42:52Z — **the app-stop guard restarted what it stopped** ✓. All 12 running on the same images at **+245 s**. AdventureLog's trip read back intact (3 places, visits). | | torn copy | on disk the adventurelog and bookstack units hold a **19:40:07 SQL dump beside 19:00 volume tars and a 19:02:34 manifest**; both restore points are dated 19:40:07Z → **R-519 (P2)** | | alarms fired | `backup_failed (error) — an app-data backup (volume dump) was interrupted by a controller restart — 1 app(s) were left stopped and have been restarted` + **operator e-mail** — **true**; `controller_started (info)` — true. No false dead-app alarm at +10 min | | should have fired, did not | a customer-side notice on the backups page (R-519) | ### F3 — power cut during a guarded Update (`phase5/F3-power-cut-during-update.txt`, `phase5/F3/`) | | | |---|---| | how | `POST /api/stacks/nextcloud/update` 19:52:02Z (no newer version exists after the Phase 4 revert, so a same-version guarded update — the only Update the catalog allows tonight); a live log follower cut power on the line `update nextcloud: phase pulling` (19:52:03Z) → **`qm stop` 19:52:05.8Z**; 30 s off; `qm start` 19:52:36Z | | journal before the cut | checking → precondition met (Tier 2 copy 49 m old) → **safety dump written** (`…/db-dumps/pre-restore-20260914T195203Z-nextcloud-mariadb.sql`) → pinning („pin advanced to the catalog's current definition", unchanged images) → pulling | | what the box did alone | all 12 running on the same images at **+230 s**; at boot the gate recreated nextcloud onto the live bind („live bind confirmed — recreating drive-backed app nextcloud"); `pin adoption: 0 pinned, 12 already pinned`. **No log line resumes, aborts or even mentions the interrupted update**; `app.yaml` has `pinned_images` = `installed_images`, no hold, no verdict field | | pin state | consistent (same definition) | | the page's message | „Fut · Naprakész"; no word of the interrupted update (fragments `frissít, megszak, vissza, hiba, tartva` absent) | | data | Nextcloud `status.php` installed, not in maintenance; 4 files read back **sha equal**; 200 photos listed | | alarms | (below) | **Verdict: recovered, but the journal's honesty is UNPROVEN for a real version change** — with the images unchanged there is nothing to resume or abort, and nothing records that an update was cut. Not filed as a defect (no wrong state was produced); recorded as a gap in what tonight could test. F3 alarms (re-read 20:03:09Z, +10 min): `controller_started (info)` only — **true**; no false alarm. ### F4 — the data drive is unplugged while apps run (`phase5/F4-drive-unplugged.txt`, `phase5/F4/`) | | | |---|---| | how | `qm set 333 --delete scsi1` at **19:57:27Z** on the running VM (the disk file stays as `unused0`) | | how fast the drive apps stop | Nextcloud's front door already **503 at +5 s**; controller `storage_disconnected` at 19:58:02Z (+35 s); **all four drive apps (nextcloud, immich, jellyfin, paperless-ngx) `stopped` by +38 s**, Nextcloud front door then 404 | | do system-disk apps keep running | **yes** — bookstack, vaultwarden, mealie, docmost (and the other system-disk apps) running throughout; BookStack front door 200 | | what the customer saw (20:04:54Z) | banner on every page: „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort · Rendszermonitor →" and „**Meghajtó leválasztva: Adatlemez (/mnt/felhom-drives/hdd_1)**" (**printed twice**, once with „Beállítások →", once with „Rendszermonitor →"). Tárhely: „Adatlemez · **Leválasztva** · Leválasztva: **2026-09-14T19:58:02Z** · Leállított alkalmazások: immich, jellyfin, nextcloud, paperless-ngx · Csatlakoztatás". App page: „Nextcloud · Leállítva · Naprakész · **Hiányzó tárhely: Adatlemez** — Ennek az alkalmazásnak az adattárolója jelenleg nem elérhető, ezért le van állítva. Csatlakoztasd újra a meghajtót, vagy helyezd át az adatokat egy másik tárhelyre." Dashboard: „12 Futó · 4 Leállítva · Adatlemez — Leválasztva". App list: each drive app „Leállítva · Naprakész · Hiányzó tárhely: Adatlemez". **Honest, Hungarian, and says what to do.** Blemishes: a raw ISO UTC timestamp; the duplicated banner; formal „nézze meg" | | alarms fired | 21:58:02 CEST `storage_disconnected (error) — Meghajtó váratlanul leválasztva: Adatlemez` + operator mail — **true**; 21:58:15 `app_start_failed (warning)` × 4 (Paperless-ngx, Jellyfin, Immich, Nextcloud) + **four more operator mails** — **true but redundant**: one unplug produced five operator e-mails | | should have fired, did not | nothing reached the **customer's** mailbox (`to:tester1@felhom.eu` unchanged) — the household learns of it only by opening the dashboard | Fail-closed path, +8 min (`phase5/F4/fail-closed-path-8min.txt`): on the host and in the guest `/mnt/felhom-drives/hdd_1` **is still a mountpoint on the dead device** — every read answers `Input/output error`; `df` still shows `/dev/sdb 98G 3.6G`. Writes cannot land on the system disk underneath (0 files visible). Rows R-520, R-521 filed; F4 copy blemishes added to R-516. ### F5 — the drive stays out 30 min, then comes back (`phase5/F5-drive-back.txt`, `phase5/F5/`) | | | |---|---| | how | F4's evidence pulled 20:27:31Z; `qm set 333 --scsi1 …vm-333-disk-2.qcow2` at **20:27:31Z** (out **1 804 s**) | | re-bind | the disk returns as **`/dev/sdc`** (the old `sdb` mount stayed `shutdown` the whole time); `EXT4-fs (sdc): recovery complete`, mounted by filesystem UUID `c4b530fd…` at `/mnt/felhom-drives/hdd_1` on host and guest | | restart | the box restarted the four drive apps by itself: immich / jellyfin `starting` at +36 s, **Nextcloud front door 200 at +58 s**, **all four `running` at +91 s**; system-disk apps never stopped | | the badge | Tárhely: „Adatlemez · Alapértelmezett · **Aktív** · 3.5 GB / 97.9 GB · ext4 · /dev/sdc · 4 alkalmazás használja (Immich 529.1 MB, Jellyfin, Nextcloud 389.5 MB, Paperless-ngx 0 B)"; dashboard „16 Futó · 0 Leállítva · Adatlemez 3.52 GB / 97.9 GB". **But the banner „Meghajtó leválasztva: Adatlemez" (×2) was still on every page at 20:29:23Z** — re-checked below | | alarm clears | hub 22:28:25 CEST `storage_reconnected (info) — Meghajtó újra csatlakoztatva: Adatlemez` (no mail) — **true** | | written to the fail-closed path meanwhile? | **no** — while out, the path stayed a mountpoint on the dead device and every access failed with `Input/output error` (host and guest, 20:05Z); no write could land on the system disk beneath. After return, **11 of 11 sampled Nextcloud files sha equal**, „Közös mappa" lists 398 entries | F5 banner re-check (`phase5/F5/banner-10min.txt`, 20:37:54Z): „Meghajtó leválasztva" count **0** on dashboard and storage page (control „Tárhely" present) — the stale banner cleared within 10 minutes of the drive's return. ### F6 — the drive is unplugged during a backup (`phase5/F6-drive-out-during-backup.txt`, `phase5/F6/`) | | | |---|---| | how | `POST /api/backup/run` 20:37:58Z; unplugged at **20:37:59Z**, during `DB dump: nextcloud-db` (1 s in); re-attached 20:42:09Z | | what the backup did | DB dumps of all six databases completed (immich's to a `.tmp` at 20:38:00 that was never promoted); system-disk apps' volume dumps continued normally (adventurelog, bookstack, docmost, gokapi, mealie, privatebin, … each stopped and restarted); at 20:38:33Z `[gate] drive ABSENT … stopped+blocked 4 app(s)`; then **„Skipping volume dump for immich / jellyfin / nextcloud / paperless-ngx — drive disconnected"**; run end `db_dump {"count":6,"duration":"55.7s","success":true}` | | torn copy handled? | **partly.** Immich's point stayed honestly at 19:02:34Z (the `.tmp` not promoted). Nextcloud's point jumped to 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars. The run that skipped four apps says `success:true`; `/backups/apps` shows no skipped app (fragments `kihagy/sikertelen/részleges` = 0) → **added to R-519** | | drive back | all four drive apps running again at **+122 s** after re-attach | | the next run's honesty | `POST /api/backup/run` 20:47:24Z → 2 m 11 s, `success:true`; **all four drive apps dumped** (immich, jellyfin, nextcloud, paperless volume dumps 20:48–20:49Z); points nextcloud 20:48:59Z, immich 20:48:13Z, jellyfin 20:48:26Z, paperless 20:49:15Z; immich's torn `.tmp` gone. **Honest.** One leftover: F3's `pre-restore-20260914T195203Z-nextcloud-mariadb.sql.tmp` still in nextcloud's unit | | alarms | `storage_disconnected (error)` + 4 `app_start_failed` — **all five operator mails suppressed by cooldown** (a second, separate drive loss 40 min after the first, a reconnect in between) → **added to R-521**; `health_degraded (warning)` mailed — true | ### F7 — the system disk fills up (`phase5/F7-disk-full.txt`, `phase5/F7/`) | | | |---|---| | how | `fallocate -l 41 591 249 306` in the guest at `/mnt/sys_drive/felhom-data/userdata/import/F7-nagy-fajl.bin` at **20:50:52Z** (sys_drive 69 G: 36 % → **95 %**, 3.5 G left). The box's thin pool did not grow (fallocate on ext4: `vm-9201-disk-1` data 36.94 %) | | what the customer saw | dashboard from +3 s: „Rendszer (/) · **61.8 GB / 68.7 GB (90%)** · Kritikusan kevés hely" — **90 %, while `df` says 95 %** (the tile ignores the filesystem's reserved blocks; and the tile says „/" for the data volume `/mnt/sys_drive`). From ≈ +5 min a banner on every page: **„SSD disk usage high: 90%" — English** · „Rendszermonitor →". `/backups`: no warning. Deploy page (homebox): memory line only, **nothing about the disk** | | does the box stay reachable | yes — dashboard and BookStack 200 throughout | | what refuses first | nothing refused: a 4 MiB BookStack attachment still uploaded (200); a backup was started on the full disk (below) | | alarms | hub 22:54:46 CEST `health_degraded (warning)` — **operator mail suppressed by cooldown** (the F6 `health_degraded` 15 min earlier). **No disk-specific event** reached the hub | | a backup on the full disk | `POST /api/backup/run` 20:58:5xZ → finished 21:00:57Z, `success:true`; sys_drive stayed 95 % (units are replaced in place; the 3.5 G left sufficed); no `no space`/`ENOSPC` line; controller `[monitor] Disk (SSD) threshold breached: 90% (limit: 80%)`; apps running after (vaultwarden briefly `starting` from its own volume dump) | | recovery | file removed 21:01:19Z → sys_drive 36 %; the English banner gone at +212 s (21:04:53Z); hub 23:04:45 CEST `health_recovered (info) — Rendszer állapot helyreállt: ok (volt: warn)` | | alarms | see above: `health_degraded` (mail suppressed by cooldown) → `health_recovered`; **no disk-specific alarm** → R-521 amended; the English banner → R-516 amended | Verdict **PASS for the box** (reachable, nothing refused or broke at 95 %, the backup still completed); **the operator was not told**. ### F8 — the internet goes away for 20 minutes (`phase5/F8-internet-gone.txt`, `phase5/F8/`) **Harness slip, attempt 1 (21:05:19Z):** the bridge-filter chain was named `fwd`, a reserved word in nft; the chain was never created, so **nothing was cut** (guest → hub still 302). Stopped after 1 minute; evidence kept as `F8-attempt1-HARNESS-SLIP-rule-not-applied.txt`; the half-made empty table removed. The rule was rewritten as a file, dry-run checked (`nft -c -f` → OK) and applied for attempt 2. Scope: only `tap333i0` (VM 333); 9201/9202 untouched. **Attempt 2 (21:06:25Z):** rule applied, counters 0 → debugged with counting tables: VM 333's frames do pass the bridge forward hook; the real reason the guest still reached „the internet" is that **`hub.felhom.eu` resolves to `192.168.0.192` on this LAN** (the hub runs on DooPlex's ingress), so the probe never left the LAN. Stopped; tables removed. **Attempt 3 (the measured run), cut at 21:08:36Z** — VM 333 may reach the LAN (DNS 192.168.0.250 / .1, household devices) but **not** the internet **and not** the hub's LAN ingress 192.168.0.192, which is what a household loses when its internet goes. The LAN dashboard is probed from demo-hp (a household laptop) instead of DooPlex. Drop counters at 21:09:26Z: 27 packets to the hub, 126 to the internet, 6 IPv6 — the cut is real. Guest probe (fixed script, `phase5/F8/guest-wan-probe.txt`) at 21:09:48Z: `guest->1.1.1.1=fail guest->hub=fail cloudflared(last120s registered=1 errors=45)`. (The F8 runner's own guest column prints a shell quoting error — harness; the separate probe file is the observable.) Hub at 21:08:42Z (just after the cut): „Last report: 14 min ago · **warn**" — **not a stall**: the controller's `hub-report` job runs every 15 min (pushed 20:39:46Z and 20:54:46Z, `Hub report pushed successfully`); the „warn" is the state that last report carried (F7's disk). The next push, due 21:09:46Z, fell inside the cut. **Phase 6 pre-note (written 21:13Z, before the phase):** the brief asks to restore one DB-backed app **from off-site onto 9202**. On this record there is no off-site copy to restore from — restic Tier 3 was off (and not ticked: money), and the DR tier never provisioned (R-511); `/backup/offbox/status` → `snapshots 0`. **Not followed, one line why:** the promise „restore from off-site" cannot be walked without an off-site tier, and faking one (copying a unit by hand onto 9202) would prove a path no customer has. Phase 6 instead restores a DB-backed app from its **local** copy on the box itself, through the product's restore endpoint, and reads the data back — recorded as such. **F8 result (attempt 3)** (`phase5/F8-internet-gone.txt`, `phase5/F8/`): | | | |---|---| | how | bridge rule on demo-hp, `tap333i0` only: 21:08:36 → **21:26:07Z** — **17 m 31 s**, not the brief's 20 (the runner's loop count was short — harness). Dropped at removal: 236 packets to the hub, 1 393 to the internet, 54 IPv6 | | LAN dashboard still works? | **yes** — 200 on every probe from demo-hp (every 26 s); polled pages (every 2 min) showed **no banner at all** and the tile „Cloudflare Tunnel … **Fut** · Védett" the whole time → **R-522 (P3)** | | what the box did | cloudflared ≈ 20 errors per 2 min (`failed to dial to edge with quic: timeout`), retrying; controller `[report] Push failed … context deadline exceeded` 21:11:26Z, `Job hub-report failed … after 3 attempts`, wait channel backing off; public name 502 → **530** (no connector) | | on return | report pushed **at 21:26:07Z, the same second** the rule went; tunnel re-registered **+9 s** (21:26:16Z), all four connections **+58 s** (21:27:05Z) — **reconnected by itself** ✓; public name back to 502 (R-510) | | hub shows DOWN, and when | never DOWN (threshold 1 h). Status „warn · Last report N min ago" climbing 14 → 31 min; **`node_stale` 21:25:43Z** (31 min after the last report, 17 min into the cut) + operator mail; **`node_recovered` 21:26:43Z** + mail | | alarms true? | `node_stale` true; `node_recovered` true — a cut shorter than ≈ 13 min after a report would raise nothing, by design | | should have fired, did not | a customer-side notice that the box is offline (R-522) | ### F9 — the controller dies in the middle of a deploy (`phase5/F9-controller-killed-mid-deploy.txt`, `phase5/F9/`) | | | |---|---| | how | deploy of a throwaway app (homebox) `POST` 21:34:35Z → „Telepítés elindítva" (with a memory warning: „Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normál használat mellett ez nem okoz problémát.") → **`docker kill felhom-controller` at 21:34:39Z** (4 s in) | | what the box did by itself | **nothing brought the controller back**: `Exited (137)`, `restart=unless-stopped` — Docker does not restart a container stopped by `kill`; the in-guest `felhom-controller-bootstrap` unit is a one-shot („active (exited)" since boot) and does not watch it. 5 min later still dead, dashboard **502** on the LAN | | the deploy | the app itself came up: container `homebox Up 5 minutes (healthy)`; `app.yaml` and `applied-compose.yml` written at 21:34 | | what the customer saw | the dashboard answers 502 — no page at all | | 23 min later (21:57:39Z) | still `Exited (137) 22 minutes ago`, LAN health **502**; hub shows only `app_deployed (info) — Homebox` (the controller's last report reached the hub at **21:34:38Z**, 3 s before the kill, so the 30-minute staleness is due ≈ 22:04:38Z). The host agent's journal meanwhile: drive reconcile and stale-lock scans every 20 s, **nothing about the controller** | | 33 min later (22:08:04Z) | still `Exited (137) 33 minutes ago`, LAN 502. Hub **22:04:43Z `node_stale` — operator mail suppressed by cooldown** (key `tester-1:node_stale`, set by F8 at 21:25:43Z) | **STOP RULE MET (22:08Z).** F9 left the box in a state the product did not recover from on its own (33 min) and that a customer cannot recover from any screen (there is no dashboard). Filed **R-523 (P1)** before acting. Per the brief: **no further faults are injected — F10, F11 and F12 are not run**; the box is recovered by the household's only lever, a power-cycle, and the night moves to Phase 6. **F9 recovery — the household's only lever, a power-cycle** (`phase5/F9/recovery-power-cycle.txt`, `controller-after-reboot.log`): `qm reset 333` 22:08:17Z → the in-guest bootstrap started the controller at boot (22:10:23Z, `restarts=0`); dashboard 200 at **+131 s**; all 12 apps `running` at **+229 s** (paperless again started by the boot reconciler); **homebox `running`, `deploying false`, `deployed true`, „Fut · Naprakész" — the interrupted deploy is not stuck**. Hub: `controller_started (info)` 22:10:32Z, `node_recovered` 22:10:43Z — **mail suppressed by cooldown**. So a reboot heals it; nothing short of a reboot does, and no screen tells a household to reboot. ## Phase 6 — the morning after (`phase6/`) **Every app healthy? Every label true?** (`phase6/morning-after-check.txt`, 22:13:14Z): all 12 apps `running`, no container down or unhealthy. Labels: 11 × „Fut · Naprakész" with installed == catalog — true. **privatebin: „Frissítés elérhető — ma" while installed 2.0.6 and catalog 2.0.5** (after the Phase 4 revert) — **false: the offered Update is a downgrade** → **R-524 (P2)**. (The check script's own „label true" verdict for privatebin was wrong — it compared for difference, exactly as the product does.) **Every backup page honest?** (`phase6/backups-overview-morning.txt`): „DB mentések 6 fájl" ✓; „Távoli rendszermentés — nincs beállítva" ✓ (honest again after the reboot); whole-system tile „– · Utolsó teljes mentés – · Méret / cél · **Naprakész**" — no backup shown, labelled current (R-517); „Következő mentés — 3 órája" (R-500 class); „Pillanatkép-mód: … csak néhány másodpercre" (R-518). `/backups/restore` lists 13 apps incl. Homebox ✓. **Restore from off-site onto 9202: not walked** — no off-site copy exists on this record (see the pre-note). Instead a DB-backed app (BookStack, MariaDB) is restored from its local copy on the box through `POST /backup/restore`, after a page is deleted and the recycle bin emptied (below). **Local restore of a DB-backed app** (`phase6/local-restore-bookstack.txt`, `…-readback.txt`): before — 5 pages 200, 2 attachments sha equal, „Wifi jelszó helye" text present. Accident 22:14:36Z: the page deleted (302 → 404); the harness's „empty recycle bin" call used a wrong route (404), so the page sat in BookStack's bin — the restore does not depend on that. Point offered: `2026-09-14T20:59:20Z helyi · Belső SSD (rendszer)`. `POST /backup/restore stack_name=bookstack snapshot_id=helyi` 22:14:37Z → finished 22:15:01Z (**24 s**): „A(z) bookstack: 2 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult." After: **all 5 pages 200, the deleted page's text back, both attachments sha equal.** PASS. (Harness: the poll loop missed the status's `last` object and ran on; the read-back was taken separately at 22:24:57Z.) Phase 6 evidence pulled off the box 22:26Z (`box-logs-phase6/`), before any teardown. ## Phase 7 — teardown, three layers **Layer 1 — the machine** (`teardown-before.txt`, `teardown-layer1-machine.txt`), 22:25:58Z: `qm stop 333`, `qm destroy 333 --purge 1 --destroy-unreferenced-disks 1` (system and data qcow2 on `nvme-scratch` removed; `/mnt/hdd_1/images/` now holds only `9202`); ISO 1.27.1 removed from `local:iso`; `/root/bn` harness files removed; nft: only `inet felhom_oob` (the drill's bridge table was already deleted at the end of F8). `pvesm status` before → after: `nvme-scratch` 59 436 396 → **10 140 556 KiB** (before the drill 10 134 820); `local` 24 728 328 → **23 071 188 KiB** (before the drill 23 041 736); `local-lvm` 44.17 % unchanged all night. `qm list` empty; **9201 and 9202 running**. `pct fstrim`: not applied — the drill used only qcow2 files on a dir storage, now deleted; 9201 and 9202 were not touched and are not trimmed. **ep0 read-only before the host delete** (`teardown-ep0-before-host-delete.txt`, 22:26:02Z): WireGuard peer `10.77.0.5/32` present; namespaces `demo-felhom demo-hp tester-1`; `tester-1/ct/9201` with 1 snapshot directory (the doorstep walk's data). Nothing was written to ep0 tonight. **Layer 3 — host record and ep0 peer** (`teardown-layer3-hub.txt`): the destroyed box's last host report reached the hub at 22:23:45Z; delete-impact turned `stale · deletable` at ≈ 22:54Z. **Harness slip:** the first `POST /hosts/tester-1-a61396/delete` (22:54:40Z) sent no fields → **400**, hub `host delete refused: … confirm mismatch`; the host record and the ep0 peer stayed (ep0 read at 22:58:42Z: peer 10.77.0.5 present, tester-1 snapshot dirs 1). The handler wants `confirm_host_id` = the host id (its test D3). Attempt 2 below. **Attempt 2 (22:59:21Z)** with `confirm_host_id=tester-1-a61396` → **303 `/hosts`**; hub `host deleted: tester-1-a61396 (escrow deleted: false)`; host page **404**; not on `/hosts`; **customer `tester-1` page 200 — kept**. The first ep0 re-read (23:00:23Z) still showed the peer, because the last peer sync (22:59:13Z) predates the delete by 8 s — harness read too early; re-read after the next sync below. **ep0 after the next sync** (23:04:13Z `wgsync: pushed 4 peers`; read-only 23:04:33Z): **peer `10.77.0.5` gone (0)**, control peer `10.77.0.3` present; namespaces `demo-felhom demo-hp tester-1`; `tester-1` still holds **1 snapshot directory — kept, stated** for the operator. **Teardown complete: machine, host, hub.** The night's secrets in the session scratchpad (hub password copy, break-glass password, passphrase, codes, app passwords, bind token) shredded at 23:04Z; a search of the scratchpad for leftover tokens found none.