63 KiB
Journal — BIGNIGHT 2026-09-14 (a household's first month in one night)
Times UTC unless marked. Hub logs print CEST (UTC+2). Brief: drills/BIGNIGHT-2026-09-14.md.
Phase 1 — baselines (read 17:29–17:40Z from live source)
| repo | main == origin/main | version |
|---|---|---|
| felhom-controller | 406755fa8fba |
v0.242.0 (CHANGELOG top) |
| felhom-agent | 4586f0f7f6d1 |
v0.130.0 (tree: one untracked scripts/__pycache__/, no tracked change) |
| felhom.eu | a4d684412b49 |
hub v0.113.0 (CHANGELOG top; live image felhom-hub:0.113.0) |
| app-catalog-felhom.eu | 6d6eec307939 |
— |
Register: highest id R-508; 221 table rows (grep ^| R-n).
ISO 1.27.1: /mnt/5_hdd/felhom.eu/felhom-iso/out/felhom-installer-1.27.1-pve9.2-1.iso, 1 705 322 496 B,
sha256 25637007d5a7120ff9faa6b5b7ead3e33c0a361ac2d67e9fd4e0ee77c034c053 = .sha256 = manifest
output-sha256; manifest repo-commit 27e8ec86…, mode release, answer-file NONE. Found, not rebuilt.
DooPlex headroom: /mnt/5_hdd 37 %, / 52 %.
demo-hp (17:29Z): qm list empty; 9201 + 9202 running (each memory: 25898, cores: 7);
host RAM 29 994 MB total, 24 652 MB available; 8 threads (Ryzen V1756B); storages local 56.96 %,
local-lvm 44.17 %, nvme-scratch (dir /mnt/hdd_1, is_mountpoint yes) 1.03 %.
Gmail connector: authenticates as felhom.eu@gmail.com — the catch-all for @felhom.eu
(threads to admin@felhom.eu and drill0242@felhom.eu land there). Read only; nothing sent.
Phase 1 — the Tester 1 record (hub GET /customers/tester-1, Basic auth, 17:30Z)
Customer ID tester-1 · Name „Tester 1" · Domain enkicsifelhom.hu · Email tester1@felhom.eu ·
Config MANAGED. Status tile still reads down, „Last report 1h ago · Controller 0.242.0" — the stale
report of the destroyed VM 331 (host record deleted 16:46:52Z). Claim: „Claimed 1h ago · generation 1".
Edit tab: dr_tier ON (all four DR steps „waiting — no host enrolled yet"); offsite_enabled
OFF („Enable offsite (provisions a Hetzner Storage Box on save)"). Controller floor v0.242.0.
Mailbox: to:tester1@felhom.eu → 0 threads before tonight.
Brief vs record — off-site. The brief says "Off-site (Tier 3) is ON … its own namespace on ep0". On the record, the ep0 namespace is the DR tier (PBS whole-guest); the Tier-3 restic off-site is OFF, and turning it on provisions a Hetzner Storage Box (money, and a new external resource). Not followed: I do not tick it — money is fenced by the rules file §1. "Off-site" tonight = the DR tier to ep0 as configured. Consequence, stated up front: the Phase 4 off-site integrity check and the Phase 6 app restore "from off-site" have no restic tier to act on; each is recorded as such when reached.
Harness slip: the first read of the customer page printed the page's text into this session's tool output, which includes the customer's hub API key (not the retrieval passphrase — that is masked). It stayed on DooPlex; not written to any file.
Tunnel before any box (17:31:19Z, from DooPlex) — tunnel-before-no-box.txt
DNS resolves to Cloudflare (both resolvers); https://felhom.enkicsifelhom.hu → 530, error code: 1033
(no connector). Expected with no box.
Venue — VM 333 bignight-household on demo-hp (created 17:31:52Z)
q35 / OVMF (pre-enrolled-keys=0), 4 cores, cpu=host, 16 384 MB, virtio NIC
BC:24:11:F2:E2:95 on vmbr0, scsi0 200 G qcow2 on nvme-scratch (dir at /mnt/hdd_1, its root),
ISO 1.27.1 on ide2 (sha verified on the HP = 25637007…). One disk at install; the data disk is
added after.
Why these numbers. System disk = doorstep VM 331's 200 G. Memory: the HP has 29 994 MB; 9201 and 9202 together used ≈ 5.3 GB at 17:29Z with 24 652 MB available. Both LXCs carry a 25 898 MB limit, so no VM size leaves both limits whole; 16 GB leaves ≈ 8 GB free plus 8 GB swap for 9201's real load. Cores: the 0242 drill's 4 of 8 threads.
Power-on 17:32:13Z. GRUB (s01, s02): the two Felhom entries; text mode chosen (H4).
Phase 2.1 — install from 1.27.1 (text mode), one disk
| UTC | screen | what it said / what was done |
|---|---|---|
| 17:34:5x | s03 |
English Proxmox EULA → „I agree" |
| 17:35:11 | s04 |
„Target harddisk: /dev/sda (QEMU HARDDISK) (200.00 GiB)" — one disk, no choice to make → Next |
| 17:35:24 | s05 |
Country Hungary · Timezone Europe/Budapest · Keyboard layout Hungarian (H2: changed to U.S. English, s06–s09; one dropped keystroke landed on „Turkish" first, corrected) |
| 17:37:00 | s11 |
„Root password [at least 8 characters]" · Confirm · „Administrator email" prefilled mail@example.invalid → 20-char generated password (0600 file on DooPlex), tester1@felhom.eu (the guide: „a saját e-mail címedet") |
| 17:39:01 | s13 |
nic0 bc:24:11:f2:e2:95, virtio_net · Hostname pve.example.invalid · IP 192.168.0.136/24 (the DHCP lease, frozen static) · GW 192.168.0.1 · DNS 192.168.0.250 · „[X] Pin network interface names" → hostname set felhom.enkicsifelhom.hu per the guide (s15) |
| 17:40:36 | s16 |
Summary: ext4 · /dev/sda · Europe/Budapest · U.S. English · tester1@felhom.eu · nic0 · felhom.enkicsifelhom.hu · 192.168.0.136/24 · 192.168.0.1 · 192.168.0.250 · „[X] Automatically reboot after successful installation" (H3: unticked, s17) |
| 17:41:02 | — | Install pressed |
| ≤17:43:47 | s18 |
„Success — Installation finished - reboot now?" (≤ 2 m 45 s copying; qcow2 4 058 MB, settled) |
| 17:44:xx | first boot | H3: qm stop, --delete ide2, --boot order=scsi0, qm start |
Seeds prepared on DooPlex (scratchpad, tmpfs): 200 JPEG 1600×1200 noise + caption (264 MB), 20 three-page PDFs (2.3 MB), a 50 MB random file, a 150 s 1280×720 H.264 video (98.9 MB).
Phase 2.2 — first screen, and waiting for the mail
17:44:53Z — hub lists Unclaimed appliance 28 (38 s after power-on): smbios uuid 1ecd1c40…, pairing
code ***-*** (redacted), MAC bc:24:11:f2:e2:95, „Standard PC (Q35 + ICH9, 2009)", 15.6 GB, three SSH
host keys. Bind list offers „Tester 1 (0 hosts)".
The console, 17:45:31Z (s19, code redacted), verbatim:
Felhom otthoni szerver
Ezen a gépen most nincs dolgod, és bejelentkezni sem kell.
A beállításhoz kövesd a Felhomtól kapott útmutatót.
felhom login:
==============================================
Felhom — a doboz készen áll, és a párosításra vár.
Párosító kód: ***-***
Nyisd meg az e-mailben kapott linket, és add meg
ezt a kódot és a Tulajdonosi jelmondatodat
(az 5 szót a Felhom üzemeltetőjétől kaptad).
Ez a képernyő magától frissül — nincs teendő a
doboznál, és nyugodtan itt hagyhatod bekapcsolva.
==============================================
No Proxmox :8006 line on the first boot (the 1.27.1 fix holds on a fresh install). Text says
„otthoni szerver"; the guide §3 quotes „Ezen a képernyőn nincs teendőd" — the guide's quote does not
match the screen word for word (meaning is the same).
Mailbox, 17:46Z: to:tester1@felhom.eu → 0 threads. The console asks for „az e-mailben kapott linket".
Hub source: the self-bind link is auto-sent only at customer creation (configs.go:725) and at
RESET completion (customer_reset.go:162); tester-1 was created 32 days ago with no e-mail, so no
link was ever sent to tester1@felhom.eu. Waiting the brief's 10 minutes (to 17:55Z) before acting.
17:55:11Z — ten minutes, 0 messages to tester1@felhom.eu. Filed R-509 (P1) at 17:56Z, before acting.
Intervention I1: the operator's „Send self-bind link" button on the customer's Setup tab is pressed
(a volunteer cannot press it).
17:55:42Z POST /customers/tester-1/selfbind-link → 303 flash=selfbind-sent; hub: self-bind link (hash effe179d…, valid 7 days) emailed to the registered address of tester-1.
The mail, as the volunteer reads it (Gmail connector, message 1a0a10f8cd07d8f6, Date 17:55:43Z — 1 s
after the press; Gmail threads it under the older drill0242 mail of the same subject, a mailbox quirk):
[Felhom] Kösd össze a Felhom dobozodat — from
monitoring@felhom.eutotester1@felhom.euKedves Ügyfél! Elkészült a Felhom dobozod, és készen áll az összekötésre. Az alábbi hivatkozáson tudod te magad összekötni a fiókoddal — nincs szükség bejelentkezésre:
https://hub.felhom.eu/bind/<64 hex, redacted>A hivatkozás megnyitása után két adatot kell megadnod:
- A párosító kódot, amely a doboz képernyőjén (a monitoron) látható.
- A tulajdonosi jelmondatodat (az 5 szóból álló kifejezést), amelyet a Felhom üzemeltetőjétől kaptál — személyesen vagy telefonon, e-mailben soha. Ez igazolja, hogy a fiók a tiéd. A hivatkozás 7 napig érvényes. Biztonsági okból 5 sikertelen próbálkozás után zárolódik — ilyenkor vedd fel a kapcsolatot az ügyfélszolgálattal. Ha nem te kérted ezt, hagyd figyelmen kívül ezt az e-mailt. Üdvözlettel, Felhom.eu
The Tulajdonosi jelmondat is taken from the hub customer page into a 0600 file (the operator's hand-over, per the guide's prerequisite 3 — not a volunteer act, not counted).
Phase 2.2 (cont.) — following the link (the self-bind page, first time exercised by any walk)
17:56:11Z GET https://hub.felhom.eu/bind/<token> → 200 (screen-bind-page.txt), verbatim:
„Felhom — Doboz összekötése" · „Kösd össze a most telepített Felhom dobozodat a fiókoddal. Add meg a doboz
képernyőjén látható párosító kódot és a tulajdonosi jelmondatodat." · Párosító kód („A doboz monitorán
jelenik meg, a telepítés után.", placeholder ABC-234) · Tulajdonosi jelmondat („Az öt szóból álló
kifejezés, amelyet a beállításkor kaptál.", placeholder „öt szó, kötőjellel vagy szóközzel") ·
Összekötés · „Biztonsági okból 5 sikertelen próbálkozás után a hivatkozás zárolódik. …"
Observation: the page still says the phrase was received „a beállításkor" (at setup); the mail (hub v0.113.0) says „a Felhom üzemeltetőjétől kaptál — személyesen vagy telefonon". Two wordings for one hand-over.
17:56:29Z POST /bind/<token> (pairing code + Tulajdonosi jelmondat) → 200 (screen-bind-result.txt):
„Sikeres összekötés. A doboz kb. egy percen belül folytatja a telepítést. Ezt az oldalt bezárhatod — a
beállítás a háttérben befejeződik, és a vezérlőpultod hamarosan elérhető lesz."
The self-bind link works end to end — first time any walk exercised it. One try, no error.
Hub (CEST): 19:56:30 self-bind SUCCESS: appliance 28 bound to customer tester-1 by customer self-service ·
19:56:47 credentials DELIVERED once · 19:57:23 host enrolled: tester-1-a61396 · 19:57:24 [claim] reenroll code (gen 2) emailed … reset code re-issued (gen 2) for tester-1 on box re-enrollment (clean-slate reinstall) ·
19:57:25 manifest agent 0.130.0 golden 0.242.0 · 19:57:42 wg registered … ip=10.77.0.5/32 ·
19:57:42 [ERROR] pbsdr auto-provision for tester-1 (WG-registration hook): the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action ·
19:59:06 controller_started (info) — Controller elindult (0.242.0) · 19:59:43 tester-1 down → ok.
Console after bind (s21, codes redacted): the pairing banner is printed three times with the same code;
it never changes to "bound" (R-214 class).
Which claim path — the brief asked "fresh claim or reset flow". Neither a fresh claim nor the RESET:
the hub treated the box as a re-enrolment of a claimed customer and sent the reinstall mail (gen 2).
Quoted (Gmail 1a0a111205fa7fe8, 17:57:24Z, 54 s after the bind):
[Felhom] Új beállító kód — újratelepült a szervered Kedves Ügyfél! A Felhom szervered újratelepült, ezért a vezérlőpultod belépését újra be kell állítani. A korábbi jelszavad már nem érvényes. Beállító kód:
<3 words, redacted>A kód 72 óráig érvényes, és egyszer használható fel. Nyisd meg a vezérlőpultot — "A szerver beállítása" oldal fogad —, add meg a kódot, majd válassz új jelszót: https://felhom.enkicsifelhom.hu Ha nem te telepítetted újra a szervered, vedd fel a kapcsolatot az üzemeltetővel.
For a volunteer on their first box this mail is wrong in two places: the guide §5 tells them to look for „Elindult a Felhom szervered — beállító kód", and this one says their server was reinstalled and their previous password no longer works — a password they never had. A careful stranger reads „Ha nem te telepítetted újra … vedd fel a kapcsolatot" and stops. (Cause: the record was claimed once today by VM 331.)
Tunnel after bind, before claim — 18:07:48Z from DooPlex (tunnel-after-bind.txt): 502 ×3
(server: cloudflare, body error code: 502). Changed from 530/1033 (no connector) to 502 (connector present,
origin failing).
Phase 2.4 — the gate: the dashboard through the tunnel — FAIL (P1)
box-logs-phase2/cloudflared-before-claim.txt, tunnel-origin-diagnosis.txt. Guest 9201 on the box:
192.168.0.150; containers felhom-controller:0.242.0 (healthy), traefik:v3.6.7, cloudflared:2026.6.0,
filebrowser 1.3.3-stable. cloudflared: 4 connections registered (17:59:02Z, vie06 …), pre-checks all PASS.
Then every request: Request failed … tls: failed to verify certificate: x509: certificate is valid for 544346c4….traefik.default, not traefik … ingressRule=0 originService=https://traefik.
Traefik on 127.0.0.1:443 with SNI felhom.enkicsifelhom.hu serves CN=*.enkicsifelhom.hu, Let's Encrypt YR1,
valid 2026-09-14 17:01 → 12-13 — the right certificate exists. Control (demo-hp 9201, read only): its route
{"hostname":"*.enkisfelhom.hu","originRequest":{"noTLSVerify":true},"service":"https://traefik"}.
Conclusion: the operator's route reaches the box now (530 → 502) but lacks „No TLS Verify". Filed R-510 (P1)
at 18:1xZ before acting. Intervention I2: from here the claim and all dashboard traffic go to
192.168.0.150:443 with the name forced (--resolve), from DooPlex on the same household LAN as the box.
Phase 2.2 (cont.) — the claim, with the mailed code (I2 route)
screen-claim-page.txt. 18:10:06Z GET / → 302 /claim; the page: „A szerver beállítása — Tester 1 —
Add meg az e-mailben kapott beállító kódot, majd válassz saját jelszót a vezérlőpult védelméhez." · Beállító kód ·
Új jelszó (min. 12 karakter) · Új jelszó megerősítése · Beállítás és belépés · „Nem kaptad meg a kódot? Új kód
kérése". (The page does not mention the reinstall the mail talked about — the two agree on the code, not the story.)
Harness slip: the first POST /claim (18:10:06Z) sent both of the page's _csrf values joined (two forms, two
tokens) → 200 „Érvénytelen űrlap — töltsd újra az oldalt." (form refused, not a code attempt). Re-fetched, one token.
18:10:28Z POST /claim (the mailed code, a generated 24-char password twice) → 302 / + felhom_session.
The mailed setup code works — first walk to use the real mail instead of a box-printed code.
/launcher 200: „Indítópult — Felhom.eu", tile Filebrowser, sharing off; menu Vezérlőpult · Alkalmazások ·
Tárhely (Meghajtók, Hálózati tárhely) · Biztonsági mentés (Áttekintés, Távoli mentés, Alkalmazások, Visszaállítás) ·
Megosztás (Hálózati megosztás) · Rendszermonitor · Debug · Beállítások (Rendszer, Értesítések, Biztonság és
hozzáférés) · 0.242.0.
Power-on → claimed dashboard: 38 m 15 s (17:32:13 → 18:10:28), of which 10 min were the brief's mail wait.
Phase 2.3 — version
Box: felhom-controller:0.242.0 (healthy), felhom-agent 0.130.0. Hub: „Controller 0.242.0 · Registry latest
v0.242.0 — up to date · Effective floor v0.242.0 — at/above floor". PASS — landed on golden 0.242.0, reports
current, no self-update needed (golden == floor). Bind → controller_started 2 m 36 s.
Phase 2.4 — the gate after the claim — FAIL, still R-510
tunnel-after-claim.txt: 18:10:41Z from DooPlex https://felhom.enkicsifelhom.hu/login → 502 ×6; box cloudflared
at 18:10:42Z: the same x509 … traefik.default, not traefik for each. Continued on the LAN address (I2).
Phase 2.1 (cont.) — the data disk (100 G scsi1, hot-attached 17:46:42Z, before the bind)
- Box
lsblk:sdb 100G disk, no partitions. - Dashboard (
screen-storage-dashboard-after-claim.txt): „Lemezek állapota — QEMU QEMU HARDDISK · Nincs adat · 0 °C" (one line, which disk is not said); nothing mentions a new disk. - Tárhely → Meghajtók: „Nincs regisztrált adattároló. Adjon hozzá egyet az alábbi űrlappal." · „Új meghajtó inicializálása" · „Meglévő meghajtó csatolása" · „Nem regisztrált meghajtók — az ügynök által észlelt, még nem regisztrált adatmeghajtók" · „Rendszermeghajtók — védett …" · „Már csatlakoztatott tárhely hozzáadása kézzel — Elérési út — Pl. /mnt/hdd_1 …". Formal „Adjon" (the rest of the product says „te").
GET /api/disks/candidates:initialize: [/dev/sdb 107374182400 B, QEMU HARDDISK, data_bearing false];attach: [/dev/mapper/pve-vm--9201--disk--1, mount_source /mnt/sys_drive, already_mounted true](the guest's own system volume offered for attach — observed in the 0242 drill too).- Would a household know to enrol it? No. Nothing on the dashboard, the launcher or any mail says a disk appeared; the volunteer guide has no step for it; the page that lists it is two menus deep and speaks of „inicializálás" and „ügynök". A volunteer who plugs in a second disk has to go looking.
The DR tier on this box — stuck (operator act O1, not a customer step)
Hub 19:57:42 CEST: pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit "Re-issue PBS credentials" action. 18:11:57Z operator POST /configs/tester-1/ pbsdr-reissue → 400 No provisioned PBS DR tier for this customer. Neither path provisions it. Filed R-511 (P2).
Not forced further: removing the old ep0 token is a destructive act on ep0 tenancy, and RESET would also remove the
tunnel (fenced by the brief: the customer is kept).
Consequence for tonight, stated once: this box has no off-site tier of either kind — restic Tier-3 is off
(money) and the PBS DR tier cannot provision. Nothing is written to ep0 tonight. Phase 4's off-site integrity check
and Phase 6's restore „from off-site" are recorded as not possible when reached, with the local tiers walked instead.
No „A vezérlőpultod mostantól jelszóval védett" mail arrived for the claim (the 0242 drill's box received one at 13:31:23Z, 12 s after its claim); at 18:12Z the mailbox holds only the bind mail and the reinstall mail.
Phase 2.1 (cont.) — enrolling the data drive, as a household could from the screens
18:12:57Z POST /api/storage/init (the „Új meghajtó inicializálása" wizard's call) {device /dev/sdb, ext4, mount_name hdd_1, label Adatlemez, set_default true} → formatting → done at 18:12:59Z (2 s), where
/mnt/felhom-drives/hdd_1 (drive-init.txt). The wizard's „Csatlakoztatási név" field has no value, only the
placeholder hdd_1 — a person must type a name. Tárhely now: „Adatlemez · /mnt/felhom-drives/hdd_1 ·
Alapértelmezett · Aktív · 0.0 GB / 97.9 GB · ext4 · /dev/sdb[/felhom-data] · QEMU HARDDISK · Nincs alkalmazás ezen a
tárolón". /api/disks: role user-data, durable id uuid:c4b530fd….
Phase 2 evidence pulled off the box 18:14Z: box-logs-phase2/ (bootstrap + agent journals, controller log,
box state, cloudflared). Secret patterns searched (passphrase=|retrieval_password|password:): 0 hits.
Phase 2 interventions: 2
- I1 (R-509): the self-bind mail never came for an existing customer; the operator's send button was pressed.
- I2 (R-510): the tunnel answers 502; claim and dashboard over the LAN address with the name forced. Operator acts that are prerequisites, not counted: passphrase hand-over; PBS re-issue attempt (O1, refused).
Phase 3 — a family moves in (phase3/)
Deploy pages captured before any deploy (phase3/deploy-pages-before.txt); the memory line each showed at 18:14Z
with 3 infra containers running: bookstack 1525 / 11444 MB · docmost 1568 · privatebin 1404 · gokapi 1408 ·
nextcloud 1633 · immich 3422 · vaultwarden 1427 · paperless-ngx 1880 · jellyfin 1892 · mealie 1578 · uptime-kuma
1431 · adventurelog 1476 (guest limit 11 444 MB of the VM's 16 GB). Every deploy goes through the page's own call,
POST /api/stacks/<app>/deploy {"values":…} with the form's server-generated values; a required empty password
(gokapi, nextcloud admin) is typed as a customer would — a generated value in a 0600 file.
| app | deploy | seed through the front door | household use |
|---|---|---|---|
| bookstack | 18:14:35 → running 80 s | default login → admin e-mail + password changed (old default refused); book „Családi tudástár BIGNIGHT"; 5 Hungarian pages (Töltött káposzta, Wifi jelszó helye, Kutya oltási naptár, Nyaralás Balaton 2026, Autó szerviz); 2 attachments 256 KiB + 512 KiB, sha equal ×2 | 2nd user „Kovács Péter" (Editor) created, logs in, edits Balaton page (edit visible); Péter's attachment upload returned an HTML page, not JSON — inconclusive (harness: page id hard-coded); Anna deletes „Kutya oltási naptár" → 404 → recycle bin → restore → 200 with its text |
| docmost | 18:16:10 → running ≈ 90 s | /api/auth/setup workspace „Kovács család"; space „Háztartás"; 5 docs via Markdown import (titles taken from each file's # heading: Rezsi, Iskola, Lista, Orvos, Kert); read back „ELMŰ", „árvíztűrő" present |
renamed „Bevásárlólista — szombat"; deleted „Kert" → trash → restored → listed; invited peter@ (200) |
| privatebin | 18:18:2x → running 15 s | 5 encrypted pastes (browser v2 format): 1week, never, 1month, never, 10min (the expiring one); every one decrypts equal; TTL meta 604800 / [] / 2592000 / [] / 600. Harness: PrivateBin's 10-second post limit refused one post (status 1), retried | — (file-only; pastes are the use) |
| gokapi | 18:18:43 → running 20 s | deploy page asks „Admin jelszó *" (customer types it); login admin; 3 uploads via the dropzone's chunk calls: 50 MB (2 chunks, 1.6 s), a PDF, a JPEG; admin lists 3 | a stranger downloads all three, sha equal ×3; the PDF deleted through the UI's API → its link no longer serves it (78-byte page) |
App pages: „Fut · Naprakész" (title „Ez az alkalmazás a legfrissebb elérhető változatot futtatja."). Gokapi's password: app page „Első lépések" says „a jelszó a Beállítások oldalon található"; Beállítások shows „Admin jelszó 👁 — Telepítéskor beállított kezdeti jelszó — ha az alkalmazásban megváltoztattad, az itt nem frissül." Findable. BookStack page „Első lépések": „Nyisd meg a wiki.DOMAIN címet" (R-498, still).
nextcloud 18:20:14 deploy (HDD_PATH offered: „Adatlemez — 92.9 GB szabad (alapértelmezett)") → +50 s starting →
+121 s unhealthy (the harness stopped polling there — its own mistake) → at 18:23:43Z running, all three
containers (healthy). So the card reads „unhealthy" for a while on a first start; recorded as observed, not a defect.
nextcloud seed (WebDAV, 4 parallel, as the admin the deploy page asked for): folders Fotók/Balaton 2026, Számlák, Közös mappa; 200 JPEG + 20 PDF, 220 × 201, 86 s (264 MB of photos); 2nd user „peter" (OCS 100), „Közös mappa" shared with him (OCS 200); Péter uploads into it and Anna reads it back (sha equal); foto-123 read back sha equal. Household use: a child deletes foto-007 (204 → 404); trashbin lists it; restore (201); photo back, sha equal. Admin used 343 161 297 B. Nextcloud 34.0.1.
immich 18:22:57 deploy → +116 s degraded → +126 s starting … (continues below).
immich running after 146 s (degraded 116 s on the way). Seed through its API: admin sign-up (201), 200 JPEG uploads, 200 × 201 in 16 s; stats photos 200, 275 682 444 B; ML queues draining (faceDetection / smartSearch / ocr).
vaultwarden 18:25:37 → running 15 s. Seed with the Bitwarden client crypto (PBKDF2 600 000, AES-CBC + HMAC):
register (200), token (200), 10 login entries (200 × 10), /api/sync → 10 items, all names decrypt to the
originals. Attachment: the v2 slot's returned url is /ciphers/<id>/attachment/<id> — posting it to the web root
404s, under /api 200 (a client detail, the harness's first try left one empty attachment slot on „Ügyfélkapu+").
The retried attachment downloads 200 with its recorded size; byte-equality not proven (harness lost the key).
Finding R-512 (P2): registration stays open and the page's instruction to close it points at a read-only
field; a stranger registered with no invite (200).
SECURITY — R-513 (P1), found 18:30Z while looking for a way to put a video on the drive
The launcher's Filebrowser tile opens files.enkicsifelhom.hu: password login, no signup. No screen gives the
customer a FileBrowser login. A customer's first guess, admin / admin, works (200 + token); wrong password
401. With it, both sources list (Adatlemez → documents, …; Beolvasás → paperless). The same login works on demo-hp's
9201 and 9202, and 9201's login page answers 200 through Cloudflare from the internet (GET only, no login
attempted remotely). Filed R-513 before anything else; nothing changed on any box.
immich (cont.) ML queues empty at 18:32:09Z (≈ 6 min after upload); asset 42 original read back sha equal.
paperless-ngx 18:27:xx deploy → running. Token with the deploy-generated admin password (the app card's
„admin / admin" is not what was set — the form generates a 16-char password). Tags Számla / Garancia / Adó (201 ×3);
20 PDFs posted 18:29:57Z → 0 documents. Kernel: the container's cgroup OOM-killed gs and the celery worker at
18:31:22Z; tasks 11 FAILURE (WorkerLostError) · 1 STARTED · 8 PENDING, unchanged at 18:44Z. App still „Fut". R-514 (P2).
jellyfin 18:29:34 → running 50 s. Startup wizard through its REST calls (Hungarian UI culture, user „anna"),
login 200; libraries none; the container sees /media/{audiobooks,books,comics,movies,music,photos,podcasts,tv}
(the drive's userdata/media, read-only).
mealie → running 85 s. The card's default login changeme@example.com / MyPassword works (200); password
changed (old now 401); 5 Hungarian recipes with ingredients + steps (201/200 ×5); meal plan 15–19 Sept (201 ×5) reads
back all five; Gulyásleves reads back its accents.
uptime-kuma → running 50 s. Its socket.io endpoint answers the polling transport with the SPA page (websocket only); DooPlex has no websocket library — a raw client is written next.
uptime-kuma (cont.) The first screen of a fresh Uptime Kuma 2.4 is an English „Which database would you like to
use?" page (SQLite / Embedded MariaDB / …, „Next") — the controller's card says nothing about it; the socket API is
not up until it is answered. SQLite chosen through its own call (POST /setup-database → {"ok":true}), then over a
websocket: setup „anna" (successAdded), login ok, 3 monitors (the box's own dashboard, the family wiki, the router
by ping). Heartbeats after 90 s: router UP (0.5 ms); dashboard and wiki DOWN — „Request failed with status code
502": from inside the box, its own public names go out to Cloudflare and hit R-510. A household that adds its own
apps as monitors gets red for every one.
jellyfin media via the product's own route — „Hálózati megosztás" (SMB): enable (303 „Beállítás mentve…"),
password (first try used wrong field names → „legalább 8 karakter" flash, harness; second 303 „Megosztási jelszó
beállítva"), share „Filmek" = existing folder userdata/media/movies (303 „A megosztás létrehozva."). State „fut";
list shows beolvasas (rendszer) and Filmek (Felhőmentés bekapcsolva). From demo-hp on the LAN (a household
laptop): smbclient //192.168.0.150/Filmek -c put 98 898 905 B in 2.6 s. The drive root itself is refused for
sharing („Ez a mappa nem osztható meg.").
jellyfin (cont.) the video, copied through \FELHOM\Filmek, is visible inside the container at
/media/movies/Csaladi-nyaralas-2026.mp4; library „Családi videók" (movies, /media/movies) added (204); the scan finds
„Csaladi-nyaralas" (RunTimeTicks 1 500 000 000 = 150 s) in 5 s; static stream of the first MiB → 206, 1 048 576 B.
adventurelog → running 161 s. Account creation took five harness attempts (evidence kept in order): the
headless /auth/browser/v1/auth/signup through the frontend refuses without a usable CSRF token (the frontend clears
csrftoken on every response); the backend's own /accounts/signup/ form works (302) — a household would use the
frontend's sign-up page, which this harness did not drive. Then the frontend login form (POST /login) → sessionid
(Domain .enkicsifelhom.hu) and through it: collection „Balaton 2026 nyár" 201; Tihany, Badacsony, Keszthely 201 with a
visit each 201; Badacsony edited (rating 4) 200; read back 3 locations with their visits. Photos: 502 then 500 —
see below. (An anonymous POST to /api/collections/ makes the backend raise ValueError … AnonymousUser → 500: the
app's own behaviour, noted, not filed.)
Version labels (phase3/version-labels.txt, 18:56:27Z): all 12 app pages „Fut · Naprakész" (ASCII fragment
Naprak = 1 each, control 0).
The memory guard never refused. All twelve deployed on the default box; the last deploy page before immich read
3 422 / 11 444 MB and before vaultwarden 3 734 / 11 444 MB; free in the guest at 18:45Z: 3 520 MB used of 11 828.
How many apps fit on a default box: at least these twelve, with room left; the one app that ran out of memory did
so inside its own container cap (R-514), which the guard does not see.
adventurelog photos — the frontend proxy logs RequestContentLengthMismatchError: Request body length does not match content-length header for the multipart POST: the R-483 class („photo upload fails from any non-browser client";
the operator confirmed the browser path on 2026-09-13). Not re-filed; the harness cannot prove the browser path.
Hub during Phase 3: 12 × app_deployed (info) for tester-1, 20:14:35 → 20:45:42 CEST; nothing else — no alarm
for the Paperless OOM (R-514).
Phase 3 interventions: 0 (every seed through the app's own front door or the product's own SMB share;
harness retries are recorded, none needed an act a customer could not make). Rows: R-512, R-513, R-514, R-515, R-516.
Evidence pulled off the box 19:00Z: box-logs-phase3/.
Phase 4 — a month of routines (phase4/)
Backup pages read before (phase4/backup-pages-before.txt). The nightly's legs, each through its own endpoint, in
the nightly's order (window: db-dump W, Tier-2 W+60 m, off-site W+105 m, whole-system after):
| leg | endpoint | start → end | result |
|---|---|---|---|
| Tier 1 — DB dumps + recovery units („Mentés most" on Alkalmazások) | POST /api/backup/run |
19:00:26 → 19:02:38 | db_dump count 6, 2m8s, success true |
| Tier 2 — off-drive copies | POST /api/backup/tier2 (no page calls it; the nightly does) |
19:02:5x → (below) | |
| Tier 3 — off-site | none: not configured on this record (restic off; PBS DR stuck, R-511) | — | not run |
| whole-system (Plane-2 local) | POST /api/guest-backup/trigger („Mentés most" on Áttekintés) |
(below) |
Catalog: no app had a newer version (12 × template == catalog, all catalog_since 2026-07-18). So a real one-step
bump: privatebin 2.0.5 → 2.0.6 (upstream Docker Hub 2026-08-08), catalog commit d5d91e0 pushed 18:59:36Z
with catalog_since 2026-09-14, gates OK. At 19:02:41Z the box still reads 2.0.5 (the 15-minute sync).
R-499 still present on /stacks/bookstack/backup („már szerepelnek a teljes rendszermentésben (PBS)").
Tier 2 result (controller log, phase4/controller-log-t1-t2.txt): Tier 2 run complete: 12 app(s) processed at
19:03:06Z, 17 s. The eight system-disk apps copied to the data drive (hdd_1/backups/secondary/…, e.g. bookstack
154.9 MB); the four drive apps copied to the internal SSD marked [SSD: state-only] (immich 1.7 GB, nextcloud
1.3 GB, paperless 68.8 MB, jellyfin 0.6 MB) — definitions and databases, not the photos, as the Nextcloud page said
(„csak az adatbázis és a konfiguráció másolódik"). The status API's running flag covers only the DB leg, so it read
false during Tier 2 (the harness loop exited on it; the log is the observable).
Whole-system „Mentés most" (POST /api/guest-backup/trigger 19:03:23Z; phase4/guest-backup-run.txt,
guest-backup-quiesce-log.txt): the controller quiesced all 12 apps („backup due on 2 tier(s)"); tier local
19:03:49 → 19:09:59Z, 8 877 619 753 B, 362 s, success; then tier felhom-pbs started with the apps still stopped,
failed (storage 'felhom-pbs' does not exist — the DR tier that never provisioned, R-511), unquiesce 19:10:09Z, last app
started 19:11:12Z → apps down ≈ 7 m 45 s under a page line promising „csak néhány másodpercre" → R-518 (P2).
Backoff logged: next PBS attempt in 15 min „so the apps are not stopped again for a backup that cannot succeed".
Harness: the status API kept the previous job's done until the new job id appeared; the first poll stopped on it.
Backup pages, read as a first-timer after the run (phase4/backup-pages-after.txt, 19:15:36Z):
- Áttekintés: „✗ · Utolsó teljes mentés 2026-09-14 21:09 (5 perce) · 0 B · Biztonsági szerver – külön hardver (PBS) · Naprakész" and „✓ Távoli rendszermentés — külön hardveren (PBS)" — false on both counts; the real 8.9 GB local backup is not shown → R-517 (P1). Dates are local time (21:09 CEST) ✓. „DB mentések 6 fájl" ✓. „Következő mentés — 0 órája" (reads backwards, seen in the 0242 drill). Both honest warnings about one disk remain although a data drive is now enrolled — one of them offers „Adatlemez · Kijelölöm" as the whole-system target ✓.
- Alkalmazások: every app „1. mentés Auto helyi Utolsó: …"; drive apps' second copy state-only as said.
/stacks/bookstack/backupstill claims PBS coverage (R-499).- Off-site integrity check: not possible on this record — „Távoli mentés": „Még nincs beállítva távoli mentési cél";
/backup/offbox/status→snapshots 0, status "". Verdict and duration: none — no tier to check (recorded, not run).
Hub after the routines (phase4/hub-check.txt, 19:15:31Z): „Tester 1 · ok · Controller 0.242.0 · Last report
2 min ago · Containers 24/24" (all 12 apps' containers listed), storage „SSD 33 % · Adatlemez 4 %", „Last DB dump
13 min ago", „No off-site data reported … this page cannot say". 13 × crossdrive_completed (info); 21:11:12
whole_guest_backup_failed (error) + operator e-mail — a TRUE alarm (the tier really cannot succeed), with no word of
the local success. The hub page shows no whole-system backup line at all.
Catalog travel: the box synced at 18:59:00Z (36 s before the push) and picked the bump up at 19:14:01Z („Sablonok frissítve — frissítve: privatebin") — 14 m 25 s after the push. App page: „Frissítés elérhető — ma" (title „Újabb változat érhető el …"). PrivateBin before the update: 2 pastes decrypt equal; the 10-minute paste answers „Document does not exist, has expired or has been deleted." — honest expiry.
The guarded Update, privatebin 2.0.5 → 2.0.6 (phase4/privatebin-update.txt): POST /api/stacks/privatebin/update
19:16:33Z → „Frissítés elindult – az állapot a kártyán követhető". Phases: checking → precondition met — Tier 2
(second drive) copy from 19:03:05Z (13 m old, limit 24 h) → safety-dump (no database, no-op) → pinning („pin
advanced to the catalog's current definition") → pulling → starting → verifying → „healthy after 5s" → installed-images
recorded privatebin/pdo:2.0.6 (sha256:4c141b23…) → DONE in 11 s. Container privatebin/pdo:2.0.6 (healthy). App
page „Fut · Naprakész". Data read back after: both surviving pastes decrypt equal; the expired one still
expired. PASS. (Observation: template_images in the API still read 2.0.5 for ≈ 15 s after updating went false.)
Catalog revert a161ccb pushed 19:17:55Z (image back to 2.0.5; catalog_since stays 2026-09-14 by the gate).
Phase 4 evidence pulled off the box 19:18Z (box-logs-phase4/).
Phase 5 — the accidents (phase5/)
Method, the same for every fault: a family loop (family.py) keeps three apps busy (BookStack page loads each
second, a Nextcloud note written every 2 s, Mealie API each second) and logs ok/fail per 5 s; watch_recovery.py
records from the fault's T0 when the dashboard's /api/health answers and when each of the 12 apps is running again
with the same images as the Phase 4 baseline; alarms.sh reads the hub's events, staleness changes and operator
mails for tester-1; the Gmail connector is read for what reached the customer's and the operator's mailbox. For each
fault: what the customer saw · what the box did alone · time to steady state · which alarm fired and was it true ·
which alarm should have fired and did not. Evidence off the box before the next fault.
Before F1 — does the failed PBS tier stop the apps again? The controller deferred the tier 15 min at 19:11:12Z („so the apps are not stopped again for a backup that cannot succeed"). Read at 19:27:40Z (below).
Quiesce retry check (19:27:40Z): no quiesce or stack stop between 19:18 and 19:27Z — the deferred PBS tier did
not stop the apps again. The agent's own attempt still failed at 21:27:43 CEST (hub WARN host … backup FAILED: target=felhom-pbs … storage 'felhom-pbs' does not exist), no operator mail for it.
F1 — power cut while the family uses three apps (phase5/F1-power-cut.txt, phase5/F1/, family-activity-F1.txt)
| how | qm stop 333 19:28:08Z (hard) · 61 s off · qm start 19:29:09Z |
| what the customer saw | from 19:28:12Z every request to wiki / cloud / recipes fails (connection refused/timeouts) for ≈ 3 m 30 s; BookStack answers again 19:31:37Z, Nextcloud 19:31:42Z, Mealie after 19:32:02Z; the dashboard asks for login again |
| what the box did alone | dashboard /api/health 200 at +133 s; privatebin, vaultwarden, jellyfin +134–135 s … immich, adventurelog +191 s; paperless-ngx left „boot-orphaned" and started by the boot reconciler at 19:32:26Z („1 app(s) recovered in 1 attempt(s)"); all 12 running at +243 s, every one on the same images as before |
| time to steady state | 4 m 03 s from power-on |
| alarms fired | hub: controller_started (info) — Controller elindult (0.242.0) only — true. Re-read at +10 min (past the 90 s dead-app grace): no app_start_failed, no false alarm |
| should have fired, did not | none — a 61 s outage is under the hub's 30-minute staleness threshold, by design |
Verdict PASS. Jellyfin's health probe answered 503 / refused for ≈ 45 s after its start (WARN lines), then healthy.
F2 — power cut during the nightly backup (phase5/F2-power-cut-during-backup.txt, phase5/F2/)
| how | POST /api/backup/run 19:40:03Z; the harness waited for the first „Stopping … for safe volume dump" (adventurelog, 19:40:07Z) and cut power at 19:40:08Z; 33 s off; qm start 19:40:41Z |
| what the customer saw | every app down ≈ 3½ min, as F1. Afterwards /backups/apps: „Utolsó adatbázis mentés 2026-09-14 21:40 (5 perce)", six databases „21:40 · OK", every app „Utolsó: 5 perce"; /backups and /dashboard: no word of an interruption; dashboard tile „Utolsó mentés: 2026-09-14 19:40" (UTC beside local times — R-500); whole-system tile „– · Naprakész" with no backup listed (added to R-517) |
| what the box did alone | controller 19:42:48Z: [appstop] crash recovery: an app-data backup (volume dump) … was interrupted and left 1 app(s) stopped — restarting them: [adventurelog] → restarted 19:42:52Z — the app-stop guard restarted what it stopped ✓. All 12 running on the same images at +245 s. AdventureLog's trip read back intact (3 places, visits). |
| torn copy | on disk the adventurelog and bookstack units hold a 19:40:07 SQL dump beside 19:00 volume tars and a 19:02:34 manifest; both restore points are dated 19:40:07Z → R-519 (P2) |
| alarms fired | backup_failed (error) — an app-data backup (volume dump) was interrupted by a controller restart — 1 app(s) were left stopped and have been restarted + operator e-mail — true; controller_started (info) — true. No false dead-app alarm at +10 min |
| should have fired, did not | a customer-side notice on the backups page (R-519) |
F3 — power cut during a guarded Update (phase5/F3-power-cut-during-update.txt, phase5/F3/)
| how | POST /api/stacks/nextcloud/update 19:52:02Z (no newer version exists after the Phase 4 revert, so a same-version guarded update — the only Update the catalog allows tonight); a live log follower cut power on the line update nextcloud: phase pulling (19:52:03Z) → qm stop 19:52:05.8Z; 30 s off; qm start 19:52:36Z |
| journal before the cut | checking → precondition met (Tier 2 copy 49 m old) → safety dump written (…/db-dumps/pre-restore-20260914T195203Z-nextcloud-mariadb.sql) → pinning („pin advanced to the catalog's current definition", unchanged images) → pulling |
| what the box did alone | all 12 running on the same images at +230 s; at boot the gate recreated nextcloud onto the live bind („live bind confirmed — recreating drive-backed app nextcloud"); pin adoption: 0 pinned, 12 already pinned. No log line resumes, aborts or even mentions the interrupted update; app.yaml has pinned_images = installed_images, no hold, no verdict field |
| pin state | consistent (same definition) |
| the page's message | „Fut · Naprakész"; no word of the interrupted update (fragments frissít, megszak, vissza, hiba, tartva absent) |
| data | Nextcloud status.php installed, not in maintenance; 4 files read back sha equal; 200 photos listed |
| alarms | (below) |
Verdict: recovered, but the journal's honesty is UNPROVEN for a real version change — with the images unchanged there is nothing to resume or abort, and nothing records that an update was cut. Not filed as a defect (no wrong state was produced); recorded as a gap in what tonight could test.
F3 alarms (re-read 20:03:09Z, +10 min): controller_started (info) only — true; no false alarm.
F4 — the data drive is unplugged while apps run (phase5/F4-drive-unplugged.txt, phase5/F4/)
| how | qm set 333 --delete scsi1 at 19:57:27Z on the running VM (the disk file stays as unused0) |
| how fast the drive apps stop | Nextcloud's front door already 503 at +5 s; controller storage_disconnected at 19:58:02Z (+35 s); all four drive apps (nextcloud, immich, jellyfin, paperless-ngx) stopped by +38 s, Nextcloud front door then 404 |
| do system-disk apps keep running | yes — bookstack, vaultwarden, mealie, docmost (and the other system-disk apps) running throughout; BookStack front door 200 |
| what the customer saw (20:04:54Z) | banner on every page: „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort · Rendszermonitor →" and „Meghajtó leválasztva: Adatlemez (/mnt/felhom-drives/hdd_1)" (printed twice, once with „Beállítások →", once with „Rendszermonitor →"). Tárhely: „Adatlemez · Leválasztva · Leválasztva: 2026-09-14T19:58:02Z · Leállított alkalmazások: immich, jellyfin, nextcloud, paperless-ngx · Csatlakoztatás". App page: „Nextcloud · Leállítva · Naprakész · Hiányzó tárhely: Adatlemez — Ennek az alkalmazásnak az adattárolója jelenleg nem elérhető, ezért le van állítva. Csatlakoztasd újra a meghajtót, vagy helyezd át az adatokat egy másik tárhelyre." Dashboard: „12 Futó · 4 Leállítva · Adatlemez — Leválasztva". App list: each drive app „Leállítva · Naprakész · Hiányzó tárhely: Adatlemez". Honest, Hungarian, and says what to do. Blemishes: a raw ISO UTC timestamp; the duplicated banner; formal „nézze meg" |
| alarms fired | 21:58:02 CEST storage_disconnected (error) — Meghajtó váratlanul leválasztva: Adatlemez + operator mail — true; 21:58:15 app_start_failed (warning) × 4 (Paperless-ngx, Jellyfin, Immich, Nextcloud) + four more operator mails — true but redundant: one unplug produced five operator e-mails |
| should have fired, did not | nothing reached the customer's mailbox (to:tester1@felhom.eu unchanged) — the household learns of it only by opening the dashboard |
Fail-closed path, +8 min (phase5/F4/fail-closed-path-8min.txt): on the host and in the guest /mnt/felhom-drives/hdd_1
is still a mountpoint on the dead device — every read answers Input/output error; df still shows /dev/sdb 98G 3.6G. Writes cannot land on the system disk underneath (0 files visible). Rows R-520, R-521 filed; F4 copy blemishes
added to R-516.
F5 — the drive stays out 30 min, then comes back (phase5/F5-drive-back.txt, phase5/F5/)
| how | F4's evidence pulled 20:27:31Z; qm set 333 --scsi1 …vm-333-disk-2.qcow2 at 20:27:31Z (out 1 804 s) |
| re-bind | the disk returns as /dev/sdc (the old sdb mount stayed shutdown the whole time); EXT4-fs (sdc): recovery complete, mounted by filesystem UUID c4b530fd… at /mnt/felhom-drives/hdd_1 on host and guest |
| restart | the box restarted the four drive apps by itself: immich / jellyfin starting at +36 s, Nextcloud front door 200 at +58 s, all four running at +91 s; system-disk apps never stopped |
| the badge | Tárhely: „Adatlemez · Alapértelmezett · Aktív · 3.5 GB / 97.9 GB · ext4 · /dev/sdc · 4 alkalmazás használja (Immich 529.1 MB, Jellyfin, Nextcloud 389.5 MB, Paperless-ngx 0 B)"; dashboard „16 Futó · 0 Leállítva · Adatlemez 3.52 GB / 97.9 GB". But the banner „Meghajtó leválasztva: Adatlemez" (×2) was still on every page at 20:29:23Z — re-checked below |
| alarm clears | hub 22:28:25 CEST storage_reconnected (info) — Meghajtó újra csatlakoztatva: Adatlemez (no mail) — true |
| written to the fail-closed path meanwhile? | no — while out, the path stayed a mountpoint on the dead device and every access failed with Input/output error (host and guest, 20:05Z); no write could land on the system disk beneath. After return, 11 of 11 sampled Nextcloud files sha equal, „Közös mappa" lists 398 entries |
F5 banner re-check (phase5/F5/banner-10min.txt, 20:37:54Z): „Meghajtó leválasztva" count 0 on dashboard and
storage page (control „Tárhely" present) — the stale banner cleared within 10 minutes of the drive's return.
F6 — the drive is unplugged during a backup (phase5/F6-drive-out-during-backup.txt, phase5/F6/)
| how | POST /api/backup/run 20:37:58Z; unplugged at 20:37:59Z, during DB dump: nextcloud-db (1 s in); re-attached 20:42:09Z |
| what the backup did | DB dumps of all six databases completed (immich's to a .tmp at 20:38:00 that was never promoted); system-disk apps' volume dumps continued normally (adventurelog, bookstack, docmost, gokapi, mealie, privatebin, … each stopped and restarted); at 20:38:33Z [gate] drive ABSENT … stopped+blocked 4 app(s); then „Skipping volume dump for immich / jellyfin / nextcloud / paperless-ngx — drive disconnected"; run end db_dump {"count":6,"duration":"55.7s","success":true} |
| torn copy handled? | partly. Immich's point stayed honestly at 19:02:34Z (the .tmp not promoted). Nextcloud's point jumped to 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars. The run that skipped four apps says success:true; /backups/apps shows no skipped app (fragments kihagy/sikertelen/részleges = 0) → added to R-519 |
| drive back | all four drive apps running again at +122 s after re-attach |
| the next run's honesty | POST /api/backup/run 20:47:24Z → 2 m 11 s, success:true; all four drive apps dumped (immich, jellyfin, nextcloud, paperless volume dumps 20:48–20:49Z); points nextcloud 20:48:59Z, immich 20:48:13Z, jellyfin 20:48:26Z, paperless 20:49:15Z; immich's torn .tmp gone. Honest. One leftover: F3's pre-restore-20260914T195203Z-nextcloud-mariadb.sql.tmp still in nextcloud's unit |
| alarms | storage_disconnected (error) + 4 app_start_failed — all five operator mails suppressed by cooldown (a second, separate drive loss 40 min after the first, a reconnect in between) → added to R-521; health_degraded (warning) mailed — true |
F7 — the system disk fills up (phase5/F7-disk-full.txt, phase5/F7/)
| how | fallocate -l 41 591 249 306 in the guest at /mnt/sys_drive/felhom-data/userdata/import/F7-nagy-fajl.bin at 20:50:52Z (sys_drive 69 G: 36 % → 95 %, 3.5 G left). The box's thin pool did not grow (fallocate on ext4: vm-9201-disk-1 data 36.94 %) |
| what the customer saw | dashboard from +3 s: „Rendszer (/) · 61.8 GB / 68.7 GB (90%) · Kritikusan kevés hely" — 90 %, while df says 95 % (the tile ignores the filesystem's reserved blocks; and the tile says „/" for the data volume /mnt/sys_drive). From ≈ +5 min a banner on every page: „SSD disk usage high: 90%" — English · „Rendszermonitor →". /backups: no warning. Deploy page (homebox): memory line only, nothing about the disk |
| does the box stay reachable | yes — dashboard and BookStack 200 throughout |
| what refuses first | nothing refused: a 4 MiB BookStack attachment still uploaded (200); a backup was started on the full disk (below) |
| alarms | hub 22:54:46 CEST health_degraded (warning) — operator mail suppressed by cooldown (the F6 health_degraded 15 min earlier). No disk-specific event reached the hub |
| a backup on the full disk | POST /api/backup/run 20:58:5xZ → finished 21:00:57Z, success:true; sys_drive stayed 95 % (units are replaced in place; the 3.5 G left sufficed); no no space/ENOSPC line; controller [monitor] Disk (SSD) threshold breached: 90% (limit: 80%); apps running after (vaultwarden briefly starting from its own volume dump) |
| recovery | file removed 21:01:19Z → sys_drive 36 %; the English banner gone at +212 s (21:04:53Z); hub 23:04:45 CEST health_recovered (info) — Rendszer állapot helyreállt: ok (volt: warn) |
| alarms | see above: health_degraded (mail suppressed by cooldown) → health_recovered; no disk-specific alarm → R-521 amended; the English banner → R-516 amended |
Verdict PASS for the box (reachable, nothing refused or broke at 95 %, the backup still completed); the operator was not told.
F8 — the internet goes away for 20 minutes (phase5/F8-internet-gone.txt, phase5/F8/)
Harness slip, attempt 1 (21:05:19Z): the bridge-filter chain was named fwd, a reserved word in nft; the chain was
never created, so nothing was cut (guest → hub still 302). Stopped after 1 minute; evidence kept as
F8-attempt1-HARNESS-SLIP-rule-not-applied.txt; the half-made empty table removed. The rule was rewritten as a file,
dry-run checked (nft -c -f → OK) and applied for attempt 2. Scope: only tap333i0 (VM 333); 9201/9202 untouched.
Attempt 2 (21:06:25Z): rule applied, counters 0 → debugged with counting tables: VM 333's frames do pass the bridge
forward hook; the real reason the guest still reached „the internet" is that hub.felhom.eu resolves to
192.168.0.192 on this LAN (the hub runs on DooPlex's ingress), so the probe never left the LAN. Stopped; tables removed.
Attempt 3 (the measured run), cut at 21:08:36Z — VM 333 may reach the LAN (DNS 192.168.0.250 / .1, household
devices) but not the internet and not the hub's LAN ingress 192.168.0.192, which is what a household loses when
its internet goes. The LAN dashboard is probed from demo-hp (a household laptop) instead of DooPlex. Drop counters at
21:09:26Z: 27 packets to the hub, 126 to the internet, 6 IPv6 — the cut is real. Guest probe (fixed script,
phase5/F8/guest-wan-probe.txt) at 21:09:48Z: guest->1.1.1.1=fail guest->hub=fail cloudflared(last120s registered=1 errors=45). (The F8 runner's own guest column prints a shell quoting error — harness; the separate
probe file is the observable.)
Hub at 21:08:42Z (just after the cut): „Last report: 14 min ago · warn" — not a stall: the controller's
hub-report job runs every 15 min (pushed 20:39:46Z and 20:54:46Z, Hub report pushed successfully); the „warn" is the
state that last report carried (F7's disk). The next push, due 21:09:46Z, fell inside the cut.
Phase 6 pre-note (written 21:13Z, before the phase): the brief asks to restore one DB-backed app from off-site
onto 9202. On this record there is no off-site copy to restore from — restic Tier 3 was off (and not ticked: money),
and the DR tier never provisioned (R-511); /backup/offbox/status → snapshots 0. Not followed, one line why: the
promise „restore from off-site" cannot be walked without an off-site tier, and faking one (copying a unit by hand onto
9202) would prove a path no customer has. Phase 6 instead restores a DB-backed app from its local copy on the box
itself, through the product's restore endpoint, and reads the data back — recorded as such.
F8 result (attempt 3) (phase5/F8-internet-gone.txt, phase5/F8/):
| how | bridge rule on demo-hp, tap333i0 only: 21:08:36 → 21:26:07Z — 17 m 31 s, not the brief's 20 (the runner's loop count was short — harness). Dropped at removal: 236 packets to the hub, 1 393 to the internet, 54 IPv6 |
| LAN dashboard still works? | yes — 200 on every probe from demo-hp (every 26 s); polled pages (every 2 min) showed no banner at all and the tile „Cloudflare Tunnel … Fut · Védett" the whole time → R-522 (P3) |
| what the box did | cloudflared ≈ 20 errors per 2 min (failed to dial to edge with quic: timeout), retrying; controller [report] Push failed … context deadline exceeded 21:11:26Z, Job hub-report failed … after 3 attempts, wait channel backing off; public name 502 → 530 (no connector) |
| on return | report pushed at 21:26:07Z, the same second the rule went; tunnel re-registered +9 s (21:26:16Z), all four connections +58 s (21:27:05Z) — reconnected by itself ✓; public name back to 502 (R-510) |
| hub shows DOWN, and when | never DOWN (threshold 1 h). Status „warn · Last report N min ago" climbing 14 → 31 min; node_stale 21:25:43Z (31 min after the last report, 17 min into the cut) + operator mail; node_recovered 21:26:43Z + mail |
| alarms true? | node_stale true; node_recovered true — a cut shorter than ≈ 13 min after a report would raise nothing, by design |
| should have fired, did not | a customer-side notice that the box is offline (R-522) |
F9 — the controller dies in the middle of a deploy (phase5/F9-controller-killed-mid-deploy.txt, phase5/F9/)
| how | deploy of a throwaway app (homebox) POST 21:34:35Z → „Telepítés elindítva" (with a memory warning: „Az alkalmazások csúcsterhelése meghaladhatja a rendelkezésre álló memóriát. Normál használat mellett ez nem okoz problémát.") → docker kill felhom-controller at 21:34:39Z (4 s in) |
| what the box did by itself | nothing brought the controller back: Exited (137), restart=unless-stopped — Docker does not restart a container stopped by kill; the in-guest felhom-controller-bootstrap unit is a one-shot („active (exited)" since boot) and does not watch it. 5 min later still dead, dashboard 502 on the LAN |
| the deploy | the app itself came up: container homebox Up 5 minutes (healthy); app.yaml and applied-compose.yml written at 21:34 |
| what the customer saw | the dashboard answers 502 — no page at all |
| 23 min later (21:57:39Z) | still Exited (137) 22 minutes ago, LAN health 502; hub shows only app_deployed (info) — Homebox (the controller's last report reached the hub at 21:34:38Z, 3 s before the kill, so the 30-minute staleness is due ≈ 22:04:38Z). The host agent's journal meanwhile: drive reconcile and stale-lock scans every 20 s, nothing about the controller |
| 33 min later (22:08:04Z) | still Exited (137) 33 minutes ago, LAN 502. Hub 22:04:43Z node_stale — operator mail suppressed by cooldown (key tester-1:node_stale, set by F8 at 21:25:43Z) |
STOP RULE MET (22:08Z). F9 left the box in a state the product did not recover from on its own (33 min) and that a customer cannot recover from any screen (there is no dashboard). Filed R-523 (P1) before acting. Per the brief: no further faults are injected — F10, F11 and F12 are not run; the box is recovered by the household's only lever, a power-cycle, and the night moves to Phase 6.
F9 recovery — the household's only lever, a power-cycle (phase5/F9/recovery-power-cycle.txt,
controller-after-reboot.log): qm reset 333 22:08:17Z → the in-guest bootstrap started the controller at boot
(22:10:23Z, restarts=0); dashboard 200 at +131 s; all 12 apps running at +229 s (paperless again started by
the boot reconciler); homebox running, deploying false, deployed true, „Fut · Naprakész" — the interrupted deploy
is not stuck. Hub: controller_started (info) 22:10:32Z, node_recovered 22:10:43Z — mail suppressed by cooldown.
So a reboot heals it; nothing short of a reboot does, and no screen tells a household to reboot.
Phase 6 — the morning after (phase6/)
Every app healthy? Every label true? (phase6/morning-after-check.txt, 22:13:14Z): all 12 apps running, no
container down or unhealthy. Labels: 11 × „Fut · Naprakész" with installed == catalog — true. privatebin: „Frissítés
elérhető — ma" while installed 2.0.6 and catalog 2.0.5 (after the Phase 4 revert) — false: the offered Update is a
downgrade → R-524 (P2). (The check script's own „label true" verdict for privatebin was wrong — it compared for
difference, exactly as the product does.)
Every backup page honest? (phase6/backups-overview-morning.txt): „DB mentések 6 fájl" ✓; „Távoli rendszermentés —
nincs beállítva" ✓ (honest again after the reboot); whole-system tile „– · Utolsó teljes mentés – · Méret / cél ·
Naprakész" — no backup shown, labelled current (R-517); „Következő mentés — 3 órája" (R-500 class);
„Pillanatkép-mód: … csak néhány másodpercre" (R-518). /backups/restore lists 13 apps incl. Homebox ✓.
Restore from off-site onto 9202: not walked — no off-site copy exists on this record (see the pre-note). Instead a
DB-backed app (BookStack, MariaDB) is restored from its local copy on the box through POST /backup/restore, after a
page is deleted and the recycle bin emptied (below).
Local restore of a DB-backed app (phase6/local-restore-bookstack.txt, …-readback.txt): before — 5 pages 200,
2 attachments sha equal, „Wifi jelszó helye" text present. Accident 22:14:36Z: the page deleted (302 → 404); the
harness's „empty recycle bin" call used a wrong route (404), so the page sat in BookStack's bin — the restore does not
depend on that. Point offered: 2026-09-14T20:59:20Z helyi · Belső SSD (rendszer). POST /backup/restore stack_name=bookstack snapshot_id=helyi 22:14:37Z → finished 22:15:01Z (24 s): „A(z) bookstack: 2 adatkötet és
az adatbázis visszaállítva — az alkalmazás újraindult." After: all 5 pages 200, the deleted page's text back, both
attachments sha equal. PASS. (Harness: the poll loop missed the status's last object and ran on; the read-back was
taken separately at 22:24:57Z.)
Phase 6 evidence pulled off the box 22:26Z (box-logs-phase6/), before any teardown.
Phase 7 — teardown, three layers
Layer 1 — the machine (teardown-before.txt, teardown-layer1-machine.txt), 22:25:58Z: qm stop 333, qm destroy 333 --purge 1 --destroy-unreferenced-disks 1 (system and data qcow2 on nvme-scratch removed; /mnt/hdd_1/images/
now holds only 9202); ISO 1.27.1 removed from local:iso; /root/bn harness files removed; nft: only inet felhom_oob (the drill's bridge table was already deleted at the end of F8). pvesm status before → after:
nvme-scratch 59 436 396 → 10 140 556 KiB (before the drill 10 134 820); local 24 728 328 → 23 071 188 KiB
(before the drill 23 041 736); local-lvm 44.17 % unchanged all night. qm list empty; 9201 and 9202 running.
pct fstrim: not applied — the drill used only qcow2 files on a dir storage, now deleted; 9201 and 9202 were not
touched and are not trimmed.
ep0 read-only before the host delete (teardown-ep0-before-host-delete.txt, 22:26:02Z): WireGuard peer
10.77.0.5/32 present; namespaces demo-felhom demo-hp tester-1; tester-1/ct/9201 with 1 snapshot directory (the
doorstep walk's data). Nothing was written to ep0 tonight.
Layer 3 — host record and ep0 peer (teardown-layer3-hub.txt): the destroyed box's last host report reached the
hub at 22:23:45Z; delete-impact turned stale · deletable at ≈ 22:54Z. Harness slip: the first POST /hosts/tester-1-a61396/delete (22:54:40Z) sent no fields → 400, hub host delete refused: … confirm mismatch; the
host record and the ep0 peer stayed (ep0 read at 22:58:42Z: peer 10.77.0.5 present, tester-1 snapshot dirs 1). The
handler wants confirm_host_id = the host id (its test D3). Attempt 2 below.
Attempt 2 (22:59:21Z) with confirm_host_id=tester-1-a61396 → 303 /hosts; hub host deleted: tester-1-a61396 (escrow deleted: false); host page 404; not on /hosts; customer tester-1 page 200 — kept. The first ep0
re-read (23:00:23Z) still showed the peer, because the last peer sync (22:59:13Z) predates the delete by 8 s — harness
read too early; re-read after the next sync below.
ep0 after the next sync (23:04:13Z wgsync: pushed 4 peers; read-only 23:04:33Z): peer 10.77.0.5 gone (0),
control peer 10.77.0.3 present; namespaces demo-felhom demo-hp tester-1; tester-1 still holds 1 snapshot
directory — kept, stated for the operator. Teardown complete: machine, host, hub. The night's secrets in the
session scratchpad (hub password copy, break-glass password, passphrase, codes, app passwords, bind token) shredded at
23:04Z; a search of the scratchpad for leftover tokens found none.