The automatic connect e-mail is proven with a real mailbox: the host record was deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second later, with selfbind_link_sent (host delete) on the timeline. The requirement was two minutes. The hub refuses to delete an ONLINE host with no override, so the record had to fall stale first — that wait is part of the proof. Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box. The verdict is still no, for a new reason: a one-drive box with no off-site tier keeps none of the household's own files in any backup, the page says otherwise, and the restore that should save them makes it worse (R-537, R-538). Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing on the off-site server written or removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
18 KiB
DRILL — prove the P1 fixes on a fresh box (2026-09-16)
Interventions: 0 (O1 the operator's self-bind press and O2 the PBS re-issue press were pre-declared and are counted apart).
Ready for a volunteer: NO — and the reason is new, not one of the old ones. Every P1 fix this drill set out to prove did hold on a fresh box. But a one-drive box with no off-site tier — the state every fresh install starts in — keeps none of the household's own files in any backup, while the backup page says it does, and a restore then reports success and leaves the app listing photos it cannot open (R-537, R-538).
The automatic connect e-mail: PASSED. The host record was deleted at 12:22:59Z and the mail „Kösd össze a Felhom dobozodat” reached tester1@felhom.eu at 12:23:00Z — one second later, with selfbind_link_sent … (host delete) on the customer timeline. The requirement was two minutes.
Baselines at start (re-verified against live Gitea): controller
383a30b3c07bv0.243.0 (Unreleased: the „0 B" tile fix), agente98b857684f4v0.131.0, felhom.eu351296114c4dhub v0.114.0, catalog94bc5febaca2. Golden before this run: 0.242.0 (2026-09-14). Customertester-1, domainenkicsifelhom.hu, e-mailtester1@felhom.eu, no host, DR tier ticked, ep0 token held with no descriptor, namespace EMPTY (phase3-ep0-before.txt).
Phase 0 — rulings, the golden, the N100
Ruling 1 (2026-09-16, operator). The signing keys stay on DooPlex, owner-only
(/mnt/5_hdd/felhom.eu/felhom-op-operational, felhom-rec-recovery, felhom_op_ed25519, mode 0600);
CC may sign agent_update jobs with them until the first PAYING customer — testers excluded. Recorded in
CONTEXT.md and 04-control-plane-authorization.md §3.1. R-533 closes (path + mode);
R-530 narrows to "a fleet rollout step is still undesigned".
N100 signed and updated (R-530). felhom-opsign -op agent_update -host demo-felhom-8363b5 -key-id felhom-op-1 … -agent-version 0.131.0 -sha256 1118b552…c9c, queued 08:51:34Z, the box took it at
09:04:43Z (its own poll, ~13 min), dwelled 60 s and committed; controller-supervisor: started
(interval 30s, confirm 2, crash-loop 3/15m) in its journal. Peti's box untouched. The N100 keeps
controller 0.242.0; its floor was not moved by this run.
Ruling 2 — hub v0.115.0 (R-529). host_stale / host_down / host_recovered join the node_*
cooldown bypass with the same 5-minute dedupe. Red-proof: with the three host types removed the test fails
at "host_stale 39 min after the previous one was suppressed (sent=1)". Recorded in 08-alarm-ladder.md §6.2
beside the 2026-09-15 node ruling. Deployed by GitOps: ArgoCD Synced, deploy/hub image 0.115.0.
Golden 0.243.0 baked and vouched. First attempt FAILED and is recorded: the template picker took the
LAST debian-13 line, which is arm64, and the container would not start ("Detected container
architecture: arm64"); no golden was produced, the scratch guest was destroyed and the disk reverted.
Re-run with the amd64 template: GOLDEN_VERSION=0.243.0,
GOLDEN_SHA256=e2d1843c8b648910ddde7cef4fe9f9ba2ee25002cdb58fbc42543a3e8967c10a. Markers from the saved log
(phase0-bake-full.log, 326 lines): docker OK (overlay2 ×1, including mount point ×2 (rootfs + mp0),
upload OK (HTTP 201) ×1, FATAL 0, excluding 0. Token leak checks: control copy 1, committed log 0.
Vouched as a three-field change — golden 0.243.0 + agent 0.131.0 + min agent 0.131.0 — hub logged
Artifact manifest set: agent=0.131.0 golden=0.243.0 min_agent="0.131.0" wrapper_sha=true.
Phase 1 — the walk, as a volunteer
| # | step | what happened | time (UTC) |
|---|---|---|---|
| 1 | download from felhom.eu/letoltes |
the page names felhom-installer-1.27.1-pve9.2-1.iso, 1 705 322 496 B, sha256 25637007…c053; the downloaded file matched byte-for-byte |
08:55–08:56 |
| 2 | install | boot menu default is the graphical entry (15 s). It IS drivable by keyboard (Tab moves focus, Enter presses), but each step needs a screenshot to see where focus is, so the walk switched to the text-mode entry — the other entry a volunteer is offered — and says so. Screens: EULA (English Proxmox), disk target /dev/sda 32 GiB only (the 100 GB data disk is never offered as the target), Hungary/Europe-Budapest/Hungarian prefilled (the keyboard was set to U.S. English for typing only), password + mail@example.invalid prefilled, host name prefilled pve.example.invalid (replaced with tester1.enkicsifelhom.hu), DHCP 192.168.0.128/24, summary with [X] Automatically reboot after successful installation |
09:00–09:12 start |
| 2b | the reboot trap, measured | the automatic reboot re-entered the installer: a guest reboot reuses the running process's boot order, so a disk-first change made during the install does not apply. A volunteer with the stick still in sees the installer again. Cold stop + remove the stick + start → boots the installed system | 09:57 |
| 3 | first console screen | Felhom-only, Hungarian, no admin URL (8006 0 occurrences): „Felhom otthoni szerver … Ezen a gépen most nincs dolgod … Párosító kód: 37S-NFE … Nyisd meg az e-mailben kapott linket". The hub's Hosts page lists the same box as an unclaimed appliance with the same code |
09:58 |
| 4 | the connect mail + bind | O1 (pre-declared, not counted): the operator's „Send self-bind link" pressed once — the customer was already waiting when the automatic trigger shipped. Mail arrived at 09:59:56Z to tester1@felhom.eu, subject „Kösd össze a Felhom dobozodat", naming the two things to enter and the 7-day / 5-attempt limits. The bind page took the pairing code + owner passphrase and answered „Sikeres összekötés." |
09:59:55 → 10:01:14 |
| 5 | „ready" | the hub served agent 0.131.0 / golden 0.243.0 (this run's own bake) at 09:01:59Z; the guest came up and the controller announced itself: controller_started (0.243.0) at 10:03:30Z. The box's first host-report carrying a guest: 10:17:16Z (1 guest, 2 storage targets, 1 backup) | 10:01–10:17 |
| 6 | the tunnel from outside (R-510) | three GETs from DooPlex to https://felhom.enkicsifelhom.hu → 302 ×3 (cloudflare), following → 200 on „A szerver beállítása" — no 502, no 530. R-510 CLOSED | 10:04:19–25 |
| 6b | claim | the setup-code mail („Új beállító kód — újratelepült a szervered", 10:01:57Z, 72 h) + a new dashboard password → 302 → /. First try refused with „Érvénytelen űrlap": the field names are code / new_password / confirm_password | 10:06:18 |
| 7 | the off-site tier (R-511 → R-534) | the WG hook refused exactly as R-511 describes; O2 pressed → hub v0.114.0's ADOPT path ran and the ENDPOINT refused: status 255 … missing Datastore.Modify on /datastore/felhom-offsite → 502, nothing written (fail-closed). So the adopt fix is sound and inert until the ep0 grant is fixed — new row R-534 (P1). The box therefore has no whole-guest off-site tier | 10:02:15 / 10:03:29 |
| 8 | the file manager (R-513 on a FRESH box) | the app page shows „Kezdeti belépési adatok — Felhasználónév admin … A jelszót a Felhom állította be ezen a gépen"; reveal → 200, 16 chars; through the tunnel files.enkicsifelhom.hu: revealed 200, admin/admin 401, wrong 401 — and admin/admin was already 401 on a probe taken BEFORE any dashboard action | 10:06:43 |
| 8b | the data drive | „Nincs regisztrált adattároló" on arrival; the wizard formatted and registered /dev/sdb as „Adatlemez" (/mnt/felhom-drives/adatlemez, default, 97.9 GB) in ~2 s. Two refusals worth recording: the mount name rejects accented letters („érvénytelen csatlakoztatási név"), which is the first thing a Hungarian volunteer types, and the API needs mount_name, not label | 10:20 |
| 11 | the backup page (R-517 on a FRESH box) | per tier: „Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-16 12:08 · 624.3 MB · Naprakész"; „Biztonsági szerver – külön hardver (PBS) nincs beállítva"; the remote tile „Távoli rendszermentés nincs beállítva" (not ticked); the button says „A mentés alatt az alkalmazások leállnak — általában néhány perc, nagyobb adatnál több." | 10:22 |
| 9 | four apps | deployed through the dashboard API with the new drive as their data path: privatebin and vaultwarden running by 10:23:29Z, paperless-ngx and nextcloud by 10:24:29Z (2–3 min each); the hub logged four app_deployed events. Cards: no admin / admin anywhere; Vaultwarden's „Első lépések" is invite-first („Users" → „Invite User", „a regisztráció alapból le van zárva"); Paperless points at „Beállítások → Automatikusan generált értékek" and at the file manager's own login page | 10:21–10:24 |
| 9b | Vaultwarden, R-512 on a fresh box | over the public address, a stranger's send-verification-email → 400 „Registration not allowed or user already exists"; the app answers otherwise (its /api/config 200) | 10:25:05 |
| 9c | Paperless, R-514 on a fresh box | 20 three-page PDFs posted at once through the public address → 20 SUCCESS, 20 documents, no OOM; the container carries the new catalog limits (1280M, 1 worker × 1 thread) | 10:25:26–10:27 |
| 10–11 | „Mentés most" and the page | the button's promise measured: privatebin answered 200, went unreachable 10:27:17→10:27:43Z = 26 s, then answered again while the dump continued — inside „általában néhány perc". The page kept „Helyi tároló (local) ✓ … 624.3 MB · Naprakész" and „PBS — nincs beállítva" throughout, and showed „Mentés folyamatban… (fázis: snapshotted)" while running. The box's own first scheduled backup had already proven R-518 live: backup_tier_skipped (warning) at 12:08:25 CEST — „tier felhom-pbs skipped: its storage does not exist … No app was stopped for it" + operator mail | 10:27 |
| 12 | which golden it landed on | MEASURED, not inferred: the day-0 install left the golden archive on the box (vzdump-lxc-9100-…tar.zst, 653 997 919 B) whose sha256 is e2d1843c…c10a — byte-identical to this drill's own bake; the guest runs controller 0.243.0, so no self-update was needed | 12:02 |
New finding on the walk: R-535 (P2) — 25 minutes after a successful bind AND claim, with apps deploying, the box's console still showed „a doboz készen áll, és a párosításra vár" with the pairing code, and still promised „Ez a képernyő magától frissül".
Phase 2 — the faults
Five faults, each recorded the same way: what the customer saw · what the box did · time to steady ·
did an alarm fire and was it true · did an alarm that should have fired stay silent. Evidence per
fault in evidence-drill-0243-2026-09-16/phase2-*.txt.
| fault | customer saw | box did | steady | alarm |
|---|---|---|---|---|
| F9' — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | 37 s | controller_restarted_by_agent — true. But app_deployed for Mealie had already been sent at accept time → R-536 |
| F9'' — three more kills, 20 min apart, at idle | dashboard gone 30–90 s each time; the apps themselves never stopped | restarted every time: 61 s / 41 s / 61 s back to 200 | 61/41/61 s | none — and that is the finding: the restarts NEVER accumulated (window 15 min), so the 30-minute brake was never armed. A controller dying every 20 minutes is restarted forever, traced only by an info event that mails nobody → R-531, an operator ruling, measured not changed |
| F10 — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is back and lists all five, and none opens (Sabre\DAV\Exception\NotFound) |
tier-1 restore replayed 3 volumes + the database in 35 s and reported plain success | 35 s | none fired, and none exists for "restored database points at files that are not there" → R-537, R-538 |
| F11 — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then "Túl sok próbálkozás — próbáld újra 15 perc múlva" from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | claim_lockout (warning) at 13:11:42 CEST, operator mail 1 s later — fired and true |
| F12 — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | 124 s after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm |
| M1 — memory pressure (R-528) | app restarted 9–10 times | all three OOM signals silent: OOMKilled=false, docker events oom empty, the cgroup invisible inside the guest, dmesg unreadable |
— | no OOM alarm is possible on this box, same as on scratch 9202 |
F10 is the finding of this drill. On a one-drive box with no off-site tier — the state every fresh
install starts in — the household's own files are in no backup at all: the whole-guest tiers exclude
the data drive by design (07-backup-architecture.md, "[FACT] What the whole-guest tiers do NOT carry",
confirmed live: excluding bind mount point mp8 … (not a volume)), and the app's file leg lives at
tier 2 / tier 3, both unset. The design is not the defect. The defects are that the page calls tier 1
„DB + Konfig + Adatok" and prints the drive size beside it (R-537), and that a restore reports
success while leaving the app listing files it cannot open — and wipes the app's own trash, which still
held every byte (R-538).
Phase 3 — the morning after
Every app healthy, labels true. Front doors at 11:15Z: PrivateBin 200, Vaultwarden 200, Paperless
302 (its login redirect), Nextcloud status.php 200, file manager 200 at files.enkicsifelhom.hu.
All seven stacks read running. Controller 0.243.0 on the page, agent 0.131.0 on the hub, the
customer row reads Version 0.243.0 / Floor v0.242.0 — a floor is a minimum, so that is correct, and the
box runs the golden this drill baked.
Off-site restore onto 9202 — NOT WALKED, and not for time. There is nothing to restore from. The
off-site tier was never provisioned on this box: the re-issue fails on the endpoint token's missing
Datastore.Modify grant (R-534), so no descriptor and no upload path were ever created. Read-only
listing of ep0 confirms it: ns/tester-1/ct is empty both before (09:12Z) and after (11:14Z) the
drill, while the neighbouring ns/demo-hp/ct lists group 9201 — the positive control proving the
listing method shows groups when they exist. Nothing on ep0 was written, removed or pruned by this
session.
Interventions — counted, with the reason for each verdict
Pre-declared and counted apart: O1 the operator's „Send self-bind link" press (the customer was already waiting when the automatic trigger shipped) and O2 the „Re-issue PBS credentials" press.
Counted interventions: 0. Nothing on this walk needed a shell, an operator, or knowledge a household does not have. The four moments that could be mistaken for one, and why each is not:
| moment | why it is not an intervention |
|---|---|
| the automatic reboot re-entered the installer | a volunteer removes the USB stick and powers the box off and on — the same act. It is a real trap and is recorded as one, not as help from outside |
| the text-mode installer entry was used | the graphical entry IS keyboard-drivable; the walk switched because I cannot see the screen without a screenshot. A volunteer looking at a monitor has no such problem. A deviation of method, declared |
| the claim form and the drive form each refused once | both refusals were mine: I drove the API and guessed field names (code/new_password/confirm_password, and mount_name not label). A volunteer uses the browser form and never meets them |
| Paperless had to be repaired mid-drill | my own damage — I re-ran docker compose up -d by hand during the memory measurement and lost the controller-injected environment. Repaired through the controller's own API. Not a product fault |
What a volunteer WOULD have hit, as defects rather than interventions: the console keeps showing the pairing banner long after the box is bound and claimed (R-535), the drive wizard rejects an accented mount name — the first thing a Hungarian types (recorded on the walk), and the backup story in Phase 2 (R-537, R-538).
Teardown — three layers, stated
Evidence first (R-320). The agent journal (2000 lines), the controller log (91 lines) and the box's
own state (pct config, qm list, pvesm status, docker ps) were copied to DooPlex before anything
was destroyed. Token-leak control: 0 hits in all three files.
Layer 1 — the machine. Nested VM 334 (tester1-drill-0243) stopped and destroyed with
--purge --destroy-unreferenced-disks 1. After: qm list empty, /mnt/hdd_1/images/334 absent, no file
matching *334* left on the drive. Fence check: demo-hp's own containers 9201 and 9202 are both still
running and were not touched.
Layer 2 — the host. Nothing to remove: the drill box WAS the nested VM. demo-hp keeps its standing apps
and bentopdf.
Layer 3 — the hub. The host record tester-1-652049 is deleted; the customer tester-1 stays,
with its e-mail, domain and tunnel. RESET was never used — the customer-delete form on the hub is the
one that resets everything, and it was deliberately not touched. The host delete is refused while a host
is ONLINE (409, with no override by design, because a live agent would be permanently 401'd), so the
record had to fall stale first — which it does only after the machine is really gone. That wait is part of
the proof, not an obstacle.