Nothing on the walk needed a shell or an operator. The four moments that could be mistaken for help are listed with the reason each is not one — two of them were my own errors driving the API, and one was my own damage during the memory test. The alarm table is now measured from two independent sides: the hub's own log lines and the inbox. The one-hour operator cooldown is proven to suppress AND to release (backup_tier_skipped mailed 12:08, suppressed 12:37 and 12:58, mailed again 13:18). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
16 KiB
DRILL — prove the P1 fixes on a fresh box (2026-09-16)
Interventions: pending (O1/O2 pre-declared, counted apart). Ready for a volunteer: pending. The automatic connect e-mail: pending.
Baselines at start (re-verified against live Gitea): controller
383a30b3c07bv0.243.0 (Unreleased: the „0 B" tile fix), agente98b857684f4v0.131.0, felhom.eu351296114c4dhub v0.114.0, catalog94bc5febaca2. Golden before this run: 0.242.0 (2026-09-14). Customertester-1, domainenkicsifelhom.hu, e-mailtester1@felhom.eu, no host, DR tier ticked, ep0 token held with no descriptor, namespace EMPTY (phase3-ep0-before.txt).
Phase 0 — rulings, the golden, the N100
Ruling 1 (2026-09-16, operator). The signing keys stay on DooPlex, owner-only
(/mnt/5_hdd/felhom.eu/felhom-op-operational, felhom-rec-recovery, felhom_op_ed25519, mode 0600);
CC may sign agent_update jobs with them until the first PAYING customer — testers excluded. Recorded in
CONTEXT.md and 04-control-plane-authorization.md §3.1. R-533 closes (path + mode);
R-530 narrows to "a fleet rollout step is still undesigned".
N100 signed and updated (R-530). felhom-opsign -op agent_update -host demo-felhom-8363b5 -key-id felhom-op-1 … -agent-version 0.131.0 -sha256 1118b552…c9c, queued 08:51:34Z, the box took it at
09:04:43Z (its own poll, ~13 min), dwelled 60 s and committed; controller-supervisor: started
(interval 30s, confirm 2, crash-loop 3/15m) in its journal. Peti's box untouched. The N100 keeps
controller 0.242.0; its floor was not moved by this run.
Ruling 2 — hub v0.115.0 (R-529). host_stale / host_down / host_recovered join the node_*
cooldown bypass with the same 5-minute dedupe. Red-proof: with the three host types removed the test fails
at "host_stale 39 min after the previous one was suppressed (sent=1)". Recorded in 08-alarm-ladder.md §6.2
beside the 2026-09-15 node ruling. Deployed by GitOps: ArgoCD Synced, deploy/hub image 0.115.0.
Golden 0.243.0 baked and vouched. First attempt FAILED and is recorded: the template picker took the
LAST debian-13 line, which is arm64, and the container would not start ("Detected container
architecture: arm64"); no golden was produced, the scratch guest was destroyed and the disk reverted.
Re-run with the amd64 template: GOLDEN_VERSION=0.243.0,
GOLDEN_SHA256=e2d1843c8b648910ddde7cef4fe9f9ba2ee25002cdb58fbc42543a3e8967c10a. Markers from the saved log
(phase0-bake-full.log, 326 lines): docker OK (overlay2 ×1, including mount point ×2 (rootfs + mp0),
upload OK (HTTP 201) ×1, FATAL 0, excluding 0. Token leak checks: control copy 1, committed log 0.
Vouched as a three-field change — golden 0.243.0 + agent 0.131.0 + min agent 0.131.0 — hub logged
Artifact manifest set: agent=0.131.0 golden=0.243.0 min_agent="0.131.0" wrapper_sha=true.
Phase 1 — the walk, as a volunteer
| # | step | what happened | time (UTC) |
|---|---|---|---|
| 1 | download from felhom.eu/letoltes |
the page names felhom-installer-1.27.1-pve9.2-1.iso, 1 705 322 496 B, sha256 25637007…c053; the downloaded file matched byte-for-byte |
08:55–08:56 |
| 2 | install | boot menu default is the graphical entry (15 s). It IS drivable by keyboard (Tab moves focus, Enter presses), but each step needs a screenshot to see where focus is, so the walk switched to the text-mode entry — the other entry a volunteer is offered — and says so. Screens: EULA (English Proxmox), disk target /dev/sda 32 GiB only (the 100 GB data disk is never offered as the target), Hungary/Europe-Budapest/Hungarian prefilled (the keyboard was set to U.S. English for typing only), password + mail@example.invalid prefilled, host name prefilled pve.example.invalid (replaced with tester1.enkicsifelhom.hu), DHCP 192.168.0.128/24, summary with [X] Automatically reboot after successful installation |
09:00–09:12 start |
| 2b | the reboot trap, measured | the automatic reboot re-entered the installer: a guest reboot reuses the running process's boot order, so a disk-first change made during the install does not apply. A volunteer with the stick still in sees the installer again. Cold stop + remove the stick + start → boots the installed system | 09:57 |
| 3 | first console screen | Felhom-only, Hungarian, no admin URL (8006 0 occurrences): „Felhom otthoni szerver … Ezen a gépen most nincs dolgod … Párosító kód: 37S-NFE … Nyisd meg az e-mailben kapott linket". The hub's Hosts page lists the same box as an unclaimed appliance with the same code |
09:58 |
| 4 | the connect mail + bind | O1 (pre-declared, not counted): the operator's „Send self-bind link" pressed once — the customer was already waiting when the automatic trigger shipped. Mail arrived at 09:59:56Z to tester1@felhom.eu, subject „Kösd össze a Felhom dobozodat", naming the two things to enter and the 7-day / 5-attempt limits. The bind page took the pairing code + owner passphrase and answered „Sikeres összekötés." |
09:59:55 → 10:01:14 |
| 5 | „ready" | the hub served agent 0.131.0 / golden 0.243.0 (this run's own bake) at 09:01:59Z; the guest came up and the controller announced itself: controller_started (0.243.0) at 10:03:30Z. The box's first host-report carrying a guest: 10:17:16Z (1 guest, 2 storage targets, 1 backup) | 10:01–10:17 |
| 6 | the tunnel from outside (R-510) | three GETs from DooPlex to https://felhom.enkicsifelhom.hu → 302 ×3 (cloudflare), following → 200 on „A szerver beállítása" — no 502, no 530. R-510 CLOSED | 10:04:19–25 |
| 6b | claim | the setup-code mail („Új beállító kód — újratelepült a szervered", 10:01:57Z, 72 h) + a new dashboard password → 302 → /. First try refused with „Érvénytelen űrlap": the field names are code / new_password / confirm_password | 10:06:18 |
| 7 | the off-site tier (R-511 → R-534) | the WG hook refused exactly as R-511 describes; O2 pressed → hub v0.114.0's ADOPT path ran and the ENDPOINT refused: status 255 … missing Datastore.Modify on /datastore/felhom-offsite → 502, nothing written (fail-closed). So the adopt fix is sound and inert until the ep0 grant is fixed — new row R-534 (P1). The box therefore has no whole-guest off-site tier | 10:02:15 / 10:03:29 |
| 8 | the file manager (R-513 on a FRESH box) | the app page shows „Kezdeti belépési adatok — Felhasználónév admin … A jelszót a Felhom állította be ezen a gépen"; reveal → 200, 16 chars; through the tunnel files.enkicsifelhom.hu: revealed 200, admin/admin 401, wrong 401 — and admin/admin was already 401 on a probe taken BEFORE any dashboard action | 10:06:43 |
| 8b | the data drive | „Nincs regisztrált adattároló" on arrival; the wizard formatted and registered /dev/sdb as „Adatlemez" (/mnt/felhom-drives/adatlemez, default, 97.9 GB) in ~2 s. Two refusals worth recording: the mount name rejects accented letters („érvénytelen csatlakoztatási név"), which is the first thing a Hungarian volunteer types, and the API needs mount_name, not label | 10:20 |
| 11 | the backup page (R-517 on a FRESH box) | per tier: „Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-16 12:08 · 624.3 MB · Naprakész"; „Biztonsági szerver – külön hardver (PBS) nincs beállítva"; the remote tile „Távoli rendszermentés nincs beállítva" (not ticked); the button says „A mentés alatt az alkalmazások leállnak — általában néhány perc, nagyobb adatnál több." | 10:22 |
| 9 | four apps | deployed through the dashboard API with the new drive as their data path: privatebin and vaultwarden running by 10:23:29Z, paperless-ngx and nextcloud by 10:24:29Z (2–3 min each); the hub logged four app_deployed events. Cards: no admin / admin anywhere; Vaultwarden's „Első lépések" is invite-first („Users" → „Invite User", „a regisztráció alapból le van zárva"); Paperless points at „Beállítások → Automatikusan generált értékek" and at the file manager's own login page | 10:21–10:24 |
| 9b | Vaultwarden, R-512 on a fresh box | over the public address, a stranger's send-verification-email → 400 „Registration not allowed or user already exists"; the app answers otherwise (its /api/config 200) | 10:25:05 |
| 9c | Paperless, R-514 on a fresh box | 20 three-page PDFs posted at once through the public address → 20 SUCCESS, 20 documents, no OOM; the container carries the new catalog limits (1280M, 1 worker × 1 thread) | 10:25:26–10:27 |
| 10–11 | „Mentés most" and the page | the button's promise measured: privatebin answered 200, went unreachable 10:27:17→10:27:43Z = 26 s, then answered again while the dump continued — inside „általában néhány perc". The page kept „Helyi tároló (local) ✓ … 624.3 MB · Naprakész" and „PBS — nincs beállítva" throughout, and showed „Mentés folyamatban… (fázis: snapshotted)" while running. The box's own first scheduled backup had already proven R-518 live: backup_tier_skipped (warning) at 12:08:25 CEST — „tier felhom-pbs skipped: its storage does not exist … No app was stopped for it" + operator mail | 10:27 |
| 12 | which golden it landed on | MEASURED, not inferred: the day-0 install left the golden archive on the box (vzdump-lxc-9100-…tar.zst, 653 997 919 B) whose sha256 is e2d1843c…c10a — byte-identical to this drill's own bake; the guest runs controller 0.243.0, so no self-update was needed | 12:02 |
New finding on the walk: R-535 (P2) — 25 minutes after a successful bind AND claim, with apps deploying, the box's console still showed „a doboz készen áll, és a párosításra vár" with the pairing code, and still promised „Ez a képernyő magától frissül".
Phase 2 — the faults
Five faults, each recorded the same way: what the customer saw · what the box did · time to steady ·
did an alarm fire and was it true · did an alarm that should have fired stay silent. Evidence per
fault in evidence-drill-0243-2026-09-16/phase2-*.txt.
| fault | customer saw | box did | steady | alarm |
|---|---|---|---|---|
| F9' — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | 37 s | controller_restarted_by_agent — true. But app_deployed for Mealie had already been sent at accept time → R-536 |
| F10 — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is back and lists all five, and none opens (Sabre\DAV\Exception\NotFound) |
tier-1 restore replayed 3 volumes + the database in 35 s and reported plain success | 35 s | none fired, and none exists for "restored database points at files that are not there" → R-537, R-538 |
| F11 — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then "Túl sok próbálkozás — próbáld újra 15 perc múlva" from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | claim_lockout (warning) at 13:11:42 CEST, operator mail 1 s later — fired and true |
| F12 — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | 124 s after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm |
| M1 — memory pressure (R-528) | app restarted 9–10 times | all three OOM signals silent: OOMKilled=false, docker events oom empty, the cgroup invisible inside the guest, dmesg unreadable |
— | no OOM alarm is possible on this box, same as on scratch 9202 |
F10 is the finding of this drill. On a one-drive box with no off-site tier — the state every fresh
install starts in — the household's own files are in no backup at all: the whole-guest tiers exclude
the data drive by design (07-backup-architecture.md, "[FACT] What the whole-guest tiers do NOT carry",
confirmed live: excluding bind mount point mp8 … (not a volume)), and the app's file leg lives at
tier 2 / tier 3, both unset. The design is not the defect. The defects are that the page calls tier 1
„DB + Konfig + Adatok" and prints the drive size beside it (R-537), and that a restore reports
success while leaving the app listing files it cannot open — and wipes the app's own trash, which still
held every byte (R-538).
Phase 3 — the morning after
Every app healthy, labels true. Front doors at 11:15Z: PrivateBin 200, Vaultwarden 200, Paperless
302 (its login redirect), Nextcloud status.php 200, file manager 200 at files.enkicsifelhom.hu.
All seven stacks read running. Controller 0.243.0 on the page, agent 0.131.0 on the hub, the
customer row reads Version 0.243.0 / Floor v0.242.0 — a floor is a minimum, so that is correct, and the
box runs the golden this drill baked.
Off-site restore onto 9202 — NOT WALKED, and not for time. There is nothing to restore from. The
off-site tier was never provisioned on this box: the re-issue fails on the endpoint token's missing
Datastore.Modify grant (R-534), so no descriptor and no upload path were ever created. Read-only
listing of ep0 confirms it: ns/tester-1/ct is empty both before (09:12Z) and after (11:14Z) the
drill, while the neighbouring ns/demo-hp/ct lists group 9201 — the positive control proving the
listing method shows groups when they exist. Nothing on ep0 was written, removed or pruned by this
session.
Interventions — counted, with the reason for each verdict
Pre-declared and counted apart: O1 the operator's „Send self-bind link" press (the customer was already waiting when the automatic trigger shipped) and O2 the „Re-issue PBS credentials" press.
Counted interventions: 0. Nothing on this walk needed a shell, an operator, or knowledge a household does not have. The four moments that could be mistaken for one, and why each is not:
| moment | why it is not an intervention |
|---|---|
| the automatic reboot re-entered the installer | a volunteer removes the USB stick and powers the box off and on — the same act. It is a real trap and is recorded as one, not as help from outside |
| the text-mode installer entry was used | the graphical entry IS keyboard-drivable; the walk switched because I cannot see the screen without a screenshot. A volunteer looking at a monitor has no such problem. A deviation of method, declared |
| the claim form and the drive form each refused once | both refusals were mine: I drove the API and guessed field names (code/new_password/confirm_password, and mount_name not label). A volunteer uses the browser form and never meets them |
| Paperless had to be repaired mid-drill | my own damage — I re-ran docker compose up -d by hand during the memory measurement and lost the controller-injected environment. Repaired through the controller's own API. Not a product fault |
What a volunteer WOULD have hit, as defects rather than interventions: the console keeps showing the pairing banner long after the box is bound and claimed (R-535), the drive wizard rejects an accented mount name — the first thing a Hungarian types (recorded on the walk), and the backup story in Phase 2 (R-537, R-538).