dfd854474e
gates / gates (push) Successful in 21s
The automatic connect e-mail is proven with a real mailbox: the host record was deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second later, with selfbind_link_sent (host delete) on the timeline. The requirement was two minutes. The hub refuses to delete an ONLINE host with no override, so the record had to fall stale first — that wait is part of the proof. Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box. The verdict is still no, for a new reason: a one-drive box with no off-site tier keeps none of the household's own files in any backup, the page says otherwise, and the restore that should save them makes it worse (R-537, R-538). Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing on the off-site server written or removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
149 lines
18 KiB
Markdown
149 lines
18 KiB
Markdown
# DRILL — prove the P1 fixes on a fresh box (2026-09-16)
|
||
|
||
**Interventions: 0** (O1 the operator's self-bind press and O2 the PBS re-issue press were pre-declared and are counted apart).
|
||
**Ready for a volunteer: NO — and the reason is new, not one of the old ones.** Every P1 fix this drill set out to prove did hold on a fresh box. But a one-drive box with no off-site tier — the state every fresh install starts in — keeps **none of the household's own files in any backup**, while the backup page says it does, and a restore then reports success and leaves the app listing photos it cannot open (R-537, R-538).
|
||
**The automatic connect e-mail: PASSED.** The host record was deleted at 12:22:59Z and the mail „Kösd össze a Felhom dobozodat” reached `tester1@felhom.eu` at **12:23:00Z — one second later**, with `selfbind_link_sent … (host delete)` on the customer timeline. The requirement was two minutes.
|
||
|
||
> Baselines at start (re-verified against live Gitea): controller `383a30b3c07b` v0.243.0 (`Unreleased`: the
|
||
> „0 B" tile fix), agent `e98b857684f4` v0.131.0, felhom.eu `351296114c4d` hub v0.114.0, catalog `94bc5febaca2`.
|
||
> Golden before this run: 0.242.0 (2026-09-14). Customer `tester-1`, domain `enkicsifelhom.hu`, e-mail
|
||
> `tester1@felhom.eu`, **no host**, DR tier ticked, ep0 token held with no descriptor, namespace EMPTY
|
||
> (`phase3-ep0-before.txt`).
|
||
|
||
## Phase 0 — rulings, the golden, the N100
|
||
|
||
**Ruling 1 (2026-09-16, operator).** The signing keys stay on DooPlex, owner-only
|
||
(`/mnt/5_hdd/felhom.eu/felhom-op-operational`, `felhom-rec-recovery`, `felhom_op_ed25519`, mode 0600);
|
||
CC may sign `agent_update` jobs with them until the first PAYING customer — testers excluded. Recorded in
|
||
`CONTEXT.md` and `04-control-plane-authorization.md` §3.1. **R-533 closes** (path + mode);
|
||
**R-530 narrows** to "a fleet rollout step is still undesigned".
|
||
|
||
**N100 signed and updated (R-530).** `felhom-opsign -op agent_update -host demo-felhom-8363b5 -key-id
|
||
felhom-op-1 … -agent-version 0.131.0 -sha256 1118b552…c9c`, queued 08:51:34Z, **the box took it at
|
||
09:04:43Z** (its own poll, ~13 min), dwelled 60 s and committed; `controller-supervisor: started`
|
||
(interval 30s, confirm 2, crash-loop 3/15m) in its journal. Peti's box untouched. The N100 keeps
|
||
controller 0.242.0; its floor was not moved by this run.
|
||
|
||
**Ruling 2 — hub v0.115.0 (R-529).** `host_stale` / `host_down` / `host_recovered` join the `node_*`
|
||
cooldown bypass with the same 5-minute dedupe. Red-proof: with the three host types removed the test fails
|
||
at "host_stale 39 min after the previous one was suppressed (sent=1)". Recorded in `08-alarm-ladder.md` §6.2
|
||
beside the 2026-09-15 node ruling. Deployed by GitOps: ArgoCD Synced, `deploy/hub` image 0.115.0.
|
||
|
||
**Golden 0.243.0 baked and vouched.** First attempt FAILED and is recorded: the template picker took the
|
||
LAST `debian-13` line, which is **arm64**, and the container would not start ("Detected container
|
||
architecture: arm64"); no golden was produced, the scratch guest was destroyed and the disk reverted.
|
||
Re-run with the amd64 template: `GOLDEN_VERSION=0.243.0`,
|
||
`GOLDEN_SHA256=e2d1843c8b648910ddde7cef4fe9f9ba2ee25002cdb58fbc42543a3e8967c10a`. Markers from the saved log
|
||
(`phase0-bake-full.log`, 326 lines): `docker OK (overlay2` ×1, `including mount point` ×2 (rootfs + mp0),
|
||
`upload OK (HTTP 201)` ×1, `FATAL` 0, `excluding` 0. Token leak checks: control copy 1, committed log 0.
|
||
Vouched as a three-field change — golden 0.243.0 + agent 0.131.0 + min agent 0.131.0 — hub logged
|
||
`Artifact manifest set: agent=0.131.0 golden=0.243.0 min_agent="0.131.0" wrapper_sha=true`.
|
||
|
||
## Phase 1 — the walk, as a volunteer
|
||
|
||
| # | step | what happened | time (UTC) |
|
||
|---|---|---|---|
|
||
| 1 | download from `felhom.eu/letoltes` | the page names `felhom-installer-1.27.1-pve9.2-1.iso`, 1 705 322 496 B, sha256 `25637007…c053`; the downloaded file matched **byte-for-byte** | 08:55–08:56 |
|
||
| 2 | install | boot menu default is the **graphical** entry (15 s). It IS drivable by keyboard (Tab moves focus, Enter presses), but each step needs a screenshot to see where focus is, so the walk switched to the **text-mode** entry — the other entry a volunteer is offered — and says so. Screens: EULA (English Proxmox), disk target **`/dev/sda 32 GiB` only** (the 100 GB data disk is never offered as the target), Hungary/Europe-Budapest/Hungarian prefilled (the keyboard was set to U.S. English **for typing only**), password + `mail@example.invalid` prefilled, host name prefilled `pve.example.invalid` (replaced with `tester1.enkicsifelhom.hu`), DHCP `192.168.0.128/24`, summary with **[X] Automatically reboot after successful installation** | 09:00–09:12 start |
|
||
| 2b | the reboot trap, measured | the automatic reboot **re-entered the installer**: a guest reboot reuses the running process's boot order, so a disk-first change made during the install does not apply. A volunteer with the stick still in sees the installer again. Cold stop + remove the stick + start → boots the installed system | 09:57 |
|
||
| 3 | first console screen | **Felhom-only, Hungarian, no admin URL** (`8006` 0 occurrences): „Felhom otthoni szerver … Ezen a gépen most nincs dolgod … Párosító kód: 37S-NFE … Nyisd meg az e-mailben kapott linket". The hub's Hosts page lists the same box as an unclaimed appliance with the same code | 09:58 |
|
||
| 4 | the connect mail + bind | **O1 (pre-declared, not counted):** the operator's „Send self-bind link" pressed once — the customer was already waiting when the automatic trigger shipped. Mail arrived at **09:59:56Z** to `tester1@felhom.eu`, subject „Kösd össze a Felhom dobozodat", naming the two things to enter and the 7-day / 5-attempt limits. The bind page took the pairing code + owner passphrase and answered **„Sikeres összekötés."** | 09:59:55 → 10:01:14 |
|
||
|
||
| 5 | „ready" | the hub served **agent 0.131.0 / golden 0.243.0** (this run's own bake) at 09:01:59Z; the guest came up and the controller announced itself: `controller_started (0.243.0)` at 10:03:30Z. The box's first host-report carrying a guest: 10:17:16Z (1 guest, 2 storage targets, 1 backup) | 10:01–10:17 |
|
||
| 6 | **the tunnel from outside** (R-510) | three GETs from DooPlex to `https://felhom.enkicsifelhom.hu` → **302 ×3 (cloudflare)**, following → **200** on „A szerver beállítása" — no 502, no 530. **R-510 CLOSED** | 10:04:19–25 |
|
||
| 6b | claim | the setup-code mail („Új beállító kód — újratelepült a szervered", 10:01:57Z, 72 h) + a new dashboard password → **302 → /**. First try refused with „Érvénytelen űrlap": the field names are `code` / `new_password` / `confirm_password` | 10:06:18 |
|
||
| 7 | **the off-site tier** (R-511 → **R-534**) | the WG hook refused exactly as R-511 describes; **O2 pressed** → hub v0.114.0's ADOPT path ran and the ENDPOINT refused: `status 255 … missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing written (fail-closed). So the adopt fix is sound and **inert until the ep0 grant is fixed** — new row **R-534 (P1)**. The box therefore has **no** whole-guest off-site tier | 10:02:15 / 10:03:29 |
|
||
| 8 | **the file manager** (R-513 on a FRESH box) | the app page shows „Kezdeti belépési adatok — Felhasználónév admin … A jelszót a Felhom állította be ezen a gépen"; reveal → 200, 16 chars; through the tunnel `files.enkicsifelhom.hu`: revealed **200**, `admin`/`admin` **401**, wrong **401** — and admin/admin was already 401 on a probe taken BEFORE any dashboard action | 10:06:43 |
|
||
| 8b | the data drive | „Nincs regisztrált adattároló" on arrival; the wizard formatted and registered `/dev/sdb` as „Adatlemez" (`/mnt/felhom-drives/adatlemez`, default, 97.9 GB) in ~2 s. **Two refusals worth recording:** the mount name rejects accented letters („érvénytelen csatlakoztatási név"), which is the first thing a Hungarian volunteer types, and the API needs `mount_name`, not `label` | 10:20 |
|
||
| 11 | **the backup page** (R-517 on a FRESH box) | per tier: „Helyi tároló (local) ✓ Utolsó sikeres mentés: 2026-09-16 12:08 · 624.3 MB · Naprakész"; „Biztonsági szerver – külön hardver (PBS) **nincs beállítva**"; the remote tile „Távoli rendszermentés **nincs beállítva**" (not ticked); the button says „A mentés alatt az alkalmazások leállnak — általában néhány perc, nagyobb adatnál több." | 10:22 |
|
||
|
||
| 9 | four apps | deployed through the dashboard API with the new drive as their data path: privatebin and vaultwarden **running by 10:23:29Z**, paperless-ngx and nextcloud **by 10:24:29Z** (2–3 min each); the hub logged four `app_deployed` events. Cards: **no `admin / admin` anywhere**; Vaultwarden's „Első lépések" is invite-first („Users" → „Invite User", „a regisztráció alapból le van zárva"); Paperless points at „Beállítások → Automatikusan generált értékek" and at the file manager's own login page | 10:21–10:24 |
|
||
| 9b | **Vaultwarden, R-512 on a fresh box** | over the public address, a stranger's `send-verification-email` → **400 „Registration not allowed or user already exists"**; the app answers otherwise (its `/api/config` 200) | 10:25:05 |
|
||
| 9c | **Paperless, R-514 on a fresh box** | 20 three-page PDFs posted at once through the public address → **20 SUCCESS, 20 documents**, no OOM; the container carries the new catalog limits (1280M, 1 worker × 1 thread) | 10:25:26–10:27 |
|
||
| 10–11 | **„Mentés most" and the page** | the button's promise measured: privatebin answered 200, went unreachable **10:27:17→10:27:43Z = 26 s**, then answered again while the dump continued — inside „általában néhány perc". The page kept „Helyi tároló (local) ✓ … 624.3 MB · Naprakész" and „PBS — nincs beállítva" throughout, and showed „Mentés folyamatban… (fázis: snapshotted)" while running. **The box's own first scheduled backup had already proven R-518 live**: `backup_tier_skipped (warning)` at 12:08:25 CEST — „tier felhom-pbs skipped: its storage does not exist … No app was stopped for it" + operator mail | 10:27 |
|
||
| 12 | which golden it landed on | **MEASURED, not inferred:** the day-0 install left the golden archive on the box (`vzdump-lxc-9100-…tar.zst`, 653 997 919 B) whose sha256 is **e2d1843c…c10a** — byte-identical to this drill's own bake; the guest runs controller **0.243.0**, so no self-update was needed | 12:02 |
|
||
|
||
**New finding on the walk: R-535 (P2)** — 25 minutes after a successful bind AND claim, with apps deploying, the box's console still showed „a doboz készen áll, és a párosításra vár" with the pairing code, and still promised „Ez a képernyő magától frissül".
|
||
|
||
|
||
## Phase 2 — the faults
|
||
|
||
Five faults, each recorded the same way: what the customer saw · what the box did · time to steady ·
|
||
did an alarm fire and was it true · did an alarm that should have fired stay silent. Evidence per
|
||
fault in `evidence-drill-0243-2026-09-16/phase2-*.txt`.
|
||
|
||
| fault | customer saw | box did | steady | alarm |
|
||
|---|---|---|---|---|
|
||
| **F9'** — controller killed 5 s into a deploy, empty budget | dashboard gone ~37 s; the interrupted app simply not installed | agent saw it on sweep 1, confirmed and restarted on sweep 2 (12:31:45 → 12:32:15/16 CEST) | **37 s** | `controller_restarted_by_agent` — true. **But** `app_deployed` for Mealie had already been sent at accept time → **R-536** |
|
||
| **F9''** — three more kills, 20 min apart, at idle | dashboard gone 30–90 s each time; the apps themselves never stopped | restarted every time: **61 s / 41 s / 61 s** back to 200 | 61/41/61 s | none — and that is the finding: the restarts NEVER accumulated (window 15 min), so the 30-minute brake was never armed. A controller dying every 20 minutes is restarted forever, traced only by an `info` event that mails nobody → **R-531**, an operator ruling, measured not changed |
|
||
| **F10** — a child deletes the photo folder | folder and 5 photos gone; after the restore the folder is **back and lists all five**, and **none opens** (`Sabre\DAV\Exception\NotFound`) | tier-1 restore replayed 3 volumes + the database in **35 s** and reported plain success | 35 s | none fired, and **none exists** for "restored database points at files that are not there" → **R-537, R-538** |
|
||
| **F11** — forgotten passwords ×5 | Nextcloud: 401 five times, no lock-out, correct password still works. The box's setup code: wrong twice, then **"Túl sok próbálkozás — próbáld újra 15 perc múlva"** from the third try; "Új beállító kód kérése" still works | claim counter locks the page for 15 min | immediate | `claim_lockout` (warning) at 13:11:42 CEST, operator mail **1 s later** — fired and **true** |
|
||
| **F12** — two reboots inside two minutes | dashboard gone ~2 min, then everything back | boot reconciler restarted all seven stacks; no double start; nothing stuck "telepítés folyamatban" | **124 s** after the second reset | none fired; correct — a reboot inside the liveness window is not an alarm |
|
||
| **M1** — memory pressure (R-528) | app restarted 9–10 times | all three OOM signals **silent**: `OOMKilled=false`, `docker events oom` empty, the cgroup invisible inside the guest, `dmesg` unreadable | — | no OOM alarm is possible on this box, same as on scratch 9202 |
|
||
|
||
**F10 is the finding of this drill.** On a one-drive box with no off-site tier — the state every fresh
|
||
install starts in — the household's own files are in **no backup at all**: the whole-guest tiers exclude
|
||
the data drive by design (`07-backup-architecture.md`, "[FACT] What the whole-guest tiers do NOT carry",
|
||
confirmed live: `excluding bind mount point mp8 … (not a volume)`), and the app's file leg lives at
|
||
tier 2 / tier 3, both unset. The design is not the defect. The defects are that the page calls tier 1
|
||
„DB + Konfig + Adatok" and prints the drive size beside it (**R-537**), and that a restore reports
|
||
success while leaving the app listing files it cannot open — and wipes the app's own trash, which still
|
||
held every byte (**R-538**).
|
||
|
||
## Phase 3 — the morning after
|
||
|
||
**Every app healthy, labels true.** Front doors at 11:15Z: PrivateBin 200, Vaultwarden 200, Paperless
|
||
302 (its login redirect), Nextcloud `status.php` 200, file manager 200 at `files.enkicsifelhom.hu`.
|
||
All seven stacks read `running`. Controller **0.243.0** on the page, agent **0.131.0** on the hub, the
|
||
customer row reads Version 0.243.0 / Floor v0.242.0 — a floor is a minimum, so that is correct, and the
|
||
box runs the golden this drill baked.
|
||
|
||
**Off-site restore onto 9202 — NOT WALKED, and not for time.** There is nothing to restore from. The
|
||
off-site tier was never provisioned on this box: the re-issue fails on the endpoint token's missing
|
||
`Datastore.Modify` grant (**R-534**), so no descriptor and no upload path were ever created. Read-only
|
||
listing of ep0 confirms it: `ns/tester-1/ct` is **empty** both before (09:12Z) and after (11:14Z) the
|
||
drill, while the neighbouring `ns/demo-hp/ct` lists group 9201 — the positive control proving the
|
||
listing method shows groups when they exist. **Nothing on ep0 was written, removed or pruned by this
|
||
session.**
|
||
|
||
## Interventions — counted, with the reason for each verdict
|
||
|
||
**Pre-declared and counted apart:** **O1** the operator's „Send self-bind link" press (the customer was
|
||
already waiting when the automatic trigger shipped) and **O2** the „Re-issue PBS credentials" press.
|
||
|
||
**Counted interventions: 0.** Nothing on this walk needed a shell, an operator, or knowledge a household
|
||
does not have. The four moments that could be mistaken for one, and why each is not:
|
||
|
||
| moment | why it is not an intervention |
|
||
|---|---|
|
||
| the automatic reboot re-entered the installer | a volunteer removes the USB stick and powers the box off and on — the same act. It is a real trap and is recorded as one, not as help from outside |
|
||
| the text-mode installer entry was used | the graphical entry IS keyboard-drivable; the walk switched because *I* cannot see the screen without a screenshot. A volunteer looking at a monitor has no such problem. A deviation of method, declared |
|
||
| the claim form and the drive form each refused once | both refusals were **mine**: I drove the API and guessed field names (`code`/`new_password`/`confirm_password`, and `mount_name` not `label`). A volunteer uses the browser form and never meets them |
|
||
| Paperless had to be repaired mid-drill | **my own damage** — I re-ran `docker compose up -d` by hand during the memory measurement and lost the controller-injected environment. Repaired through the controller's own API. Not a product fault |
|
||
|
||
**What a volunteer WOULD have hit, as defects rather than interventions:** the console keeps showing the
|
||
pairing banner long after the box is bound and claimed (**R-535**), the drive wizard rejects an accented
|
||
mount name — the first thing a Hungarian types (recorded on the walk), and the backup story in Phase 2
|
||
(**R-537**, **R-538**).
|
||
|
||
## Teardown — three layers, stated
|
||
|
||
**Evidence first (R-320).** The agent journal (2000 lines), the controller log (91 lines) and the box's
|
||
own state (`pct config`, `qm list`, `pvesm status`, `docker ps`) were copied to DooPlex **before** anything
|
||
was destroyed. Token-leak control: 0 hits in all three files.
|
||
|
||
**Layer 1 — the machine.** Nested VM **334** (`tester1-drill-0243`) stopped and destroyed with
|
||
`--purge --destroy-unreferenced-disks 1`. After: `qm list` empty, `/mnt/hdd_1/images/334` absent, no file
|
||
matching `*334*` left on the drive. **Fence check: demo-hp's own containers 9201 and 9202 are both still
|
||
running and were not touched.**
|
||
|
||
**Layer 2 — the host.** Nothing to remove: the drill box WAS the nested VM. demo-hp keeps its standing apps
|
||
and `bentopdf`.
|
||
|
||
**Layer 3 — the hub.** The **host record** `tester-1-652049` is deleted; the **customer** `tester-1` stays,
|
||
with its e-mail, domain and tunnel. **`RESET` was never used** — the customer-delete form on the hub is the
|
||
one that resets everything, and it was deliberately not touched. The host delete is **refused while a host
|
||
is ONLINE** (409, with no override by design, because a live agent would be permanently 401'd), so the
|
||
record had to fall stale first — which it does only after the machine is really gone. That wait is part of
|
||
the proof, not an obstacle.
|