# DRILL — Day-0 clean-room install / uninstall / reinstall (2026-07-03) **Class:** Clean-room drill (empirical validation of `documentation/runbooks/day0-install.md` + residue-proof of `--uninstall`). Companion deliverables: the runbook itself and `scripts/felhom-host-install.sh` **v1.9.1**. **Verdict (short):** the Day-0 guide is **executable verbatim end-to-end** on a virgin PVE 9.2 box: install → verify → app deploy → uninstall → **zero-Felhom-residue diff** → reinstall all pass. The drill found **one installer gap that mattered** (the controller never deployed without a guest reboot — fixed in v1.9.1), **five uninstall residue items** (all fixed in v1.9.1, re-proven by a second residue-diff), and produced a decisive **OQ-2 answer** (a fresh box does NOT self-update its controller — a one-time manual update is a mandatory install-day step). **Go for Peti** with the go/no-go list in §8. No secrets appear in this document. Drill credentials lived only on the build server (0600) and in the drill VM; the drill VM was reverted to its virgin snapshot afterwards. --- ## 1. Environment (Phase 0 — primary mechanism, validated) | Item | Value | |---|---| | Host | build server 192.168.0.180 (Debian 13, /dev/kvm, QEMU 10.0.8 installed for the drill) | | VM | nested **PVE 9.2.2** (unattended install: `proxmox-auto-install-assistant` 9.2.7 answer file, ext4/LVM-thin layout), 4 vCPU / 8 GiB RAM / 150 GiB sparse qcow2 | | Resulting storage | `local` 46 GiB dir + `local-lvm` **76 GiB** thin — deliberately BELOW the 120 GiB floor, so the warn path is exercised | | Network | QEMU user-mode NAT (slirp); outbound to hub.felhom.eu / gitea.dooplex.hu / felhom.eu confirmed from host AND guest; SSH via hostfwd | | Restore point | qcow2 internal snapshot `virgin` taken post-install, pre-first-boot | | Hub side | live hub v0.31.0; throwaway customer **drill-1** (domain `drill.demo-felhom.eu`, fake CF tunnel token, real Gitea read credential copied from the demo customer's row, hub-generated-format passphrase) inserted directly into `customer_configs` (the web create form is operator-password-gated — CC cannot drive it; the row is byte-identical to `SaveCustomerConfig`, spike-day0 precedent) | | Versions under test | script v1.9.0 → v1.9.1, agent v0.62.0, golden v0.85.1 (hub artifact manifest), controller floor 0.87.0, latest 0.98.3 | **Environment quirk (drill-only, not product):** slirp's DHCP pool starts at 10.0.2.15 — the same address the PVE host holds statically — so the freshly-provisioned guest's DHCP grabbed the host's IP, intermittently breaking the drill SSH (and truncating the D2 transcript). Worked around by pinning the guest to 10.0.2.40. A real LAN with a real DHCP server does not have this failure mode. ## 2. D0 — baseline Captured off-box before any Felhom artifact touched the box: sorted file list of `/etc /usr/local/bin /usr/local/sbin /var/lib` (5 076 entries), `systemctl list-unit-files` (379 units), `pveum user/role/acl/pool list`, `pvesm status`, `pct list; qm list` (empty). Stored at `~/drill/D0/` on the build server. ## 3. D1 — hub onboarding (guide Part A) - Customer-create field set verified against `hub/internal/web/configs.go` (`handleConfigCreate` + `buildConfigJSON`): customer_id/name/domain/email + `infrastructure.cf_tunnel_token`, `infrastructure.cf_api_token`, `git.username`, `git.token`; passphrase (5 words) + api-key hub-generated. **`git.token` is load-bearing for Day-0** — the installer dies at step 5/8 without it (guide Part A marks it required). - Box-facing contracts probed for drill-1: `GET /api/v1/config/drill-1` → **200**; `GET /api/v1/artifacts/drill-1` → agent 0.62.0 + golden 0.85.1; wrong passphrase → **401**. - **UI caveat:** the operator-password-gated create form itself was not click-driven (CC has no operator password); the row was inserted store-identically. Part A's field list is code-verified and its outputs are live-verified. ## 4. D2 — install (guide Parts B/C, script v1.9.0) Part B checklist ran verbatim (two guide corrections folded in: hub root answers **302**, and a cleaner node-count one-liner). Script fetched from `https://felhom.eu/scripts/felhom-host-install.sh` (text/plain, the git-synced repo file). `--dry-run` previewed all 8 steps correctly. Real run: ``` ./felhom-host-install.sh --customer-id drill-1 --vmid 120 --cores 2 --memory 4096 \ --force-gitea-golden --acl-storages "local local-lvm" --passphrase-file /root/.drill-pass ``` - pre-flight: PVE 9.2.2, node auto (single), **`local-lvm ~75 GiB < 120 GiB` warn fired as designed**, passphrase validated, grows auto-computed for the small box (docker +34G, sys +12G). - token + pool + scoped 3-role ACL applied (Base@/, Guest@/pool/felhom, Store@[local local-lvm]). - host-enroll: **201 MINTED** (`drill-1-c867f3`). - agent 0.62.0 fetched + **sha256-verified** against the hub manifest; sudo installed; non-root user/sudoers/unit installed; selftest passed; service active. - golden 0.85.1 fetched + sha256-verified from Gitea (`--force-gitea-golden` — the real fresh-box path). - provision: guest 120 up in 1m26s, cores=2/memory=4096 applied, 32/50/20G volumes, mp9 bootstrap mount, onboot=1, pool member. State file recorded `provisioned_vmid=120` (**confirms** the v0.62.0-era open question: a fresh install DOES record it). ## 5. D3/D4 — verify + app deploy - Agent: `--selftest` all-OK **including `pool read pool "felhom", 1 member`** (the v0.62.0 `pve:pool-read` capability); service active as `felhom-agent`. - Hub: host row reporting (agent 0.62.0), guest 120 `running`; events pipeline live from the fresh box (`controller_started`, `app_deployed`, `controller_updated`, 2× `host_leaf_changed` warnings from the reinstalls — correct behavior). - **Finding R6 (the big one):** after provision the controller was NOT running — `felhom-controller-bootstrap.service` was skipped (`ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json` unmet at boot): the agent back-half **hot-plugs** the bootstrap mount into the already-running guest and its own output says *"next: reboot the guest"*, but the v1.9.0 script never rebooted. Fast hardware (demo) usually won this race; the slow clean-room box lost it deterministically. **Fixed in v1.9.1** (post-provision `pct reboot` + bounded verify wait); re-proven in D6b (§7). - Dashboard: the controller sits behind traefik — correct probe is `curl -k -H "Host: felhom." https://127.0.0.1/` → **200** (Hungarian dashboard). The guide's draft `:8080` probe was wrong and was corrected. - **D4 app deploy:** `bentopdf` deployed through the exact UI pipeline (page fetch → the same `POST /api/stacks/bentopdf/deploy` + empty-CSRF headers the dashboard JS sends — endpoint-proxy method per the live-validation rule; no browser bridge was available in this session). Container healthy; app answers **200** at `pdf.drill.demo-felhom.eu` via traefik. Public Cloudflare leg not drill-testable (fake tunnel token) — that leg is live-proven daily on the demo box. - **OQ-2 evidence:** fresh box landed controller **0.85.1** (golden-baked) and did NOT self-update (floor 0.87.0 ignored — 0.85.1 predates the v0.86.0 floor code; its one report ACK carried no consequence). Manual update (`/api/selfupdate/check` + `/update`, the settings-page button flow) first **failed with the known fresh-install local-API 401** (`POST /controller/swap: HTTP 401`, stale token map); after `systemctl restart felhom-agent` the same trigger pulled 0.98.3 and the agent swap landed it: **0.85.1 → 0.98.3, healthy**. Both the failure and the recovery are now guide steps (D.1b). ## 6. D5 — uninstall + residue diff (v1.9.0 → the gap list) `--uninstall --vmid 120 --remove-golden` ran with the typed-vmid confirmation; guest/agent/pveum/ pool/state/golden all removed. Full re-capture + diff vs D0 classified every delta: | Delta (v1.9.0 uninstall) | Class | |---|---| | `/etc/felhom-agent/agent.json` (+dir) — **holds the live hub api_key** | **RESIDUE R1** → fixed v1.9.1 | | `felhom-shared-parent.service` (+wants links) + `/usr/local/sbin/felhom-shared-parent.sh` + `/mnt/felhom-drives` self-bind still mounted | **RESIDUE R2** → fixed v1.9.1 | | `/usr/local/sbin/felhom-mkfs-guarded` | **RESIDUE R3** (suspected by the spec — confirmed) → fixed v1.9.1 | | `/var/lib/vz/snippets/felhom-guest-hook.sh` | **RESIDUE R4** (would remain where the snippet install succeeded; on this box it never installed — see B2) → fixed v1.9.1 | | `/etc/dnsmasq.d/felhom-resolver-base.conf` | **RESIDUE R5** → fixed v1.9.1 | | `sudo` + `dnsmasq` packages (+ their files/units) | expected remnant (documented in the guide; dnsmasq newly noted) | | hub host record + report/guest/event history | expected remnant (no hub delete path — B4) | | LVM archive/backup metadata, dpkg lists, PVE rrd + `/var/lib/lxc/120` scratch, `.pve-staged-mounts` | system noise (PVE/dpkg-owned) | ## 7. D6 / D6b — reinstall proofs - **D6 (v1.9.0, post-uninstall box):** full reinstall succeeded — host-enroll **200 REUSED** (mint-once-reuse proven), fresh leaf minted (box had no guests → the populated-host guard correctly stayed silent), golden re-fetched, guest provisioned. Uninstall leaves a reinstallable box. (Controller again required the reboot — the R6 evidence repeated deterministically, incl. the agent's own "next: reboot the guest" line and the guesthook warn, see B2.) - **v1.9.1 re-verification (all fixes):** - `bash -n` + `-h` (header now v1.9.1) + full `--dry-run` previews of the new removal steps. - **Uninstall re-drill:** v1.9.1 `--uninstall --remove-golden` → re-diff vs D0: **zero `felhom`-named entries** in the file diff, zero pveum objects, no mounts, both dirs gone. - **D6b install re-drill:** full v1.9.1 install on the clean box → the script rebooted the guest itself → verify reported `controller: Up … (after ~0s)`, cloudflared up — **no manual intervention**. R6 closed at the installer. ## 8. OQ answers + go/no-go for Peti (sajatfelhom.hu) **OQ-1 (offsite backup): recommendation stands — go live with local backups only** (`--acl-storages "local local-lvm"`, exactly the drilled canonical command). The PBS box is LAN-only until Headscale/WireGuard lands; the ACL/`--rescope-acl` path retrofits the offsite grant later without reinstalling. The guide documents this as the current standard. (Operator holds the pen — if offsite-first is wanted instead, install day waits on the VPN spike.) **OQ-2 (fresh-install controller version): answered with evidence** — lands 0.85.1, never self-updates (pre-floor code). Raising the floor does NOT fix a fresh box. Install-day step D.1b (restart agent → trigger update → verify 0.98.3) is mandatory and drill-proven. Structural fix = rebuild + re-vouch the golden with a ≥ 0.86.0 controller (operator follow-up, not a Peti blocker). **OQ-3 (pre-install facts):** the guide's Part B checklist is the collection instrument (PVE 9.x, single-node vs cluster → `--node`, ≥120 GiB local-lvm, free vmid from `pct list`+`qm list`, host cores/RAM for the cap, root SSH, outbound reach incl. the three exact URLs+expected codes, python3). **Go/no-go list:** | # | Item | State | |---|---|---| | 1 | Install path proven verbatim on a virgin box (twice + reinstall) | ✅ GO | | 2 | Uninstall residue-clean (v1.9.1) — safe on a shared box | ✅ GO (deploy/serve v1.9.1 BEFORE install day; drilled from the served URL at v1.9.0 + locally at v1.9.1) | | 3 | D.1b one-time controller update executed on install day | ⚠️ REQUIRED STEP (guide covers it; 2 minutes) | | 4 | Day-0 manifest vouches agent **0.62.0** + golden **0.85.1** | ✅ current; **but hub Day-0 manifest bump to agent 0.62.0 was already flagged as operator follow-up — CONFIRMED DONE (drill fetched 0.62.0)** | | 5 | Real CF tunnel token created + entered in Part A (drill used a fake) | ⚠️ operator does A.1 for sajatfelhom.hu; tunnel leg is demo-proven, not drill-proven | | 6 | `git.username`/`git.token` set in the customer config | ⚠️ REQUIRED (installer dies without it) | | 7 | Multi-node? If Peti's box is a cluster, add `--node` | per Part B answer | | 8 | Known non-blockers on the box: B2 (no pre-start hook on fresh boxes), B3 (401-until-agent-restart) | **both FIXED in agent v0.63.0** (2026-07-03, Gitea-published); effective for fresh installs once the Day-0 manifest vouches ≥ 0.63.0 — until then the D.1b restart step covers B3 | **Verdict: GO** — with items 3/5/6 executed on/before install day. ## 9. Findings ledger (blockers + follow-ups, per the §9 gap policy) | ID | Where | Severity | What | Status | |---|---|---|---|---| | R1–R5 | installer | — | uninstall residue (config w/ secrets, shared-parent unit+script+mount, mkfs wrapper, hook snippet, dnsmasq snippets) | **FIXED v1.9.1**, residue-diff-proven | | R6 | installer | — | controller-bootstrap unit skipped: no post-provision reboot (agent hot-plugs the mount; unit condition is boot-time) | **FIXED v1.9.1** (reboot + bounded verify wait), D6b-proven. NOTE (2026-07-03, golden-098 task): goldens ≥ 0.98.3 bake a `.path` unit that makes the reboot redundant — the reboot is RETAINED as a belt; its removal is a recorded candidate cleanup, not done | | B1 | agent/golden | LOW | structural fix for R6 belongs in the product too (path unit in the golden, or agent-side restart after back-half) — installer reboot is a correct but external crutch | **FIXED** (build-golden.sh v2.0.0, golden 0.98.3): baked `felhom-controller-bootstrap.path` starts the service on the bootstrap-mount hot-plug — isolated + full-install proven, `DRILL-golden-098-2026-07-03.md` §4/§5 | | B2 | agent | LOW | `guesthook.InstallSnippet` fails on a fresh box: `/var/lib/vz/snippets` doesn't exist and `install` won't create it → no pre-start self-heal hook, warn-only (`install: cannot create regular file … No such file or directory`, agent journal 14:58:40) | **FIXED agent v0.63.0** (fenced `mkdir -p` precedes the install + the one sudoers grant; red-proofed — felhom-agent REPORT 2026-07-03) | | B3 | agent | MED | fresh-install local-API **401 until `systemctl restart felhom-agent`** — root cause: the daemon's TokenStore index is built once at open, while provisioning is a SEPARATE one-shot process minting into the shared file; NEW consequence found: it blocks the controller self-update agent-swap | **FIXED agent v0.63.0** (`Lookup` reload-on-miss, red-proofed + run on felhom-pve — felhom-agent REPORT 2026-07-03). Guide D.1b restart-step narrowed to "agent < v0.63.0" — applies until the Day-0 manifest vouches ≥ 0.63.0 | | B4 | hub | LOW | **no host-delete path** (UI or API) — drill host rows removed via direct SQL (spike precedent); customer-delete exists but leaves hosts/reports/guests/events orphaned | RECORDED (candidate small hub follow-up) | | B5 | golden/fleet | MED | golden bakes a pre-floor controller (0.85.1) → every fresh install needs D.1b manually | **FIXED** (build-golden.sh v2.0.0: mandatory controller tag; golden **0.98.3** baked, published, vouched): fresh install lands 0.98.3 on first boot, self-update reports up-to-date, guide D.1b retired — `DRILL-golden-098-2026-07-03.md` §5/§6 | ## 10. Cleanup confirmation - **Hub:** all `drill-1` rows deleted (customer_configs / hosts / host_reports / reports / guests / events / notification_log / customer_notifications → remaining **0**); `demo-felhom` confirmed intact. (The web customer-delete would have left the host/report rows — B4 — hence SQL, spike precedent.) - **Drill VM:** powered off; disk reverted to the `virgin` snapshot and kept (**virgin-snapshotted-and-stopped**, not removed — reusable for future drills). Footprint ~6.7 GiB at `~/drill/` on the build server (ISO + qcow2 + evidence logs + keys, keys 0600); the baked auto-install ISO variant was deleted. - Evidence archived off-box at `~/drill/{D0,D5,D5b}/` + `D2-install.log`, `D5-uninstall.log`, `D5b-uninstall.log`, `D6-install.log`, `D6b-install.log` on the build server. - QEMU/`proxmox-auto-install-assistant` packages remain installed on the build server (useful for re-drills; noted).