DRILL-day0-vm 2026-07-12: full report + runbook corrections (F-4/F-6 headline findings)

Audit doc for the Day-0 VM drill: appliance install, floor-at-first-report,
escrow ceremony + auto-confirm FIRST LIVE FIRING (~7.5 min, zero clicks),
offsite backup + restore round-trip, PBS-DR/WG fork (Peti-sequence rehearsal).
Corrects day0-install.md A.2 (git creds optional since v1.11.2, anonymous
fetch is the normal shape) and RUNBOOK-escrow-ceremony.md (identity-only mode
does NOT exist — F-6; age prereq — F-10). REPORT.md overwritten; CONTEXT.md
one-liner added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NptTCFtu7dz2Ru89qHRagN
This commit is contained in:
2026-07-12 17:16:25 +02:00
parent 0ed87f5dc4
commit 3e949bc513
5 changed files with 207 additions and 19 deletions
+11
View File
@@ -3,6 +3,17 @@
> Created with the REUSE.md rollout (2026-07-03). Authoritative history: `hub/CHANGELOG.md` (hub),
> `website/CHANGELOG.md`, `scripts/CHANGELOG.md`; end-of-task detail in `REPORT.md`.
- **2026-07-12 — Day-0 VM DRILL COMPLETE (auto-confirm FIRST LIVE FIRING): full arc proven on a
fresh nested-PVE box** — appliance Day-0 → floor-at-first-report → ceremony → **auto-confirm
pending→escrowed in ~7.5 min, zero clicks** → offsite backup + restore round-trip. Two HIGH gaps:
**F-4 no operator password-set path exists (G10 unclosable, dashboards born OPEN — blocks tester
gate)** and **F-6 identity-only ceremony never implemented (no-PBS appliance can't escrow — drill
forked to PBS DR tier = full Peti-sequence rehearsal, all green)**. Installer fresh-box gaps:
felhom-pbs-apply not shipped (F-7), `age` missing (F-10), root-owned guests/ parents (F-3 —
check demo for the latent copy), silent root@pam rotation UX (F-8). Runbook fixes committed
(day0 A.2 anonymous-fetch; escrow-ceremony identity-only claim CORRECTED + age prereq). Report:
`documentation/audits/DRILL-day0-vm-2026-07-12.md`. Drill VM qm 300 kept (3 snapshots) for
re-drills; teardown list in report §9.
- **2026-07-12 — CAMPAIGN-3 Task A SHIPPED: agent v0.85.0 boot/recovery plane + appliance self-heal.**
Fixes F12 (CRITICAL boot ordering cycle — templates drop network-online, `MigrateNetworkUnits`
repairs installed units), F11/F10/F9 (reassert: fstype-driven classify, reset-failed+rearm, per-unit
+7 -12
View File
@@ -2,19 +2,14 @@
> **Overwrite** this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in [hub/CHANGELOG.md](hub/CHANGELOG.md); the scripts history lives in [scripts/CHANGELOG.md](scripts/CHANGELOG.md).
## CAMPAIGN-3 — unattended "no mercy" night run (NAS · deploy · backup · chaos · observability) — 2026-07-11/12
## DRILL — Day-0 on the Demo-VM (nested PVE) + escrow/auto-confirm first firing — 2026-07-12
**Full report: [`documentation/audits/CAMPAIGN-3-2026-07-11.md`](documentation/audits/CAMPAIGN-3-2026-07-11.md).** Run 22:09 → 04:27 CEST against the demo box (controller 0.117.0 / agent 0.84.0), operator-unattended, real surfaces only, no code changes. Ledger (final incl. morning RCA/recovery): **31 PASS · 17 FAIL · 12 FINDING · 1 DISCREPANCY** (70+ scenario entries, evidence at `180:~/campaign3/`).
**Full report: [`documentation/audits/DRILL-day0-vm-2026-07-12.md`](documentation/audits/DRILL-day0-vm-2026-07-12.md).** Supervised drill (Viktor + CC), ~15:1517:09 CEST, on a throwaway nested PVE 9.2.2 VM (qm 300 on felhom-pve) against hub customer `demo-vm-felhom` / `enkisfelhom.hu`. No code changes; two runbook docs corrected (day0-install.md A.2 anonymous-fetch wording; RUNBOOK-escrow-ceremony.md's false "identity-only ≥0.80.0" claim + `age` prereq).
### Headlines
- **CRITICAL F12 — the overnight host loss, RCA closed (hardware exonerated):** the agent's network-storage **automount** template (`After=`+`Wants=network-online.target`, implicitly `Before=local-fs.target`) creates a boot **ordering cycle**; systemd breaks it by deleting an arbitrary job. Boot at 23:31 sacrificed `networking.service` → host up **7 hours with no network**; the 06:45 power-cycle boot hit the same cycle and sacrificed the **automount** instead (NAS dead, healed manually). **Every boot of a host with an enrolled share is a coin flip until the template drops the network-online ordering** (`_netdev` on the `.mount` suffices). 4e caught exactly what it was designed to catch.
- **CRITICAL F10 + HIGH F11/F9 — the reboot/recovery plane around NAS automounts is broken:** `mount-start-limit-hit` is never re-armed by any heal path (once even blocked guest start → guest DOWN); guest reboot with an idle share leaves the autofs trigger unpropagated into the container — the post-start reassert **logs its own WARNING and then skips** the automount restart that provably heals ("skip-active" branch). Every such reboot = 4 NAS apps dead-at-boot. Reproduced on 3 of 3 guest reboots.
- **HIGH F7 — backup dumps are written in place (no tmp+rename):** a mid-backup NAS cut left a 0-byte tar *replacing* the last good 247M dump; in that window restore = empty volume. Next run self-heals; run-level `success:false` is the only signal.
- **The data plane held:** all 5 refusal categories ×2 correct + fast (25 s, retry=0), verify/rollback/single-flight/orphan flows clean, deploy-view truth holds, restore round-trips **byte-identical**, EIO same-second under outage, organic stub → badge + deploy-409 live-validated, hardlinks work on NFSv4.1.
- **Fix-6 answered with numbers:** ring cap horizon = ~55 min idle but **~6.5 min under load**; every restart/reboot wipes both rings — persistence, not just size, is the gap.
- **Policy discovery (docs):** tier-1 = volumes+config only (NAS media userdata excluded by design); NAS apps' tier-1 lands *on the NAS*, tier-2 is what gets it off; volume-only apps back up to sys_drive with blank drive label and **no tier-2 copy**.
### Box state / cleanup (final, 06:53)
DooPlex NAS restored **md5-identical to baseline** (exports + smb.conf; campaign user/share/dirs/creds removed); services never touched (exportfs-only rail held). Host recovered post-power-cycle; guest 9201 in **defined state**: 6 wave apps deployed + healthy with data, `privatebin` stop+removed via the real flow, verification backup `success:true`, nas-media `ok`, stub 0. Hub untouched throughout. ⚠ Next host reboot re-rolls the F12 dice until the agent template is fixed (interim: systemd drop-in on the automount units).
- **The whole arc ran to completion on a fresh box:** appliance-mode Day-0 (agent 0.85.0 + golden 0.120.0, both sha-verified, "Day-0 provision SUCCESS"), guest at floor 0.120.0 at FIRST report with zero manual steps, ActualBudget deployed through the tunnel, escrow ceremony sealed (R with Viktor only), **AUTO-CONFIRM FIRST LIVE FIRING: pending → escrowed in ~7.5 min with zero clicks**, offsite backup (1 snapshot, 14 s) **and the verification-restore round-trip proven** (48K payload decrypted back from the Storage Box).
- **F-4 (HIGH): G10 is unclosable** — no operator-set dashboard password path exists in shipped code (hub has no per-customer `password_hash` UI/API; controller's open-state page defers to the operator; Day-0 preseeded setup skips the wizard's password form). Every fresh box's dashboard stays OPEN on the internet. Blocks the tester onboarding gate.
- **F-6 (HIGH): identity-only escrow ceremony was never implemented** — on a no-PBS box (the documented appliance standard) the offsite escrow chain can never complete. Mid-drill fork (Viktor): attached the **PBS DR tier (ep0)** instead, which live-rehearsed Peti's exact pending sequence — wrapper+`age` prep → `wg_tunnel.enabled` → hands-free WG peer registration → hub PBS-DR enable (WG-first dependency guard works) → apply-bridge fresh path (consume-once → K born → escrow seed) → ceremony → auto-confirm.
- Fresh-box installer gaps found live: `felhom-pbs-apply` binary not shipped (F-7), `age` not installed (F-10), root-owned `guests/` parents break the non-root agent's lanresolver (F-3, live-fixed; **demo host likely has the same latent state**), silent root@pam rotation surprises the operator (F-8), version strings disagree ×3 (F-1).
- Snapshots on qm 300 for re-drills: `pre-day0-clean` / `post-install` / `post-drill`. Demo 9201, Peti's hub entry, demo offbox untouched.
@@ -0,0 +1,174 @@
# DRILL — Day-0 on the Demo-VM (nested PVE on felhom-pve), 2026-07-12
Executed per `RUNBOOK` (Day-0 drill on the Demo-VM). Actors: Viktor (console install, dry-run
go/no-go, fork decisions, the ceremony/R-moment, dialog clicks) + CC (everything scriptable over
SSH + the hub/dashboard browser tracks). **Outcome: the full arc ran to completion — Day-0
appliance install, first live escrow ceremony + AUTO-CONFIRM FIRST FIRING, offsite backup +
restore round-trip — at the cost of one mid-drill fork (PBS DR tier attached) forced by the
drill's headline findings (F-4, F-6).**
## 1. Baselines recorded (§1 of the runbook, verified LIVE before Phase 1)
| Item | Recorded |
|---|---|
| Day-0 artifact manifest | Agent **0.85.0** sha256 `31babb2c4f8fa5a1f961428519a0dfbb3a19b6e536fba01c2664afca339da93d`; Golden **0.120.0** sha256 `f7d7d02c76b49891719df9cf68624c9ae193e438e2a3a63b948a2a9d12887596`; MinAgent **0.81.0** — matches the 0.85/0.120 publish train, no bump needed |
| Global controller floor | Effective **v0.120.0**, source DB (hub_settings); env fallback also v0.120.0 |
| Customer `demo-vm-felhom` | exists, domain `enkisfelhom.hu`, CF tunnel + API tokens saved (API token perms incl. Zone WAF:Edit); git credentials EMPTY (correct per G3; hub edit-form already labels Git Sync "Opcionális" — only day0-install.md A.2 was stale, fixed in this commit); DEBUG MÓD was ON |
| Wildcard DNS | `*.enkisfelhom.hu` → CF proxied edge (104.21.3.175 / 172.67.130.252 + AAAA), verified via 1.1.1.1 on a real and a random subdomain |
| felhom-pve capacity | local-lvm thin pool 348.82g @ 7.47%; RAM available 12.9 GiB; nested virt Y |
| Installer | felhom-host-install.sh **v1.14.0** per `-h` (but see F-1) |
| Golden-baked controller | 0.120.0 (= floor; see 4.3 note) |
## 2. Per-phase gates
| Phase | Gate | Result | Evidence (abridged) |
|---|---|---|---|
| 0 — build VM | P0 | **PASS** | qm 300 `drill-day0` (8G/4c/host/250G thin); PVE **9.2.2** ISO 9.2-1; key SSH proven; snapshot `pre-day0-clean` 15:12:41. ISO detached + boot→scsi0 BEFORE the snapshot (clean config) |
| 1 — hub verify | — | **PASS** | §1 table above; nothing created/changed in the hub |
| 2 — box prereqs | P2 | **PASS** | single node; nested local-lvm **149.88 GiB** free (≥120 ✓ — note: a 250G VM disk yields ~150G nested pool after the installer's root/swap split); vmid 9201 absent; hub 302 / gitea 200 / felhom.eu 200 |
| 3 — install | P3 | **PASS** | dry-run reviewed by Viktor → GO → real run 15:2515:28: **"Day-0 provision SUCCESS — vmid=9201 host_id=demo-vm-felhom-2482b0"**; both artifacts sha-verified vs the hub manifest; both expected ANONYMOUS-fetch warns; 4b break-glass vaulted; passphrase file 0600, shredded after (never in CC's transcript — Viktor wrote it himself) |
| 4 — post-install | D.1D.4 | **PASS** (1 finding) | see §3 |
| 5 — hub tracks | G9/G10 | **PREMISE COLLAPSED** — F-4/F-5; geo exercised via the shipped path | see §4 |
| 6 — escrow/auto-confirm/offsite | fork-4 | **PASS after fork** (PBS DR attached) — F-6/F-7/F-10/F-11 | see §5 |
## 3. Phase 4 detail (post-install verification)
- **D.1:** agent active as non-root `felhom-agent`; full read-only selftest ALL-OK incl. pool read
(pool `felhom`, member 9201); guest running, onboot=1, mounts mp0 docker 50G / mp1 sys 20G /
mp8 felhom-drives / rootfs 32G; in-guest containers controller 0.120.0 (healthy) + filebrowser +
cloudflared + traefik; dashboard 200 in-guest (traefik https + Host header; :80 → 301).
- **4.2 appliance gating (first live appliance box):** agent.json `deployment_mode="appliance"`;
journal `selfheal: node watchdog starting mode=appliance interval_s=60` + storage watchdog
armed; felhom-mgmt-watchdog.timer enabled+active. NOT byo-defaulted.
- **4.3 floor self-update (tester-recruitment gate):** guest landed **0.120.0 = floor at its FIRST
report, zero manual steps, elapsed ≈ 0**. LIMITATION: baked == floor on this train, so the
floor-driven *update* path was not stressed — re-prove on the next train where golden < floor.
- **4.4 hub:** customer PENDING→ok ~2 min after provision; host ONLINE agent 0.85.0; guest 9201
appears at the next agent report (900 s cadence — the empty Guests panel in between is timing,
not a bug); only-degraded capabilities = the 3 `pbsdr-*` ("binary not found") — expected no-PBS
shape (and see F-7).
- **4.5 customer-visible via the REAL edge** (method: `curl --resolve` on both CF anycast IPs —
the split-horizon-proof variant): felhom.enkisfelhom.hu → 200 "Vezérlőpult", **OPEN, no auth**
(G10 "before", ~15:30).
- **4.6 first-app smoke:** ActualBudget via the open dashboard → "Telepítés sikeres"; real-edge
200 + `<title>Actual</title>` on budget.enkisfelhom.hu. (The old POST-via-public-URL no-op
gotcha did NOT reproduce on 0.120.0 for normal forms.)
- Snapshot `post-install` 15:35:37 (live, no fs-freeze — no qemu-guest-agent in the drill VM).
## 4. Phase 5 detail — the G9/G10 premise vs shipped code
- **5.1 operator password-set: IMPOSSIBLE (F-4).** No hub UI/API sets per-customer
`web.password_hash`; the controller's security page in the open state says "Kérd az
üzemeltetőt"; `settingsPasswordHandler` requires a current-password match (no initial-set);
the hub-preseeded Day-0 setup path skips the only wizard form having a password field.
- **5.2 geo-restriction:** apply lives in the CUSTOMER dashboard (controller
`settings_security.html` + `api/geo.go`), not the hub (hub has only `handleGeoDisable`) — F-5.
Exercised via the shipped path: HU-only enabled → controller created WAF rule
**"[felhom-geo] Global"** (`(not ip.src.country in {"HU"})`, action block) on the zone, sync
reported 1 active rule; HU access still 200 through the real edge. Non-HU block not testable
from an HU vantage.
- **5.3 G10 statement:** BEFORE — dashboard OPEN (proven in-guest + real-edge, ~15:30). AFTER —
**STILL OPEN; closure impossible until F-4 ships.** The tester-agreement onboarding gate that
points at this line CANNOT currently be satisfied.
## 5. Phase 6 detail — offsite → ceremony → auto-confirm → tier proof
Pre-ceremony state (all verified): offbox pre-provisioned at customer-create
(u629488-sub3@…your-storagebox.de:23, /home/felhom-repo, pinned host fingerprint, 0/50 GB);
controller `EscrowState=pending` — the /backups banner "a mentés addig nem fut" (the F6
no-single-copy guard, LIVE); staged `escrow-stage/restic_repo_password` on the agent (0600,
staged at provision — fork-4 stage-FIRST ordering held); ActualBudget toggled for NAS.
**The ceremony blocked → the drill's second headline (F-6):** `--selftest=escrow-create --upload`
refuses without a PBS storage/key; **identity-only mode does not exist in any shipped agent**
(the ≥0.80.0 claim in RUNBOOK-escrow-ceremony.md was false — fixed in this commit). On a no-PBS
appliance box the offsite arc can NEVER complete. **Fork decision (Viktor): attach the PBS DR
tier (ep0)** — which converted the drill into a full rehearsal of Peti's pending sequence:
1. Ship `configs/felhom-pbs-apply` → /usr/local/sbin (the binary is missing from host-install —
F-7; the FELHOM_PBSDR sudoers alias DOES ship). Agent restart → zero capability-DEGRADED.
2. Hub PBS-DR enable failed correctly: "host has not reported a WG key yet — the tunnel peer must
exist before the PBS DR tier" — the dependency is hub-enforced (good), and the drill runbook's
"no WG" scope was inconsistent with any PBS path.
3. WG: `wg_tunnel.enabled=true` in agent.json (installer never sets it — decide the appliance
default together with the F-6 spec) → keygen → **hands-free hub registration** (10.77.0.3/32,
gen 1, no operator vouch) → conf applied → handshake + ping 10.77.0.1 (33 ms).
4. Hub PBS-DR enable → "Configuration updated" → descriptor gen 2; apply-bridge first 403'd
(`Datastore.Allocate` on /storage/felhom-pbs) — consequence of the drill's narrowed
`--acl-storages "local local-lvm"`; the installer DEFAULT includes felhom-pbs exactly for
this. Retrofit dual-grant → next tick: **token consumed (single-use) → entry + K created →
`escrow.pbs_storage_id` seeded → state=applied** (16:17). pvesm ACTIVE; felhom-pbs.{enc,pw}
present.
5. **Ceremony** (Viktor, R on paper, nothing in CC's transcript): attempt 1 failed — `age`
missing (F-10, `apt-get install -y age` → 1.2.1); attempt 2 SUCCESS ~16:35; the bundle
auto-captured `+wg_private_key +restic_repo_password`; ceremony wiped the staged secret; hub
host page flipped to **DR RECIPE: present / KEY ESCROW: present**.
6. **AUTO-CONFIRM FIRST LIVE FIRING:** hands off, "Letét megerősítése" untouched → 16:42:36
controller log: *"hub-verified: the escrow covers the current repo password (hash
99c16e8d84e7…) — EscrowState auto-confirmed escrowed; offsite runs enabled."*
**Elapsed ≈ 7.5 min, zero clicks.** /backups gate banner gone.
7. **Tier proof:** "NAS-mentés most" → repo initialized on the Storage Box → 1 snapshot, 14 s,
"✓ Rendben". Restore-to-verify (native confirm() blocked automation twice and Viktor missed
the popup — F-11; also live operator confusion: he first launched the FULL local tier-1
restore, which itself completed healthy in 8.7 s) → offbox verification restore →
*"restored actualbudget → …/offbox-restore/actualbudget"*, 48K payload (app.yaml, .felhom.yml,
docker-compose.yml, manifest.json) decrypted from the box — **the offsite round-trip is
proven**. Bonus: the full local tier-1 restore was ALSO proven the same afternoon.
## 6. Headline metrics
| Metric | Value |
|---|---|
| Wall-clock Phase 2 → 6 complete | ~15:15 → ~17:09 (**~1 h 55 m**, including the mid-drill fork, 3 fix-and-continue stops, and two supervised STOPs) |
| Install run itself (Phase 3) | ~3 min |
| Floor self-update (4.3) | at-floor at FIRST report, 0 manual steps (baked == floor; update path not stressed) |
| **Auto-confirm (6.4)** | **~7.5 min ceremony→escrowed, zero manual clicks (FIRST LIVE FIRING)** |
| Offsite backup / restore | 14 s backup (1 snapshot) / verification restore round-trip proven |
## 7. Findings
| # | Sev | Phase | Finding | Disposition |
|---|---|---|---|---|
| F-1 | LOW | 3 | Version-string mismatches: `-h` v1.14.0 vs run banner v1.13.0 vs hub Setup-tab copy "1.12.0" | fix strings (installer + hub template) |
| F-2 | COSMETIC | 3 | dry-run prints `curl -u <git>` on the anonymous-fetch branch | fix placeholder |
| F-3 | MEDIUM | 4 | Root-run provision leaves `/var/lib/felhom-agent/guests{,/9201}` root:root 0700 inside the agent-owned state dir → non-root agent lanresolver "permission denied". LIVE-FIXED (chown the two parent dirs; the guest-root-owned bootstrap subtree untouched) | agent/installer: create parents agent-owned at provision; **check demo/felhom-pve for the same latent state** |
| **F-4** | **HIGH** | 5 | **No operator-set dashboard password path exists anywhere** (hub has no UI/API for per-customer `password_hash`; controller open-state page defers to the operator; Day-0 preseeded path skips the wizard's password form) → **G10 unclosable; every fresh box's dashboard stays OPEN on the internet** | hub feature task (operator set → config-delivered hash → controller re-pull); blocks tester onboarding gate |
| F-5 | MEDIUM | 5 | Geo-restriction APPLY is customer-dashboard-side; hub only disables. Runbook premise stale; also: the open dashboard (F-4) exposes the geo toggle unauthenticated | doc/design decision; folded into F-4's arc |
| **F-6** | **HIGH** | 6 | **Identity-only escrow ceremony was never implemented** (`escrow-create` hard-requires PBS storage + key; the ≥0.80.0 runbook claim was false) → on a no-PBS box (the documented appliance standard!) the offsite escrow chain can never complete | agent feature task (identity-only mode); ceremony runbook corrected in this commit |
| F-7 | MEDIUM | 6 | host-install ships the FELHOM_PBSDR sudoers alias but NOT the `felhom-pbs-apply` binary → pbsdr capabilities born DEGRADED on every fresh box | installer: ship the wrapper (like mkfs/selfupdate wrappers) |
| F-8 | LOW/UX | 3 | Step 4b rotates root@pam + vaults silently — operator surprised by 401 at the PVE GUI (live: Viktor) | installer: print "root@pam rotated + vaulted — retrieve at hub → host page" |
| F-9 | NOTE | 6 | Installer never sets `wg_tunnel.enabled`; WG registration itself is hands-free once enabled | decide appliance default alongside the F-6 spec |
| F-10 | MEDIUM | 6 | `age` (ceremony identity-wrap dependency) not installed by host-install — fresh-box ceremony dies; demo host masked it (spike-era install) | installer: add `age` to package set; runbook prereq added in this commit |
| F-11 | LOW | 6 | Native `confirm()` on the offbox restore-verify form: blocks browser automation, easy to miss (operator missed it twice live; meanwhile launched the full tier-1 restore from the adjacent form — the two restore controls invite confusion) | convert to the design-system inline confirm pattern; consider renaming |
Observations (not F-numbered): hub "REGISTRY LATEST v0.120.0 — up to date" vs the dashboard's own
"Új controller verzió elérhető: 0.121.0" banner disagree on "latest"; the hub dashboard row for
Peti shows "OK · minutes ago" (controller-derived, via his proxmox2 migration) while his agent
HOST is DOWN 23h — the roll-up masks a dead host; retrofit-ACL note: adding the PBS tier to a box
installed with narrowed `--acl-storages` needs the /storage/<id> dual-grant (documented default
avoids it).
## 8. What this proved for Peti — and what it did not
**Proven live on a fresh box (his exact pending sequence):** wrapper+age prep → WG enable →
hands-free peer registration → hub PBS-DR enable (dependency guard works) → apply-bridge fresh
path (consume → K → grant → escrow seed) → ceremony (with K) → **auto-confirm** → gated offsite
run → restore round-trip. Prep list for his box: `felhom-pbs-apply` + `age` + `wg_tunnel.enabled`
+ (if his ACL set was narrowed) the /storage dual-grant.
**Deliberately NOT covered:** WG OOB operator peer (separate arc); S5 DR restore from total loss
(identity-consume — queued as its own drill); byo-mode re-run from `pre-day0-clean` (queued);
non-HU geo-block verification (needs a non-HU vantage).
## 9. Snapshot inventory (qm 300 on felhom-pve, at drill end)
| Snapshot | When | State |
|---|---|---|
| `pre-day0-clean` | 15:12 | PVE 9.2.2 + keys, nothing Felhom — the universal re-drill zero |
| `post-install` | 15:35 | Day-0 SUCCESS + ActualBudget + guests-dir chown |
| `post-drill` | 17:11 | full drill end state (WG + PBS-DR + escrowed + geo HU + dashboard OPEN per F-4) |
Blast radius honored: guest demo 9201 on felhom-pve, Peti's hub entry, and the demo offbox were
never touched. Drill leftovers to tidy at teardown (deliberately kept for re-drills now): hub
customer `demo-vm-felhom` (its api_key + CF tokens surfaced in the operator UI during the drill —
delete/rotate at teardown), DEBUG MÓD on, geo HU rule on the drill zone, VM 300 running.
@@ -18,8 +18,15 @@
## Prerequisites (check BEFORE scheduling with the customer)
- **Agent version:** ≥ v0.79.0 (the ceremony records `restic_pw_sha256` — older agents produce a blob
auto-confirm can never match). **No-PBS hosts** (BYO without the PBS tier): ≥ **v0.80.0**
(identity-only mode; below that the ceremony hard-requires the PBS key and refuses).
auto-confirm can never match).
- **⚠ No-PBS hosts: the ceremony CANNOT run.** The previously documented "identity-only mode
(≥ v0.80.0)" was NEVER implemented — v0.80.0's actual feature was seeding `escrow.pbs_storage_id`
on PBS hosts. `escrow-create` hard-requires a PBS storage id + its key file (drill-proven
2026-07-12, finding F-6 of DRILL-day0-vm-2026-07-12.md). Until identity-only ships, a box MUST
have the PBS DR tier (which itself requires the WG tunnel peer first) before any escrow/offsite
arc can complete.
- **Host packages:** `age` must be installed (identity wrap dependency; NOT installed by
host-install as of v1.14.0 — drill finding F-10). `apt-get install -y age`.
- **K gate (PBS hosts):** `escrow.pbs_storage_id` set and the key file present
(`cfg.Backup.PBSEncKeyPath(<id>)`).
- **Staged secret (offsite):** offsite enabled → `EscrowState="pending"` on the controller and the staged
@@ -31,8 +38,9 @@
```bash
felhom-agent --selftest=escrow-create --upload
```
- `--storage <pbs-storage-id>` only if `escrow.pbs_storage_id` isn't configured. No-PBS hosts (agent
≥0.80.0): omit — identity-only engages automatically.
- `--storage <pbs-storage-id>` only if `escrow.pbs_storage_id` isn't configured (the PBS DR
apply-bridge seeds it automatically). No-PBS hosts: see the prerequisite warning above — the
ceremony refuses without a PBS key; identity-only mode does not exist yet.
- What it does, in order: generates a fresh **R** (EFF-wordlist passphrase; entropy printed) → seals K
(if present, via proxmox-backup-client re-key) and the IdentityBundle (age-under-R; the staged restic
password auto-injected; the live WG key auto-captured if present) → **self-verifies by recovering its
+3 -3
View File
@@ -63,8 +63,8 @@ Hub UI (`https://hub.felhom.eu`, operator password) → **Customers → New**:
| Email | customer's email | used for customer-tier notifications (Hungarian) |
| CF tunnel token | from A.1 | → `infrastructure.cf_tunnel_token` |
| CF API token | from A.1 (optional) | → `infrastructure.cf_api_token` (geo rules) |
| Git username | Gitea read account | → `git.username`**required for Day-0** |
| Git token | Gitea read token | → `git.token`**required for Day-0**: the install script and the controller fetch artifacts from Gitea with this credential; without it the install dies at step 5/8 |
| Git username | Gitea read account | → `git.username`**optional** (only for a private app catalog) |
| Git token | Gitea read token | → `git.token`**optional** since installer v1.11.2 (the G3 ruling): artifacts are world-readable and fetched ANONYMOUSLY, sha256-verified against the hub manifest. Empty is the normal customer shape; expect the installer's "fetching artifacts ANONYMOUSLY" warn line in steps 5/8 (drill-verified 2026-07-12 on the VM Day-0 drill). |
On save the hub generates two credentials:
@@ -405,7 +405,7 @@ recovery credential).
| byo dies: "acl storage(s) not found on this box" | `--acl-storages` names a storage the box lacks | pass the box's real storages (check `pvesm status`) |
| step 1 dies: "this is a N-node cluster" | multi-node cluster | re-run with `--node <name>` |
| step 5 dies: "hub artifact manifest has no agent version" | Day-0 manifest unset/incomplete | Part A.3 — set it in the operator UI |
| step 5 dies: "no git token in controller.yaml" | customer created without git credentials | Part A.2 — add `git.username`/`git.token`, regenerate config |
| step 5 warns: "no git credential — fetching artifacts ANONYMOUSLY" | customer has no git credentials — the NORMAL shape since installer v1.11.2 | nothing to do; sha256 verification is unchanged. (Pre-v1.11.2 installers die here instead — upgrade the script.) |
| step 1: passphrase REJECTED (401) | typo / wrong customer | re-check with the hub UI's printed curl command |
| step 8 fails: "CT <vmid> already exists" | vmid collision with a hub-invisible guest | pick from `pct list` + `qm list` (Part B); the agent destroys nothing on collision — re-run with a free vmid and `--resume` |
| controller container missing in-guest after provision (docker ps empty) | golden older than 0.98.3 (no baked path unit) + pre-v1.9.1 script: the controller-bootstrap unit's boot-time condition lost the race with the bootstrap-mount attach | `pct reboot <VMID>` — the unit runs on the next boot (v1.9.1 reboots itself; goldens ≥ 0.98.3 bake a path unit that starts the controller on the mount hot-plug, no reboot needed) |