Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
13 KiB
DRILL — Golden rebuild 0.98.3: current-controller baseline (B5) + hot-plug path unit (B1) (2026-07-03)
Class: implementation + clean-room drill. Companion deliverables: felhom-agent/configs/build-golden.sh
v2.0.0 (commit ceca355) and the golden archive 0.98.3 (published + operator-vouched).
Source findings: DRILL-day0-cleanroom-2026-07-03.md §9 B5 (golden bakes a pre-floor controller →
mandatory manual D.1b on every fresh install) and B1 (controller-bootstrap only fires at boot; the
hot-plugged config mount needed the installer's reboot crutch). Also closes the stale backlog note
documentation/backlog/FOLLOWUP-golden-default-controller-tag.md.
Verdict (short): both findings are FIXED in the product. A fresh Day-0 install now lands controller
0.98.3 on first boot and self-manages from there (no D.1b), and the baked
felhom-controller-bootstrap.path unit starts the controller the moment the agent hot-plugs the
bootstrap mount — no reboot needed (the installer's v1.9.1 reboot is retained as a redundant belt).
Everything was proven clean-room on the drill VM BEFORE publish; the vouch is the only production change.
No secrets appear in this document. The bake's registry credential was passed via env from a 0600 file
on the build server; the script logs out + removes /root/.docker/config.json in the build guest before
archiving (build-golden.sh — verified present in the bake transcript path).
1. Environment & baselines
| Item | Value |
|---|---|
| Drill VM | the DRILL-day0-cleanroom nested PVE 9.2.2 on the build server (192.168.0.180), restored from the virgin qcow2 snapshot before the bake and again before each install scenario |
| VM launch quirk fixed | slirp DHCP pool moved with dhcpstart=10.0.2.30 (+ explicit hostfwd=…-10.0.2.15:22) so nested guests can't grab the PVE host's static 10.0.2.15 — the prior drill's IP-collision quirk is fully avoided, no guest IP pinning needed |
| felhom-agent base | c9f963d → script commit ceca355 (build-golden.sh v2.0.0; no Go change, no agent version bump) |
| felhom-controller | v0.98.3 reconfirmed as the current released tag (CHANGELOG top entry; manifest fetch HTTP 200 with the read credential) — the tag baked |
| Hub manifest at start | agent 0.63.0 (the v0.63.0 REPORT's operator follow-up was already done) + golden 0.85.1 |
| Debian template | debian-13-standard_13.1-2_amd64.tar.zst (pveam-downloaded into the virgin VM — not present on a virgin box) |
| Throwaway hub customer | drill-1 re-created store-identically (SQL insert matching SaveCustomerConfig: plaintext retrieval passphrase reused from the prior drill's 0600 file, RandomHex(32)-format api_key, demo customer's config_json with fake CF tunnel + CF API tokens, real Gitea read credential); probes: config 200 with passphrase / 401 wrong; deleted after (§7) |
2. Script changes (Phase 1) + red-proof
build-golden.sh v2.0.0 (see felhom-agent CHANGELOG for the full rationale):
- CONTROLLER_IMAGE (arg 6) mandatory — no default; die-with-usage names the convention.
felhom-controller-bootstrap.pathbaked + enabled (PathExists=/etc/felhom-bootstrap/bootstrap.json,WantedBy=multi-user.target); the service (oneshot/RemainAfterExit + ConditionPathExists) unchanged.GOLDEN_SCRIPT_VERSION=2.0.0+ a[golden]provenance line (script version + baked tag) in every transcript.
Gates: bash -n clean (Windows + in-VM after CRLF strip). B5 red-proof: the no-arg invocation dies
with usage, exit 1, before any pct op:
[golden] FATAL: CONTROLLER_IMAGE (argument 6) is required — pass the released controller tag explicitly.
Usage: build-golden.sh [VMID] [TEMPLATE_VOLID] [ROOTFS_STORAGE] [ARCHIVE_STORAGE] [BRIDGE] CONTROLLER_IMAGE
e.g.: build-golden.sh 9100 local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst local-lvm local vmbr0 gitea.dooplex.hu/admin/felhom-controller:0.98.3
exit=1
(Re-run in the drill VM: rc=1 as well.)
3. Scenario A — bake integrity (PASS)
Bake on the virgin-restored drill VM as root@pam:
build-golden-new.sh 9100 local:vztmpl/debian-13-standard_13.1-2_amd64.tar.zst local-lvm local vmbr0 gitea.dooplex.hu/admin/felhom-controller:0.98.3
with REGISTRY_USER/REGISTRY_TOKEN from env (0600 file). Transcript (bake-A.log, archived at
~/drill/ on the build server):
[golden] build-golden.sh v2.0.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.98.3
…
[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) …
…
INFO: including mount point rootfs ('/') in backup
INFO: including mount point mp0 ('/var/lib/docker') in backup
INFO: including mount point mp1 ('/mnt/sys_drive') in backup
INFO: archive file size: 583MB
[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_07_03-18_01_21.tar.zst
- mp0 + mp1 both INCLUDED (the load-bearing vzdump guard confirmed; no exclusion lines).
- Build guest 9100 torn down (
pct destroy 9100 --purge, volumes removed). - Archive sha256
b9a02ef1b6f02b9b58babc4c6aad9cf6c053ebdfba116c78c8e7830de757fd01(612 002 238 bytes), identical after copy to the build server. - Deviation caught + corrected: the script's opt-in Gitea publish block keys on the same
REGISTRY_*vars the pull needs, so the bake ALSO published immediately — before Scenario C. The premature package was DELETED (HTTP 204, anon GET 404 verified) and re-published only after Scenario C passed (§6). Recorded as an observation (§8-O2), script not changed (no-"while-here" rule).
4. Scenario B — isolated hot-plug proof (PASS — the B1 STOP gate)
Scratch guest 9300 pct restored from the NEW archive (the bring-up mechanism, minus the bootstrap
mount — the exact provision race window), booted 16:05:58 UTC.
Before attach (the R6 failure state, reproduced):
felhom-controller-bootstrap.service Active: inactive (dead)
Condition: start condition unmet … ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json was not met
TriggeredBy: ● felhom-controller-bootstrap.path
felhom-controller-bootstrap.path Active: active (waiting)
docker ps: (no containers) ; /etc/felhom-bootstrap: No such file or directory
Attach to the RUNNING guest — the back-half's exact op (host dir + 0600 bootstrap.json +
chown -R 100000:100000 + pct set 9300 -mp9 <dir>,mp=/etc/felhom-bootstrap,ro=1) at 16:07:04.
After (~1 s later, NO reboot — uptime -s unchanged at 16:05:58):
Jul 03 16:07:05 systemd[1]: Starting felhom-controller-bootstrap.service …
Jul 03 16:07:05 felhom-controller-bootstrap.sh[521]: [ctrl-bootstrap] deploying gitea.dooplex.hu/admin/felhom-controller:0.98.3 …
Jul 03 16:07:05 systemd[1]: Finished felhom-controller-bootstrap.service
felhom-controller-bootstrap.service Active: active (exited) … status=0/SUCCESS
docker ps: felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.98.3 Up (healthy)
The systemd-semantics assumption holds empirically: a .path unit DOES see a hot-plugged LXC bind
mount and DOES start a Condition-gated oneshot service; RemainAfterExit=yes showed no re-trigger.
(Note: the hand-crafted minimal bootstrap.json's "id" didn't survive shell quoting for the
--hostname sed parse → hostname=<unset> fallback, by design fail-safe; Scenario C exercises the
real agent-rendered document.) Guest 9300 + its host dir removed after.
5. Scenario C — full Day-0 proof, LOCAL golden (PASS — the publish gate)
Virgin-restored VM; the new archive placed as the only local vzdump of GOLDEN_VMID 9100; installer
fetched from https://felhom.eu/scripts/felhom-host-install.sh (served v1.9.1); the guide's
canonical command WITHOUT --force-gitea-golden:
./felhom-host-install.sh --customer-id drill-1 --vmid 120 --cores 2 --memory 4096 \
--acl-storages "local local-lvm" --passphrase-file /root/.drill-pass
- Resolution order proof:
7/8 golden archive → [SKIP] using local golden: local:backup/vzdump-lxc-9100-2026_07_03-18_01_21.tar.zst; final verdictDay-0 provision SUCCESS — vmid=120 host_id=drill-1-20e7a3 … golden=local:backup/…. - First boot = 0.98.3, NO D.1b performed: in-guest
felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.98.3 Up (healthy); hubreportsrow showsdrill-1 | 0.98.3; host rowdrill-1-20e7a3 | agent 0.63.0. - Self-manage proof:
POST /api/selfupdate/check→{"current_version":"0.98.3","latest_version":"0.98.3","update_available":false}— the box is current and the floor mechanism owns it from here (0.98.3 ≥ the v0.86.0 floor-honoring code). - Agent
--selftestall-OK includingpool read pool "felhom", 1 member. - Base infra: traefik + filebrowser healthy; cloudflared restarting on the FAKE tunnel token (expected drill behavior, public leg demo-proven — same as the prior drill).
- D4 smoke:
bentopdfdeployed via the dashboard's own endpoint (POST /api/stacks/bentopdf/deploywith{"values":{}}— endpoint-proxy method, no browser bridge in this session) →Up (healthy), andpdf.drill.demo-felhom.euanswers 200 through traefik in-guest. - In-flow B1 note: the installer's v1.9.1 belt-reboot still ran (kept by design), so the container journal of the pre-reboot boot was not retrievable (volatile journal); the no-reboot claim rests on the §4 isolated proof.
6. Publish + vouch + Scenario D — published-artifact proof
- Publish (after Scenario C):
PUT …/api/packages/admin/generic/felhom-golden/0.98.3/golden.tar.zst(exactly the installer'sstep_goldenURL convention) → HTTP 201; round-trip GET sha256 matchesb9a02ef1…fd01. Publish creds = the build server's Gitea admin credential (out-of-band); version string = the baked controller tag per the 0.85.1 precedent. - Vouch (operator, password-gated UI): hub Day-0 manifest → golden 0.98.3 / sha
b9a02ef1…fd01(agent 0.63.0 was already vouched). Confirmed viaGET /api/v1/artifacts/drill-1:{"agent":{"version":"0.63.0",…},"golden":{"version":"0.98.3","sha256":"b9a02ef1…fd01"}}. - Scenario D (PASS): virgin-restored VM, NO local golden (dump dir +
pvesm listboth empty, verified), same canonical command +--force-gitea-golden:
[INFO] golden: none local — will fetch + verify from Gitea in step 7/8
[INFO] manifest: agent v0.63.0 (sha b4a89c81c33d51ea…), golden v0.98.3
[OK] verified sha256 b4a89c81c33d51ea… matches the hub manifest (agent)
[INFO] fetching golden v0.98.3 from Gitea → /var/lib/vz/dump/vzdump-lxc-9100-2026_07_03-18_31_49.tar.zst
[OK] verified sha256 b9a02ef1b6f02b9b… matches the hub manifest (golden)
[OK] golden imported + verified: local:backup/vzdump-lxc-9100-2026_07_03-18_31_49.tar.zst
[OK] Day-0 provision SUCCESS — vmid=120 host_id=drill-1-20e7a3 customer=drill-1
First boot: felhom-controller:0.98.3 Up (healthy); selfupdate/check →
{"current_version":"0.98.3","latest_version":"0.98.3","update_available":false}. Host-enroll
REUSED the Scenario C host id (mint-once-reuse held). This is the exact fetch+verify path a real
customer box (Peti's) takes.
7. Cleanup confirmation
- Hub: all
drill-1rows deleted AFTER Scenario D (customer_configs / hosts / host_reports / reports / guests / events / notification_log / customer_notifications — every count verified 0);demo-felhom+peti-felhomintact. The vouch (golden 0.98.3 in the Day-0 manifest) is the one intended production change and stays. - Drill VM: powered off; disk reverted to the
virginsnapshot and kept (reusable). Local archives and staged files on the VM gone with the revert. - Build server:
~/drill/golden-0.98.3.tar.zst+bake-A.logretained as evidence;.bake-cred.env(0600) removed after the run.
8. Observations (recorded, not acted on)
- O1 (SECURITY, operator follow-up): the demo customer's
git.token— the credential every customer box holds for artifact fetch — is a Gitea admin token with package-WRITE rights (the Scenario A bake proved it by successfully publishing with it). A customer-side compromise could overwrite published artifacts. The manifest-sha trust chain means installs would reject a tampered artifact, but the write capability should not exist on customer boxes at all → issue a scoped read-only Gitea account/token for customer configs and rotate the current one. - O2:
build-golden.sh's publish block auto-fires whenREGISTRY_USER/REGISTRY_TOKENare set (they are also needed for the pull) — a validate-before-publish flow must either use a read-only cred (blocked on O1) or expect the early publish and delete it. Candidate cleanup: a separateGOLDEN_PUBLISH=1opt-in flag. - O3: the installer's post-provision reboot (v1.9.1, R6 fix) is now a redundant belt — the path unit deploys the controller before the reboot even happens. Removal is a candidate cleanup for a future installer version; NOT done here (rule: keep the belt).
- O4:
journalctl -b -1inside the LXC guest returns nothing (volatile journal) — pre-reboot boot evidence must be captured before a reboot, or the journal made persistent, in future drills.