§7.2 answer: YES. A real day-0 from the existing v1.25.0 ISO reached a claimable,
app-serving box in ~10 minutes unattended, and an app's data came back from the
drive with the guest's app.yaml gone — proven readable by the application over
its own TCP path, with a discriminator (PRE-BACKUP row = 1, POST-BACKUP row = 0).
Part 0: NO ISO rebuild needed, verified against the ISO on disk rather than from
source. It bakes only felhom-bootstrap.sh, its unit and the secret-free pairing
env (full-base64 match, 1 hit each) and 0 hits for any installer, controller or
golden marker. The installer is fetched at run time; the live URL is byte-identical
to repo HEAD (v1.22.0, six days newer than the ISO) and the fresh box ran it.
Part 1: baked 0.188.0 rather than the brief's 0.187.0 — 0.187.0 lacks D5, which
is the very claim Part 2 step 6 tests. Published (404 pre-gate with a 200 control;
anonymous download, 649310288 bytes, sha match), vouched, and consumed by a real
box. R-120's gate exercised BOTH ways: 0.185.1 refused with no write, 0.188.0
allowed — evaluated, not silently skipped.
Part 3: RUNBOOK-manual-build.md cited a "RECORDED" qemu line that is itself
labelled reconstructed and whose source says it was never saved. The real
invocation is now captured from this bake as §4.0, with the bake/publish/teardown
steps; the old entry is marked SUPERSEDED.
Teardown all three layers, hub disposition stated: VM destroyed, scratch storage
removed with space returned exactly, customer sess-g DELETED via full cascade.
sess-f deliberately left (R-131) with its command recorded.
Filed, none fixed: R-128 (false ISO_VERSION invariant comment), R-129 (demo-hp's
"no baked SSH key" is stale — key auth works), R-130 (HARD_MIN_LVM_GIB warns and
proceeds), R-131 (fourth orphaned scratch customer), R-132 (curl's %{redirect_url}
printed the hub operator password into a transcript — HUB_PW needs rotating).
12 KiB
RUNBOOK — manual build / deploy / publish (agent · controller · golden · hub)
Audience: the operator, on DooPlex (192.168.0.180, Debian 13) — the same environment CC uses.
Builds are local commands; felhom-pve is one ssh hop.
Written: 2026-07-11, from the verified felhom.eu/skills/felhom-build-deploy/SKILL.md command set
- the publish-train procedure (
documentation/pilot/RUNBOOK-publish-0.79-0.110-2026-07-10.md) +documentation/runbooks/publish-train-rules.md. If this doc and the skill ever disagree, the skill wins.
0. The mental model (why the hub dropdown "lags")
There is no CI — nothing builds on push, ever. Every artifact moves through two separate, deliberate verbs:
| Verb | Means | Who sees it |
|---|---|---|
| DEPLOY | build + install on the DEMO (felhom-pve / guest 9201 / k3s) | only the demo box |
| PUBLISH | upload the artifact to Gitea packages + vouch it in the hub Day-0 manifest | Day-0 installs, the manifest dropdowns, and the remote-rollout machinery (signed agent ops, controller floor) |
A version can be live-on-demo for days without being published (agent 0.82–0.84 right now). The hub Configuration dropdowns list published artifacts only — that screen showing 0.81.0/0.113.0 is correct, not stale.
Remote rollout to customer boxes is a third, separate step and has its own runbook pattern:
agent = operator-signed agent_update op (per box); controller = the global floor (DB row —
acts immediately; save it LAST). Rules: publish-train-rules.md.
1. Session setup (every session, first)
FELHOM_ROOT=/mnt/5_hdd/felhom.eu # working root: ALL felhom repos/build/drill/iso live HERE
Clean-tree gate before any build: git status --porcelain must be empty and git rev-parse HEAD
must equal git rev-parse origin/main in the repo being built. An unpushed change does not exist.
| Host | Access | Role |
|---|---|---|
| DooPlex (this host) | local | builds ($FELHOM_ROOT/git/felhom-agent, $FELHOM_ROOT/build/felhom-controller, $FELHOM_ROOT/build/felhom-hub), kubectl |
| Demo PVE host | ssh felhom-pve (root@192.168.0.162) |
agent install, pct exec 9201 |
| Hub UI | hub.felhom.eu → Configuration | manifest vouch, MinAgent, floor (operator password) |
Housekeeping note: $FELHOM_ROOT/build/felhom-agent on 180 is a stale pre-June-23 leftover — agent
builds live in $FELHOM_ROOT/git/felhom-agent now. Safe to remove the old dir.
2. Agent (felhom-agent binary → felhom-pve, then optionally publish)
Always commit+push to main first (an unpushed change does not exist).
# BUILD on 180 (the explicit git pull is load-bearing)
cd $FELHOM_ROOT/git/felhom-agent && git pull && go build -ldflags '-X main.version=<VER>' -o /tmp/felhom-agent-<VER> ./cmd/felhom-agent
# PUSH to the PVE host — ONE hop, the binary is already local
scp /tmp/felhom-agent-<VER> felhom-pve:/tmp/
# DEPLOY with backup + restart + verify
ssh felhom-pve "cp /usr/local/bin/felhom-agent /usr/local/bin/felhom-agent.bak-\$(/usr/local/bin/felhom-agent --version | awk '{print \$2}') && install -m0755 /tmp/felhom-agent-<VER> /usr/local/bin/felhom-agent && systemctl restart felhom-agent && sleep 3 && /usr/local/bin/felhom-agent --version && journalctl -u felhom-agent -n 20 --no-pager"
Gotchas (earned): if configs/ changed in the repo, ship the sudoers + guarded wrapper WITH the
binary (several Go guards only exist when the deployed configs match); beware CRLF when scp-ing
configs through Windows; after restart the journal must show a clean ReassertGuestBinds and (since
0.84) the network-mount reassert, with no capability degradations.
PUBLISH (makes it real for the fleet):
# from the agent repo, with REGISTRY_* creds set; use the LIVE-DEPLOYED bytes, sha-verified across hops
scripts/publish-agent.sh <VER> <path-to-binary>
Pre-gate: the package GET for <VER> must be 404 before (published artifacts are immutable —
never republish over an existing version). The script prints the sha256 — record it: the same sha
goes into the hub manifest AND any signed agent_update op. One sha, three places, byte-identical.
Then hub → Configuration → Day-0 artifacts: pick the new Agent version (sha auto-read from Gitea),
set Min agent if the paired controller release declares MinAgent: in its CHANGELOG, Save
artifact manifest. (This save does NOT move the floor — that's a separate card since hub v0.45.)
3. Controller image (felhom-controller → guest 9201, remote via floor)
9201 is golden/bootstrap-managed — no compose file. The bootstrap service runs whatever tag is in
/etc/felhom-controller-image (anonymous pull).
# BUILD+PUSH the image (build.sh does NOT pull — the explicit pull is load-bearing)
cd $FELHOM_ROOT/build/felhom-controller && git -C $FELHOM_ROOT/git/felhom-controller pull && ./build.sh <VER> --push
# DEPLOY on the demo guest
ssh felhom-pve "pct exec 9201 -- bash -c 'docker pull gitea.dooplex.hu/admin/felhom-controller:<VER> && echo gitea.dooplex.hu/admin/felhom-controller:<VER> > /etc/felhom-controller-image && systemctl restart felhom-controller-bootstrap.service'"
# VERIFY
ssh felhom-pve "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
Remote rollout = the hub floor (Configuration → Managed updates). Order rules (train rules doc): the manifest vouches the target FIRST; any MinAgent must be satisfied fleet-wide (the hub now HOLDS boxes below it automatically, and flags them); save the floor LAST — the DB row acts immediately on every box below it, on their next report.
4. Golden image (fresh Day-0 installs)
The golden is a pre-baked controller-era guest image built in the drill VM on 180
(/mnt/5_hdd/felhom.eu/drill/drill.qcow2, internal snapshot virgin).
4.0 The canonical drill-VM launch — captured from a real bake, not reconstructed
Until 2026-07-31 this section told the reader to "use the RECORDED qemu launch line" from
RUNBOOK-publish-0.79-0.110-2026-07-10.md Phase C — while that line is itself labelled
"PASS (reconstructed — DEVIATION)" and its own §Deviations says "the canonical qemu-system
one-liner was never saved". The document forbade improvising and pointed at an improvisation. The
lines below were captured verbatim from the 0.188.0 bake on 2026-07-31 and are now the canonical
invocation. Re-capture them (do not retype from memory) if the bake host or disk layout ever changes.
# 1. revert the disk to virgin (non-destructive to the snapshot; succeeding also proves no qemu holds the qcow2)
qemu-img snapshot -a virgin /mnt/5_hdd/felhom.eu/drill/drill.qcow2
# 2. COLD-boot it (the snapshot is disk-only, 0 B VM_SIZE — never -loadvm)
qemu-system-x86_64 -enable-kvm -cpu host -smp 4 -m 8192 \
-drive file=/mnt/5_hdd/felhom.eu/drill/drill.qcow2,format=qcow2,if=virtio,cache=writeback \
-netdev user,id=n0,dhcpstart=10.0.2.30,hostfwd=tcp::2222-10.0.2.15:22 \
-device virtio-net-pci,netdev=n0 -display none -daemonize \
-pidfile /mnt/5_hdd/felhom.eu/drill/qemu.pid
if=virtio is load-bearing (the guest expects /dev/vda). SSH answers on :2222 in ~40 s
(ssh -i /mnt/5_hdd/felhom.eu/drill/id_drill -p 2222 root@localhost); pveversion read
pve-manager/9.2.2 on 2026-07-31. Liveness check: ps -eo comm | grep qemu-system-x86 —
pgrep -f qemu-system-x86_64 self-matches your own command line and reports a false "still running".
4.1 Bake + publish
-
Revert + boot per §4.0.
-
The debian template is absent on
virginand the exact point release rots — list the current one (pveam available --section system | grep debian-13) andpveam download local <that>. It wasdebian-13-standard_13.6-1_amd64.tar.zston 2026-07-31. -
scpinbuild-golden.sh(agent repoconfigs/) + the Gitea token (~/.gitea-tokenon 180, 0600), then run it as a transient unit so it survives a session close, reading the token from the file inside the VM so it never reaches a command line:GT=$(cat /root/.gitea-token); systemd-run --unit=golden-bake --collect \ --setenv=GITEA_USER=admin --setenv=GITEA_TOKEN=$GT \ --setenv=REGISTRY_USER=admin --setenv=REGISTRY_TOKEN=$GT \ bash -c "/root/build-golden.sh 9100 local:vztmpl/<template> local-lvm local vmbr0 \ gitea.dooplex.hu/admin/felhom-controller:<VER> > /root/bake.log 2>&1"CONTROLLER_IMAGEis a required argument (a hand-bumped default rotted twice) andGOLDEN_VERSIONis derived from it — the golden's version IS the controller it bakes. The script is the publisher: it uploads to Gitea and printsGOLDEN_VERSION+GOLDEN_SHA256. Pass markers:overlay2 OK,including mount pointfor rootfs and mp0 and mp1 with noexcluding/FATAL,upload OK (HTTP 201). The 404 pre-gate applies to the package URL. -
Teardown:
pct destroy 9100 --purge,shred -uthe token/script/log after copying the log out for evidence,poweroff, wait for qemu to exit,qemu-img snapshot -a virgin. Token-leak grep on the saved log =grep -c -F "$(cat ~/.gitea-token)"(the literal value — a broad[a-f0-9]{40}pattern false-hits image shas). -
Hub → Configuration → Day-0 artifacts: pick the new Golden, Save. The R-120 gate lives on this save (
hub/internal/web/configs.go:1165) and REFUSES a golden older than the newest controller the fleet reports. It does not run on a controller image deploy — it is not a general drift net.
A golden is only needed when a publish train wants fresh installs current — demo deploys never need it.
The full 0.188.0 run, with the observables: documentation/audits/tester-gate-golden-0.188.0-2026-07-31.md.
5. Hub (felhom.eu/hub → k3s, GitOps)
The manifest is the truth — a built image deploys NOTHING until manifests/hub.yaml's image:
tag changes in git and the ArgoCD app is deliberately synced (auto-sync is OFF; never
kubectl set image, never :latest).
cd $FELHOM_ROOT/build/felhom-hub && ./build.sh <VER> --push
# edit manifests/hub.yaml image tag → <VER>; commit; push
sudo kubectl -n argocd annotate application felhom argocd.argoproj.io/refresh=hard --overwrite; sleep 8; sudo kubectl -n argocd get application felhom -o jsonpath='{.status.sync.status} {.status.sync.revision}{"\n"}'
sudo kubectl -n argocd patch application felhom --type merge -p '{"operation":{"initiatedBy":{"username":"op"},"sync":{"syncStrategy":{"apply":{}}}}}'
# verify: Synced/Healthy + rollout + live image tag + logs
sudo kubectl -n argocd get application felhom -o jsonpath='sync={.status.sync.status} health={.status.health.status}{"\n"}'; sudo kubectl -n felhom-system rollout status deploy/hub --timeout=90s; sudo kubectl -n felhom-system get deploy hub -o jsonpath='{.spec.template.spec.containers[0].image}'; echo; sudo kubectl -n felhom-system logs -l app=hub --tail 10
6. "What's live right now?" one-liners
ssh felhom-pve "/usr/local/bin/felhom-agent --version" # agent on demo
ssh felhom-pve "pct exec 9201 -- cat /etc/felhom-controller-image" # controller on demo
sudo kubectl -n felhom-system get deploy hub -o jsonpath='{.spec.template.spec.containers[0].image}' # hub
# published = the hub Configuration dropdowns (they read Gitea packages live)
# fleet = hub Dashboard per-host rows (agent_version + controller version per box)
7. State snapshot as of 2026-07-11 (so the screens make sense)
| Artifact | Live on demo | Published / vouched | Peti |
|---|---|---|---|
| agent | 0.84.0 | 0.81.0 | 0.81.0 |
| controller | 0.117.0 | golden 0.113.0, floor 0.113.0 (DB = env, aligned) | 0.113.0 |
| hub | 0.46.0 | n/a (central) | n/a |
The gap between columns 1 and 2 is the pending publish train (agent 0.84 + golden/floor 0.117 + MinAgent 0.81 + Peti's journal one-liner + temp-creds deletion) — its runbook follows the 0.81/0.113 pattern with these numbers.
8. Iron rules (recap)
Never :latest; never republish over an existing package version (404 pre-gate); never
kubectl set image / bare kubectl apply against GitOps surfaces; manifest before floor, floor
LAST; one sha in three places, byte-identical; secrets never in transcripts or commits; an unpushed
change does not exist.