Half 1, the artifact: golden 0.186.0 baked, published, vouched, and proven on a REAL day-0 on demo-hp (not the fixture, per the rule committed in Part 1). With the target detached, the fresh box's endpoint returned the TargetAbsent copy -- "A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod..." -- with offer_path absent entirely. The day-old read on the 0.185.1 golden had returned the false system-disk message plus an offer of the other drive. That is the customer-visible defect closed. Half 2, the mechanism: operator ruled REFUSE, shipped as hub v0.82.0 and DEPLOYED. Proven live by re-attempting the original mistake -- vouching the stale 0.185.1 golden now yields HTTP 303 flash=golden_behind_fleet plus [WARN] artifact vouch REFUSED, and the manifest reads back unchanged at 0.186.0. Refused AND unwritten, against the real fleet signal rather than a unit fixture. Recorded on R-29's audit list as the first ENFORCED gate beside its three orphans, so the contrast is kept rather than lost. The orphans are unchanged -- this proves the pattern is available, not that the backlog moved. Teardown all three layers: VM 9402 purged, r120-images removed with the space measured back, hub layer gate-blocked on ONLINE with the command recorded. Last session's sess-e was deleted this run, discharging its recorded layer 3.
16 KiB
R120-golden-rebake-2026-07-30 — the golden is current again, and the class now has a gate that refuses
Run: R-120, CC on DooPlex, 2026-07-30. CLOSED on both halves. Golden 0.186.0 baked, published, vouched and proven on a real day-0; and the class now has a mechanism — hub v0.82.0's vouch-path gate, which refuses a golden the fleet has already outrun, deployed and proven live by re-attempting the original mistake.
1. Baselines, established separately
| Thing | Value | Established from |
|---|---|---|
controller on main |
0.186.0 | felhom-controller/CHANGELOG.md:1 |
| controller the golden baked before this task | 0.185.1 | the golden's own record: drill/bake-0.185.1.log:1 (baking controller …:0.185.1) and :330 (GOLDEN_VERSION=0.185.1) |
| controller running, felhom-pve | 0.186.0 | pct exec 9201 -- docker ps |
| controller running, demo-hp | 0.185.1 | same |
| agent published + vouched | 0.116.0, sha b47c5c4d… |
registry GET + manifest read-back |
| golden vouched before | 0.185.1, sha dba00f3e… |
manifest read-back |
felhom.eu HEAD at start |
772956d |
git rev-parse |
R-120 confirmed at source, not inferred: build-golden.sh:345 derives
GOLDEN_VERSION="${GOLDEN_VERSION:-${CONTROLLER_IMAGE##*:}}" — the golden version is the controller
tag it bakes. So golden 0.185.1 ⇒ every fresh box installed controller 0.185.1, which predates R-114 +
R-112 (shipped in 0.186.0), while a deployed box was already on 0.186.0. One release behind, and the
missing release is the one that tells the truth about an absent backup target.
2. The bake procedure — it IS written down
Contrary to the risk the task flagged, the procedure exists and was followed, not reconstructed:
documentation/runbooks/RUNBOOK-manual-build.md:101-115§4 "Golden image (fresh Day-0 installs)" — the five steps: revert the drill VM tovirgin, boot it, runbuild-golden.sh <controller VER>, teardown per GL-1, then vouch in the hub.felhom-agent/configs/build-golden.sh— the tool, with its usage in its own header ([VMID] [TEMPLATE_VOLID] [ROOTFS_STORAGE] [ARCHIVE_STORAGE] [BRIDGE] CONTROLLER_IMAGE), and it publishes to Gitea itself (:335-371) whenGITEA_USER+GITEA_TOKENare set. There is no separatepublish-golden.sh— the bake script is the publisher.documentation/runbooks/publish-train-rules.md— rules 1–4. Rule 2 (save the floor field LAST, separate save) and rule 3 (MinAgent) were consulted; see §4 for how they applied.
One documentation-integrity finding, small but worth recording. RUNBOOK-manual-build.md:103-105
says to use the "RECORDED qemu launch line; do not reconstruct from memory" and points at
RUNBOOK-publish-0.79-0.110-2026-07-10.md Phase C — but that line is itself labelled
PASS (reconstructed — DEVIATION) at :29, with :77-79 stating the canonical one-liner was never
saved. So the runbook forbids reconstruction while citing a reconstruction. The line at :29 is
nonetheless correct — it booted first try here, and independently earlier today. A trap learned today
and worth adding to it: adding any -device at launch shifts PCI slot assignment and renames the
guest NIC, so vmbr0 loses its port and SSH never comes up while the console shows a healthy PVE login.
Boot the recorded line unchanged and hot-plug extras over QMP.
3. The bake — controller 0.186.0
Run in the DooPlex nested drill VM (its documented purpose — target-selection.md names it a bake
fixture, and this is a bake), reverted to virgin first, build-golden.sh v2.1.0 as a transient unit:
[golden] build-golden.sh v2.1.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.186.0
docker OK (overlay2; data-root /var/lib/docker)
INFO: including mount point rootfs ('/')
INFO: including mount point mp0 ('/var/lib/docker')
INFO: including mount point mp1 ('/mnt/sys_drive')
INFO: archive file size: 618MB
[golden] upload OK (HTTP 201)
GOLDEN_VERSION=0.186.0
GOLDEN_SHA256=b760ac6a33e70700f427b4207afeebc8694bd6bb9478973c30901d9f059bde17
Pass markers, counted on the saved log: overlay2 present, FATAL 0, exclusions 0, 3 mounts
included, upload 201. Template: debian-13-standard_13.6-1_amd64.tar.zst (the pinned version in
the script's default has rotted away; the current one was resolved with pveam available).
Log retained out-of-VM at drill/bake-0.186.0.log (330 lines); token-leak grep against the literal
token value on the saved log = 0.
GL-1 teardown: build guest 9100 pct destroy --purge; .gitea-token, the script and the in-VM log
shred -u'd (0 residual); VM powered off; drill.qcow2 reverted to virgin (snapshot intact, ID 1,
2026-07-03) — the fixture is exactly as found.
4. The three publish/vouch observables, quoted as returned
(1) PUBLISHED — anonymous GET, i.e. what the installer actually does, sha over the bytes returned:
HTTP 200 bytes=648930639
b760ac6a33e70700f427b4207afeebc8694bd6bb9478973c30901d9f059bde17 /tmp/golden-0.186.0.tar.zst
(identical to the bake's GOLDEN_SHA256)
(2) VOUCHED — manifest read BACK after the POST, never the 303:
agent_version: vouched=['0.116.0'] agent_sha256 = b47c5c4dab641ee5…
golden_version: vouched=['0.186.0'] golden_sha256 = b760ac6a33e70700…
min_agent = 0.113.0 min_controller_version = 0.156.0
wrapper_sha256 = 104db0a4401f65bb… (re-checked against configs/felhom-pbs-apply — NO drift)
(3) A CONSUMER RESOLVES IT — the hub telling a real fresh box what to install:
[INFO] Artifact manifest served for customer sess-f (agent=0.116.0 golden=0.186.0)
Publish-train rules as applied: rule 1 (manifest before floor) — satisfied trivially, no floor moved.
Rule 2 (floor field LAST, separate save) — the floor was not touched at all; min_controller_version
read back 0.156.0, unchanged, and the floor is a separate form (/configuration/global-floor) which
was never posted. Rule 3 (MinAgent) — left at 0.113.0, because controller 0.186.0's CHANGELOG
header declares MinAgent 0.113.0 (unchanged); raising it would hold demo-hp (agent 0.113.0) for no
reason. The agent vouch was carried through untouched.
5. §2.1 — the fresh-box proof
Machine: demo-hp (Tier 0, the designated drill+build VM host). A real day-0, not the fixture —
per the rule committed in Part 1, this is a path claim, so a fixture could not prove it. VM 9402
r120-drill, disks on a dir storage r120-images at /mnt/nvme-1tb root (never local-lvm),
installed from the v1.25.0 ISO: self-register (appliance 14) → operator bind to sess-f → credentials
delivered once → guest provisioned from the new golden.
R-120's fix, proven at the install path:
hub → Artifact manifest served for customer sess-f (agent=0.116.0 golden=0.186.0)
hub → Event from sess-f: controller_started (info) — Controller elindult (0.186.0)
box → felhom-agent --version → felhom-agent 0.116.0
box → pct exec 9201 -- docker ps → gitea.dooplex.hu/admin/felhom-controller:0.186.0
A freshly installed box now runs 0.186.0. Before this task it ran 0.185.1. That is the defect closed.
Target drive enrolled through the real endpoints and detached (real hot-detach), so the box sits in R-120's exact state — and agent 0.116.0's R-116 fix is visible on it:
/disks felhom-backup row: state=disconnected role=system mount_path=''
guest_path='/mnt/felhom-drives/cel' backup_target=True
pvesm: felhom-backup dir inactive 0 0 0
The capture — R-120's symptom, inverted
Operator-relayed claim code, then GET /api/storage/backup-target with the target still detached, on
controller 0.186.0 from the new golden. Verbatim:
{"data":{"degraded":true,"known":true,"label":"",
"message":"A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod, a teljes rendszermentés nem készül el.",
"target":"felhom-backup"},"ok":true}
- ✅ the
TargetAbsentcopy — the backup drive is unreachable; until you reconnect it the full system backup will not be made. Reconnect that drive. - ✅
offer_pathis absent from the payload entirely — no drive is offered as a remedy. - ✅ not the generic "A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer".
For contrast, the same read on the OLD 0.185.1 golden one day earlier
(audits/R116-v0116-2026-07-30.md §6a caveat) returned the generic system-disk message plus
offer_label: "Adat meghajto" and offer_path: "/mnt/felhom-drives/adat". The false message and the wrong
offer are both gone. That is R-120 closed at the customer-visible layer.
The claim gate had to be cleared by an operator-relayed code — the endpoint first returned
{"ok":false,"error":"dashboard not yet claimed"}. The claim gate runs before auth
(controller/internal/web/auth.go:52-67) and exempts only /claim, /claim/request-new-code,
/api/health, /static/ (claim.go:221-228) — no Bearer escape, and regen-password
(hub/internal/web/configs.go:965-980) regenerates the retrieval passphrase, not the dashboard
password. So an operator-relayed code is the only route. This is R-119's fourth instance.
6. Part 3 — the ruling, and the recommendation with its wiring
Recommendation: the gate, and it must REFUSE, not warn. The record is the argument: this gap has
opened three times (R-111, R-115, R-120) and the first two were closed by re-baking and
remembering. Remembering has now failed twice. And R-29 is the standing proof that a check nobody
runs is worse than none, because it reads as coverage — hostinstall_gates.py sat red and uninvoked
across three version bumps and hub_confirm_gate.py has never run at all. So the distinguishing
property is not does a check exist but does it block.
Where it wires — established, not proposed loosely. hub/internal/web/configs.go:1116
handleSetArtifacts is the vouch path: the only UI writer of store.SetArtifactManifest
(:1142), reached from server.go:484. It already normalises the versions (:1121-1125) and resolves
the SHAs (:1128-1131), so a refusal belongs immediately before :1142. Nothing can vouch a golden
without passing through it — which is precisely what makes it a mechanism rather than a fourth orphan
in scripts/.
The comparison signal is already local. The hub stores controller_version per guest
(hub/internal/store/store.go:294), populated from reports. A gate comparing the submitted
golden_version against the newest controller version any box has reported would have caught this
exact case: felhom-pve was reporting 0.186.0 while the manifest vouched a 0.185.1 golden. No
outbound registry call, no new credential.
Its honest blind spot, stated up front: if no box has ever run the newer controller, the hub cannot know it exists, so a golden baked behind an unreleased-to-the-fleet controller would still pass. That is a real limit — and it does not touch the failure that actually happened three times, which is a golden left behind a controller already running in the field.
The test must go through the production wiring, not an injected seam: drive handleSetArtifacts over
httptest with a guest row reporting 0.186.0 and a submitted golden of 0.185.1, then assert both
that the response is the refusal and that the manifest was not written. Three shipped defects in
this project were fully green with the seam disconnected; a gate that can be inert is the thing this
gate exists to prevent.
RULING: REFUSE — approved by the operator, 2026-07-30. Built, deployed, proven live.
Shipped as hub v0.82.0: the gate in handleSetArtifacts immediately before the only write;
store.NewestReportedControllerVersion() over reports.controller_version (the column SaveReport
denormalises, store.go:903), semver-compared in Go; an operator-legible refusal flash naming the
remedy; fail-open in exactly two deliberate cases — an empty golden field, and an unknown fleet version.
A near-miss worth keeping. The first draft read guests.controller_version — a column that exists
in the schema (store.go:294) and that nothing writes. That gate would always have seen "" and
failed open: inert, i.e. this gate's own failure shape, and one grep from shipping. Caught by looking
for a writer before trusting the column.
Tests: 4, through the production handler over httptest, never an injected seam. The refusal asserts
both the flash and that the manifest was not written — a gate that redirects and saves anyway reads
as enforcement while providing none. Plus the allow case, both fail-open cases, and the semver-ordering
case (0.99.0 vs 0.186.0, a pair this fleet has shipped). Red-proof: deleting the block makes the stale
golden vouchable and both refusal assertions fail.
PROVEN LIVE on the deployed hub — the original mistake, re-attempted and blocked
Hub 0.82.0 built, pushed, manifest pinned, ArgoCD hard-refreshed + synced → Synced / Healthy, pod on
felhom-hub:0.82.0. Then the exact R-120 mistake was re-attempted — vouch the stale 0.185.1 golden:
POST /configuration/artifacts → HTTP 303 …/configuration?flash=golden_behind_fleet
hub log: [WARN] artifact vouch REFUSED: golden 0.185.1 is older than the newest controller the fleet
reports (0.186.0) — a fresh install would land on stale application code (R-120)
manifest read back afterwards: golden still 0.186.0, sha still b760ac6a… — the refused POST did NOT write
Refused and unwritten, on the live hub against the real fleet signal rather than a unit fixture.
R-29's audit list now records it (ROADMAP.md) as the first enforced gate beside its three
orphans, so the contrast is kept rather than lost. The orphans are unchanged — this proves the pattern is
available, not that the backlog moved.
7. Teardown
| layer | item | disposition |
|---|---|---|
| 1 — machine | DooPlex bake fixture: build guest 9100 purged, drill.qcow2 reverted to virgin |
DONE |
| 1 — machine | demo-hp VM 9402 r120-drill + all volumes |
DESTROYED qm destroy 9402 --purge; qm list shows only drill-r50 |
| 2 — host | r120-images dir storage at /mnt/nvme-1tb |
REMOVED; felhom-backup available 928787076 KiB after == before, local-lvm 38.84 % unchanged — space returned |
| 3 — hub | customer sess-f, host sess-f-2670b5, appliance 14 |
GATE-BLOCKED — command recorded below. Cascade attempted, correctly refused HTTP 409 (host still ONLINE, customer_delete.go:220-228). Deletable once it ages ONLINE→DOWN, exactly as sess-e was this run |
pvesm status on demo-hp before this run, for the layer-2 comparison at teardown:
felhom-backup dir active 983379700 4566012 928787076 0.46%
local dir active 40516856 14980024 23446440 36.97%
local-lvm lvmthin active 56545280 21962186 34583093 38.84% ← the fence figure
Teardown commands, recorded now:
ssh demo-hp 'qm stop 9402; qm destroy 9402 --purge; pvesm remove r120-images; pvesm status'
POST /configs/sess-f/delete ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=sess-f expect_hosts=1
Previous session's layer 3 is now DISCHARGED. sess-e was gate-blocked on ONLINE last run with the
command recorded; it aged to STALE and the cascade completed this run — /customers/sess-e → 404, 0
hosts, and the hub logged customer DELETE cascade COMPLETE for sess-e (journal #7) — full teardown
including residue purged (reports=4 app_telemetry=4 … appliance_registrations=1). That is the §13
layer-3 discipline working as intended: recorded, then executed.
Fences held: drill-r50 untouched · guest 9201 on both demo boxes untouched · neither demo box
re-targeted · nothing on local-lvm · Peti untouched · no existing box received the new golden ·
felhom-agent/felhom-controller source untouched · the hub DB copy taken for the break-glass
credential was shred -u'd immediately.