12 KiB
REPORT — hub v0.73.1: allowlist disk_health_degraded (2026-07-24)
Part of the disk-health train (agent v0.94.0 + controller v0.169.0/.1 + hub v0.73.1). Appended above the previous report (kept below).
Baseline: felhom.eu be9edb4 → code 19dbb02 (allowlist) → manifest 5902703 (deploy). Clean at start.
Change: added "disk_health_degraded": true to allowedEventTypes (hub/internal/api/handler.go)
so the controller v0.169.0 per-disk SMART degradation notification is ingested, not 400-rejected. Like
offbox_enlarge_blocked, no customerMessages entry — the controller sends a dynamic Hungarian
message (disk label + attribute names) that the FormatCustomerEmail fallback preserves; a static entry
would discard the specifics. Test TestHandleEvent_DiskHealthDegradedAccepted; red-proof: drop the
allowlist line → 400 (confirmed). Full hub suite: 17 packages pass, 0 fail; build + vet clean.
Deploy (GitOps): built felhom-hub:0.73.1 (digest 79dc3082…), bumped manifests/hub.yaml tag,
ArgoCD sync → Synced/Healthy, pod hub-59857b486f-hzw48 running :0.73.1, clean startup.
§13(3) end-to-end notification proof (the exact /api/v1/event POST the controller's notifier makes,
customer demo-felhom, via curl — no committed fixture):
- BEFORE (v0.73.0):
400 Invalid event_type: disk_health_degraded— the allowlist was genuinely required (without it the customer's disk-degradation email is silently dropped). - AFTER (v0.73.1):
200 {"ok":true}; hub log:Event from demo-felhom: disk_health_degraded (warning) — Lemez állapot romlás: teszt-lemez — romló érték: függőben lévő szektorokandOperator email sent for demo-felhom/disk_health_degraded→ hub accepted + dispatched through the Resend notification pipeline. Nothing to revert (curl only); the demo's real disks are all "Nincs adat", so normal operation stays quiet.
REPORT — ISO train v1.25.0 + hub v0.73.0 (2026-07-23)
Task: the four-ruling ISO train (F9 belt-everywhere, apt no-subscription, R-63 console ő/ű,
R-71 golden-lift mitigation) + Part 6 root-vault rider + Part 7 hub offsite_stale fix.
scripts: host-install 1.17.0→1.18.0, build-felhom-iso 1.24.0→1.25.0, felhom-bootstrap R-63;
hub v0.72.0→v0.73.0. Commits: b620435 (scripts code), hub 0.73.0 train, docs follow.
Status by part
| Part | State |
|---|---|
7 — hub offsite_stale newborn anchor |
✅ SHIPPED + deployed (hub v0.73.0, Synced/Healthy); red-proofed; live: demo-hp/demo-felhom silent |
| 1 — OOB belt (F9) | ✅ code (default appliance leg) + LIVE on BOTH boxes + login PROVEN on felhom-pve |
| 2 — apt no-subscription | ✅ code + LIVE aligned on both boxes (apt-get update clean) |
| 3 — R-63 console font | ✅ code + live ő/ű capture on the HP drill (képernyő/teendő correct) |
| 4 — R-71 build-gate + rule | ✅ code + red-proofed + LIVE-demonstrated blocking the real gap; standing rule 5 added |
| 4 — golden republish | ✅ built + published + vouched — golden 0.161.0 (sha 77624408…), floor stays 0.156.0; the LIVE gap is CLOSED |
| 5 — ISO v1.25.0 build | ✅ built (…v1.25.0-nested-vm-generic-mkimage.iso, sha 99b38420…); gate passed live; rootpw+R-63 present |
| 5 — nested drill | ✅ RUN on the HP (KVM after SVM-enable): R-63 ő/ű captured, pairing registered, rootpw opens console; belt/apt fresh-install needs a bind |
| 6 — felhom-pve root vault | ✅ rotated + vaulted + verified authenticating over :22 |
CRITICAL FINDING — the R-71 gap was LIVE, now CLOSED
RESOLVED 2026-07-23: golden republished at controller 0.161.0 (built on felhom-pve via
build-golden.sh, published to Gitea 618 MB, sha 77624408bd86ec21f3de25eb57f0d9c30d9c374f6e2c19bfcea0816ccb347188,
Gitea-served copy sha-verified), operator-vouched in the manifest (golden 0.153.0 → 0.161.0;
floor unchanged at 0.156.0; agent 0.93.0). The gate now passes live and the v1.25.0 ISO built. The
gap below is the finding as discovered; it no longer exists.
The hub's current artifacts were golden 0.153.0 < managed floor 0.156.0. Every fresh install
today boots a controller BELOW the floor, so the day-0 managed update fires within minutes —
the exact update-vs-bridge race that burned demo-hp's offsite credential (DIAG-f10 / R-71). The
new build-gate correctly refuses to build a v1.25.0 ISO in this state (demonstrated live:
R-71 BUILD GATE FAILED: golden 0.153.0 < managed floor 0.156.0). The fix is Part 4's golden
republish at controller 0.161.0 + vouch — until then no valid pairing ISO can be built, by
design. The vouch (handleSetArtifacts) is operator-password-gated (no admin API), so it needs
the operator.
Part detail
Part 1 (belt). host-install: appliance mode installs the H1 belt by default (--no-oob opts
out; byo still refuses --enable-oob — deliberate exclusion, flagged not changed), and the config
writer now sets oob.enabled=true so the agent activates the belt instead of leaving the units
inert (the pre-v1.25.0 state that left it unusable — F9). Lockout-safe with no operator peer
(empty sets + no key, :22 untouched). Live: belt installed on felhom-pve + demo-hp (both
8822, sets filled {10.77.0.250}/{8822}, felhom-op key installed, :22 pid unchanged). Login
PROVEN on felhom-pve end-to-end (felhom-op@demo-felhom) — see the operator-identity note below.
Operator identity re-established (latent gap fixed). The registered oob_operator_ssh_pubkey
was ORPHANED — it matched no key the operator held (an OOB credential the operator can't reproduce
is no better than the lockout it prevents). Re-anchored to the operator's real machine: registered
its WireGuard pubkey as the operator peer 10.77.0.250 + its felhom-op-1 SSH key, built the
operator wg-felhom tunnel config (ep0.felhom.eu:443), added + persisted ep0's forward rules
(10.77.0.250 → 10.77.0.2/.3 in /etc/nftables.conf). Along the way: a corp-network UDP/443 block
was diagnosed (ep0 tcpdump — packets arriving on 51820 but not 443) and the login completed from a
home network on 443. All secrets file→file, shredded.
Part 2 (apt). align_apt_repos (appliance-only, idempotent, pre-apt): disable enterprise
deb822 .sources (Enabled: no) / legacy .list, add pve-no-subscription.sources (suite+keyring
derived), apt-get update clean. Live: both boxes were on enterprise (401 on update); aligned →
enterprise+ceph disabled, no-subscription added, apt-get update exit 0 on both.
Part 3 (R-63). felhom-bootstrap loads a Latin-2 console font (Lat2-Terminus16 → Lat2-Fixed16
→ Lat2-Terminus14) once before the first paint (idempotent, best-effort, never blocks boot). Font
names verified present in the trixie console-setup package. Renders ő/ű in the R-59 network screen
(élő/telepítő/ellenőrizze) + pairing banner (képernyő/teendő). LIVE-captured on the HP drill (see the drill section + R-63 evidence PNG).
Part 4 (build-gate + rule). assert_golden_ge_floor + ver_ge in build-felhom-iso; runs before
workspace/mint; dies with both versions on golden < floor; FELHOM_ASSERT_GOLDEN/FLOOR inputs;
unset warns loudly. Red-proofed both directions (floor>golden → exit 1; golden≥floor → OK; equal →
OK; unset → WARN). Standing rule 5 in publish-train-rules.md. Live-demonstrated against the real
production gap (see above).
Part 6 (root vault). On felhom-pve, generated a fresh strong root password (never printed),
chpasswd, vaulted via PUT /api/v1/hosts/demo-felhom-8363b5/recovery-credential (HTTP 200, the
same day-0 step_break_glass mechanism, host api_key). Verified: retrieved via the operator
endpoint (global key) and it AUTHENTICATED root over :22 (VAULT-AUTH-OK: demo-felhom); copy
shredded. felhom-pve now has G1 vault parity with demo-hp. (Note: the vault's set_at reflects the
original day-0 stamp — not bumped on rotation; cosmetic, the SECRET is the rotated one, proven by
the auth test.)
Nested drill — RUN on the HP (2026-07-23)
The N100 attempt was aborted (install crawled, contended by live guest 9201). Per the operator's
steer the drill moved to the HP (demo-hp). Blocker found + fixed: the HP had AMD SVM disabled in
BIOS (SVM disabled (by BIOS) in MSR_VM_CR) so KVM was unavailable — the operator enabled
Security → System Security → Virtualization Technology = Enable; on reboot /dev/kvm appeared
(kvm_amd auto-loads via the persisted /etc/modules-load.d/felhom-kvm.conf) and the live guest
9201 auto-recovered.
The KVM-accelerated install reached first-boot in ~12 min (vs the N100's 40+ that never finished) and PROVED, live:
- R-63 — ő/ű console capture: the first-boot pairing banner renders „képernyő"/„teendő" with
correct double-acute ő under the Lat2 font (evidence:
documentation/audits/R-63-console-font-drill-2026-07-23.png). This is the spec's R-63 closure evidence. - v1.25.0 pairing chain: the box self-registered as an unclaimed appliance (hub log
appliance registered: new unclaimed box uuid=206c8838…). - R-61 rootpw regression: the build's
.rootpw.txtopens the installed console (ROOTPW-OK: felhom-host). - Bonus:
apt-get install sshpasssucceeded on demo-hp — confirms the Part-2 no-subscription alignment works for real installs, not justupdate.
Still requiring an operator bind (password-gated) to exercise on the ISO-driven fresh install: the belt-default + apt legs inside host-install and the golden-0.161.0 pull (host-install only runs after a box is bound). Those legs are already proven live on both real boxes this session, and the no-floor-update outcome is structurally guaranteed post-vouch (golden 0.161 ≥ floor 0.156). The unclaimed drill box (uuid 206c8838) can be bound to a throwaway customer to close that last leg, or discarded.
Nested drill — (earlier N100 attempt, superseded)
Golden republish + vouch + the v1.25.0 ISO build are DONE. The nested drill was attempted on the
N100 (felhom-pve): 1.7 GB ISO transferred + sha-verified, VM 9313 created (OVMF SB-off for the
mkimage loader, disk-first boot), auto-install started. It ran KVM-accelerated (confirmed:
/usr/bin/kvm, cpu host,+kvm_pv_eoi…, q35+pve0 — not TCG, not memory-starved) but crawled
pathologically — ~40+ min pegged at 100% of one core, stuck in the kernel/grub/initramfs
configure at ~87% — because the N100 was simultaneously running the live customer guest 9201
plus the golden build scratch. Aborted to protect the live guest (a runaway install must not
starve a production customer workload); VM 9313 + the golden scratch (9100 + local vzdump) + the
drill ISO were destroyed/removed, 9201 untouched, 70 GB freed.
Operator steer (2026-07-23): future drills/builds move to the HP (demo-hp) — more RAM, faster CPU, and it carries no live customer guest, so a nested install there will be far quicker.
What the drill would have added, and its current standing:
- R-63 console ő/ű capture — the one leg genuinely PENDING a live boot. The font code is
verified (Lat2-Terminus16/Fixed16/Terminus14 confirmed present in the trixie console-setup
package;
set_console_fontcorrect + embedded in the built ISO's bootstrap payload). The visual screendump rides the next HP-based drill. - belt installed+active + apt=no-subscription (fresh-install path) — the legs themselves are already proven live on both real boxes this session; the ISO-driven path additionally needs an operator bind (password-gated, like the v1.24.0 rehearsal's drill customer).
- golden 0.161.0 boots with NO managed-update firing — now structurally guaranteed (golden 0.161.0 ≥ floor 0.156.0 after the vouch; the update only fires below the floor); a live boot would merely confirm it.
.rootpw.txtopens the console — R-61 regression; rides the HP drill.- build-gate red-proof in the nest — already demonstrated live on DooPlex against the real gap
(
golden 0.153 < floor 0.156→ build refused).
Observed / not acted on
- The live golden(0.153)<floor(0.156) gap predates this train (an old golden bake vs a later floor lift) — the new gate surfaces it; the golden republish fixes it.
- ep0 root is reachable from DooPlex (used for the operator-tunnel + forward-rule work); the forward rules are persisted to survive reboot.