Golden 0.146.0 baked on the drill VM and published to gitea: felhom-golden/0.146.0/golden.tar.zst sha256 4834c703162c5437467a329144b1a523019bf5693ab9d439558be7323587e955 612696588 B (584 MB archive), controller 0.146.0 confirmed baked in All pass markers green: Result=success/ExecMainStatus=0, 0 FATAL/exclusions, docker OK (overlay2), ALL THREE mounts included (rootfs + mp0 /var/lib/docker + mp1 /mnt/sys_drive), pre-delete HTTP 404 (the pre-gate — version did not exist), upload HTTP 201. Integrity verified independently of the build host: anonymous GET | sha256sum matches byte-for-byte, ranged GET 206, content-length matches the bake's bytes. The version now appears in the hub dropdown (0.136.0, 0.143.0, 0.146.0). Teardown per GL-1: log copied out as evidence first (180:/mnt/5_hdd/felhom.eu/drill/bake-0.146.0.log), guest 9100 purged, token + script + log shredded in-VM, VM off, qemu confirmed gone via `ps -eo comm` (not the self-matching pgrep -f), drill disk reverted to the virgin snapshot exactly as found. Token-leak grep = 0 against the LITERAL token value, on the bake log and both ISO build logs from this session. REMAINING is operator-only and password-gated: Day-0 manifest Golden -> 0.146.0 (Agent stays 0.90.0, MinAgent stays 0.90.0 — v0.146.0 declares no new agent coupling), then the floor -> v0.146.0 saved LAST. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
12 KiB
felhom.eu — task reports
Overwrite this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in hub/CHANGELOG.md; the scripts history lives in scripts/CHANGELOG.md.
Pre-travel train — R-39 heal · scripts v1.21.0 + ISOs · nav polish · (golden deferred) — 2026-07-18
Commits (this repo): bcdb042 scripts v1.21.0 · ROADMAP touch (R-39 diagnosis + R-33 collapse).
Sibling commits: felhom-agent 9596d5a (v0.90.1) + f22f70c (report) · felhom-controller
24d23b8 (v0.146.0) + fd93020 (accordion tests).
**Phases 1–3 shipped. Phase 4 skipped cleanly (its own "time-permitting"). Phase 5 (golden 0.146.0
- publish) NOT started — see the closing section.**
Phase 1 — R-39: diagnosis, and the brief's hypothesis refuted
The conditional hub fix was NOT shipped, because its condition proved false. The brief said to ship a generation-bump fix "only if step 1–2 pin the mechanism to re-mint fails to bump the generation". It does not:
store.SetHostDesiredbumpsdesired_generationunconditionally — it went 2 → 3 on the re-issue.web/configs.go'sapplyPBSDRis exonerated: its "idempotent … no re-key, no second secret, no spurious generation bump" comment at ~L618 is accurate and guarded by thecur != nil && cur.Namespace != ""early return. The hub log shows mint #2 came from the re-issue path, not from an Edit-tab Save. The comment-vs-behaviour contradiction the brief expected does not exist.
The real mechanism is a signal mismatch between two tiers. The hub's re-consume signal is a
generation bump + a poke. The agent's re-apply trigger is a change in the descriptor content hash
(felhom-agent internal/pbsdr/manager.go ~L235):
if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
return // idempotent: this exact descriptor already converged
}
An ep0 credential re-issue re-keys the secret of an existing token, so token_id and
fingerprint never change and the descriptor stays byte-identical — only the side-table
host_pbs_secrets row rotates. Same hash → converged agent short-circuits → the fresh secret is
never consumed → the box keeps presenting a revoked credential → 401 forever. Proof in one line:
consumed-failed.json carries hash a4e5424…, identical to the marker.json written two
minutes before the re-issue. The comment at hub/internal/web/pbsdr.go:320 asserts the reissue
refreshes the descriptor "with the NEW token_id/fingerprint" — false for this op.
Timeline (hub log is CEST; the hub DB is UTC — a split within one service):
| CEST | Event |
|---|---|
| 18:30:51 | pbsdr provisioned … gen 2; secret stored consume-once — mint #1, via the WG-registration hook, with a generation bump |
| 18:45:51 | agent consumes mint #1 → converged state=applied |
| 18:47:52 | pbsdr credentials **re-issued** … fresh consume-once secret stored — mint #2, consumed_at stayed NULL |
A second, independent defect, found while healing. configs/felhom-pbs-apply's reconcile
passed --server to pvesm set; PVE treats server as create-only and rejects the whole call
even when the value is byte-identical. So every re-apply exited 255 — and because the agent
consumes the one-time secret before invoking the wrapper, each re-issue burned a credential.
Proven live before writing code: with --server → rejected; without → rc 0. Fixed in agent
v0.90.1 (one argv line + red-proof TestReconcileNeverPassesServerToPvesmSet, verified red then
green; it handles two vacuous-pass traps — CRLF line endings, and the WHY comment quoting the very
flag under test).
Cost I incurred: proving the mechanism consumed the pending secret against the still-unfixed
wrapper, so it burned. The box was already 401 before and after — no functional regression — but the
recoverable state was gone until an operator re-issue. The agent parked correctly in
consumed-failed.json with NOT retrying silently: no burn loop, the fail-safe worked.
HEALED — Viktor's re-issue click closed the chain in 9 s: hub re-issued 20:28:44 → agent
consumed 20:28:51 → converged state=applied 20:28:53, with the patched wrapper.
| Check | Before | After |
|---|---|---|
pvesm status |
401 Unauthorized / inactive |
active |
Direct token probe /api2/json/version |
401 |
200 |
| Real backup | none possible | felhom-pbs:backup/ct/9201/2026-07-18T18:31:06Z, 9 744 319 312 B, 13m36s |
Encrypted under fingerprint 7e:a6:af:f7:ea:6d:3e:d9 — the escrowed key, the one customer zero
holds the recovery code for. The DR tier's first real backup on the reborn box. Nothing was
destroyed: .pw, .enc (K) and the storage.cfg entry verified intact (PVE rejects atomically, so
the set-only law held).
Left for the fleet spec, deliberately not improvised: (a) make a fresh unconsumed secret actually
un-converge the agent; (b) fix the verify loop's read path — it reads /etc/pve/priv/storage/<id>.pw
directly as non-root, a file it can only ever write through the root wrapper (/etc/pve/priv is
0700 root:www-data; sudoers exposes create|reconcile|grant, no read verb); (c) an auth probe
so applied can never mean 401.
Phase 2 — scripts v1.21.0 + fresh ISOs
run_pairing() now loops inside the script (30s sleep — hub-side rate unchanged) instead of
exiting non-zero per poll, so the unit sits in activating and systemd prints nothing on the
customer's console. Registration split into register_appliance() whose transient failures the loop
retries. Journal quiet but not dark: logged once on entry, then a 10-minute heartbeat; 410 still
exits non-zero on purpose. Console banner every 5 min, single accented spelling, plus the missing
reassurance („Ez a képernyő magától frissül").
The load-bearing half is TimeoutStartSec=infinity — a Type=oneshot ExecStart is killed at 90s,
so without it systemd would kill the new wait and Restart=on-failure would silently reinstate the
exact spam this removes, after appearing to work for the first three polls.
Verified behaviourally, in a container against a stub hub answering 204 five times then
delivering: one log line plus one heartbeat, zero exits between polls, then a clean fall-through
to the direct install and exit 0. The old design produced 5 unit invocations and 5 Failed to start console lines for that same sequence.
ISOs rebuilt (both --pairing, --loader mkimage, same PVE input proxmox-ve_9.2-1.iso
sha 4e88fe41…), secret-bearing: no:
| ISO | sha256 |
|---|---|
felhom-pve-9.2-1-v1.21.0-n100-generic-mkimage.iso (safety) |
b1b25fd412b779bcacbfaa4c59002ee80a8f35e4c0bd1248cc35d997f967cb90 |
felhom-pve-9.2-1-v1.21.0-n100-demo-generic-mkimage.iso (real) |
90a0fb7da3f3f11d315bf1cdf55a7da84ab9e6d75e06b2942041470f7864031d |
Both at 180:/mnt/5_hdd/felhom.eu/felhom-iso/out/. Verified the fix actually shipped inside the
artifact, not just in git: extracted the embedded first-boot payload with xorriso and decoded it —
the shipped felhom-bootstrap.sh is byte-identical to the committed source, carries the new
cadence constants and the while true loop, the old will poll again in 30s exit line is gone,
and the embedded unit carries TimeoutStartSec=infinity. Viktor flashes the stick.
Phase 3 — controller v0.146.0 nav polish
Built, pushed and deployed to guest 9201 (0.146.0 Up (healthy)).
- Scrollbars: thin + hairline-coloured;
scrollbar-width/scrollbar-colorfor Firefox and::-webkit-scrollbar(8px, thumb--line, hover--text-3,--radius) for WebKit/Blink, since neither alone covers the browsers customers use..sidebar→--bg-2track,html→--bg-0. Tokens only. - Collapsible groups: Tárhely / Biztonsági mentés / Megosztás as accordions, chevron, exactly one
open. Header is a real
<button>witharia-expanded+aria-controls+:focus-visible, so keyboard/AT reachability is real rather than simulated. Nothing became unreachable — checked first: every group's landing page is also its first sub-item. Progressive enhancement — the active group is opened server-side, so it is correct before any JS runs. No layout jump —grid-template-rows: 0fr → 1frrather thanmax-height, animating to the content's real height with no magic number to drift; the toggle reserves its active border as transparent; both transitions off underprefers-reduced-motion.
All design-v2 gates PASS (template_id_gate, emoji_gate, native_confirm_gate,
offbox_rename_gate, mojibake_gate, app_row_dedup_gate); build/vet/tests green.
docker_run_volume_path_gate still fails on estimate.go:179 — that is R-29, pre-existing,
verified to fail identically on the untouched tree, and deliberately not bundled.
Screenshot leg NOT done, and here is the honest reason. The demo controller's password is
customer-owned since the claim flow — Viktor set it during the rehearsal — so the credentials on
the build server are stale and a curl-login returns the Bejelentkezés page. Instead of asserting
nothing, four render tests (internal/web/nav_accordion_test.go) pin the server-side half through
the real shared layout: every sub-page opens its own group with aria-expanded=true and an .active
toggle and exactly one group open (the count is asserted, not just the expected group); a page
outside any group opens nothing; every group's landing page still exists as a sub-link; the toggle is
a real button whose aria-controls targets a real element. Red-proofed — removing the is-open
marker fails two assertions on both storage pages. The visual leg still wants Viktor's browser.
Phase 4 — skipped cleanly
Its own instruction was "time-permitting; skip cleanly if not". Nothing was started, so nothing is half-done. Note that its (a) auto-mint self-bind link, (b) post-RESET health card and (c) unprovisioned-offsite flash correspond to R-36 / R-37 / R-36 and remain open as written. The conditional Phase-1 hub fix is not part of any v0.67.0 train, because its condition was refuted.
Phase 5 — golden 0.146.0 baked + published (STOP → Viktor)
Run on the nested drill VM on 180, following the recorded Phase-C procedure
(pilot/RUNBOOK-publish-0.85-0.120-2026-07-12).
Artifact: felhom-golden/0.146.0/golden.tar.zst
| GOLDEN_SHA256 | 4834c703162c5437467a329144b1a523019bf5693ab9d439558be7323587e955 |
| Size | 612 696 588 B (archive 584 MB) |
| Controller baked in | 0.146.0 (confirmed in the log, not assumed) |
Every pass marker green: unit Result=success / ExecMainStatus=0 · 0 FATAL or exclusions ·
docker OK (overlay2; data-root /var/lib/docker) · all three mounts included — rootfs /,
mp0 /var/lib/docker, mp1 /mnt/sys_drive · pre-delete HTTP 404 (the pre-gate — the version did
not previously exist) · upload HTTP 201.
Integrity verified independently of the build host — anonymous GET | sha256sum matches
4834c703…e955 byte-for-byte, ranged GET returns 206, and content-length matches the bytes
the bake reported. The golden now appears in the hub's dropdown alongside 0.136.0 and 0.143.0.
Teardown per GL-1 discipline: bake log copied out first as evidence
(180:/mnt/5_hdd/felhom.eu/drill/bake-0.146.0.log), then build guest 9100 pct destroy --purge
(0 guests remain), token + script + log shred -u in-VM, VM powered off, qemu confirmed gone (via
ps -eo comm, not the self-matching pgrep -f), drill disk reverted to the virgin snapshot
exactly as found. Token-leak grep = 0 against the literal token value, on the bake log and on
both ISO build logs from this session.
STOP — the two remaining saves are yours (password-gated)
- Day-0 manifest → Golden 0.146.0 (sha auto-reads as
4834c703…e955). Agent stays 0.90.0, MinAgent stays 0.90.0 — the v0.146.0 CHANGELOG declares no new agent coupling, so nothing justifies moving either. Save.- Then the floor → v0.146.0, saved LAST.
The traveling box converges over the tunnel — that would be the second live floor lift, this time from the road. Worth noting in CONTEXT when it lands.