# felhom.eu — task reports > **Overwrite** this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in [hub/CHANGELOG.md](hub/CHANGELOG.md); the scripts history lives in [scripts/CHANGELOG.md](scripts/CHANGELOG.md). ## Pre-travel train — R-39 heal · scripts v1.21.0 + ISOs · nav polish · (golden deferred) — 2026-07-18 **Commits (this repo):** `bcdb042` scripts v1.21.0 · ROADMAP touch (R-39 diagnosis + R-33 collapse). **Sibling commits:** `felhom-agent` `9596d5a` (v0.90.1) + `f22f70c` (report) · `felhom-controller` `24d23b8` (v0.146.0) + `fd93020` (accordion tests). **Phases 1–3 shipped. Phase 4 skipped cleanly (its own "time-permitting"). Phase 5 (golden 0.146.0 + publish) NOT started — see the closing section.** --- ## Phase 1 — R-39: diagnosis, and the brief's hypothesis refuted **The conditional hub fix was NOT shipped, because its condition proved false.** The brief said to ship a generation-bump fix "only if step 1–2 pin the mechanism to *re-mint fails to bump the generation*". It does not: - `store.SetHostDesired` bumps `desired_generation` **unconditionally** — it went **2 → 3** on the re-issue. - `web/configs.go`'s `applyPBSDR` is **exonerated**: its "idempotent … no re-key, no second secret, no spurious generation bump" comment at ~L618 is accurate and guarded by the `cur != nil && cur.Namespace != ""` early return. The hub log shows mint #2 came from the **re-issue** path, not from an Edit-tab Save. The comment-vs-behaviour contradiction the brief expected does not exist. **The real mechanism is a signal mismatch between two tiers.** The hub's re-consume signal is *a generation bump + a poke*. The agent's re-apply trigger is *a change in the descriptor content hash* (`felhom-agent internal/pbsdr/manager.go` ~L235): ```go if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) { return // idempotent: this exact descriptor already converged } ``` An ep0 credential re-issue re-keys the **secret of an existing token**, so `token_id` and `fingerprint` never change and the descriptor stays **byte-identical** — only the side-table `host_pbs_secrets` row rotates. Same hash → converged agent short-circuits → the fresh secret is never consumed → the box keeps presenting a revoked credential → **401 forever**. Proof in one line: `consumed-failed.json` carries hash `a4e5424…`, **identical** to the `marker.json` written two minutes before the re-issue. The comment at `hub/internal/web/pbsdr.go:320` asserts the reissue refreshes the descriptor "with the NEW token_id/fingerprint" — false for this op. **Timeline (hub log is CEST; the hub DB is UTC — a split *within one service*):** | CEST | Event | |---|---| | 18:30:51 | `pbsdr provisioned … gen 2; secret stored consume-once` — mint #1, via the WG-registration hook, **with** a generation bump | | 18:45:51 | agent consumes mint #1 → `converged state=applied` | | 18:47:52 | `pbsdr credentials **re-issued** … fresh consume-once secret stored` — mint #2, `consumed_at` stayed NULL | **A second, independent defect, found while healing.** `configs/felhom-pbs-apply`'s `reconcile` passed `--server` to `pvesm set`; PVE treats `server` as **create-only** and rejects the whole call even when the value is byte-identical. So *every* re-apply exited 255 — and because the agent consumes the one-time secret **before** invoking the wrapper, each re-issue **burned a credential**. Proven live before writing code: with `--server` → rejected; without → **rc 0**. Fixed in **agent v0.90.1** (one argv line + red-proof `TestReconcileNeverPassesServerToPvesmSet`, verified red then green; it handles two vacuous-pass traps — CRLF line endings, and the WHY comment quoting the very flag under test). **Cost I incurred:** proving the mechanism consumed the pending secret against the still-unfixed wrapper, so it burned. The box was already 401 before and after — no functional regression — but the recoverable state was gone until an operator re-issue. The agent parked correctly in `consumed-failed.json` with `NOT retrying silently`: **no burn loop**, the fail-safe worked. **HEALED — Viktor's re-issue click closed the chain in 9 s:** hub re-issued 20:28:44 → agent consumed 20:28:51 → `converged state=applied` 20:28:53, with the patched wrapper. | Check | Before | After | |---|---|---| | `pvesm status` | `401 Unauthorized` / `inactive` | **`active`** | | Direct token probe `/api2/json/version` | `401` | **`200`** | | Real backup | none possible | **`felhom-pbs:backup/ct/9201/2026-07-18T18:31:06Z`, 9 744 319 312 B, 13m36s** | Encrypted under fingerprint `7e:a6:af:f7:ea:6d:3e:d9` — the **escrowed** key, the one customer zero holds the recovery code for. The DR tier's **first real backup on the reborn box**. Nothing was destroyed: `.pw`, `.enc` (K) and the `storage.cfg` entry verified intact (PVE rejects atomically, so the set-only law held). **Left for the fleet spec, deliberately not improvised:** (a) make a fresh unconsumed secret actually un-converge the agent; (b) fix the verify loop's read path — it reads `/etc/pve/priv/storage/.pw` directly as non-root, a file it can only ever *write* through the root wrapper (`/etc/pve/priv` is `0700 root:www-data`; sudoers exposes `create|reconcile|grant`, **no read verb**); (c) an auth probe so `applied` can never mean `401`. --- ## Phase 2 — scripts v1.21.0 + fresh ISOs `run_pairing()` now loops **inside** the script (30s sleep — hub-side rate unchanged) instead of exiting non-zero per poll, so the unit sits in `activating` and systemd prints nothing on the customer's console. Registration split into `register_appliance()` whose transient failures the loop retries. Journal quiet but not dark: logged once on entry, then a 10-minute heartbeat; `410` still exits non-zero on purpose. Console banner every 5 min, single accented spelling, plus the missing reassurance („Ez a képernyő magától frissül"). **The load-bearing half is `TimeoutStartSec=infinity`** — a `Type=oneshot` ExecStart is killed at 90s, so without it systemd would kill the new wait and `Restart=on-failure` would silently reinstate the exact spam this removes, *after appearing to work for the first three polls*. **Verified behaviourally**, in a container against a stub hub answering `204` five times then delivering: **one log line plus one heartbeat, zero exits between polls**, then a clean fall-through to the direct install and `exit 0`. The old design produced 5 unit invocations and 5 `Failed to start` console lines for that same sequence. **ISOs rebuilt (both `--pairing`, `--loader mkimage`, same PVE input `proxmox-ve_9.2-1.iso` sha `4e88fe41…`), `secret-bearing: no`:** | ISO | sha256 | |---|---| | `felhom-pve-9.2-1-v1.21.0-n100-generic-mkimage.iso` (safety) | `b1b25fd412b779bcacbfaa4c59002ee80a8f35e4c0bd1248cc35d997f967cb90` | | `felhom-pve-9.2-1-v1.21.0-n100-demo-generic-mkimage.iso` (real) | `90a0fb7da3f3f11d315bf1cdf55a7da84ab9e6d75e06b2942041470f7864031d` | Both at `180:/mnt/5_hdd/felhom.eu/felhom-iso/out/`. **Verified the fix actually shipped inside the artifact**, not just in git: extracted the embedded first-boot payload with xorriso and decoded it — the shipped `felhom-bootstrap.sh` is **byte-identical to the committed source**, carries the new cadence constants and the `while true` loop, the old `will poll again in 30s` exit line is **gone**, and the embedded unit carries `TimeoutStartSec=infinity`. Viktor flashes the stick. --- ## Phase 3 — controller v0.146.0 nav polish Built, pushed and **deployed to guest 9201** (`0.146.0 Up (healthy)`). - **Scrollbars:** thin + hairline-coloured; `scrollbar-width`/`scrollbar-color` for Firefox **and** `::-webkit-scrollbar` (8px, thumb `--line`, hover `--text-3`, `--radius`) for WebKit/Blink, since neither alone covers the browsers customers use. `.sidebar` → `--bg-2` track, `html` → `--bg-0`. Tokens only. - **Collapsible groups:** Tárhely / Biztonsági mentés / Megosztás as accordions, chevron, exactly one open. Header is a **real `