Files
felhom.eu/REPORT.md
T
admin 1f4702fe50 docs: golden 0.146.0 baked + published (Phase 5); STOP for the operator saves
Golden 0.146.0 baked on the drill VM and published to gitea:
  felhom-golden/0.146.0/golden.tar.zst
  sha256 4834c703162c5437467a329144b1a523019bf5693ab9d439558be7323587e955
  612696588 B (584 MB archive), controller 0.146.0 confirmed baked in

All pass markers green: Result=success/ExecMainStatus=0, 0 FATAL/exclusions,
docker OK (overlay2), ALL THREE mounts included (rootfs + mp0 /var/lib/docker +
mp1 /mnt/sys_drive), pre-delete HTTP 404 (the pre-gate — version did not exist),
upload HTTP 201.

Integrity verified independently of the build host: anonymous GET | sha256sum
matches byte-for-byte, ranged GET 206, content-length matches the bake's bytes.
The version now appears in the hub dropdown (0.136.0, 0.143.0, 0.146.0).

Teardown per GL-1: log copied out as evidence first
(180:/mnt/5_hdd/felhom.eu/drill/bake-0.146.0.log), guest 9100 purged, token +
script + log shredded in-VM, VM off, qemu confirmed gone via `ps -eo comm` (not
the self-matching pgrep -f), drill disk reverted to the virgin snapshot exactly
as found. Token-leak grep = 0 against the LITERAL token value, on the bake log
and both ISO build logs from this session.

REMAINING is operator-only and password-gated: Day-0 manifest Golden -> 0.146.0
(Agent stays 0.90.0, MinAgent stays 0.90.0 — v0.146.0 declares no new agent
coupling), then the floor -> v0.146.0 saved LAST.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
2026-07-18 21:20:50 +02:00

12 KiB
Raw Blame History

felhom.eu — task reports

Overwrite this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in hub/CHANGELOG.md; the scripts history lives in scripts/CHANGELOG.md.

Pre-travel train — R-39 heal · scripts v1.21.0 + ISOs · nav polish · (golden deferred) — 2026-07-18

Commits (this repo): bcdb042 scripts v1.21.0 · ROADMAP touch (R-39 diagnosis + R-33 collapse). Sibling commits: felhom-agent 9596d5a (v0.90.1) + f22f70c (report) · felhom-controller 24d23b8 (v0.146.0) + fd93020 (accordion tests).

**Phases 13 shipped. Phase 4 skipped cleanly (its own "time-permitting"). Phase 5 (golden 0.146.0

  • publish) NOT started — see the closing section.**

Phase 1 — R-39: diagnosis, and the brief's hypothesis refuted

The conditional hub fix was NOT shipped, because its condition proved false. The brief said to ship a generation-bump fix "only if step 12 pin the mechanism to re-mint fails to bump the generation". It does not:

  • store.SetHostDesired bumps desired_generation unconditionally — it went 2 → 3 on the re-issue.
  • web/configs.go's applyPBSDR is exonerated: its "idempotent … no re-key, no second secret, no spurious generation bump" comment at ~L618 is accurate and guarded by the cur != nil && cur.Namespace != "" early return. The hub log shows mint #2 came from the re-issue path, not from an Edit-tab Save. The comment-vs-behaviour contradiction the brief expected does not exist.

The real mechanism is a signal mismatch between two tiers. The hub's re-consume signal is a generation bump + a poke. The agent's re-apply trigger is a change in the descriptor content hash (felhom-agent internal/pbsdr/manager.go ~L235):

if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
    return // idempotent: this exact descriptor already converged
}

An ep0 credential re-issue re-keys the secret of an existing token, so token_id and fingerprint never change and the descriptor stays byte-identical — only the side-table host_pbs_secrets row rotates. Same hash → converged agent short-circuits → the fresh secret is never consumed → the box keeps presenting a revoked credential → 401 forever. Proof in one line: consumed-failed.json carries hash a4e5424…, identical to the marker.json written two minutes before the re-issue. The comment at hub/internal/web/pbsdr.go:320 asserts the reissue refreshes the descriptor "with the NEW token_id/fingerprint" — false for this op.

Timeline (hub log is CEST; the hub DB is UTC — a split within one service):

CEST Event
18:30:51 pbsdr provisioned … gen 2; secret stored consume-once — mint #1, via the WG-registration hook, with a generation bump
18:45:51 agent consumes mint #1 → converged state=applied
18:47:52 pbsdr credentials **re-issued** … fresh consume-once secret stored — mint #2, consumed_at stayed NULL

A second, independent defect, found while healing. configs/felhom-pbs-apply's reconcile passed --server to pvesm set; PVE treats server as create-only and rejects the whole call even when the value is byte-identical. So every re-apply exited 255 — and because the agent consumes the one-time secret before invoking the wrapper, each re-issue burned a credential. Proven live before writing code: with --server → rejected; without → rc 0. Fixed in agent v0.90.1 (one argv line + red-proof TestReconcileNeverPassesServerToPvesmSet, verified red then green; it handles two vacuous-pass traps — CRLF line endings, and the WHY comment quoting the very flag under test).

Cost I incurred: proving the mechanism consumed the pending secret against the still-unfixed wrapper, so it burned. The box was already 401 before and after — no functional regression — but the recoverable state was gone until an operator re-issue. The agent parked correctly in consumed-failed.json with NOT retrying silently: no burn loop, the fail-safe worked.

HEALED — Viktor's re-issue click closed the chain in 9 s: hub re-issued 20:28:44 → agent consumed 20:28:51 → converged state=applied 20:28:53, with the patched wrapper.

Check Before After
pvesm status 401 Unauthorized / inactive active
Direct token probe /api2/json/version 401 200
Real backup none possible felhom-pbs:backup/ct/9201/2026-07-18T18:31:06Z, 9 744 319 312 B, 13m36s

Encrypted under fingerprint 7e:a6:af:f7:ea:6d:3e:d9 — the escrowed key, the one customer zero holds the recovery code for. The DR tier's first real backup on the reborn box. Nothing was destroyed: .pw, .enc (K) and the storage.cfg entry verified intact (PVE rejects atomically, so the set-only law held).

Left for the fleet spec, deliberately not improvised: (a) make a fresh unconsumed secret actually un-converge the agent; (b) fix the verify loop's read path — it reads /etc/pve/priv/storage/<id>.pw directly as non-root, a file it can only ever write through the root wrapper (/etc/pve/priv is 0700 root:www-data; sudoers exposes create|reconcile|grant, no read verb); (c) an auth probe so applied can never mean 401.


Phase 2 — scripts v1.21.0 + fresh ISOs

run_pairing() now loops inside the script (30s sleep — hub-side rate unchanged) instead of exiting non-zero per poll, so the unit sits in activating and systemd prints nothing on the customer's console. Registration split into register_appliance() whose transient failures the loop retries. Journal quiet but not dark: logged once on entry, then a 10-minute heartbeat; 410 still exits non-zero on purpose. Console banner every 5 min, single accented spelling, plus the missing reassurance („Ez a képernyő magától frissül").

The load-bearing half is TimeoutStartSec=infinity — a Type=oneshot ExecStart is killed at 90s, so without it systemd would kill the new wait and Restart=on-failure would silently reinstate the exact spam this removes, after appearing to work for the first three polls.

Verified behaviourally, in a container against a stub hub answering 204 five times then delivering: one log line plus one heartbeat, zero exits between polls, then a clean fall-through to the direct install and exit 0. The old design produced 5 unit invocations and 5 Failed to start console lines for that same sequence.

ISOs rebuilt (both --pairing, --loader mkimage, same PVE input proxmox-ve_9.2-1.iso sha 4e88fe41…), secret-bearing: no:

ISO sha256
felhom-pve-9.2-1-v1.21.0-n100-generic-mkimage.iso (safety) b1b25fd412b779bcacbfaa4c59002ee80a8f35e4c0bd1248cc35d997f967cb90
felhom-pve-9.2-1-v1.21.0-n100-demo-generic-mkimage.iso (real) 90a0fb7da3f3f11d315bf1cdf55a7da84ab9e6d75e06b2942041470f7864031d

Both at 180:/mnt/5_hdd/felhom.eu/felhom-iso/out/. Verified the fix actually shipped inside the artifact, not just in git: extracted the embedded first-boot payload with xorriso and decoded it — the shipped felhom-bootstrap.sh is byte-identical to the committed source, carries the new cadence constants and the while true loop, the old will poll again in 30s exit line is gone, and the embedded unit carries TimeoutStartSec=infinity. Viktor flashes the stick.


Phase 3 — controller v0.146.0 nav polish

Built, pushed and deployed to guest 9201 (0.146.0 Up (healthy)).

  • Scrollbars: thin + hairline-coloured; scrollbar-width/scrollbar-color for Firefox and ::-webkit-scrollbar (8px, thumb --line, hover --text-3, --radius) for WebKit/Blink, since neither alone covers the browsers customers use. .sidebar--bg-2 track, html--bg-0. Tokens only.
  • Collapsible groups: Tárhely / Biztonsági mentés / Megosztás as accordions, chevron, exactly one open. Header is a real <button> with aria-expanded + aria-controls + :focus-visible, so keyboard/AT reachability is real rather than simulated. Nothing became unreachable — checked first: every group's landing page is also its first sub-item. Progressive enhancement — the active group is opened server-side, so it is correct before any JS runs. No layout jumpgrid-template-rows: 0fr → 1fr rather than max-height, animating to the content's real height with no magic number to drift; the toggle reserves its active border as transparent; both transitions off under prefers-reduced-motion.

All design-v2 gates PASS (template_id_gate, emoji_gate, native_confirm_gate, offbox_rename_gate, mojibake_gate, app_row_dedup_gate); build/vet/tests green. docker_run_volume_path_gate still fails on estimate.go:179 — that is R-29, pre-existing, verified to fail identically on the untouched tree, and deliberately not bundled.

Screenshot leg NOT done, and here is the honest reason. The demo controller's password is customer-owned since the claim flow — Viktor set it during the rehearsal — so the credentials on the build server are stale and a curl-login returns the Bejelentkezés page. Instead of asserting nothing, four render tests (internal/web/nav_accordion_test.go) pin the server-side half through the real shared layout: every sub-page opens its own group with aria-expanded=true and an .active toggle and exactly one group open (the count is asserted, not just the expected group); a page outside any group opens nothing; every group's landing page still exists as a sub-link; the toggle is a real button whose aria-controls targets a real element. Red-proofed — removing the is-open marker fails two assertions on both storage pages. The visual leg still wants Viktor's browser.


Phase 4 — skipped cleanly

Its own instruction was "time-permitting; skip cleanly if not". Nothing was started, so nothing is half-done. Note that its (a) auto-mint self-bind link, (b) post-RESET health card and (c) unprovisioned-offsite flash correspond to R-36 / R-37 / R-36 and remain open as written. The conditional Phase-1 hub fix is not part of any v0.67.0 train, because its condition was refuted.

Phase 5 — golden 0.146.0 baked + published (STOP → Viktor)

Run on the nested drill VM on 180, following the recorded Phase-C procedure (pilot/RUNBOOK-publish-0.85-0.120-2026-07-12).

Artifact: felhom-golden/0.146.0/golden.tar.zst

GOLDEN_SHA256 4834c703162c5437467a329144b1a523019bf5693ab9d439558be7323587e955
Size 612 696 588 B (archive 584 MB)
Controller baked in 0.146.0 (confirmed in the log, not assumed)

Every pass marker green: unit Result=success / ExecMainStatus=0 · 0 FATAL or exclusions · docker OK (overlay2; data-root /var/lib/docker) · all three mounts included — rootfs /, mp0 /var/lib/docker, mp1 /mnt/sys_drive · pre-delete HTTP 404 (the pre-gate — the version did not previously exist) · upload HTTP 201.

Integrity verified independently of the build host — anonymous GET | sha256sum matches 4834c703…e955 byte-for-byte, ranged GET returns 206, and content-length matches the bytes the bake reported. The golden now appears in the hub's dropdown alongside 0.136.0 and 0.143.0.

Teardown per GL-1 discipline: bake log copied out first as evidence (180:/mnt/5_hdd/felhom.eu/drill/bake-0.146.0.log), then build guest 9100 pct destroy --purge (0 guests remain), token + script + log shred -u in-VM, VM powered off, qemu confirmed gone (via ps -eo comm, not the self-matching pgrep -f), drill disk reverted to the virgin snapshot exactly as found. Token-leak grep = 0 against the literal token value, on the bake log and on both ISO build logs from this session.

STOP — the two remaining saves are yours (password-gated)

  1. Day-0 manifest → Golden 0.146.0 (sha auto-reads as 4834c703…e955). Agent stays 0.90.0, MinAgent stays 0.90.0 — the v0.146.0 CHANGELOG declares no new agent coupling, so nothing justifies moving either. Save.
  2. Then the floor → v0.146.0, saved LAST.

The traveling box converges over the tunnel — that would be the second live floor lift, this time from the road. Worth noting in CONTEXT when it lands.