docs(v1.24.0): R-59/R-60/R-61 SHIPPED — CHANGELOG, README, ROADMAP (+R-62), runbook, capability map, drill evidence, REPORT

Virgin-ISO nested drill closed the train: dead-NIC install baked the
fallback (incl. the dead default gateway), the R-59 screen painted
(capture committed beside the spike doc), the cable move healed +
registered at the hub in 23s unaided, and the build's rootpw file
matched the installed box's shadow hash. R-59 SHIPPED with the recorded
deviation (first-boot gate; installer-initrd abort out of scope by
operator ack). R-60 SHIPPED (spike + drill cited; F-P9 route-flush fix
included). R-61 slice 1 SHIPPED. New R-62 row (hub delete-dialog
cosmetics, XS). Capability map: new PROVEN-LIVE row (nested != metal,
said so). Cleanup verified: felhom-pve interfaces byte-identical,
bridge/VMs/ISO removed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UuFPHmHNrCJj1VhY6QdDMU
This commit is contained in:
2026-07-22 11:45:55 +02:00
parent 699325bd8a
commit a12c6f9730
8 changed files with 218 additions and 87 deletions
+4 -3
View File
@@ -72,9 +72,10 @@
| R-56 | **[P3] Apps do not say how technical they are, so a beginner can be ambushed by a config-heavy one.** The catalog presents every app as equally approachable — one Telepítés button, the same Hungarian copy — but they are not. Glance needs a hand-written `glance.yml` before it does anything; some apps need a reverse-proxy or API concept to configure; others genuinely are install-and-use. A tester who picks the wrong first app concludes the PRODUCT is broken, not that they picked an advanced app. | S | **idea (filed 2026-07-21)** | Origin: TASK-E Part 3 — **filed, deliberately not implemented**. Shape: a `difficulty:` field in `.felhom.yml` (`kezdő` / `haladó` / `technikás`) surfaced as a catalog-card badge and repeated on the deploy screen. Cheap and incremental: one optional metadata field plus a badge, classifiable app-by-app with no migration — an app with no `difficulty:` simply shows no badge. **This is the constructive half of the glance ruling**: glance STAYS in the catalog (operator ruling 2026-07-21 — it is a legitimate app, not a broken one; its missing seeded `glance.yml` is a known pre-existing finding), and the honest fix is to LABEL it rather than hide it. Pairs with R-41: that gate proves an app CAN still deploy; this field tells a customer whether THEY should be the one deploying it. **Badge plumbing is ALREADY BUILT (controller v0.158.0)**`web.MetaBadge` + the `meta_badge` template partial + the `lifecycleBadge` funcmap entry were written generic for exactly this: a `difficultyBadge` funcmap function returning the same `*MetaBadge`, plus a `difficulty:` field on `stacks.Metadata`, is the whole remaining job. No new markup, no new CSS. Re-sized accordingly |
| R-57 | **An app can be withdrawn from the catalog without orphaning the customers already running it**`.felhom.yml` `lifecycle: available / hidden / abandoned`. | S | **SHIPPED 2026-07-21 — controller v0.158.0 (+ v0.158.1 fix), LIVE-PROVEN** | **Motivating case: plant-it.** Earlier the same day it was withdrawn by moving its directory to `retired/` — which un-offers the app but ALSO makes the controller's orphan detector see the template as GONE for anyone running it, flagging their working install `Elavult` and offering a Törlés button. Withdrawing an app must never take a working app away from a customer, so the directory move was replaced by metadata. **Operator requirements, verbatim (ruling 2026-07-21):** states `available` / `hidden` / `abandoned`; abandoned apps are NOT offered to new installs (no badge-but-installable middle state); deployed instances of hidden/abandoned apps keep full function; an abandoned app shows a permanent notice that *„Az alkalmazás fejlesztője felhagyott a fejlesztéssel. A telepített verzió továbbra is használható, de frissítések és biztonsági javítások már nem érkeznek hozzá."* **Design points that matter beyond this feature:** (a) the deploy gate is server-side and fail-CLOSED before any mutation — hiding a button is not a gate, and a stale link or direct POST must be refused; (b) an unknown lifecycle value fails OPEN (→ available + one WARN), deliberately opposite, because a typo or a state from a newer catalog must never pull a working app out of every customer's list — both read the same `EffectiveLifecycle`, so they cannot disagree; (c) lifecycle NEVER reaches orphan detection, red-proofed. **LIVE-PROVEN 2026-07-21 on 9201** through the real endpoints: `lifecycle: abandoned` arrived via the normal catalog sync; plant-it renders **0 times** on the Alkalmazások page while the control app renders 10; a direct `POST /api/stacks/plant-it/deploy` returns **HTTP 409 `{"ok":false,"error":"Ez az alkalmazás jelenleg nem telepíthető."}`**; the app page carries the notice and no Telepítés button. **v0.158.1 is a shipped-and-caught defect worth remembering:** the three predicates were declared with POINTER receivers, and html/template cannot call those on the non-addressable value the handler passes — every `/apps/<slug>` returned 500, for every app, while compiling cleanly with a fully green suite, because no test rendered `app_info`. A template method call is only checked when the template runs. Follow-on: R-56's difficulty badge reuses this plumbing |
| R-58 | **[P2] Assisted disk-picker install mode — the installer should let the operator CHOOSE the target disk instead of requiring the serial up front.** Today an install is either unattended (the answer file pins one `ID_SERIAL_SHORT`, which you can only know by first booting the machine) or match-nothing safety (aborts by design). That forces a two-boot dance for every new box: boot the safety ISO to read the serial, rebuild the ISO armed, boot again. | SM | **idea — operator ruling 2026-07-21** | **Operator's argument, verbatim:** *"the installer should list the available storage devices (excluding the installation media) and let us select one, and continue."* **Shape:** a THIRD ISO mode alongside the two that exist — unattended-serial and match-nothing-safety. It enumerates candidate disks with **size / model / serial**, excludes the installation media itself, takes a selection plus a confirm, and proceeds. **Unattended+serial REMAINS the appliance/factory mode** — it is the right shape when the machine is provisioned in bulk and nobody is standing there; the picker is for the case where somebody is. **Slice 1 (cheap, same code surface, do this first):** improve the abort screen. On filter-no-match the installer currently just fails safe and says nothing useful — it should print the candidate table (size/model/serial) plus the one-line hint naming which serial to put in the profile. That alone collapses the two-boot dance from "boot, guess, go read docs, rebuild" to "boot, copy the serial off the screen, rebuild", and it is the same enumeration code the full picker needs. **Why it matters beyond convenience:** it is the BYO / reinstall flow — a customer's existing hardware, or a rebuild of a box whose disk layout nobody recorded, is exactly where the serial is unknown and a wrong guess is destructive. The current fail-safe is correct but mute. Origin: TASK-G, arming the HP install ISO — the serial had to be read off the board by hand between two boots |
| R-59 | **[P1] A no-DHCP install must HARD-ABORT — instead it bakes the installer's fallback address as a STATIC config and completes, producing a box that can never call home.** | S | **idea (found live 2026-07-21, demo-hp)** | **This is the worst silent onboarding failure shape there is:** the install *succeeds*, the box looks finished, and it is permanently unreachable — no hub check-in, no pairing, no way in except a keyboard and monitor. Found on the HP t740's first install: the 4-port NIC got no DHCP lease at this site (see the t740 gotcha in `scripts/iso/README.md`), and rather than refusing, the installer wrote its **192.168.100.2 fallback as a static `vmbr0` address** into `/etc/network/interfaces` and carried on. **The philosophy is already established one layer over — the disk filter refuses loudly and touches nothing when it cannot identify its target** (spike S5c, proven twice on real boards). Networking deserves the identical treatment: no lease on any carrier-bearing NIC ⇒ **abort with a legible screen**, never invent an address. Slice: detect "DHCP produced no lease" in the answer/first-boot path and fail with the candidate NIC table (name / MAC / carrier / link speed) plus the one-line remedy, exactly as R-58 slice 1 does for disks — same refuse-loudly grammar, same screen shape. Pairs with **R-60**, which is the self-heal for the case where the cable simply moved |
| R-60 | **[P2] First-boot NIC sweep self-heal: if the hub is unreachable, try DHCP across every carrier-bearing NIC before settling.** | S | **idea (found live 2026-07-21, demo-hp)** | `felhom-bootstrap` currently accepts whatever addressing the installer left behind and, if the hub cannot be reached, simply stays broken. On demo-hp the fix was a human moving one cable from the 4-port card to the onboard port — **a sweep would have healed it unaided**: enumerate NICs with `carrier=1`, DHCP each in turn, and keep the first that reaches the hub. Cheap because the box has nothing to lose at first boot (no customer data, no running guests) and the failure it repairs is total. Deliberately scoped to FIRST BOOT and to the hub-unreachable condition only — a running box must never re-shuffle its own networking. Complements **R-59**: that one refuses to produce an unreachable box, this one repairs the case where the truth changed after the install (cable moved, switch port died, the installer guessed the wrong port) |
| R-61 | **[P1] The baked root password must be knowable by the operator — the recurring console lockout.** | S | **idea (found live 2026-07-21; recurring)** | The ISO mints a **fresh throwaway crypt hash per build** and the plaintext is discarded, so nobody — including the person holding the machine — can log into the console of a box they just installed. Today that meant reaching demo-hp only through the G1 break-glass credential vaulted in the hub, which is the right mechanism for a *lost* password and the wrong one for a *never-known* password: it requires a working hub, a working network, and operator tooling, at exactly the moment the likely reason you need the console is that one of those is broken. **Slice 1 (do this):** the ISO build emits the baked root password into the build REPORT and the operator cheat-sheet alongside the sha256 — it is already a per-build value, so surfacing it costs nothing and closes the lockout. **Follow-up (appliance-grade):** keep it per-build random and treat the build output as the record of truth. **A fixed well-known password is explicitly REJECTED (operator ruling 2026-07-21)** — a pre-pairing box sits on a stranger's LAN with a predictable root credential, which is a far worse exposure than the lockout it would fix. Relates to G1 break-glass (the vault stays; this is about the window before/without it) |
| R-59 | **[P1] A no-DHCP install must HARD-ABORT — instead it bakes the installer's fallback address as a STATIC config and completes, producing a box that can never call home.** | S | **SHIPPED v1.24.0 (2026-07-22) — as a FIRST-BOOT refuse-loudly gate, with a RECORDED DEVIATION: the install-time abort is out of scope (the 192.168.100.2 fallback is baked inside the Proxmox auto-installer itself, unreachable without an installer-initrd hook; operator-acked, not silently dropped). The screen + gate proven on the nested drill (SPIKE-firstboot-nic-sweep-2026-07-22.md); nested ≠ metal — metal proof rides the next real install** | **This is the worst silent onboarding failure shape there is:** the install *succeeds*, the box looks finished, and it is permanently unreachable — no hub check-in, no pairing, no way in except a keyboard and monitor. Found on the HP t740's first install: the 4-port NIC got no DHCP lease at this site (see the t740 gotcha in `scripts/iso/README.md`), and rather than refusing, the installer wrote its **192.168.100.2 fallback as a static `vmbr0` address** into `/etc/network/interfaces` and carried on. **The philosophy is already established one layer over — the disk filter refuses loudly and touches nothing when it cannot identify its target** (spike S5c, proven twice on real boards). Networking deserves the identical treatment: no lease on any carrier-bearing NIC ⇒ **abort with a legible screen**, never invent an address. Slice: detect "DHCP produced no lease" in the answer/first-boot path and fail with the candidate NIC table (name / MAC / carrier / link speed) plus the one-line remedy, exactly as R-58 slice 1 does for disks — same refuse-loudly grammar, same screen shape. Pairs with **R-60**, which is the self-heal for the case where the cable simply moved |
| R-60 | **[P2] First-boot NIC sweep self-heal: if the hub is unreachable, try DHCP across every carrier-bearing NIC before settling.** | S | **SHIPPED v1.24.0 (2026-07-22) — spike + nested drill proven (SPIKE-firstboot-nic-sweep-2026-07-22.md): cable move → sweep → heal + hub registration unaided in <1 min; sweep is structurally first-boot-only (state.json gate + the unit's done-flag condition); drill also surfaced and fixed the baked-fallback-default-route trap (flush before the bounded dhclient)** | `felhom-bootstrap` currently accepts whatever addressing the installer left behind and, if the hub cannot be reached, simply stays broken. On demo-hp the fix was a human moving one cable from the 4-port card to the onboard port — **a sweep would have healed it unaided**: enumerate NICs with `carrier=1`, DHCP each in turn, and keep the first that reaches the hub. Cheap because the box has nothing to lose at first boot (no customer data, no running guests) and the failure it repairs is total. Deliberately scoped to FIRST BOOT and to the hub-unreachable condition only — a running box must never re-shuffle its own networking. Complements **R-59**: that one refuses to produce an unreachable box, this one repairs the case where the truth changed after the install (cable moved, switch port died, the installer guessed the wrong port) |
| R-61 | **[P1] The baked root password must be knowable by the operator — the recurring console lockout.** | S | **slice 1 SHIPPED v1.24.0 (2026-07-22): the build emits the plaintext into a 0600 sibling `<iso>.rootpw.txt` (single record of truth — never logged/manifested/committed); drill-verified against the installed box's shadow hash. Follow-up (appliance-grade record-keeping) stays open** | The ISO mints a **fresh throwaway crypt hash per build** and the plaintext is discarded, so nobody — including the person holding the machine — can log into the console of a box they just installed. Today that meant reaching demo-hp only through the G1 break-glass credential vaulted in the hub, which is the right mechanism for a *lost* password and the wrong one for a *never-known* password: it requires a working hub, a working network, and operator tooling, at exactly the moment the likely reason you need the console is that one of those is broken. **Slice 1 (do this):** the ISO build emits the baked root password into the build REPORT and the operator cheat-sheet alongside the sha256 — it is already a per-build value, so surfacing it costs nothing and closes the lockout. **Follow-up (appliance-grade):** keep it per-build random and treat the build output as the record of truth. **A fixed well-known password is explicitly REJECTED (operator ruling 2026-07-21)** — a pre-pairing box sits on a stranger's LAN with a predictable root credential, which is a far worse exposure than the lockout it would fix. Relates to G1 break-glass (the vault stays; this is about the window before/without it) |
| R-62 | **[P3] Hub delete dialog: show the customer-id the operator must type, and reword the three acks for the ghost shape.** | XS | **idea (operator, 2026-07-22)** | Cosmetic, hub-only, docs-only in the v1.24.0 train. The delete confirmation asks the operator to type the customer-id, but the id appears NOWHERE on the Edit page the dialog opens from — the operator has to fish it out of the URL or another tab. Also: for a GHOST customer (host already gone) the three acknowledgement checkboxes describe teardown steps that cannot happen; **wording only** — the server MUST keep requiring all three (the render-gate lesson of v0.70.1 stands: reachability and requirements are separate concerns). |
| R-53 | **`app_export.html` substituted the CSRF token where the customer domain belongs** - the open-in-browser link was wrong for every app with a subdomain, and a session CSRF token landed in a URL. | XS | **SHIPPED (controller v0.150.0, 2026-07-20)** | One template token (`{{$.CSRFToken}}` -> `{{$.Domain}}`) plus the `Domain` key in `exportPageHandler`'s data map - that handler does not go through `baseData`, which is where every other page gets it, so the template had no domain to read. Render tests assert the joined `<sub>.<domain>` and that the token appears nowhere in that line; red-proofed against the pre-fix template. Origin: `audits/AUDIT-vacation-remote-ops-2026-07-20.md` (F7) |
## P3 — post-alpha