diff --git a/documentation/audits/evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt b/documentation/audits/evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt index 82863928..494dbd36 100644 --- a/documentation/audits/evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt +++ b/documentation/audits/evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt @@ -76,3 +76,14 @@ ## „Indítás 14 másodperc múlva..." (the 15 s countdown, running) ## No Proxmox entry, no automated/answer-file entry, no admin URL. This is the proof-install half of ## the release gate beginning; the install itself follows on the same image. + +## PUBLISHED 2026-09-16T18:2xZ, on the operator's explicit yes (the one STOP of this task). +## uploaded: felhom-installer-1.28.0-pve9.2-1.iso (1 705 322 496 B) to the private bucket via the +## env-only rclone container (no credential file is ever written; the credentials were read +## file→file, never `source`d — sourcing that file leaks hyphenated keys as shell errors). +## ROUND TRIP — the published bytes, not the local file: +## downloaded 1 705 322 496 B, sha256 a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635 +## built 1 705 322 496 B, sha256 a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635 +## checksum file beside it: HTTP 200 +## The bytes a stranger receives are the bytes that were gate-checked and installed tonight. +## 1.27.1 stays in the bucket; nothing was overwritten or removed. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index e35253cc..7c25f4c1 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -719,7 +719,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-522** | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-524** | **[P2-MEDIUM] When the catalog moves an app back to an older version, a box that already updated shows „Frissítés elérhető" — and the offered Update is a downgrade.** MEASURED 2026-09-15 (BIGNIGHT Phase 6, VM 333): privatebin was updated 2.0.5 → 2.0.6 through the guarded Update after the drill bump; the catalog was then reverted to 2.0.5 (`a161ccb`). At 22:13:37Z the box reads `installed privatebin/pdo:2.0.6`, `catalog privatebin/pdo:2.0.5`, `catalog_since 2026-09-14`, and the app page tag „**Frissítés elérhető — ma**" with the title „Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot." The label compares for *difference*, not for *newer* (`09-update-architecture.md` §5.4 render table); the guarded Update would advance the pin „to the catalog's current definition" — 2.0.6 → 2.0.5. The same state follows any real upstream yank. **Not pressed tonight.** **Fix shape:** compare versions (or `catalog_since` against the installed record) and render „Naprakész" / „a katalógusnál újabb" when the box is ahead; refuse a pin move to an older tag without an operator word. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **R-534** | **[P1-HIGH] The off-site tier cannot be provisioned or adopted for a rebuilt box: the hub's endpoint token lacks `Datastore.Modify`, so every re-issue fails.** MEASURED 2026-09-16 on the drill box (`tester-1-652049`, fresh install from published ISO 1.27.1): the WG-registration hook refused as R-511 describes („the endpoint already holds a PBS token … use the explicit Re-issue PBS credentials action"); the operator pressed exactly that, hub v0.114.0's ADOPT path ran, and the endpoint answered **`Process exited with status 255 (stderr: Error: permission check failed - missing Datastore.Modify on /datastore/felhom-offsite)`** → HTTP 502, no descriptor written, no secret stored (fail-closed, correct). So the code fix of 2026-09-15 is sound and INERT: a rebuilt box has no whole-guest off-site tier, and the customer's „Távoli rendszermentés" stays absent. **Not run by hand on ep0** (fenced). **Fix shape (operator):** grant the hub's tenantsync user `Datastore.Modify` on `/datastore/felhom-offsite` (it already holds the create/delete grants the provision path uses), or give the script a token-only re-key op that needs no Modify. Evidence: `audits/evidence-drill-0243-2026-09-16/phase1-pbsdr.txt`. **THE GRANT IS GIVEN 2026-09-16, and the narrowest role was MEASURED rather than recalled:** `DatastorePowerUser` carries Datastore.Backup + Datastore.Prune only, so it does not help; PBS has no role-create command and no custom roles, so the narrowest role that carries `Datastore.Modify` is `DatastoreAdmin`. Applied for the hub's `felhom@pbs` on `/datastore/felhom-offsite` ONLY; the per-customer `DatastoreBackup` entries are untouched and nothing else on ep0 changed. Effective permissions after: Audit, Backup, Modify, Prune, Read, Verify at that path. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt`. **CLOSED 2026-09-16 — the grant works, proven END TO END on a fresh box.** After the narrow grant (DatastoreAdmin for the hub's `felhom@pbs` on `/datastore/felhom-offsite` only — `DatastorePowerUser` was measured to carry Backup+Prune and PBS has no custom roles), a newly installed box for the same rebuilt customer hit the very refusal this row describes, by itself: „pbsdr auto-provision … the endpoint already holds a PBS token for tester-1 but the hub has no descriptor — use the explicit Re-issue PBS credentials action" (19:01 CEST). Pressing that action then SUCCEEDED: „tenantsync: reissue ok … token_id=felhom@pbs!tester-1", „pbsdr ADOPTED for tester-1 (host tester-1-33b6a9, gen 2; fresh consume-once secret stored)", „pbs token secret consumed by host … (single-use)" — no permission error. This morning the identical action returned „missing Datastore.Modify … status 255" → 502. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseC-ep0-grant.txt` and `phaseC-reissue.txt`. | **CLOSED 2026-09-16 — grant given and proven end to end** | -| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom. címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. | **READY — rank P2-MEDIUM; owner: CC (ISO payload `felhom-bootstrap.sh` — ships with the next ISO)** | +| **R-535** | **[P2-MEDIUM] The box's console still says „a doboz készen áll, és a párosításra vár" long after the box is bound, claimed and running apps — and it promises that the screen refreshes itself.** MEASURED 2026-09-16 on the drill box (fresh install, ISO 1.27.1, controller 0.243.0): bind succeeded 10:01:14Z, claim 10:06:18Z, four apps deploying by 10:22Z — and at 10:23Z the console still showed the pairing banner with the code `37S-NFE` and the line „Ez a képernyő magától frissül — nincs teendő a doboznál" (`audits/evidence-drill-0243-2026-09-16/screens/33-console-after-claim.png`). A volunteer watching the monitor has no way to tell the box is finished; worse, the screen says it updates itself, so waiting longer does not help. **Fix shape:** the first-boot banner unit re-renders on claim/bind state (the controller already knows: it reports `controller_started` and the hub holds `claimed`), showing „A doboz össze van kötve — a vezérlőpult a https://felhom. címen érhető el"; or at minimum stop printing the pairing code once the appliance is claimed. **CLOSED 2026-09-16 — shipped in ISO 1.28.0, published the same day on the operator's explicit yes.** `felhom-bootstrap.sh` prints `print_bound_banner` the moment the bind delivery lands: „a doboz össze van kötve", „a beállítás magától folytatódik", „ezen a gépen nincs több teendőd" — replacing the pairing code on the console. The payload in the published image is byte-identical to repo HEAD and the string is present in it; the boot menu and the install were walked on that exact file. **What it deliberately does NOT do, recorded rather than implied away:** it does not name the dashboard URL (the one-shot bind delivery carries the customer id, passphrase and mode — not the domain), and it does not reflect the later CLAIM, because this unit has exited by then. **And the new banner was never SEEN on a screen** — the box bound itself while the walk was driving it headlessly, so the proof is the shipped payload plus the gate, not a photograph. | **CLOSED 2026-09-16 — shipped in ISO 1.28.0 (published); on-screen effect not photographed** | | **R-536** | **[P2-MEDIUM] The hub is told „Alkalmazás telepítve" the moment a deploy is ACCEPTED, so an install that never finishes is recorded as a completed one.** MEASURED 2026-09-16 on the drill box: the deploy of `mealie` was accepted at 12:31:36 CEST and the hub logged `Event from tester-1: app_deployed (info) — Alkalmazás telepítve: Mealie` in the SAME second; the controller was then killed 5 s in (F9'), and after the agent restarted it the stack read `not_deployed / deployed=false / deploying=false` — i.e. the app was never installed, and nothing corrected the event. Source confirms the ordering: `internal/api/router.go` writes the 202 „Telepítés elindítva" and then calls `NotifyAppDeployed` immediately, while the comment right above it says the deploy „runs asynchronously (compose pull/up + health happen after this returns)". The event is `info`, so nobody is mailed — but the customer timeline and the hub's app history record a completed install that did not happen (the „presence is not success" class). **Also measured, same shape:** an interrupted deploy leaves `/opt/docker/stacks//app.yaml` behind (written at accept) while the stack reads not-deployed — second instance after 2026-09-15's homebox; moved aside on the box. **Fix shape:** emit `app_deployed` from the async path when the stack reaches running/healthy (or emit `app_deploy_started` at accept and `app_deployed` at completion), and remove the accept-time `app.yaml` on a failed deploy. **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0.** `app_deploy_started` is emitted beside the 202; `app_deployed` now fires from the async path's own end, and `app_deploy_failed` (warning) replaces the silence an interrupted install used to get. Both new types are registered in `allowedEventTypes` AND `customerMessages`. The accept-time `app.yaml` is deliberately NOT deleted on failure — it is the crash-safe record with `Deployed:false` and it holds the settings the customer typed; the state every surface reads is `not_deployed`. Red-proofs: the accept-time call back → `TestDeployAcceptance_DoesNotClaimTheAppIsInstalled` fails; the success hook removed → `TestDeployDoneHook_...` fails at „the deploy ended and nothing was told about it". | **CLOSED 2026-09-16 — controller v0.244.0 + hub v0.116.0** | | **R-539** | **[P3-LOW] The restart brake catches a FAST crash loop and is blind to a SLOW one — add a second, slower counter (operator ruling 2026-09-16).** MEASURED 2026-09-16 (R-531): four controller restarts 20 minutes apart, none accumulating, because the budget window is 15 minutes; the only trace is an `info` `controller_restarted_by_agent` event, which mails nobody. A box whose controller dies every 20 minutes is restarted forever and nothing tells the operator. **The ruling:** a second counter — N restarts in 24 h → `controller_slow_crashloop` (warning) — kept beside the existing 3-in-15-minutes brake, which is unchanged. **Not built in the 2026-09-16 task** (its brief said the budget is a design and this task measures it); it is the nightly's to build. Needs: the agent-side counter, the new event type in `allowedEventTypes` + `customerMessages`, and a red-proof that a 20-minute cycle raises it while a healthy box never does. | **READY — rank P3-LOW; owner: CC (agent + hub)** | | **R-540** | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** | diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index 6764063d..b34a28f3 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,4 +1,4 @@ -## ISO v1.28.0 — the console stops showing the pairing code once the box is connected (2026-09-16, R-535) — NOT PUBLISHED +## ISO v1.28.0 — the console stops showing the pairing code once the box is connected (2026-09-16, R-535) — PUBLISHED 2026-09-16 **The defect, measured on a fresh box 2026-09-16:** 25 minutes after a successful bind AND claim, with four apps deploying, the physical console still read „a doboz készen áll, és a párosításra vár" with @@ -16,7 +16,7 @@ waiting will update it. claim-aware console needs a different owner. R-535 is closed for the measured complaint — the code stays on screen after binding — and that residue is recorded rather than implied away. -## ISO v1.27.1 — the FIRST boot is Felhom's too (2026-09-14, R-496) — NOT PUBLISHED +## ISO v1.27.1 — the FIRST boot is Felhom's too (2026-09-14, R-496) — PUBLISHED 2026-09-15 (the marker read NOT PUBLISHED until 2026-09-16; it was published on the big night and the heading was never corrected) **Why 1.27.0 was not enough, measured on its proof install (VM 331, screen s20):** `pvebanner.service` ran at the first boot before `felhom-bootstrap` could mask it, so the household's first screen still carried diff --git a/website/letoltes.html b/website/letoltes.html index 1611f54a..8ae927ca 100644 --- a/website/letoltes.html +++ b/website/letoltes.html @@ -74,16 +74,16 @@

Letöltés

-

Felhom telepítő 1.27.1 (Proxmox VE 9.2 alapon)

+

Felhom telepítő 1.28.0 (Proxmox VE 9.2 alapon)

-

felhom-installer-1.27.1-pve9.2-1.iso — 1,6 GB (1 705 322 496 bájt)

+

felhom-installer-1.28.0-pve9.2-1.iso — 1,6 GB (1 705 322 496 bájt)

Ellenőrző összeg (SHA-256):

-

25637007d5a7120ff9faa6b5b7ead3e33c0a361ac2d67e9fd4e0ee77c034c053

+

a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635

-

Ugyanez letölthető fájlként is: felhom-installer-1.27.1-pve9.2-1.iso.sha256

-

Ellenőrzés Linuxon vagy Macen: sha256sum felhom-installer-1.27.1-pve9.2-1.iso. Windowson PowerShellben: Get-FileHash felhom-installer-1.27.1-pve9.2-1.iso. A kiírt számsornak meg kell egyeznie a fentivel. Ha nem egyezik, ne használd a fájlt, és töltsd le újra.

+

Ugyanez letölthető fájlként is: felhom-installer-1.28.0-pve9.2-1.iso.sha256

+

Ellenőrzés Linuxon vagy Macen: sha256sum felhom-installer-1.28.0-pve9.2-1.iso. Windowson PowerShellben: Get-FileHash felhom-installer-1.28.0-pve9.2-1.iso. A kiírt számsornak meg kell egyeznie a fentivel. Ha nem egyezik, ne használd a fájlt, és töltsd le újra.

A letöltött fájlt egy legalább 4 GB-os USB kulcsra kell írni (például a Balena Etcher vagy a Rufus programmal). Az útmutató ezt is leírja.