From 20aafc3dec2008f8f4697129bcf95bd354edef20 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 18 Sep 2026 15:37:15 +0200 Subject: [PATCH] =?UTF-8?q?ops:=20fleet=20floor=20raised=20to=200.255.0=20?= =?UTF-8?q?=E2=80=94=20and=20it=20delivered=20on=20its=20own?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit POST /configuration/global-floor with min_controller_version=0.255.0 and the declared min_agent=0.131.0 (R-472); 303 flash=floor_set, read back from the form, not from the POST. The proof is the N100: it was never hand-deployed and its own Docker reports felhom-controller:0.255.0 healthy within five minutes of the save. Three boxes remain below — all BLOCKED or DOWN, which is a floor being held, not a floor failing; each takes it on its next check-in. R-580 filed: curl's %{redirect_url} rebuilds the request URL WITH the --netrc credentials in it, so the hub password was printed into the session's own output. Nothing written to a file, nothing committed. The build-deploy skill now carries the rule. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- STATUS.md | 17 +++++-- .../D/live/floor-0.255.0.md | 49 +++++++++++++++++++ documentation/backlog/OPEN-ITEMS.md | 1 + skills/felhom-build-deploy/SKILL.md | 5 ++ 4 files changed, 67 insertions(+), 5 deletions(-) create mode 100644 documentation/audits/i18n-slice2-2026-09-18/D/live/floor-0.255.0.md diff --git a/STATUS.md b/STATUS.md index d7d662e8..9d6cf621 100644 --- a/STATUS.md +++ b/STATUS.md @@ -26,11 +26,18 @@ browser on this machine. That part is your click. **Rows.** One opened and closed the same day. -**Needs you — one decision.** **Raise the floor to 0.255.0?** -- **If you do:** every box serves the fixed pages on its next check-in, and nobody sees the broken globe. -- **If you do nothing:** boxes on 0.254.0 keep showing it to anyone whose browser cached the old style - file. Nothing is at risk; it just looks wrong. -- I would raise it — the thing it fixes is the thing you noticed. +**The floor is raised — nothing needs you.** You asked for it, so I did it: every box now must run +0.255.0 or newer. **It delivered on its own:** the N100 was never touched by hand and picked up the +new version in under five minutes. The HP box was already on it. + +Three machines are still on older versions. All three are switched off or blocked, and a box that is +not running cannot be told anything. Each takes the new version the moment it comes back. That is the +floor working, not failing. + +**One new row, mine, nothing for you to do.** While pressing the button I used a command option that +printed your hub password back into my own session log. It went nowhere else — no file, nothing +saved, nothing sent. I wrote down how to avoid it next time. **If you want to be thorough, change +that password**; nothing suggests you need to. --- diff --git a/documentation/audits/i18n-slice2-2026-09-18/D/live/floor-0.255.0.md b/documentation/audits/i18n-slice2-2026-09-18/D/live/floor-0.255.0.md new file mode 100644 index 00000000..3c45e18c --- /dev/null +++ b/documentation/audits/i18n-slice2-2026-09-18/D/live/floor-0.255.0.md @@ -0,0 +1,49 @@ +# Fleet floor raised to 0.255.0 — 2026-09-18 + +**Method: endpoint-level.** The operator UI's own form, driven with Basic auth against the hub's +ClusterIP. No hand-deploy around the floor (R-472). + +``` +before : min_controller_version = 0.254.0, declared MinAgent 0.131.0 +impact : POST /configuration/global-floor/impact?v=0.255.0 -> {"below":4} +save : POST /configuration/global-floor + min_controller_version=0.255.0 min_agent=0.131.0 + -> 303 /configuration?flash=floor_set +after : min_controller_version = 0.255.0 (re-read from the form, not from the POST) +``` + +## It delivered — the positive observable + +The N100 was NOT hand-deployed. It self-updated from the floor alone, and its own Docker says so: + +``` +demo-felhom guest 9201 : felhom-controller:0.255.0 Up 5 minutes (healthy) +demo-hp guest 9201 : felhom-controller:0.255.0 Up 26 minutes (healthy) [hand-deployed earlier] +``` + +The hub's config table agrees, from each box's own report: + +| customer | status | version | floor | +|---|---|---|---| +| demo-felhom | OK | **0.255.0** | v0.255.0 | +| demo-hp | OK | **0.255.0** | v0.243.0 (per-customer override) | +| drill-r50 | BLOCKED | 0.213.0 | v0.255.0 ● held | +| peti-felhom | DOWN | 0.115.0 | v0.255.0 ● held | +| tester-1 | DOWN | 0.245.0 | v0.255.0 ● held | + +## Three boxes are still below, and that is the floor working, not failing + +`below` went 4 -> 3 and stopped. The three are BLOCKED or DOWN: a floor is delivered on a box's +next check-in, and a box that never checks in never receives it. **`below: 3` is therefore not a +delivery failure** — it is the count of machines that have not reported since the save. Each will +take 0.255.0 on its next report. The one that WAS reachable took it in under five minutes. + +**demo-hp's own floor is a per-customer override at v0.243.0**, so the global save does not move it; +it is on 0.255.0 because this session deployed it by hand. Counting it as proof of delivery would be +the mistake — the N100 is the proof. + +## A finding, filed: the hub password was printed back by curl + +`curl -w '%{redirect_url}'` re-injects the credentials into the URL it prints, so the hub password +appeared in this session's own output. Nothing was written to a file and nothing was committed. +Filed as **R-580**. Never pair `%{redirect_url}` with `--netrc`/`-u`; print `%{http_code}` alone. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 1e2e4e04..9556d0bc 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -755,6 +755,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-577** | **[P3-LOW] A guest SHARE visitor has no way to pick a language, and the household's setting is the wrong default for them.** FOUND 2026-09-18 by localisation slice 2 release C (R-557, controller v0.254.0): every other page a person can reach now carries a language globe — the dashboard (the household's setting), and the sign-in and claim pages (the visitor's own cookie). The two guest share pages (`launcher_shared`, `launcher_share_password`) deliberately do NOT, and `TestGuestSharePagesHaveNoGlobe` pins that so it stays a decision rather than an oversight. **Why it is the operator's and not CC's:** a share visitor is a stranger the household sent a link to, and what language they are shown is a promise the SHARE FEATURE makes, not an implementation detail. The `felhom_lang` cookie already built would fit them exactly (display-only, their own browser, never the household's setting). **Fix shape, if the operator says yes:** add `{{template "lang_globe" .}}` to both shells with the anonymous form, and one render case per page per language. | **READY - rank P3-LOW; owner: operator (the decision), CC (the change)** | | **R-578** | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | | **R-579** | **[P3-LOW] Five page shells loaded `style.css` with NO cache-buster, so a browser holding an older copy kept being served CSS that did not know about the newest UI.** FOUND 2026-09-18 by the operator's screenshot of the v0.254.0 language globe: it rendered as a bare, unstyled `
` — a stray triangle and two plain words outside the card — because `login.html`, `claim.html`, `recovery.html`, `launcher_shared.html` and `launcher_share_password.html` requested `/static/style.css` with no `?v=`, while `layout.html` has used `?v={{.Version}}` since v0.166.0. **`.Version` was also absent from three of those five data maps.** FIXED AND CLOSED in the same session (controller v0.255.0): the parameter on all five, and `Version` set once in `executeTemplateLang` so a new shell cannot miss it; `TestGlobeOnAnonymousShells` now refuses an absent or EMPTY `?v=`. **The general form, which is the part worth keeping: a template that loads a versioned asset WITHOUT its version is invisible to every test that reads markup — the markup is correct and the browser fetches the wrong file.** A gate over "every stylesheet/script link in a template carries `?v=`" would catch the class; not built, because 8 first-boot-wizard templates would fail it and R-554 deletes them. | **CLOSED 2026-09-18 - controller v0.255.0** | +| **R-580** | **[P3-LOW] `curl -w '%{redirect_url}'` prints Basic-auth credentials back into the session transcript.** FOUND 2026-09-18 raising the fleet floor to 0.255.0: the hub's `/configuration/global-floor` form was driven with `--netrc-file` and `-w 'POST http=%{http_code} -> %{redirect_url}'` to read the `flash=floor_set` confirmation. curl re-injects the credentials into the redirect URL it formats, so the operator's hub password appeared in the session's own output. **Nothing was written to a file and nothing was committed** — the exposure is the transcript. **The general form, which is the part worth keeping: a curl format variable can carry a secret that never appeared in the command.** `%{redirect_url}`, `%{url_effective}` and `%{referer}` all reconstruct the request URL, and `--netrc`/`-u` credentials go back into it. **Fix shape:** print `%{http_code}` alone and read the destination from `-D-` headers, or pipe the format output through a redactor. Recorded in the `felhom-build-deploy` skill so the next hub session does not repeat it. | **READY - rank P3-LOW; owner: CC** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-538** | **[P1-HIGH] A tier-1 app restore reports plain success and leaves Nextcloud listing files whose bytes were never in the backup — and it destroys the app's own trash, the customer's last copy.** MEASURED 2026-09-16 on the drill box, F10 („a child deletes the photo folder"): the five photos were deleted through Nextcloud (DELETE 204, PROPFIND 404), then restored through the page exactly as a customer would (`POST /backup/restore` `stack_name=nextcloud` `snapshot_id=helyi` → 302, finished in **35 s**, „A(z) nextcloud: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult."). Afterwards the folder is back and **lists all five photos**, and **none of them opens**: `GET nyaralas-1..5` = 404 / 503×4 with `Sabre\DAV\Exception\NotFound`, while the positive controls at the same moment pass (`status.php` 200, WebDAV PUT 201, GET 200). Cause: the replayed MariaDB dump (11:01:45Z) knows the photos, the bytes live on `mp8` and were never captured (R-537). **Worse:** the bytes were still on the drive in Nextcloud's own trash (`appdata/nextcloud/admin/files_trashbin/files/Fotok.d1789556707/nyaralas-1..5.jpg`, all five present) and the restored database no longer references them — the trash listing comes back **empty**, so „restore from trash", the one route that would have worked, is gone. The customer is left with five unopenable photos, a success message, and no warning. **Fix shape:** before replaying a database whose app has an uncaptured file leg, refuse or warn („ennek az alkalmazásnak a fájljai nincsenek ebben a mentésben — a visszaállítás után a fájlok hiányozni fognak"); and never present a DB-only restore of a class-A app as a complete one. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** A unit restore refuses before anything is touched when the unit cannot return the app's drive-side files, and names the route that can. Fired live on demo-hp: `POST /backup/restore` for paperless-ngx → 302 with „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza …", and the app read `running` before AND after, so nothing was stopped and no trash was made unreachable. The database-and-settings-only path exists as a separately worded second step. Red-proof: disabling the guard fails `TestUnitRestore_RefusesWhenTheUnitCannotHoldTheFiles`. **RE-PROVEN 2026-09-16 on a FRESH box, and this time the refusal had somewhere to point:** after five photos were deleted, `POST /backup/restore` was refused with „…a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: … „Teljes visszaállítás (fájlok + adatbázis)"", the app read `running` before AND after, and the wastebasket was untouched. The off-site route then returned all five photos — 200 with the exact uploaded sizes and sha256 IDENTICAL to the originals, 5/5, with a negative control. Evidence: `audits/evidence-backup-promise-2026-09-16/phaseE-photos.txt`. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | | **R-525** | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** | diff --git a/skills/felhom-build-deploy/SKILL.md b/skills/felhom-build-deploy/SKILL.md index 6cdd2ab4..b904cd4c 100644 --- a/skills/felhom-build-deploy/SKILL.md +++ b/skills/felhom-build-deploy/SKILL.md @@ -130,6 +130,11 @@ curl -fsSL -o /tmp/rt.iso https://iso.felhom.eu/felhom-installer--pve. dashed focus ring, and `Enter` lands in text *fields*, not `Next`. Not checking once aborted an install. - **Proof installs register unclaimed appliances at the hub — discard them** or R-131 grows: `curl -u ":$HUB_PW" -X POST http://:8080/appliances//discard` → 303. + +> **Never pair `-w '%{redirect_url}'` with `-u` or `--netrc` (R-580).** curl rebuilds the request URL +> for that variable **with the credentials in it**, so the hub password is printed even though it +> never appeared in the command. `%{url_effective}` and `%{referer}` do the same. Print +> `%{http_code}` alone; read the destination from `-D-` headers if you need it. The verb is **`/discard`**, POST only (`hub/internal/web/server.go:345`); `/delete` 404s. - Venue: `demo-hp`, scratch `dir` storage at **`/mnt/nvme-1tb` root** (a subdirectory reads `disconnected` forever — the agent's `exactMount` check). Never `local-lvm`. Remove the storage at teardown.