diff --git a/REPORT-i18n-closing.md b/REPORT-i18n-closing.md index 616402ea..8ec0bd6c 100644 --- a/REPORT-i18n-closing.md +++ b/REPORT-i18n-closing.md @@ -15,7 +15,7 @@ Written as `REPORT-.md` because `REPORT.md` is shared in this repo. | "the mail says 'three words' … `strings.Count(code,"-")+1`" | **No claim mail states a count.** They say `Setup code: %s`. The only count wording in the product was the **bind page's** passphrase hint ("five words"); its English half is now count-free, Hungarian unchanged. | | "§8's phone-safe filter: no two words differing by one letter in the first six" | **Measured, then declined.** It removes **5270 of 7772** words — 68%, 12.92 → 11.29 bits/word — and would make this list stricter than the one the product already uses for the code a household writes on paper during a disaster. Reason and measurement recorded in source; **operator may reverse.** Replaced by an assertion: every word is 3–9 lower-case ASCII letters, no digit, no separator. | | "request a reset code for **the demo customer (`en`)**" | **There is no English customer on this hub.** All five are `hu`. A scratch customer was created, proven, and deleted. | -| "demo-hp guest 9201" | **Guest 9201 is on `felhom-pve`**, and **demo-hp is offline** on every route tried (R-601). | +| "demo-hp guest 9201" | **Guest 9201 is on `felhom-pve`.** I also claimed demo-hp was offline — **that was MY error, withdrawn the same day (R-601)**: the box had been up four and a half weeks and reporting; both of my SSH routes pointed at stale addresses. | | "`customer.language` reaches the anonymous claim page" | **TRUE**, verified at source before any edit and now **pinned by a test** rather than assumed. | | "the box checks a hash and needs no change" | **TRUE**, and pinned by `TestClaimAcceptsAnEnglishWordCode`. | | "29 633 words"; the line numbers | **Right.** (29 634 lines, 29 609 after dedup.) Every cited line number was accurate. | @@ -86,9 +86,16 @@ after the fix), with seven decoys; `05-hub-architecture.md` §15.6; `10-localisa ## 5. Rows **Closed:** R-596, R-597, R-598 — each with what it actually turned out to be, not just "fixed". -**Opened:** R-601 (demo-hp unreachable — needs a person, it is a physical machine), R-602 (a live -probe that uses a cookie on a signed-in page reports a fixed defect as unfixed), R-603 (an English -string with an apostrophe silently never matches a rendered page). +**Opened:** R-602 (a live probe that uses a cookie on a signed-in page reports a fixed defect as +unfixed), R-603 (an English string with an apostrophe silently never matches a rendered page), +**R-604 (a per-customer floor override silently excludes a box from every global raise — demo-hp had +missed four)**. +**Withdrawn as false the same day:** R-601 ("demo-hp is unreachable"). The operator looked at the hub +and said it was online; it was, and had been for four and a half weeks. Both of my routes pointed at +stale addresses — one at a tailnet peer for a box with no tailscale installed, one at an address the +box left behind at a reprovision. **The hub had carried the right address in every report.** The +lesson kept in the row: the standing rule says a "no access" claim must list what was tried; it does +not say the list makes the claim true. Six failures against one wrong assumption is one failure. --- @@ -102,4 +109,11 @@ a stranger on a fresh install**, and this project's own rule, written into the r is that **fixes are not a journey**. The next English walk is what turns this into a green row; it is also the walk that would exercise the two backup warnings, and it wants a one-drive machine. -**Needs the operator, still:** the fleet floor. It now stands at 0.257.0 against a released 0.259.0. +**The fleet floor is raised to 0.259.0** (operator asked, same session), `min_agent` 0.131.0 +declared — above the vouched golden 0.258.0, so the declaration carries it (R-472). **Both live boxes +run 0.259.0.** demo-hp took it **by itself in under four minutes** once its stale per-customer +override was cleared, and its claim page then answered **"Wrong or expired code"** in English — the +floor delivered the FIX to a box nobody hand-deployed, which is the only thing that shows a raise +worked. Evidence: `audits/i18n-closing-2026-09-21/floor-raise-0.259.0.md`. + +**Needs the operator: nothing from this session.** diff --git a/STATUS.md b/STATUS.md index 75fcff39..1cd31774 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,6 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-21 — the one screen that stopped an English speaker is fixed, and I watched it work.** +**Updated 2026-09-21 — the one screen that stopped an English speaker is fixed, the floor is raised, and both boxes have it.** > **Ready for an English-speaking tester: yes — nothing known now stands in their way.** > Ready for a Hungarian volunteer: yes, unchanged. @@ -34,15 +34,28 @@ stranger on a brand-new machine** — that is the walk that turns this from "not into "someone did it". It is worth doing before you hand a box to a real English tester, and it is also the walk that would show the two Backup warnings on a screen rather than in a test. -**Also found, needs you.** **The HP box is switched off or unplugged** — it answers on nothing I can -reach, and the network says it has been dark for a while. It is a physical machine, so it needs a -person. Nothing depends on it today; it just means we have one demo box running, not two. +**The floor is raised — both boxes have the fixes.** You asked, so I did it. Both machines now run +today's version, and the second one **updated itself** in under four minutes with nothing installed +by hand. I checked its screen afterwards: it answers *"Wrong or expired code"* in English. That is +the proof that matters — the fix reached a box I never touched. -**Rows.** Three closed, three opened. +**I was wrong about the HP box, and you were right.** I said it was switched off. It was not: it has +been running for four and a half weeks and reporting all along. **Both of my ways in were pointing at +old addresses.** The box moved to a new address on your network, and the other route goes through a +service that is not even installed on it. I fixed both, checked both, and corrected our notes. The +lesson I am keeping: I tried six ways in and all six failed, but they were six goes at one wrong +assumption, not six pieces of proof. -**Needs you — one decision, no rush.** Whether to raise the box floor to today's version (0.259.0). -Everything already works without it; it only means every box gets these fixes instead of just the -demo one. This is the same decision that was waiting yesterday, now one version further on. +**And that hunt found something real.** The HP box had a private version limit set on it during a +test five days ago, and nobody removed it. It quietly **overrode the fleet setting** — so that box +has been missing the last **four** upgrades, and nothing anywhere said so. The alarm only speaks when +something changes, so a box stuck behind a private limit stays silent for ever. I cleared it, and +that is when the box finally took today's version. Worth fixing properly: raising the fleet version +should tell you which machines it will **not** move. + +**Rows.** Three closed, one withdrawn as wrong, four opened. + +**Needs you — nothing.** The floor is done. **If you do nothing:** nothing breaks. diff --git a/documentation/audits/i18n-closing-2026-09-21/floor-raise-0.259.0.md b/documentation/audits/i18n-closing-2026-09-21/floor-raise-0.259.0.md new file mode 100644 index 00000000..d8bb8331 --- /dev/null +++ b/documentation/audits/i18n-closing-2026-09-21/floor-raise-0.259.0.md @@ -0,0 +1,73 @@ +# Floor raise to 0.259.0 — and the box it nearly missed + +**2026-09-21, at the operator's request.** Global controller floor `0.257.0` → **`0.259.0`**, +`min_agent` **`0.131.0`** declared alongside it (R-472: the floor is above the vouched golden +0.258.0, so the declaration is what carries it). + +## What the hub logged + +``` +08:58:36 Global controller-version floor set to "0.259.0" (declared MinAgent "0.131.0") +08:58:38 managed floor SERVED for demo-felhom: floor 0.259.0, agent requirement "0.131.0" + from declared (golden 0.258.0) +``` + +**And nothing for demo-hp** — which reported one second later, at 08:58:39, and stayed on 0.258.0. + +## Why: a per-customer override nobody could see (R-604) + +`customer_configs.min_controller_version` for demo-hp held **`0.243.0`** — a per-customer floor that +wins over the global one. It is a leftover from the **2026-09-16 drill**, whose golden was 0.243.0. + +**R-343 measured on 2026-08-18 that all five rows were EMPTY** and recorded that as a safety +property: *"zero — all five `customer_configs` rows carry an empty `min_controller_version`, so +nothing hides behind a lower override."* It stopped being true and nothing surfaced the change. +demo-hp had silently missed the raises to **0.253.0, 0.254.0, 0.257.0 and 0.259.0**. + +**The silence is structural, not accidental.** `managed floor SERVED` fires once per *change* +(`h.floorNotes`, `hub/internal/api/handler.go:600`) — deliberately, because a box reports every few +minutes. So a box whose override never changes is silent **for ever**, and that silence looks exactly +like the silence of a box that already logged its line. A session raises the floor, reads one SERVED +line, and reasonably concludes the fleet took it. + +## Cleared, and the result + +Override cleared (**rollback line:** POST `/customers/demo-hp/floor` with +`min_controller_version=0.243.0`, `min_agent=0.131.0`): + +``` +09:11:02 Customer demo-hp controller-version floor override set to "" (declared MinAgent "") +09:11:05 managed floor SERVED for demo-hp: floor 0.259.0, agent requirement "0.131.0" + from declared (golden 0.258.0) +``` + +demo-hp then updated **itself**, in under four minutes, with nothing deployed by hand: + +``` +07:11:18 UTC [selfupdate] SetFloor: floor "" → "0.259.0" +07:11:18 UTC [selfupdate] maybeAutoUpdate: current 0.259.0 >= floor 0.259.0 — no action +07:11:22 UTC [offsite-apply] settle-gate: GO — at/above floor 0.259.0 (we are 0.259.0) +``` + +**The floor delivered the FIX, not a version string.** On demo-hp's claim page, reached at the +controller's own address with its `Host` header and the `felhom_lang=en` cookie: + +``` +http=200 alert-error">Wrong or expired code +``` + +That is the exact screen the 2026-09-20 English drill stopped on, now in English **on a box nobody +hand-deployed** — which is the only thing that shows a floor raise did its job. + +## Fleet after + +| customer | controller | last report | override | +|---|---|---|---| +| demo-felhom | **0.259.0** | 2026-09-21 07:05 | none | +| demo-hp | **0.259.0** | 2026-09-21 07:11 | none (cleared today) | +| tester-1 | 0.245.0 | 2026-09-17 — **down** | none | +| drill-r50 | 0.213.0 | 2026-08-12 — **down** | none | +| peti-felhom | 0.115.0 | 2026-07-15 — host row deleted, cannot receive a floor | none | + +**All five overrides are now empty**, which restores the property R-343 recorded. The three boxes +that are down take the floor unattended if they ever return — untested on this version. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 21d87451..2cfb4fff 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -776,7 +776,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-598** | **[P2-MEDIUM] The Backup page's two protection warnings — the ones that say whether the household's files are safe — are Hungarian on an English dashboard.** FOUND 2026-09-20 by the slice-6 drill on a fresh box, and confirmed on the demo box. Of 73 lines on `/backups` exactly four are Hungarian: **„Csak egy másolat készül (nincs második meghajtó) — a 3-2-1 mentéshez csatlakoztasson egy második meghajtót vagy offsite tárolót"**, **„A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem"**, and the two backup-target names **„Helyi tároló (local)"** and **„Biztonsági szerver – külön hardver (PBS)"**. They come from `internal/web/backup_handlers.go` (12 Hungarian literals) and `internal/web/backup_target_offer.go` (10) — again composed sentences handed to the page, the R-573/R-596 shape. **It matters more than its line count:** those two warnings are the only place the product tells a household that one copy on one disk is not protection, and the volunteer guide's §9 sends every tester to exactly this page to read exactly these two sentences. **Fix shape:** keys + args for both files, with the retrieval-promise gate run over the English (these sentences are about what a backup does and does not protect). **CLOSED 2026-09-21, controller v0.259.0.** The row's count of `backup_handlers.go` was 12; **nine are code and three are Hungarian inside COMMENTS**. The offer file's ten is right. `degradedMessageFor` now returns a **KEY** — the decision stays language-free and in one place, the words are chosen by the caller that knows the reader — and `buildTierViews` / `backupTargetLabel` / `loadGuestBackup` take the language the way `buildDataPathCards` already did. **The English is asserted to carry the same NEGATION the Hungarian does** ("protects against corrupted files, **but not** against a disk failure"); an English sentence that promised disk-failure protection would be worse than leaving it Hungarian. **Proven LIVE on guest 9201** for the two tier names ("Local storage (felhom-backup)", "Backup server – separate hardware (PBS)"); **the two warnings themselves were NOT walked live** — that box is healthy and a healthy box renders nothing by design, and producing the state would mean un-assigning a live backup target. They are covered by render tests through the real handler in both states. **An apostrophe cost a render:** the first English absent-drive sentence never matched because `html/template` escapes `'` to `'` — caught by the test, not by review. | **CLOSED 2026-09-21 — controller v0.259.0; the two warnings proven by render test, not live** | | **R-599** | **[P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so.** FOUND 2026-09-20 tearing the slice-6 drill down. VM destroyed at 17:12Z; the hub then refused **both** `POST /configs//delete` (409, *host … is ONLINE*) and the host delete (`deletable:false`) — correctly, because an online host would receive permanent 401s. But "online" is not a liveness probe: it is a **report-staleness window**, and the window is **45 minutes** — `manifests/hub.yaml` sets `alerting.stale_threshold: "45m"`, which `hostStatus()` reads (`ok` under it, `stale` over, `down` at 2x). A machine that no longer exists therefore reads ONLINE for three quarters of an hour. **Measured the boring way, and worth recording:** this row first said 30 minutes, because `monitor/host_staleness.go`'s literal default says 30m — the DEPLOYED value is in the manifest, and the 409s kept coming after the half hour was up. Reading a default and calling it the live value is the same mistake in a smaller coat. **The consequence is not theoretical:** a session that destroys its VM and then tears down the hub side walks away believing the delete failed, or leaves the customer behind — and the 2026-09-14 drill's teardown had the same shape without recording this. **Fix shape (smallest first):** the 409 body says *how long* it will refuse ("the last report was N minutes ago; deletion opens at HH:MM"), and `runbooks/target-selection.md`'s drill section names the wait. A force flag is NOT proposed — the refusal is right, only silent about its own clock. | **READY — rank P3-LOW; owner: CC (hub)** | | **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. | **READY - rank P2-MEDIUM; owner: CC (hub)** | -| **R-601** | **[P2-MEDIUM] `demo-hp` has been unreachable since at least this morning — offline on the tailnet for 30 days by Tailscale's own count, and not on the LAN either.** FOUND 2026-09-21 while looking for guest 9201 to ship controller v0.259.0. **What was tried, in order:** `ssh demo-hp` (tailnet `100.76.96.79`) → connection timed out; `sshpass` with the hub-vaulted G1 break-glass password → same timeout; `demo-hp-lan` (`192.168.0.87` via `ProxyJump felhom-pve`) → *No route to host*; `ping 100.76.96.79` → 100% loss; `tailscale status` from the DooPlex pod → **`demo-hp … offline, last seen 30d ago, tx 6084 rx 0`**; `ip neigh` on felhom-pve → `192.168.0.87 … FAILED`. The box answers on no route this session has. **The 30-day figure is Tailscale's and should not be believed on its own** — the slice-6 drill ran its nested drill VM on demo-hp on 2026-09-20 and tore it down, which is not consistent with a box that has been dark for a month; the likelier reading is that its tailscale link has been down for 30 days while the box was reached another way, and the box itself went down more recently. **Either way it is off now.** It also means the fleet's 0.258.0 box is the one that is dark and `demo-felhom` (0.257.0 until today) is the one that is up — so **the fleet is one box, not two, until someone looks at it.** Needs a person: it is a physical machine. | **READY — rank P2-MEDIUM; owner: OPERATOR (physical)** | +| **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp` to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran `ip neigh` on felhom-pve, and `192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** | +| **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | | **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** | | **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | diff --git a/documentation/operations/nodes.md b/documentation/operations/nodes.md index e9fc7796..28f8d610 100644 --- a/documentation/operations/nodes.md +++ b/documentation/operations/nodes.md @@ -16,7 +16,7 @@ | Agent / controller | *not recorded here* — see the note below | *not recorded here* | | Control plane | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) | **island** `169.254.253.1:8443` on `vmbr9` (R-50, migrated 2026-07-25) | | SSH alias | `felhom-pve` | **`demo-hp`** | -| Tailnet | `100.70.170.35` | **`100.76.96.79`** | +| Tailnet | `100.70.170.35` | **NONE — tailscale is NOT INSTALLED on demo-hp** (checked on the box 2026-09-21: no `tailscaled`, no `tailscale` binary). The peer `100.76.96.79` still listed by `tailscale status` is a STALE entry from an earlier build and **can never answer**; it showed `offline, last seen 30d ago` while the box was up and reporting. Use the LAN address. | | Loader used to install | `mkimage` (unsigned, SB **off** — firmware workaround) | **`shim`, Secure Boot ENABLED** | > **No component versions are recorded on this page — deliberately.** Agent, controller, hub and @@ -81,10 +81,10 @@ installer baked its `192.168.100.2` fallback as a **static** `vmbr0` address and that looked installed and could never call home. Repaired on the console by bridging `vmbr0` to `enp2s0f0`. Filed as **R-59** (must hard-abort) and **R-60** (first-boot NIC sweep self-heal). -Current, post-repair: `vmbr0` static `192.168.0.87/24`, gw `192.168.0.1`, bridge-port `enp2s0f0`. +Current, read off the box 2026-09-21: `vmbr0` **static `192.168.0.104/24`**, gw `192.168.0.1`, bridge-port **`nic0`**. (It was `192.168.0.87/24` on `enp2s0f0` before a reprovision; both were stale here for long enough to cost a session a false "the box is down" — R-601.) **The hub always knows the truth:** every host report carries `addresses: [{iface, cidr}, …]`, so read it there rather than from this page. No trace of `192.168.100.2` remains. `wg-felhom` `10.77.0.3/32` up to the hub. Guest **9201 `demo-hp`** running. Agent config shape (R-50 island): `local_api` on `169.254.253.1:8443`/`vmbr9`, -guest `eth1 169.254.253.2/30`, `lan_resolver.host_ip` pinned to `192.168.0.87`. +guest `eth1 169.254.253.2/30`, `lan_resolver.host_ip` pinned to the host's LAN address. ### Designated drill + build VM host (operator ruling, 2026-07-25) @@ -139,7 +139,7 @@ Then `sshpass -e ssh root@demo-hp` (sshpass is on DooPlex, not on the nodes). the plaintext, so the console is unreachable without a working hub and network — precisely what you may be trying to fix. Slice 1 is to emit the baked password into the build report. -`demo-hp-lan` (`192.168.0.87` via `ProxyJump felhom-pve`) is the fallback while the box is away. +`demo-hp-lan` (**`192.168.0.104`** via `ProxyJump felhom-pve`) is the fallback when DooPlex is not on the home LAN. **`ssh demo-hp` now goes DIRECT to `192.168.0.104`** — DooPlex is on the same LAN, and the tailnet route for this box does not exist (see the table above). Both were verified 2026-09-21. ## OOB belt (H1) — both boxes, since 2026-07-23 (ISO train v1.25.0)