R-601 said demo-hp was unreachable. The operator looked at the hub and said it
was online. It was, and had been up four and a half weeks, reporting every few
minutes. Both of my SSH routes pointed at stale addresses: `demo-hp` at a
tailnet peer for a box that has no tailscale installed at all, and `demo-hp-lan`
at 192.168.0.87 when the box is statically on .104 since a reprovision. The hub
had carried the right address in every host report, and `ip neigh` on felhom-pve
had .104 four lines above the .87 I quoted — I searched that output for the
address I expected instead of reading it for the address that was there.
Both ssh entries repointed and verified; nodes.md corrected, including that the
tailnet route for this box does not exist.
The hunt then found R-604, which is the real defect: demo-hp carried a
per-customer floor override of 0.243.0 left over from the 2026-09-16 drill, so
it had silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0.
`managed floor SERVED` fires once per change by design, so a box behind a static
override is silent for ever and its silence is indistinguishable from a box that
already logged. Cleared; demo-hp self-updated to 0.259.0 in under four minutes
and its claim page now answers "Wrong or expired code" in English.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
@@ -15,7 +15,7 @@ Written as `REPORT-<topic>.md` because `REPORT.md` is shared in this repo.
| "the mail says 'three words' … `strings.Count(code,"-")+1`" | **No claim mail states a count.** They say `Setup code: %s`. The only count wording in the product was the **bind page's** passphrase hint ("five words"); its English half is now count-free, Hungarian unchanged. |
| "§8's phone-safe filter: no two words differing by one letter in the first six" | **Measured, then declined.** It removes **5270 of 7772** words — 68%, 12.92 → 11.29 bits/word — and would make this list stricter than the one the product already uses for the code a household writes on paper during a disaster. Reason and measurement recorded in source; **operator may reverse.** Replaced by an assertion: every word is 3–9 lower-case ASCII letters, no digit, no separator. |
| "request a reset code for **the demo customer (`en`)**" | **There is no English customer on this hub.** All five are `hu`. A scratch customer was created, proven, and deleted. |
| "demo-hp guest 9201" | **Guest 9201 is on `felhom-pve`**, and **demo-hp is offline** on every route tried (R-601). |
| "demo-hp guest 9201" | **Guest 9201 is on `felhom-pve`.** I also claimed demo-hp was offline — **that was MY error, withdrawn the same day (R-601)**: the box had been up four and a half weeks and reporting; both of my SSH routes pointed at stale addresses. |
| "`customer.language` reaches the anonymous claim page" | **TRUE**, verified at source before any edit and now **pinned by a test** rather than assumed. |
| "the box checks a hash and needs no change" | **TRUE**, and pinned by `TestClaimAcceptsAnEnglishWordCode`. |
| "29 633 words"; the line numbers | **Right.** (29 634 lines, 29 609 after dedup.) Every cited line number was accurate. |
@@ -86,9 +86,16 @@ after the fix), with seven decoys; `05-hub-architecture.md` §15.6; `10-localisa
## 5. Rows
**Closed:** R-596, R-597, R-598 — each with what it actually turned out to be, not just "fixed".
**Opened:** R-601 (demo-hp unreachable — needs a person, it is a physical machine), R-602 (a live
probe that uses a cookie on a signed-in page reports a fixed defect as unfixed), R-603 (an English
string with an apostrophe silently never matches a rendered page).
**Opened:** R-602 (a live probe that uses a cookie on a signed-in page reports a fixed defect as
unfixed), R-603 (an English string with an apostrophe silently never matches a rendered page),
**R-604 (a per-customer floor override silently excludes a box from every global raise — demo-hp had
missed four)**.
**Withdrawn as false the same day:** R-601 ("demo-hp is unreachable"). The operator looked at the hub
and said it was online; it was, and had been for four and a half weeks. Both of my routes pointed at
stale addresses — one at a tailnet peer for a box with no tailscale installed, one at an address the
box left behind at a reprovision. **The hub had carried the right address in every report.** The
lesson kept in the row: the standing rule says a "no access" claim must list what was tried; it does
not say the list makes the claim true. Six failures against one wrong assumption is one failure.
---
@@ -102,4 +109,11 @@ a stranger on a fresh install**, and this project's own rule, written into the r
is that **fixes are not a journey**. The next English walk is what turns this into a green row; it is
also the walk that would exercise the two backup warnings, and it wants a one-drive machine.
**Needs the operator, still:** the fleet floor. It now stands at 0.257.0 against a released 0.259.0.
**The fleet floor is raised to 0.259.0** (operator asked, same session), `min_agent` 0.131.0
declared — above the vouched golden 0.258.0, so the declaration carries it (R-472). **Both live boxes
run 0.259.0.** demo-hp took it **by itself in under four minutes** once its stale per-customer
override was cleared, and its claim page then answered **"Wrong or expired code"** in English — the
floor delivered the FIX to a box nobody hand-deployed, which is the only thing that shows a raise
@@ -776,7 +776,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-598** | **[P2-MEDIUM] The Backup page's two protection warnings — the ones that say whether the household's files are safe — are Hungarian on an English dashboard.** FOUND 2026-09-20 by the slice-6 drill on a fresh box, and confirmed on the demo box. Of 73 lines on `/backups` exactly four are Hungarian: **„Csak egy másolat készül (nincs második meghajtó) — a 3-2-1 mentéshez csatlakoztasson egy második meghajtót vagy offsite tárolót"**, **„A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok ellen véd, lemezhiba ellen nem"**, and the two backup-target names **„Helyi tároló (local)"** and **„Biztonsági szerver – külön hardver (PBS)"**. They come from `internal/web/backup_handlers.go` (12 Hungarian literals) and `internal/web/backup_target_offer.go` (10) — again composed sentences handed to the page, the R-573/R-596 shape. **It matters more than its line count:** those two warnings are the only place the product tells a household that one copy on one disk is not protection, and the volunteer guide's §9 sends every tester to exactly this page to read exactly these two sentences. **Fix shape:** keys + args for both files, with the retrieval-promise gate run over the English (these sentences are about what a backup does and does not protect). **CLOSED 2026-09-21, controller v0.259.0.** The row's count of `backup_handlers.go` was 12; **nine are code and three are Hungarian inside COMMENTS**. The offer file's ten is right. `degradedMessageFor` now returns a **KEY** — the decision stays language-free and in one place, the words are chosen by the caller that knows the reader — and `buildTierViews` / `backupTargetLabel` / `loadGuestBackup` take the language the way `buildDataPathCards` already did. **The English is asserted to carry the same NEGATION the Hungarian does** ("protects against corrupted files, **but not** against a disk failure"); an English sentence that promised disk-failure protection would be worse than leaving it Hungarian. **Proven LIVE on guest 9201** for the two tier names ("Local storage (felhom-backup)", "Backup server – separate hardware (PBS)"); **the two warnings themselves were NOT walked live** — that box is healthy and a healthy box renders nothing by design, and producing the state would mean un-assigning a live backup target. They are covered by render tests through the real handler in both states. **An apostrophe cost a render:** the first English absent-drive sentence never matched because `html/template` escapes `'` to `'` — caught by the test, not by review. | **CLOSED 2026-09-21 — controller v0.259.0; the two warnings proven by render test, not live** |
| **R-599** | **[P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so.** FOUND 2026-09-20 tearing the slice-6 drill down. VM destroyed at 17:12Z; the hub then refused **both**`POST /configs/<id>/delete` (409, *host … is ONLINE*) and the host delete (`deletable:false`) — correctly, because an online host would receive permanent 401s. But "online" is not a liveness probe: it is a **report-staleness window**, and the window is **45 minutes** — `manifests/hub.yaml` sets `alerting.stale_threshold: "45m"`, which `hostStatus()` reads (`ok` under it, `stale` over, `down` at 2x). A machine that no longer exists therefore reads ONLINE for three quarters of an hour. **Measured the boring way, and worth recording:** this row first said 30 minutes, because `monitor/host_staleness.go`'s literal default says 30m — the DEPLOYED value is in the manifest, and the 409s kept coming after the half hour was up. Reading a default and calling it the live value is the same mistake in a smaller coat. **The consequence is not theoretical:** a session that destroys its VM and then tears down the hub side walks away believing the delete failed, or leaves the customer behind — and the 2026-09-14 drill's teardown had the same shape without recording this. **Fix shape (smallest first):** the 409 body says *how long* it will refuse ("the last report was N minutes ago; deletion opens at HH:MM"), and `runbooks/target-selection.md`'s drill section names the wait. A force flag is NOT proposed — the refusal is right, only silent about its own clock. | **READY — rank P3-LOW; owner: CC (hub)** |
| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. | **READY - rank P2-MEDIUM; owner: CC (hub)** |
| **R-601** | **[P2-MEDIUM] `demo-hp` has been unreachable since at least this morning — offline on the tailnet for 30 days by Tailscale's own count, and not on the LAN either.** FOUND 2026-09-21 while looking for guest 9201 to ship controller v0.259.0. **What was tried, in order:**`ssh demo-hp`(tailnet `100.76.96.79`) → connection timed out; `sshpass` with the hub-vaulted G1 break-glass password → same timeout; `demo-hp-lan` (`192.168.0.87` via `ProxyJump felhom-pve`) → *No route to host*; `ping 100.76.96.79` → 100% loss; `tailscale status` from the DooPlex pod → **`demo-hp … offline, last seen 30d ago, tx 6084 rx 0`**;`ip neigh` on felhom-pve →`192.168.0.87 … FAILED`. The box answers on no route this session has. **The 30-day figure is Tailscale's and should not be believed on its own** — the slice-6 drill ran its nested drill VM on demo-hp on 2026-09-20 and tore it down, which is not consistent with a box that has been dark for a month; the likelier reading is that its tailscale link has been down for 30 days while the box was reached another way, and the box itself went down more recently. **Either way it is off now.** It also means the fleet's 0.258.0 box is the one that is dark and `demo-felhom` (0.257.0 until today) is the one that is up — so **the fleet is one box, not two, until someone looks at it.** Needs a person: it is a physical machine. | **READY — rank P2-MEDIUM; owner: OPERATOR (physical)** |
| **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp`to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran`ip neigh` on felhom-pve, and`192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** |
| **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:**`managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** |
| **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** |
| **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.****Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** |
| **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:**`07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** |
| Tailnet | `100.70.170.35` | **NONE — tailscale is NOT INSTALLED on demo-hp** (checked on the box 2026-09-21: no `tailscaled`, no `tailscale` binary). The peer `100.76.96.79` still listed by `tailscale status` is a STALE entry from an earlier build and **can never answer**; it showed `offline, last seen 30d ago` while the box was up and reporting. Use the LAN address. |
| Loader used to install | `mkimage` (unsigned, SB **off** — firmware workaround) | **`shim`, Secure Boot ENABLED** |
> **No component versions are recorded on this page — deliberately.** Agent, controller, hub and
@@ -81,10 +81,10 @@ installer baked its `192.168.100.2` fallback as a **static** `vmbr0` address and
that looked installed and could never call home. Repaired on the console by bridging `vmbr0` to
`enp2s0f0`. Filed as **R-59** (must hard-abort) and **R-60** (first-boot NIC sweep self-heal).
Current, read off the box 2026-09-21: `vmbr0`**static `192.168.0.104/24`**, gw `192.168.0.1`, bridge-port **`nic0`**. (It was `192.168.0.87/24` on `enp2s0f0` before a reprovision; both were stale here for long enough to cost a session a false "the box is down" — R-601.) **The hub always knows the truth:** every host report carries `addresses: [{iface, cidr}, …]`, so read it there rather than from this page.
No trace of `192.168.100.2` remains. `wg-felhom``10.77.0.3/32` up to the hub. Guest **9201
`demo-hp`** running. Agent config shape (R-50 island): `local_api` on `169.254.253.1:8443`/`vmbr9`,
guest `eth1 169.254.253.2/30`, `lan_resolver.host_ip` pinned to `192.168.0.87`.
guest `eth1 169.254.253.2/30`, `lan_resolver.host_ip` pinned to the host's LAN address.
### Designated drill + build VM host (operator ruling, 2026-07-25)
@@ -139,7 +139,7 @@ Then `sshpass -e ssh root@demo-hp` (sshpass is on DooPlex, not on the nodes).
the plaintext, so the console is unreachable without a working hub and network — precisely what you
may be trying to fix. Slice 1 is to emit the baked password into the build report.
`demo-hp-lan` (`192.168.0.87` via `ProxyJump felhom-pve`) is the fallback while the box is away.
`demo-hp-lan` (**`192.168.0.104`** via `ProxyJump felhom-pve`) is the fallback when DooPlex is not on the home LAN. **`ssh demo-hp` now goes DIRECT to `192.168.0.104`** — DooPlex is on the same LAN, and the tailnet route for this box does not exist (see the table above). Both were verified 2026-09-21.
## OOB belt (H1) — both boxes, since 2026-07-23 (ISO train v1.25.0)
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.