From 8db9232dea045a51af021e60caf7f1785f1a3284 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 20 Jul 2026 09:45:30 +0200 Subject: [PATCH] docs: REPORT for v0.150.0 (green gate restored, F7 export-link fix, dhclient cleanup) --- REPORT.md | 382 ++++++++++++++++++++++++++---------------------------- 1 file changed, 185 insertions(+), 197 deletions(-) diff --git a/REPORT.md b/REPORT.md index 40f9873..adb0836 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,8 +1,8 @@ -# REPORT — remote-site remediation (F1/F2/Immich) + dashboard backup-card fix (F3) +# REPORT — green-gate restoration + F7 export-page fix + remediation-arc cleanup -**Date:** 2026-07-20 · **Repo:** felhom-controller (v0.148.0 → **v0.149.0**) · Trunk, pushed to `main`. -**Origin:** `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md` (findings F1–F7). -**Run class:** implementation (Part 2) + supervised operational remediation (Part 1). +**Date:** 2026-07-20 · **Repo:** felhom-controller (v0.149.0 → **v0.150.0**) · Trunk, pushed to `main`. +**Origin:** the three open hygiene items from the 2026-07-20 remediation arc (REPORT §Green gate, +§Observations) + R-53. --- @@ -10,248 +10,236 @@ | Repo | Baseline @ start | Gate | After | |------|------------------|------|-------| -| felhom-controller | `9d00177` | `HEAD == origin/main`, tree clean ✔ | `c059fe4` (code+CHANGELOG+CONTEXT), then this docs commit | -| felhom.eu | `648ee67` | `HEAD == origin/main` ✔ | docs commit (audit addendum, ROADMAP, capability map) | +| felhom-controller | `4646be1` | `HEAD == origin/main`, tree clean ✔ | `9f436c8` (code), then this docs commit | +| felhom.eu | `30fd9d3` | `HEAD == origin/main` ✔ | `1e11819` (ROADMAP) | -Deployed controller **before**: `felhom-controller:0.148.0` (Up, healthy) — matched the baseline. -Deployed controller **after**: `felhom-controller:0.149.0` (Up, healthy). +Deployed controller **before**: `felhom-controller:0.149.0` (Up, healthy) — matched the baseline. +Deployed controller **after**: `felhom-controller:0.150.0` (Up 6 seconds, healthy). --- -## 2. Part 1 — remote-site remediation, evidence chain +## 2. Part 1 — the 7 red `internal/backup` tests -### 2.1 Pre-flight — the situation had already moved +### 2.1 Classification (1.1) — one class, not several -The router reservation had taken effect before this run began: `vmbr0` already held -**`192.168.0.162`, but via DHCP** (`dhclient.vmbr0.leases` stamped 08:13 CEST). Nothing had retried -the agent since systemd gave up at 07:25. - -The spec's duplicate-address STOP needed interpretation rather than obedience: -`ping -c1 192.168.0.162` **answered** — because the host answers *itself*. The genuine test is from -another node: ARP for `.162` seen from guest 9201 resolved to `68:1d:ef:5d:a6:64`, which **is** this -host's own vmbr0 MAC. No duplicate → the STOP did not apply. `ip neigh` on the host was empty (a host -does not ARP its own address), consistent with the same conclusion. - -### 2.2 Interfaces diff (`.bak-2026-07-20` → new) - -One stanza; `bridge-ports` / `bridge-stp` / `bridge-fd` and every other line preserved byte-for-byte: - -```diff - auto vmbr0 --iface vmbr0 inet dhcp -+iface vmbr0 inet static -+ address 192.168.0.162/24 -+ gateway 192.168.0.1 - bridge-ports enp1s0 -``` - -Gateway taken from the **live** default route, not from the spec. Tabs verified with `cat -A`. -Applied detached: `setsid ifreload -a >/tmp/ifreload-2026-07-20.log` → log contained only -`Killed old client process` (the vmbr0 dhclient exiting). - -### 2.3 Post-flip verification +The v0.149.0 REPORT's device-identity hypothesis is **confirmed**, and the code said so in Hungarian +without anyone having to infer it. `TestSharesTier2MirrorsBothClasses` failed with: ``` -inet 192.168.0.162/24 scope global vmbr0 (no "dynamic" flag → static) -default via 192.168.0.1 dev vmbr0 proto kernel onlink -curl https://gitea.dooplex.hu → 200 -tailscale: dooplex … active; direct 37.191.56.193:45127 +LastStatus:no_target +LastError:/tmp/…/hdd_1: nincs másik fizikai meghajtó — a 2. mentéshez 2. meghajtó szükséges ``` -Because the address was unchanged (.162 → .162) connectivity never blipped; cloudflared needed no -reconnect and the public dashboard kept serving. +Tier-2's whole purpose is to refuse a target on the source disk; it asks +`system.SamePhysicalDevice` → `isSameBlockDevice` → `statA.Dev == statB.Dev` +(`internal/system/mounts_linux.go:225`). On DooPlex every `t.TempDir()` lands on one filesystem, so a +fixture's `hdd_1` and `hdd_2` share an `st_dev`, the guard correctly refuses, and the test can never +reach the behaviour it exists to check. **The tests were right and the code was right — the +environment could not express the fixture's premise.** -### 2.4 Agent recovery — Scenario A satisfied +| Test | Failing assertion | Root-cause class | +|---|---|---| +| `TestSharesTier2MirrorsBothClasses` | shares not mirrored; payload + marker missing; `LastStatus != ok` | device-identity | +| `TestSharesTier2SkipsDeadMountAndContinues` | "the healthy share must still be mirrored" | device-identity | +| `TestSharesTier2ReconcilePrunesRemovedShare` | precondition: dest dir absent after first run | device-identity | +| `TestSharesTier2NotifierNeverLeaksReservedKey` | "precondition: the notifier should have fired" | device-identity | +| `TestTier2V2_MigrationAndMarkerLast` | flat `appdata/` not removed; legs not mirrored; no marker | device-identity | +| `TestTier2V2_Reconcile` | stale sibling / unrelated dirs not reconciled away | device-identity | +| `TestTier2V2_NetworkExclusion` | picked `…/sys/felhom-data` instead of the local drive | device-identity | -``` -systemctl reset-failed felhom-agent; systemctl start felhom-agent -→ active, ActiveEnterTimestamp Mon 2026-07-20 08:40:11 CEST -→ LISTEN 192.168.0.162:8443 users:(("felhom-agent",pid=72859)) -→ level=INFO msg="local-api server listening" addr=192.168.0.162:8443 -→ wg-felhom latest handshake: 1 minute, 16 seconds ago (< 3 min) -``` +**No test failed for a non-environmental reason** — no real defect was found hiding behind the red, +so there is no §1.1 STOP finding. -Controller side: **zero** `agentapi` / `no route to host` / `channel not verified` lines in the next -5 minutes — explicitly **not** the "agent active but controller still erroring" wrong outcome the -scenario names. On an authenticated fetch of the real dashboard the red -**„A tárolókezelő ügynök nem elérhető"** banner is **gone**; the only `alert-warning` strings left on -the page are two modal "this cannot be undone" confirmations. +### 2.2 Fix — tool 2 (minimal seam), once, for all seven -The one remaining agent WARN is the pre-existing F6 (`/etc/pve/priv/storage/felhom-pbs.pw: -permission denied`) — not a new failure reason, so no STOP. +Tool 1 (test-only) was considered and rejected: the only way to get two genuinely distinct devices +on this host is to place one fixture drive on `/mnt/5_hdd` and the other on `/tmp`, which hard-codes +DooPlex's disk layout into the tests — non-hermetic, and it would re-break on any other machine. Tool +3 (`t.Skip`) would have hidden seven real assertions behind an environment excuse. -### 2.5 F2 — gap closed (and it closed itself first) - -Restoring the agent was sufficient: the quiesce loop found the overdue whole-guest backup and ran it -**unprompted** — the strongest available evidence that the seam is healthy end to end. - -``` -06:40:36Z [quiesce] backup due — quiescing 3 stack(s): [bookstack calibre-web immich] -06:40:45Z backup job backup-9201-1784529645204076480 started -06:40:55Z … snapshotted — resuming app early (8B.2) -06:42:52Z … done -``` - -A manual app-data run followed through the real UI endpoint (`POST /api/backup/run`, authenticated -session + `X-CSRF-Token` read from the page's `` — the call the „Teljes mentés" button makes): - -``` -06:45:27Z DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 228ms, 41 tables) -06:45:28Z DB dump: immich-postgres → immich-postgres.sql (49.9 MB, 1.28s, 66 tables) -06:45:44Z Volume dump: calibre-web/…_config → 360.0 KB -06:45:55Z / 06:45:58Z / 06:45:59Z immich ml_cache 785.5 MB, postgres_data 241.5 MB, redis_data 6.6 MB -06:46:10Z App-data backup completed: 2 databases (50.0 MB total), 3 volume dump(s) (43.348s) -``` - -**T = 2026-07-20 06:46:10Z** (completion; newest dump artifact 06:45:28Z). All 12 containers verified -`Up (healthy)` afterwards — every stack the run stopped for a safe volume dump came back. - -### 2.6 Immich — RESOLVED, plus a correction to the audit - -Started via `POST /api/stacks/immich/start` → all four containers `Up (healthy)`. - -**Correction:** the audit's §3.3 probe used the wrong hostname. The traefik router rule is -``Host(`photos.demo-felhom.eu`)``, not `immich.demo-felhom.eu`. The container-down diagnosis stands -(and the 404 was genuinely traefik having no backend), but that specific probe URL was invalid. On -the correct host: **`https://photos.demo-felhom.eu` → 200** in 0.28 s. - ---- - -## 3. Part 2 — the F3 fix - -**File:** `controller/internal/web/handlers.go`, `dashboardHandler` backup block. - -`dashboard.html:116` branches on `{{if .BackupStatus}}` and reads `.Success` / `.LastRun`, but the -handler never set that key — so the `{{if}}` arm was unreachable and the „Még nem futott" -else-branch rendered unconditionally on **every** box. The adjacent „Adatbázisok: N mentve" row kept -working because it reads `DBDumpStatus`, which *was* passed: exactly the contradiction seen live. - -Verified before editing: `DBDumpStatus` exposes precisely the `LastRun time.Time` / `Success bool` -the template consumes, and `dashboard.html` is the **only** consumer of `BackupStatus` (the -`debug.html` hits are unrelated JS function names). Added inside the existing `if s.backupMgr != nil`: +**The seam** (`internal/backup/backup.go`) — one field + one wrapper, modelled on the package's +existing `tier2Mirror` / `tier2SSDFits` / `sharesPassdbCapture` nil-defaulted field seams: ```go -data["BackupStatus"] = fullStatus.LastDBDump +samePhysicalDevice func(a, b string) bool // nil → system.SamePhysicalDevice + +func (m *Manager) sameDevice(a, b string) bool { + if m.samePhysicalDevice != nil { return m.samePhysicalDevice(a, b) } + return system.SamePhysicalDevice(a, b) +} ``` -No template change, no new view-model, no change to what "utolsó mentés" means. `*DBDumpStatus` -nil/non-nil maps exactly onto the template's branch, so a fresh box still reads „Még nem futott". +**Seam justification (one line, per §9.4):** the off-drive predicate is the single environmental fact +the tier-2 tests cannot supply on a one-filesystem host; routing it through a nil-defaulted field +makes it injectable while leaving production on the identical `st_dev` call. -### Files changed +Nil-check lives in the wrapper rather than in `NewManager` deliberately: `newSharesEnv` builds +`&Manager{…}` by literal, so a constructor-only default would leave the field nil and panic. -| File | Change | -|------|--------| -| `controller/internal/web/handlers.go` | +1 assignment (+4 comment lines citing the audit) | -| `controller/internal/web/dashboard_backup_card_test.go` | **new** — 3 tests + 2 helpers | -| `CHANGELOG.md` | v0.149.0 entry (newest on top) | -| `CONTEXT.md` | remediation + v0.149.0 block | -| `REPORT.md` | this file (overwritten) | +Seven call sites now go through `m.sameDevice(...)`: `tier2.go` ×5 (`selectTier2TargetFrom` ×3, +`RunTier2`, `Tier2Info`), `tier2_shares.go` ×1 (`RunSharesTier2`), `backup.go` ×1 +(`hasOffDriveTarget`). Behaviour with a nil seam is byte-for-byte the previous code. + +**Test-side** (`device_seam_test.go`, new): `oneDrivePerSubtree` models one drive per directory +subtree — two paths share a device only when one contains the other, which is exactly how real +mountpoints behave. Installed in the two fixtures only (`newTier2V2`, `newSharesEnv`). + +Nothing was weakened: the guard still runs and still refuses same-device targets — the +same-drive refusal test `TestSharesTier2NeverTargetsItsOwnSourceDrive` passes under the seam, and it +is the test that would catch a fake that simply said "always different". + +**Surviving `t.Skip`s: none.** No test was skipped, deleted, renamed, or had an expected value +changed. + +### 2.3 Scenario B — mutation table + +Every one of the seven had the defect it guards re-introduced in **production** code, one at a time, +then reverted: + +| Test | Mutation | Result | +|---|---|---| +| `TestSharesTier2MirrorsBothClasses` | skip the payload mirror to the target | **FAIL** ✔ | +| `TestSharesTier2SkipsDeadMountAndContinues` | stop skipping `Disconnected/Decommissioned` drives | **FAIL** ✔ (`dead-mount share reached a mirror call`) | +| `TestSharesTier2ReconcilePrunesRemovedShare` | `tier2ReconcileRoots` → no-op | **FAIL** ✔ | +| `TestSharesTier2NotifierNeverLeaksReservedKey` | pass `SharesPseudoStack` raw on the success path | **FAIL** ✔ (`the reserved key reached the notification boundary raw: "_shares"`) | +| `TestTier2V2_MigrationAndMarkerLast` | skip the old-flat-`appdata/` migration cleanup | **FAIL** ✔ | +| `TestTier2V2_Reconcile` | `tier2ReconcileRoots` → no-op | **FAIL** ✔ | +| `TestTier2V2_NetworkExclusion` | disable the auto `sp.IsNetwork()` skip | **FAIL** ✔ | + +Honest note on method: my first attempt at two of these mutated the wrong code path (the +mirror-failure notify branch instead of the success one; the unregistered-root skip instead of the +disconnected-drive one) and both tests stayed green. That was a bad mutation, not a hollow test — the +corrected mutations bite, and the notifier one reproduces that test's own documented red-proof +verbatim. Tree verified free of mutation residue afterwards (`grep -c "// MUT:"` → 0). + +### 2.4 Scenario A — green gate, twice + +| Run | Command | Result | +|---|---|---| +| build / vet | `go build ./...`, `go vet ./...` | **OK** | +| 1 | `go test ./... -count=1` | **exit 0 — 23 packages ok, 0 FAIL** | +| 2 | `go test ./... -count=1` | **exit 0 — 23 packages ok, 0 FAIL** | + +`grep -c SKIP` over run 2 → **0**. The gate is genuinely green, not green-by-omission. --- -## 4. Tests and the red-proof +## 3. Part 2 — F7 / R-53, the export page -The tests drive the **real handler** through `ServeHTTP` (not the template alone), backed by a real -`backup.Manager` whose per-drive dump scan finds a planted `.sql` file — so they bite on the handler -wiring, which is what the red-proof requires. +`app_export.html` L93 read `var domain = '{{.Stack.Meta.Subdomain}}' ? '{{…}}.{{$.CSRFToken}}' : '';` +— the session CSRF token where the customer domain belongs. + +**A spec premise turned out to be wrong, and the fix needed one more line than planned.** §5 states +`{{$.Domain}}` is "set by `baseData`". It is (`handlers.go:120`) — but `exportPageHandler` +(`handler_export.go`) never calls `baseData`; it builds `{"Stack", "Drives"}` and `executeTemplate` +injects only `CSRFField` / `CSRFToken`. Swapping the token alone would have rendered an empty +domain — trading a wrong link for a broken one. So the fix is two lines: + +- template: `{{$.CSRFToken}}` → `{{$.Domain}}`; +- handler: `"Domain": s.cfg.Customer.Domain` added to the map, mirroring `baseData`'s own line. + +`csrfH()` and the meta-tag CSRF flow are the correct usage and are untouched. + +**Tests** (`app_export_domain_test.go`, new — v0.149.0's render harness): the `var domain` line is +isolated first so an assertion cannot match the token where it legitimately appears (the meta tag). | Scenario | Test | Result | |---|---|---| -| B — real dump ⇒ its timestamp | `TestDashboardBackupCard_ShowsLastRun` | **PASS** | -| C — no dump ⇒ still honest, no `0001-01-01` | `TestDashboardBackupCard_FreshBoxStaysHonest` | **PASS** | -| D — failed run ⇒ „Sikertelen" | `TestDashboardBackupCard_FailedRunShowsSikertelen` | **PASS** | +| C — joined from the customer domain, token absent | `TestAppExportDomainUsesCustomerDomainNotCSRFToken` | **PASS** | +| C — empty-subdomain branch still `''` | `TestAppExportDomainEmptyWithoutSubdomain` | **PASS** | -**Companion red-proof (§10).** Deleted `data["BackupStatus"] = fullStatus.LastDBDump` from -`dashboardHandler` and re-ran: +**Red-proof:** restored the pre-fix template line → both tests fail, showing the defect verbatim: ``` ---- FAIL: TestDashboardBackupCard_ShowsLastRun (0.21s) - dashboard_backup_card_test.go:118: card must show the real last run "2026-07-19 03:30", got: … - dashboard_backup_card_test.go:121: a box WITH a dump must not claim it never ran, got: … +--- FAIL: TestAppExportDomainUsesCustomerDomainNotCSRFToken + export link must be built from the customer domain, got: + var domain = 'photos' ? 'photos.deadbeefcafebabe…' : ''; + the CSRF token must NEVER appear in the export URL, got: … (same line) +--- FAIL: TestAppExportDomainEmptyWithoutSubdomain ``` -Scenario B failed on **both** assertions (C and D correctly still passed — they do not depend on the -handler line for their outcome). Line restored; suite re-run green. +Template restored; both green. -Scenario D asserts the template branch directly rather than through the handler: the failure flag -lives only in the manager's in-memory run state, unreachable from the `web` package without shelling -out to docker. That limitation is stated in the test's own comment, not hidden. +**Live verification (Scenario C)** — authenticated endpoint fetch of the real export page +`/stacks/immich/export` (no browser on DooPlex), ASCII-safe grep per the new gotcha: -**Test count: 736 → 739 (+3).** +``` +line 237: var domain = 'photos' ? 'photos.demo-felhom.eu' : ''; +meta tag: csrf-token" content="f7f928f1… ← the token lives here, and only here +``` -### Green gate — honest result - -`go build ./...` **OK** · `go vet ./...` **OK** · `go test ./...` → **22 packages ok, 1 FAIL**. - -The failing package is `internal/backup` with 7 tests (`TestTier2V2_*`, `TestSharesTier2*`). **These -are pre-existing and unrelated to this change** — verified by checking the baseline commit `9d00177` -out into a throwaway `git worktree` and running the same tests there: **identical failures with none -of my changes present**. They look environment-dependent on DooPlex (the tier-2 tests reason about -distinct physical devices, and every `t.TempDir()` here lands on the same filesystem) — the same -class as the recorded "Windows green gate silently red" lesson. `internal/web`, the package this -change touches, is **green**. - -I did not attempt to fix them: out of scope for this task, and quietly absorbing a red gate would be -worse than reporting it. +The rendered link is correct and the token appears nowhere in it. Note this also confirms the +handler-side `Domain` key, which the render tests alone could not prove. --- -## 5. Build / deploy / live verification +## 4. §3.4 — the orphaned `dhclient` + +Evidence first. Exactly one match, on an interface that does not exist: ``` -build.sh 0.149.0 --push → Build complete ✓ gitea.dooplex.hu/admin/felhom-controller:0.149.0 (145M) -pct exec 9201: docker pull … && echo … > /etc/felhom-controller-image - && systemctl restart felhom-controller-bootstrap.service -docker ps → gitea.dooplex.hu/admin/felhom-controller:0.149.0 Up 5 seconds (healthy) +BEFORE: 1922 /sbin/dhclient -pf /run/dhclient.eth0.pid -lf /var/lib/dhcp/dhclient.eth0.leases eth0 + PPID 1652, started Mon Jul 20 07:25:30 (boot) +interfaces: lo enp1s0 wlp2s0 tailscale0 vmbr0 wg-felhom veth9201i0 ← no eth0 +vmbr0: iface vmbr0 inet static / address 192.168.0.162/24 ← no dhclient should run at all +neither /run/dhclient.eth0.pid nor /var/lib/dhcp/dhclient.eth0.leases existed ``` -**Scenario B, live on the real dashboard** — authenticated endpoint fetch of the rendered page -(no browser on DooPlex; `claude-in-chrome` unavailable per CLAUDE.md): - -```html -Utolsó mentés: -2026-07-20 06:45 -Adatbázisok: -3 mentve -``` - -The identical page rendered „Még nem futott" before the deploy. The card shows 06:45 (newest dump -artifact) rather than 06:46 (run completion) because the deploy restarted the controller and cleared -the in-memory run state, so it renders from the dumps on disk — the honest value, and precisely the -path that must never fabricate a timestamp when there are no dumps. +Killed. Re-checked immediately and again ~25 min later: `pgrep -a dhclient` → **none, no respawn**. +`vmbr0` still `192.168.0.162/24`, `felhom-agent` still `active`. No other host mutation. --- -## 6. Not done / open +## 5. Files changed -| Item | Status | -|---|---| -| **Durable F1** — control plane pinned to a LAN literal | **OPEN → R-50** (island-bridge, **spike-first**). Today's static `.162` is a *window mitigation*: it still assumes the site's subnet is `192.168.0.0/24` and `.162` is free | -| **F4** — dead-primary alerting in multi-container apps | **OPEN → R-51** (constraint recorded: must NOT simply fold `unhealthy` into down — that reverts fix-3) | -| **F5** — boot desired-state reconciliation | **OPEN → R-52** (includes root-causing the `unless-stopped` non-resurrection; the evidence is gone, so it needs a fresh repro) | -| **F6** — agent cannot read the PBS `.pw` | **OPEN, already tracked** as R-39's remaining item (b). No duplicate item minted | -| **F7** — `app_export.html` CSRF-token-as-domain | **OPEN → R-53** | -| `internal/backup` 7 red tests | **OPEN, pre-existing** — not investigated (out of scope) | +| File | Change | +|------|--------| +| `internal/backup/backup.go` | seam field + `sameDevice` wrapper; 1 call site routed | +| `internal/backup/tier2.go` | 5 call sites routed through the seam | +| `internal/backup/tier2_shares.go` | 1 call site routed; unused `system` import dropped | +| `internal/backup/device_seam_test.go` | **new** — `oneDrivePerSubtree` | +| `internal/backup/tier2_v2_test.go`, `shares_test.go` | seam injected in the two fixtures (1 line each) | +| `internal/web/templates/app_export.html` | 1 token: `{{$.CSRFToken}}` → `{{$.Domain}}` | +| `internal/web/handler_export.go` | `"Domain"` added to the export page's data map | +| `internal/web/app_export_domain_test.go` | **new** — 2 tests + 2 helpers | +| `CLAUDE.md` | 2 gotchas in the live-validation section | +| `CHANGELOG.md` / `CONTEXT.md` / `REPORT.md` | v0.150.0 bookkeeping | -Capability map gained a **PARTIAL** row: *"Box survives a site/network change (relocation, different -subnet, DHCP re-lease) with the control plane intact"*, citing the audit and pointing at R-50/51/52. +**Commits:** felhom-controller `9f436c8` (code) + this docs commit; felhom.eu `1e11819` (ROADMAP). --- -## 7. Observations (noticed, not acted on) +## 6. ROADMAP corrections (felhom.eu) -1. **An orphaned `dhclient` runs on the host for a non-existent `eth0`** (`pgrep -a dhclient` → - `/sbin/dhclient … eth0`, pid 1922) while no `eth0` exists in `ip addr`. Harmless today, but it is - leftover state of exactly the kind that makes a future network change confusing to diagnose. -2. **Accented grep patterns are unreliable through the `ssh → pct exec → bash -c` chain.** A search - for „nem elérhető" returned a false `0` while the string was present — which reads exactly like - "the banner cleared". The page content is fine; the *pattern* gets mangled in transit. Use - ASCII-only substrings for any grep that gates a conclusion. This nearly produced a wrong - "banner gone" claim before an ASCII re-check caught it. -3. **The controller password contains `'` and `!`.** A helper script written via an unquoted heredoc - had its `!!` mangled by history expansion and produced one failed login (visible in the controller - log as `[WARN] [web] Failed login`). The proven-safe form is the inline `-d "password=$PW"` the - audit run used. -4. The guest journal is **volatile** (current boot only) and the controller's logs are lost on - container recreate — which is why F5's root cause could not be established after the fact. Worth - knowing before any post-mortem that depends on in-guest history. +- **R-50 factual correction.** The entry claimed an address move needs the cert's SAN set to cover + the new address. Read against source (`internal/agentapi/client.go` L105–129): the controller→agent + leg sets `InsecureSkipVerify: true` and replaces chain verification with a `VerifyPeerCertificate` + that does a raw **SHA-256 match on the leaf DER**. Hostname/SAN never enters verification on this + leg, so an address move most likely needs **no cert re-issuance** — only the endpoint the guest + dials. The entry now says so, and still requires the spike to confirm it empirically. +- **R-53 collapsed** to its one-liner + **SHIPPED (controller v0.150.0)**, noting the handler-side + `Domain` key as part of the fix. + +--- + +## 7. Not done / open + +R-50 (durable F1 — spike-first, not started), R-51, R-52, R-39(b)/F6 — all untouched, as scoped. + +--- + +## 8. Observations (noticed, not acted on) + +1. **Four files in `internal/backup` are not `gofmt`-clean at HEAD** — `offbox_3a_test.go`, + `recovery_unit.go`, `tier2.go`, `tier2_v2_test.go`. I verified this is **pre-existing** (checked + each file's HEAD blob through `gofmt -l` before my edits) and left them alone: `gofmt` is not part + of the green gate (`build`/`vet`/`test`), and reformatting them would have buried this task's real + diff in noise. Worth a one-shot `gofmt -w` pass in its own commit. +2. **`system.SamePhysicalDevice` is mount/device-granular, not disk-granular** — its own doc comment + says two partitions on one physical disk look "different" here, with the agent's durable-id as the + stronger guarantee. Unchanged by this work, but it means the tier-2 off-drive promise is + partition-level on a box that partitions one disk. +3. **The tier-2 tests silently stopped exercising their subject rather than failing loudly at the + premise.** Each failed on a downstream assertion ("share was not mirrored") instead of on "this + fixture needs two devices and this host has one". A fixture-level precondition check would have + turned a week of ambiguous red into one clear message.