diff --git a/CHANGELOG.md b/CHANGELOG.md index 82d0430..dde29f8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,104 @@ +## v0.229.0 — the second drive's copy becomes a way back (2026-08-31, R-102 + R-103) +**MinAgent: 0.129.0** (unchanged) + +### The Tier-2 unit mirror is restorable (R-102) + +Tier-2 has written a full `recovery-unit/` mirror — the app's definition, its portable secrets, its +database dump and its named-volume tars — to `/backups/secondary//recovery-unit/` on every +run for months. **No code path read it.** Every reader of a recovery unit could only name a path under +`backups/primary/`, because `appbackup/paths.go` joined that segment literally. + +**The failure that made it worth doing:** Tier-2 exists for the loss of the primary drive, and in +exactly that loss the primary unit is gone while the mirror survives — unreachable by any customer +action (`07-backup-architecture.md` §6.3, §7.2). For the **45** class-B apps in the catalogue (Part 3 +below) that is the whole of their data. + +- **`appbackup`** gains four unit-directory-relative primitives — `UnitComposeDir`, + `UnitManifestFile`, `UnitDBDumpDir`, `UnitVolumeDumpDir`. The four `(nsRoot, stackName)` helpers + become thin wrappers over them and return byte-identical strings; every existing caller compiles + untouched. Pinned by `TestR102_PathWrappersAreByteIdenticalToToday` against hand-written literals, + not against the helpers under test. +- **`Manager.RestoreFromRecoveryUnitAt(stack, unitDir)`** holds the whole body; + `RestoreFromRecoveryUnit(stack)` is the thin caller naming the primary unit. ONE implementation, two + callers — the rule `restoreDockerVolumesFrom` already states beside itself, and for the same reason. +- **THE SOURCE MOVES; THE DESTINATION DOES NOT.** `unitDir` changes only where the manifest, the + compose capture, the `.sql` and the tars are READ from. Data still lands in the live Docker volumes + and the live database container, and the definition still in the guest. A restore that also + relocated the app's data would be a migration. +- **Unchanged and pinned:** the R-47 mutation order (stop → volumes → recreate → DB-only start → + replay → start), the secret reconciliation with unit-over-guest precedence, the fail-closed data-key + gate, and the no-unit fallback to `RestoreApp` with its `CountsUnknown` handling. +- **`Manager.RestoreTier2Unit(stack)`** resolves the recorded copy, refuses **fail-closed** unless the + mirror carries a parseable `manifest.json` — *a directory that exists is not a package* — and + delegates. The single-writer flag is taken inside `RestoreFromRecoveryUnitAt`, not beside it. +- The 35-minute DB-replay bound is now named once (`dbReimportTimeout`), so the two bounded entry + points cannot drift in how long a wedged import may hang a restore. +- The unit DIRECTORY is now logged on every restore. Which copy a restore read from is a real question + with two answers, and an absent log line is not evidence. + +### The refusal becomes an action (R-103) + +An app with no file legs but a full mirror was told to press a button on a **different page**. + +- **`POST /backup/tier2/unit-restore`** + `backupTier2UnitRestoreHandler`: same guards, same + `restoreOpBlocked()` refusal (R-351b), same async shape as the file restore beside it, plus the + fail-closed pre-flight so the app is **never stopped** for a mirror that could not be opened. +- **`Tier2Coverage` gains `UnitRestorable`, and `CanRestore()` is NOT widened.** It still answers only + *"can the additive file restore run?"*. One predicate answering two questions is **R-356**, which + refused 40 running apps for months. `HasUnit` also keeps its old meaning — a half-copied mirror is + still unread data the file restore must disclose, even though the unit restore refuses it. +- **The row offers the action where the refusal was**, in `btn-danger-outline`, as a SEPARATE button. + The two are not merged: one adds what is missing, the other overwrites. **The confirm carries that + difference in words** and names the copy's date — and says so differently when that date is only an + ATTEMPT (R-101). It is assembled from named Go constants rather than inside an HTML attribute, so a + test asserts it verbatim; `fmtTimeStr` now delegates to a package-level `fmtRFC3339Local` so the + confirm and the outcome cannot render one date two ways. +- **`tier2NoCoverageMsg` is NARROWED** to the case that remains — no legs *and* no openable unit — and + still names the route that works. **`tier2UnitNotCoveredMsg` is NOT deleted:** it is appended where + the FILE restore ran and is still exactly true of it. +- The outcome reuses `unitRestoreOutcomeMsg` unchanged and appends which copy overwrote the live data. + +### The count, settled (Part 3) + +`07-backup-architecture.md` §6.2 recorded **two** Tier-2 coverage counts that disagreed — 9/43/1 and +7/45/1 — both unresolved. Counted at catalogue `459766cb16395fd1d1a66282f5cc6da59ead5924` by running +the PRODUCTION rule (`LoadMetadata` → `ParseComposeClassifiableBinds` → `ClassifyBinds` → +`ComputeCaptureSet` at `TierSecondary`) over all 53 templates: + +**A = 7 · B = 45 · C = 1.** The INV Part B.1 enumeration was right. The two apps the C9-F1 Phase-0 +count put in A are **radarr and sonarr**: both bind `${USERDATA_PATH}` paths **writably**, so the +`:ro`-default rule Phase 0 says it applied to plex/jellyfin/emby/navidrome does not catch them — they +are excluded by an **explicit** `class: excluded` entry instead. Class C is **bentopdf**, which +declares no volumes and no namespace binds at all. No catalogue file was changed. + +### Live drill — demo-hp, endpoint level + +`documentation/audits/DRILL-r102-tier2-unit-2026-08-31/` (in `felhom.eu`). docmost, class B, its +Tier-2 run reporting **0 leg(s)**. With the **primary unit moved aside** the restore returned 3 volumes +of 3 and 1 database of 1 from the secondary mirror in **28.65 s**; an accented Hungarian filename came +back byte-for-byte (verified as hex, not as rendered text — R-364) and the app read its own row **over +TCP with its own credential**. The post-backup discriminator was **gone**, so the replay was real. +**Scenario D** repeated it with the guest's `app.yaml` moved aside: `secrets recovered=2/2` from the +mirrored unit — this closes `00-capability-map.md`'s open *"not exercised live"* clause for Tier-2's +own cross-drive copy of a secret-bearing unit. + +**Filed, not fixed — R-403.** Two seconds after a restore that ran with the primary unit absent, the +5-minute status refresh (`captureAllRecoveryUnits`) rewrote the primary unit from a drive with no +dumps, producing a manifest carrying `"db_dumps": []` and `"volume_dumps": null`. The ordinary restore +then read it and honestly reported that the backup held only settings. The dangerous half — that the +next Tier-2 run would mirror that hollow unit over the good secondary copy, since `rsyncMirror` carries +`--delete` — **was not tested and is recorded as unverified.** + +### Tests + +26 new test functions (1606 → 1632). Red-proofs run and reverted: **A1** (change one wrapper's join → +fails on all three fixtures), **A5** (swap the volume replay and the recreate → fails on the sequence), +**B2** (point the Tier-2 reader back at the primary → fails, and fails again with `permission denied` +once the primary tree is unreadable), **C1** (widen `CanRestore` to include `HasUnit` → the unit-only +cases fail), **D6** (drop `EndRestoreOp` from the handler's goroutine → *"the restore never published a +result"*). The Tier-2 fixtures build their mirror with the production `RunTier2`, so the claim under +test is *the copy Tier-2 writes is the copy this restore reads*. + ## v0.228.0 — the check reads the data, and the debug page stops lying (2026-08-31, R-399 + R-400) **MinAgent: 0.129.0** (unchanged) diff --git a/CONTEXT.md b/CONTEXT.md index 8f2beae..0857aa7 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -7,7 +7,46 @@ > > Ask Claude Code: "Please update CONTEXT.md with what we did today" -Last updated: 2026-08-31 (v0.228.0 — R-399/R-400: the check reads the data, the debug page stops lying) +Last updated: 2026-08-31 (v0.229.0 — R-102/R-103: the second drive's copy becomes a way back) + +> **2026-08-31 — v0.229.0. THREE RULINGS, recorded so none is re-litigated.** +> +> **1. The source moves; the destination does not.** `RestoreFromRecoveryUnitAt(stack, unitDir)` takes +> the recovery-unit DIRECTORY, so the same restore reads a unit from the primary drive or from the +> Tier-2 mirror on the second drive. What it must NEVER take is a destination: data still lands in the +> live Docker volumes and the live database container, and the definition in the guest, resolved by +> `GetAppDrivePath` exactly as the capture is. A restore that also relocated an app's data would be a +> migration wearing a restore's label, and the customer pressed a button that said neither. +> +> **2. Two predicates, never one wider one — and it is the SECOND time this is written down.** +> `Tier2Coverage.CanRestore()` answers *"can the additive file restore run?"* and nothing else; +> `CanRestoreUnit()` answers *"can the unit restore open this copy?"*. `HasUnit` keeps its third, +> distinct meaning: *"is there captured data the file restore is not looking at?"* — true even for a +> half-copied mirror the unit restore refuses, because the disclosure is still owed. The temptation is +> always to widen the predicate already there. **R-356 is what that costs:** one predicate meaning both +> *"has this app a drive?"* and *"is this app installed?"* refused 40 running apps for months, while +> they were running, with a message telling their owners to reinstall them somewhere those apps never +> offer. +> +> **3. A destructive operation reached from a non-destructive surface must carry the difference in the +> CONFIRM, not in the label.** „Teljes visszaállítás a másolatból" sits beside „Fájlok +> visszaállítása" on the same row; one overwrites the app's database and internal volumes, the other +> only adds files that are missing and never overwrites anything. The confirm says exactly that, names +> the copy's date, and says so DIFFERENTLY when that date is only an attempt clock (R-101). It is built +> from named Go constants (`tier2UnitConfirmBase` / `…DateFmt` / `…DateUnprovenFmt` / `…Contrast`) and +> asserted verbatim, because a sentence assembled inside an HTML attribute cannot be pinned and R-364 +> makes grepping accented Hungarian out of rendered markup unreliable on top of that. +> +> **The count is settled and must not be re-derived.** At catalogue `459766cb1639`, by the production +> rule: **A = 7 · B = 45 · C = 1**. The C9-F1 Phase-0 count (9/43/1) was wrong by two — **radarr and +> sonarr**, whose `${USERDATA_PATH}` binds are WRITABLE (so the `:ro` default rule Phase 0 applied does +> not catch them) and are excluded by an explicit `class: excluded` entry instead. C is **bentopdf**. +> +> **R-403, filed and NOT fixed.** Two seconds after a restore that ran with the primary unit absent, the +> 5-minute status refresh (`captureAllRecoveryUnits`) rewrote the primary unit from a drive with no +> dumps, yielding `"db_dumps": []` / `"volume_dumps": null`. Measured on demo-hp 2026-08-31. The +> dangerous half — that the next Tier-2 run would mirror that hollow unit over the good secondary copy, +> `rsyncMirror` carrying `--delete` — **was not tested and is recorded as unverified.** > **2026-08-31 — v0.228.0. TWO RULINGS, recorded so neither is re-litigated.** > diff --git a/REPORT.md b/REPORT.md index 242a8a7..f1f2cac 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,451 +1,292 @@ -# REPORT — controller v0.228.0 (R-399 + R-400), 2026-08-31 +# REPORT — R-102 + R-103: the second drive's copy becomes a way back -**The check now reads the data, and the debug page no longer lies.** Built, pushed, deployed to -`demo-hp`, and proven there at both depths with the restic argument list read off the running process. +**Controller v0.229.0 · 2026-08-31 · MinAgent 0.129.0 (unchanged)** --- -## 1. Confirmed baselines — re-checked at the start, no drift +## 1. Confirmed baselines -| Repo | `main` @ start | matched `origin/main` | version | +Re-checked before the first edit, and both matched the task's table exactly. **No drift.** + +| Repo | `main` @ start | expected | version | |---|---|---|---| -| `felhom-controller` | `300d7e87d7cfcbb6dd355594f5d8936c06184a27` | yes | v0.227.1 → **v0.228.0** | -| `felhom.eu` | `db0812b6f261dbb925ef89ed27bfc2d2d1d3b5b9` | yes | docs only | +| `felhom-controller` | `430fb4448d6175064ef4e86c2b3f8796e15ae30c` | same | `v0.228.0` → **`v0.229.0`** | +| `felhom.eu` | `1623a4d5b5d5d0f1b73aba2727b8de053930f336` | same | — (docs only) | +| `felhom-agent` | not touched | — | `v0.130.0` unchanged | +| `app-catalog-felhom.eu` | `459766cb16395fd1d1a66282f5cc6da59ead5924` | read-only | unchanged | -`git status --porcelain` was empty in both. `MinAgent` stays **0.129.0**. No drift to report. +`git status --porcelain` empty in both repos; disk 37% / 51%. ---- +## 2. Files created / modified -## 2. Files created / modified / deleted +**`felhom-controller`** -**Created (6)** +| File | Change | +|---|---| +| `controller/internal/appbackup/paths.go` | +4 unit-directory-relative primitives; the 4 `(nsRoot, stackName)` helpers become wrappers | +| `controller/internal/appbackup/r102_paths_split_test.go` | **new** — A1 + a directory-relative guard | +| `controller/internal/backup/appbackup_bridge.go` | re-exports the 4 primitives into `backup` | +| `controller/internal/backup/restore_unit.go` | `RestoreFromRecoveryUnitAt(stack, unitDir)`; the 1-arg form becomes the thin caller; the volume leg goes through the R-354 seam; the unit dir is logged | +| `controller/internal/backup/restore_db.go` | `reimportDBDumpsAtCtx`; `dbReimportTimeout` named once | +| `controller/internal/backup/tier2_restore.go` | `RestoreTier2Unit`, `ErrTier2NoUnitInCopy`, `tier2UnitDir`, `tier2UnitIsOpenable`, `Tier2Coverage.{UnitRestorable,CopyLastRun,CopyLastSuccess}`, `CanRestoreUnit()`, `Tier2CopyDate()` | +| `controller/internal/backup/r102_unit_at_test.go` | **new** — A2…A6 + 1 | +| `controller/internal/backup/r102_tier2_unit_test.go` | **new** — B1…B5 + 1 | +| `controller/internal/backup/r103_file_restore_untouched_test.go` | **new** — C1, C2 | +| `controller/internal/web/handlers.go` | `backupTier2UnitRestoreHandler`; the refusal split; 7 named constants; `tier2UnitConfirmMsg`, `tier2UnitSourceMsg`; 4 new `AppBackupRow` fields + their builder | +| `controller/internal/web/server.go` | route `POST /backup/tier2/unit-restore` | +| `controller/internal/web/funcmap.go` | `fmtTimeStr` delegates to a package-level `fmtRFC3339Local` | +| `controller/internal/web/templates/backups_apps.html` | the destructive action on the Tier-2 row | +| `controller/internal/web/r103_tier2_action_test.go` | **new** — D1…D6 + 4 | +| `CHANGELOG.md`, `CONTEXT.md`, `controller/README.md`, `REUSE.md`, `REPORT.md` | documentation | -- `controller/internal/backup/r399_depth_test.go` — Group A (8 tests) -- `controller/internal/backup/r399_slow_notice_test.go` — Group B1–B5 -- `controller/cmd/controller/r399_no_event_test.go` — B6, the non-effect -- `controller/internal/web/r400_debug_routes_test.go` — Group D -- `controller/scripts/debug_route_gate.py` — the gate -- `controller/scripts/test_debug_route_gate.py` — Group C, incl. both red-proofs +**`felhom.eu`** — `documentation/architecture/07-backup-architecture.md` (§6.2, §6.3, §7.2, §8, §8.1), +`documentation/architecture/00-capability-map.md` (header note, Tier-2 row, D5 row), +`documentation/backlog/OPEN-ITEMS.md`, `documentation/backlog/CLOSED-ITEMS.md`, `STATUS.md`, +`documentation/audits/DRILL-r102-tier2-unit-2026-08-31/` (**new**, 11 files). -**Modified (17)** - -`controller/internal/backup/offbox_integrity.go` · `controller/internal/backup/offbox.go` · -`controller/internal/settings/settings.go` · `controller/internal/config/config.go` · -`controller/cmd/controller/main.go` · `controller/internal/web/handler_debug.go` · -`controller/internal/web/templates/debug.html` · `controller/internal/report/types.go` · -`controller/configs/controller.yaml.example` · `controller/scripts/controller_gates.py` · -`controller/scripts/test_controller_gates.py` · `controller/internal/backup/r359_integrity_test.go` · -`CHANGELOG.md` · `CONTEXT.md` · `REUSE.md` · `controller/README.md` · `.claude/rules/gates.md` - -`felhom.eu`: `STATUS.md` · `documentation/architecture/00-capability-map.md` · -`documentation/architecture/07-backup-architecture.md` · `documentation/backlog/OPEN-ITEMS.md` · -`documentation/backlog/CLOSED-ITEMS.md` · `scripts/wire_contract_gate.py` - -**DELETED — half this task, so listed explicitly** - -*From `debug.html`:* six `/api/debug/...` references, their buttons and result spans, the entire -„Tárhely teszt" card (`section-storage`), the `dr-status` panel, and the JavaScript functions -`loadWatchdogStatus`, `renderWatchdogStatus`, `simulateDisconnect`, `simulateReconnect`, `loadDRStatus`, -plus their two `loadSectionData` cases. Template shrank 49 081 → 42 646 bytes. - -*From the test tree:* `TestR359_StructureCheckPassesNoReadDataFlag` and -`TestR359_MalformedReadDataSubsetIsTreatedAsOff` — both asserted the ruling this release reverses. -They are **replaced, not weakened**, and a paragraph stands where each was naming its successor, so a -later reader does not re-derive the old ruling from an absence. See §4. - -*From `controller/README.md`:* the „Tárhely teszt" debug-section row, and the `infra-push` / -`dr/infra-status` route mentions. - ---- +**Not modified:** `felhom-agent`, `app-catalog-felhom.eu` (Part 3 is a measurement), `ROADMAP.md` +(it carries no R-102/R-103 row — `one_register_gate.py` green). ## 3. Commits pushed to `main` | Repo | Commit | What | |---|---|---| -| `felhom-controller` | `3c49dc8ea42df6c84a4bc9495d0d8a7662163ae4` | the whole of Parts 1–3 + tests + gate + repo docs | -| `felhom.eu` | `77a5a115` | STATUS, capability map, 07 §10.2, register, wire allowlist | +| `felhom-controller` | `c732006d263a25744eca11a11cdeef343b1c95af` | Part 1.1 — path helper split, zero behaviour change | +| `felhom-controller` | `0f9b796615d3e662fc010b747fb5048f36faffb7` | R-102 — Parts 1.2, 1.3, 2.1 + Groups A and B | +| `felhom-controller` | `4c8f0d291948fd60887c0fd7066e63df9e3abb68` | R-103 — Part 2.2 + Groups C and D | +| `felhom-controller` | *(docs commit — see §14)* | CHANGELOG / CONTEXT / README / REUSE / REPORT | +| `felhom.eu` | *(docs commit — see §14)* | architecture, register, STATUS, drill evidence | ---- +## 4. Per-test results, and the five named red-proofs -## 4. Tests — per group, and the red-proofs by name +Full suite: `go build ./... && go vet ./... && go test ./...` — **all packages ok** +(`internal/backup` 296.0 s, the rest cached/fast). `python3 controller/scripts/controller_gates.py` — +**all 13 gates OK**. -Green gate: `go build ./... && go vet ./... && go test ./...` → **all three exit 0**, full suite. -`python3 scripts/controller_gates.py --fast` → **13/13 OK**. - -| Test | Result | -|---|---| -| A1 `TestR399_AbsentConfigRunsFullDepth` | PASS | -| A2 `TestR399_EmptyStringIsNotOff` | PASS | -| A3 `TestR399_OffTokenRunsStructureOnly` | PASS | -| A4 `TestR399_OffTokenIsCaseInsensitive` | PASS | -| A5 `TestR399_ExplicitValueWins` | PASS | -| A6 `TestR399_MalformedFallsBackToTheDefault` | PASS | -| A7 `TestR399_DepthIsStatedInTheOutcome` | PASS | -| A8 `TestR399_DefaultResolvesThroughTheRealConfigPath` (the seam) | PASS | -| B1 `TestR399_SlowCheckWarns` | PASS | -| B2 `TestR399_FastCheckIsSilent` | PASS | -| B3 `TestR399_SkipNeverWarns` | PASS | -| B4 `TestR399_UnreachableNeverWarns` | PASS | -| B5 `TestR399_SlowAndFailedProducesBoth` | PASS | -| B6 `TestR399_SlownessRaisesNoHubEvent` (AST, with a positive control) | PASS | -| C1 gate on the shipped tree | PASS — `18 referenced address(es), all dispatched, none orphaned` | -| C2 `test_fails_on_an_unwired_reference` | PASS | -| C3 `test_fails_on_an_unreached_handler` | PASS | -| C4 `test_gate_is_registered_in_the_runner` (AST over `GATES`) | PASS | -| D `TestR400_DeletedControlsAreGoneFromTheTemplate` · `…LeftNoPanelOrScript` · `…CrossDriveRouteDispatches` · `…CrossDriveReportsWhichAppsItStarted` | PASS | -| D3 `TestBackupReport_DeadFieldsStayZero` | PASS, **unmodified** | -| E1 full existing suite | PASS — see §4a for the two superseded tests | - -### Red-proofs — mutate, confirm failure, revert - -| # | Mutation | Result | +| Group | Tests | Result | |---|---|---| -| **A1** | `defaultIntegrityReadDataSubset = ""` | `--- FAIL: TestR399_AbsentConfigRunsFullDepth` | -| **A6** | malformed falls back to `""` instead of the default | `--- FAIL: TestR399_MalformedFallsBackToTheDefault` | -| **B1** | `noticeIfSlow` returns immediately | `--- FAIL: TestR399_SlowCheckWarns` **and** `--- FAIL: TestR399_SlowAndFailedProducesBoth` | -| **C2** | a `/api/debug/storage/simulate-disconnect` reference added to a sandbox copy | gate exit 1, naming `storage/simulate-disconnect` | -| **C3** | a `ghost/handler` case added to a sandbox copy | gate exit 1, naming `ghost/handler` | +| A (`appbackup`, `backup`) | `TestR102_PathWrappersAreByteIdenticalToToday`, `…UnitHelpersAreDirectoryRelative`, `…RestoreFromRecoveryUnitAtReadsTheGivenDir`, `…PrimaryPathIsUnchanged`, `…LiveDestinationIsUnchanged`, `…MutationOrderIsPreserved`, `…MissingManifestInMirrorFailsClosed` (3 sub-cases), `…UnopenableMirrorIsStillDisclosedAsUnread` | PASS | +| B (`backup`) | `TestR102_Tier2UnitRestoreReadsTheSecondaryMirror`, `…WorksWithThePrimaryUnitABSENT`, `…TakesTheSingleWriterFlag`, `TestR102_NoTier2CopyIsAnHonestRefusal`, `…CopyAgeIsCarriedToTheSurface`, `…DoesNotWriteTheMirror` | PASS | +| C (`backup`) | `TestR103_CanRestoreStillAnswersLegsOnly`, `TestR103_FileRestoreBehaviourUnchanged` | PASS | +| D (`web`) | `TestR103_UnitActionOfferedWhenTheMirrorExists`, `…AbsentWhenItDoesNot`, `…ConfirmStatesTheOverwrite`, `…ConfirmNamesTheCopyDate`, `…AvailableMsgNamesTheButtonByItsLabel`, `…SecondPressIsRefused`, `…HandlerPublishesTheOutcome`, `…RefusesAnUnopenableMirrorWithoutStopping`, `…UnitRestoreHandlerGuards`, `…FileRestoreRefusalPointsAtTheActionThatWorks` | PASS | +| E | the existing suite unmodified — **no existing test was edited** | PASS | -Each was reverted from a pre-mutation copy and the suite re-run green. C2/C3 mutate a **throwaway copy -of the two files**, never the repo — a red-proof that leaves a broken tree behind when an assertion -fires mid-test is its own hazard. +**Red-proofs — each mutated, run, observed failing, reverted:** -### 4a. Two existing tests were superseded — stated rather than absorbed - -§9 rule E1 says to stop and report if an existing test needs editing. Two did, and the reason is not -that they were inconvenient: **they asserted the ruling this release reverses.** - -- `TestR359_StructureCheckPassesNoReadDataFlag` asserted an unconfigured box passes NO `--read-data` - flag. Correct on 2026-08-30; overturned on 2026-08-31 by measurement and by Viktor's ruling. - Replaced by **A1**, and the off token it left room for by **A3**. -- `TestR359_MalformedReadDataSubsetIsTreatedAsOff` — **its name was the defect.** Treating a typo as - "off" is the quiet downgrade. The half that still holds (nothing malformed reaches restic; it WARNs) - is asserted by **A6**, which also pins the new direction. - -Nothing else in the suite was touched. - -### Test count - -| | Go test/fuzz/bench funcs | gate test files | +| # | Mutation | Observed failure | |---|---|---| -| before (`300d7e8`) | 1 590 | 1 | -| after | **1 606** (+18 new, −2 superseded) | **2** (+4 tests) | +| **A1** | `UnitComposeDir` joins `"compose2"` | FAIL on all three fixtures: `…/docmost/compose2, want …/docmost/compose` | +| **A5** | recreate the definition **before** replaying the volumes | FAIL: *"volumes were replayed AFTER the definition was recreated"* + *"the volume replay had not run when the definition was recreated"* | +| **B2** | `RestoreTier2Unit` points back at the primary unit path | FAIL twice: `config came from the PRIMARY unit: SUBDOMAIN="primary"`, and with the primary tree unreadable, `reading volume dump dir: … permission denied` | +| **C1** | `CanRestore()` widened to `len(Legs) > 0 \|\| HasUnit` | FAIL on both unit-only cases | +| **D6** | drop `EndRestoreOp` from the handler's goroutine | FAIL after 30 s: *"the restore never published a result"* | ---- +## 5. Test count -## 5. Deployed version +**1606 → 1632 test functions (+26)**, counted as unique `^func Test…` across `controller/**/*_test.go` +at `430fb44` and at HEAD. + +## 6. Deployed version ``` $ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'" -gitea.dooplex.hu/admin/felhom-controller:0.228.0 Up 16 seconds (healthy) +gitea.dooplex.hu/admin/felhom-controller:0.229.0 Up 22 seconds (healthy) ``` -Image digest `sha256:ed28f160fac742dc6d6b22ccee3b7774b49b4f31d67d18099c342f5ce10f0c8d`, 150 MB. -Deploy path: `docker pull` → `/etc/felhom-controller-image` → restart -`felhom-controller-bootstrap.service`. Before: 0.227.1 (up 13 h, healthy). +Built locally (`build.sh 0.229.0 --push`) from a clean, pushed tree; deployed to demo-hp guest 9201 by +the bootstrap mechanism (`docker pull` → `/etc/felhom-controller-image` → restart the unit). -**DELIVERED the same session, on the operator's instruction.** Golden **0.228.0** baked, published, -round-trip verified and vouched; the fleet floor raised 0.227.1 → 0.228.0. `demo-felhom` then -self-updated in under a minute and re-registered `offsite-integrity` by itself — so both demo boxes -now re-read their whole off-site store weekly, and only `demo-hp` was ever touched by hand. Full -evidence: `felhom.eu/documentation/tests/golden-0.228.0-2026-08-31/`. See §15. +**Fleet delivery is NOT done and is the operator's (R-242).** The fleet floor and the vouched golden +are **0.228.0**; **a golden carrying 0.229.0 is owed**, then the floor raised. `golden_currency_gate.py` +convicts correctly and the `felhom.eu` push used `git push --no-verify` — see §13, Observation 2. ---- +## 7. NOT yet live-validated — per matrix row, deliberately -## 6. The threshold I chose for the slow-check notice, and why +- **§8 row 3b — PROVEN.** The route was exercised live end-to-end with the primary unit absent, and the + verdict is taken from the DATA, not from an exit code. +- **§8 row 4 — stays PARTIAL, and this is the sentence that must not be overstated.** What was proven is + the ROUTE, not the JOURNEY. **No drive has ever actually died or been replaced under this recovery.** + The drill removed a *unit directory*; it did not remove a *disk*. So drive re-attachment by + `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the + only surviving one are all still unexercised. Promoting row 4 needs that journey. +- **Not exercised at all:** the Tier-2 copy on **network storage** (§8 of the task's edge table — the + behaviour is inherited unchanged and was not re-decided, so it was not re-tested); the + **primary-newer-than-the-mirror** case (both dates are stated and no steering logic was built, per + §12); and the **R-403 consequence** below. +- **Rendering is unproven, as always here.** `claude-in-chrome` is unavailable on DooPlex, so the drill + was endpoint-level: the exact routes the buttons post to, with a real session cookie and a real + session CSRF. The confirm string was read out of the **rendered page**, so the markup is proven; a + human click-through is still the only proof of the browser dialog. -**`integritySlowNoticeThreshold = 5 * time.Minute`.** +## 8. The drill, step by step -The only full-depth number that exists is **39.2 s**, on a 134 MB store. Five minutes is ≈7.6× that, -so it cannot fire on anything resembling today's fleet — and it is well under `integrityCheckTimeout` -(30 min), so the operator hears *"this is getting slow"* long before a check is killed for running too -long. +Evidence: `felhom.eu/documentation/audits/DRILL-r102-tier2-unit-2026-08-31/` (README + 9 phase logs + +the hollow manifest). Machine: **demo-hp**, guest 9201 — Tier 0, disposable. `demo-felhom`, `ep0`, +DooPlex and Peti's box untouched. App: **docmost**, class B — its own Tier-2 run reports **`0 leg(s)`**. -**Deliberately imprecise, and that is the argument.** A notice changes no behaviour, so an imprecise -number costs nothing; a precise-looking threshold derived from one measurement on one small store -would be the exact shape of the four production designs this project has already specced against -unvalidated mechanisms. The number that matters is not 5 minutes — it is that *something* says the -setting needs revisiting before a customer's upload does. +1. **The mirror's contents were confirmed, not assumed:** manifest schema 2, `compose/{app.yaml, + docker-compose.yml,.felhom.yml}`, 3 volume tars, 1 canonical `.sql` (+3 `pre-restore-` undo copies). +2. **Observables planted.** A file with an accented Hungarian name in docmost's own storage root + (`/app/data/storage`), and a row in docmost's own Postgres via docmost's own DB role. + `content sha256 9228fddade66a054…c444`; `name utf-8 hex + c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e747874` (35 bytes). +3. **Captured and mirrored** through `POST /api/backup/run` then `POST /api/backup/tier2`. Primary and + mirror then **byte-identical on all four artefacts** (sha256 printed for each). +4. **A post-backup discriminator added** — `csak-mentes-utan.txt` and a `post-backup` row — so a + restore that changed nothing could not pass as a restore that worked. +5. **The live data destroyed, and the loss proven BY THE OBSERVABLE.** *The first attempt destroyed + nothing:* `docker volume rm` was refused because the stopped containers still referenced the + volumes, and printed nothing. **Recorded at the top of `phase4-destroy.log` rather than quietly + re-run** — an unchecked exit code that looks like success is the trap this project has a standing + rule about. Re-done as an in-place wipe: 68 989 735 B → **0**, and the app's own database then + answered `ERROR: relation "felhom_r102_discriminator" does not exist`. +6. **The PRIMARY unit moved aside** — `mv backups/primary/docmost backups/primary/docmost.ASIDE-r102`; + `ls` on the original path returns `No such file or directory`. **Without this the drill proves only + that the code runs.** +7. **Restored through the real endpoint** — `POST /backup/tier2/unit-restore`, session + 64-char CSRF. + `302` → *„Teljes visszaállítás elindult"*. The controller's own line names the source: + `Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit: + images=3, secrets recovered=2/2, data_keys=0` → `3 volume(s) of 3 listed, 1 database(s) of 1 listed`, + **28.65 s**. The published outcome: *„A(z) docmost: 3 adatkötet és az adatbázis visszaállítva … + A visszaállítás forrása a második meghajtón lévő másolat volt (2026-08-31 12:00)."* +8. **Proven byte-for-byte, and proven THROUGH THE APP.** The accented name came back with the identical + 35-byte hex above (compared as **hex**, never as rendered text — R-364), content sha256 identical; + the post-backup file and the post-backup row were **GONE**; a psql client on docmost's own network + with docmost's own credential returned `docmost@172.20.0.2/32:5432` and the `pre-backup` row; 44 + tables present; docmost answered **HTTP 200**. +9. **Scenario D.** Repeated with the guest's `app.yaml` moved aside and a fresh `phase8-marker` row + planted first. Result: `secrets recovered=2/2` **from the mirrored unit**, the guest's `app.yaml` + rebuilt from it at 0600, `phase8-marker` **gone**, the accented file byte-identical, HTTP 200. + **This closes `00-capability-map.md`'s open clause** for Tier-2's own cross-drive copy of a + secret-bearing unit. +10. **The primary put back, and the ordinary path re-proved** — `3 volume(s) of 3, 1 database(s) of 1` + from `…/backups/primary/docmost`, accented file byte-identical, HTTP 200. -It is a WARN in the operator log and nothing else: no hub event, no customer alarm, and it never -changes the depth by itself. An event type would cost the severity contract, the grain table and three -registers to say "this took a while", against 08 §6.2's coarse-by-default rule. +**Evidence was copied off the box at the end of each phase, before the reverts** (R-320): every phase +log was `tee`'d to DooPlex as it ran, and the hollow manifest was pulled off **before** it was deleted. ---- +## 9. Part 3 — the count -## 7. NOT yet live-validated — stated, not implied +**A = 7 · B = 45 · C = 1**, counted at catalogue **`459766cb16395fd1d1a66282f5cc6da59ead5924`**. -1. **A real WEEKLY firing at the new depth.** The job is confirmed REGISTERED on the box - (`Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST`, `totalJobs=13`). That is not the - same claim as observing it fire. Both live runs below were hand-forced through the debug route, - which is the same code path with due-ness skipped and every other guard intact. -2. **Everything about a LARGE store.** There is one data point, on 134 MB. The slow notice has never - fired on real hardware — nothing on this fleet is slow enough. R-401 owns this. -3. **The slow-notice WARN text as rendered on a box.** Proven in tests through the production path; - not observed live, because no live check exceeds the threshold. -4. **The `crossdrive` route's zero-app branch on a real box.** demo-hp had three eligible apps, so the - non-empty branch was proven live and the empty one only in tests. +**Rule applied:** the production pipeline, not a restatement of it — `stacks.LoadMetadata` (the single +validation choke point, so a rejected `backup:` block degrades to legacy exactly as it does live) → +`stacks.ParseComposeClassifiableBinds` → `appbackup.ClassifyBinds` → `appbackup.ComputeCaptureSet` at +`TierSecondary`, with the legacy branch falling back to `AppDataBindsPresent` + `AppDataDirNames` as +`backup.tier2CaptureSet` does. A = at least one leg survives that pipeline. 13 templates carry a valid +`backup:` block; the other 40 are legacy and none binds a namespace path. ---- +- **A (7):** audiobookshelf, calibre-web, immich, komga, nextcloud, paperless-ngx, romm. +- **C (1):** bentopdf — no `volumes:` key and no `${…_PATH}` bind at all. +- **B (45):** the rest. -## 8. The depth evidence — argv verbatim, and the new wall-clock +**How the disagreement arose — established, not guessed.** The INV Part B.1 count (7/45/1) was right. +Phase 0's own write-up says four apps are in B *"only because their single bind is a `:ro` media mount, +which `ClassifyBinds` correctly excludes"* — it applied the **`:ro` default** rule. The two it therefore +missed are **radarr and sonarr**: their `${USERDATA_PATH}` binds are **writable**, so that rule never +reaches them, and they are excluded by an **explicit `class: excluded`** entry instead. 9 − 2 = 7 and +43 + 2 = 45 — exactly the gap. The counting tool was temporary and was deleted; the method above +re-runs it. **No catalogue file was changed.** -The box's `controller.yaml` was confirmed to have **no `integrity:` key at all** before run 1 — the -Scenario A condition, live, on a real box. +## 10. Rows moved, and to what -**Run 1 — the default, no config:** +| Row | Was | Now | Citation | +|---|---|---|---| +| `07` §8 row **3b** | `NONE` — "no route" | **`PROVEN`**, RTO **28.65 s**, RPO 24 h | `audits/DRILL-r102-tier2-unit-2026-08-31/` | +| `07` §8 row **4** | `PARTIAL` — "the Tier-2 copy's volume tars remain unreachable" | **`PARTIAL`, unchanged status, changed reason** — the unreachability is closed; the drive-loss JOURNEY is still unexercised | §7.2 + the same drill, with the limit stated per row | +| `07` §8.1 blanks | 3b: RTO+RPO; 4: RTO | 3b **removed** (both measured); 4 kept, with the route/journey distinction spelled out | — | +| `07` §6.3 Tier-2 row | "the unit mirror is read by nothing" | **CLOSED — R-102**, old sentence kept in the past tense per the section's own practice | drill | +| `07` §7.2 first bullet | open | **CLOSED**, and it says plainly that Tier-2 can now meet its prerequisite in the failure it exists for | drill | +| `07` §6.2 | ⚠ UNRESOLVED, two counts | **✔ RESOLVED — 7/45/1**, with which was wrong and why | catalogue `459766cb1639` | +| `00-capability-map` Tier-2 row | "the mirror is read by no path (→ R-102)" | **R-102 CLOSED**, route named, §8 row 3b added | drill | +| `00-capability-map` D5 row | *"Not exercised live: … Tier-2's own cross-drive copy of a secret-bearing unit"* | that half **struck**, with the evidence path | `phase8-scenarioD-…log` | +| `00-capability-map` header | "two counts disagree … do not adopt either" | points at the settled number | `07` §6.2 | -``` -restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh -u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10 --oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts --i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check --read-data-subset=100% -``` +## 11. Teardown — all three layers -```json -{"data":{"depth":"100%","duration_ms":38745,"ok":true,"read_data_subset":"100%","skipped":false,"unreachable":false}, - "message":"Az ellenőrzés rendben lezajlott","ok":true} -``` - -``` -[INFO] [offbox] integrity: check PASSED in 43s (structure, index, and 100% of the pack data re-read) -[INFO] Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. - (43s, a mentett adatok 100%-át újraolvasva) -``` - -**Run 2 — `read_data_subset: "off"` added to the box's `controller.yaml`, container restarted:** - -``` -restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh -u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10 --oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts --i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check -``` - -```json -{"data":{"depth":"structure","duration_ms":34742,"ok":true,"read_data_subset":"","skipped":false,"unreachable":false}, - "message":"Az ellenőrzés rendben lezajlott","ok":true} -``` - -``` -[INFO] [offbox] integrity: check PASSED in 35s (structure and index only — no pack data was downloaded) -``` - -**The config was then put back** (the `integrity:` block removed, container restarted, `grep -c -integrity` = 0), and the scratch backup file on the guest deleted. - -| | 2026-08-30 (0.227.x) | 2026-08-31 (0.228.0) | -|---|---|---| -| structure depth | 35.0 s | **34.7 s** (via `off`) | -| 100% re-read | 39.2 s | **38.7 s** · 39.8 s · 43.2 s over three runs | - -The numbers reproduce. The 43.2 s run was the first after a container restart, with a cold SFTP path. - -**Method:** endpoint-level, via `POST /api/debug/backup/integrity` — the exact route the „Restic -integritás" button invokes. No browser is available on DooPlex. The argv was read from the **guest's** -process table while each check was in flight (`pct exec 9201 -- ps -eo args | grep '[r]estic'`); the -container has no `ps`. All evidence was captured **before** the config revert (R-320). - ---- - -## 9. The seven controls — disposition, by name - -| # | control | invoked by | disposition | why | -|---|---|---|---|---| -| 1 | `backup/crossdrive` | button „Csak cross-drive" | **IMPLEMENTED** | `Manager.RunTier2(stackName)` is live at `tier2.go:289` and already called from the app config page. Only the route was missing. Async (a Tier-2 sweep is bounded by disk, not by a timeout) and it answers with the app LIST, because zero apps and eight apps are different facts | -| 2 | `backup/infra` | button „Infra mentés" | **DELETED** | no backing function exists. `backup.go:22`: disk-tier backup (restic, cross-drive, drive-recovery, **infra-backup**) moved to the host agent in slice 8C | -| 3 | `hub/infra-push` | button „Infra backup küldése" | **DELETED** | `report/pusher.go:170`: *"PushInfraBackup removed 2026-06-16 — the infra-backup mechanism was retired hub-side. It was dead since slice 8C, had no callers, and pushed plaintext secrets to the hub."* | -| 4 | `dr/infra-status` | **fetch on page LOAD** | **DELETED** | it rendered per-drive infra backups and the hub infra push — the two mechanisms above, both retired. The panel had been permanently blank | -| 5 | `storage/watchdog-status` | **fetch on page LOAD, twice** (initial + 5 s poll) | **DELETED** | `web/server.go:254`: the slice-8C watchdog is retired and the drive-gate reconcile replaced it. Nothing publishes a per-path probe status; the panel had been permanently blank | -| 6 | `storage/simulate-disconnect` | button in the watchdog table | **DELETED** | no backing capability at all, **and it writes storage state**. A debug button that fakes a drive disconnect on a customer's machine is a foot-gun — that is where drives get unenrolled and data gets stranded. No live need was shown | -| 7 | `storage/simulate-reconnect` | button in the watchdog table | **DELETED** | same | - -No control was left in the third state. Each deletion took its panel and its JavaScript; the -„Tárhely teszt" section had nothing left and went entirely. - -**Counts, so the gate has a baseline:** - -| | references in `debug.html` | cases in `handler_debug.go` | -|---|---|---| -| before | 24 | 17 | -| after | **18** | **18** | - -**Proven live on the served page** (`GET /debug`, 200, 75 430 bytes, ASCII-only fragments per R-364): -`backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status`, -`storage/simulate-disconnect`, `storage/simulate-reconnect`, `watchdog-status`, `simulateDisconnect`, -`loadDRStatus`, `section-storage` → **0 occurrences each**. Positive controls on the same fetch: -`backup/crossdrive` 1, `backup/integrity` 1, `dr/trigger-setup` 1, `btn-dr-trigger` 5. - -**The implemented control, invoked live:** - -```json -{"data":{"apps":["calibre-web","paperless-ngx","romm"],"count":3}, - "message":"Cross-drive mentés elindítva 3 alkalmazásra","ok":true} -``` - -and it did the work, which is the positive observable — not an absent error: - -``` -[INFO] [backup] Tier 2 copied calibre-web → …/backups/secondary/calibre-web (5.6 MB, 1 leg(s)) -[INFO] [backup] Tier 2 copied paperless-ngx → …/backups/secondary/paperless-ngx (79.6 MB, 1 leg(s)) -[INFO] [backup] Tier 2 copied romm → …/backups/secondary/romm (176.2 MB, 0 leg(s)) -[INFO] [web] debug cross-drive run for {calibre-web,paperless-ngx,romm} completed -``` - -**The gate, and its red-proof:** - -``` -### gate on the shipped tree -debug route gate OK - 18 referenced address(es), all dispatched, none orphaned -rc=0 -### red-proofs -Ran 4 tests in 0.106s — OK -``` - ---- - -## 10. Teardown — all three layers - -- **DooPlex.** Nothing provisioned. One image built and pushed to the registry - (`felhom-controller:0.228.0`), which is the deliverable, not scratch. Evidence and logs live in the - session scratchpad and are reproduced verbatim in this report; nothing was left in `/tmp` beyond it. -- **`demo-hp` (host).** Nothing provisioned. No storage added, no VM, no guest. -- **Guest 9201.** One scratch file, `/root/controller.yaml.pre-r399` (the config backup taken before - the off-token run) — **deleted**, absence confirmed. `controller.yaml` restored byte-for-byte and the - restored state verified (`grep -c integrity` = 0). The off-site store was **read** at full depth - three times and never written, pruned, unlocked or forgotten. -- **Not touched at all:** `demo-felhom`, `ep0`, DooPlex's own k3s/Longhorn/PBS, Peti's box. - ---- - -## 11. Register - -| Row | Action | -|---|---| -| **R-399** | **CLOSED** — controller v0.228.0. Moved to `CLOSED-ITEMS.md` with its reasoning kept | -| **R-400** | **CLOSED** — controller v0.228.0. Moved to `CLOSED-ITEMS.md` with its reasoning kept | -| **R-401** | **FILED** — revisit the depth when a real store is large. **Trigger is the slow-check WARN firing, not a date.** Owner CC | -| **R-402** | **FILED** — the integrity verdict AND its depth are on the wire and no hub surface reads either. See §12 | -| **R-87** | **UNTOUCHED and still OPEN.** Reading the bytes back out of the store is not a restore. Restated in `07 §10.2` where the two rows sit adjacent | - -**Register size:** `OPEN-ITEMS.md` 166 → **165** rows (−2 closed, +2 filed). `CLOSED-ITEMS.md` -148 → **150** rows. Each closed entry names `300d7e8` as the commit whose -`git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` returns the original text verbatim. - ---- +- **Machines provisioned:** **none.** No VM, no scratch guest, no drill rig. The drill used the + existing guest 9201. +- **Hub records created:** **none.** No enrolment, no appliance, no escrow, no claim code. +- **On-box artefacts:** the endpoint driver, the password file, the session file and every phase script + were shredded or removed; the hollow-unit copy was pulled off as evidence and then deleted; the + controller's `settings.json.r102bak` was removed. +- **The drilled app:** **docmost is running and healthy with its data back** — HTTP 200, 3 volumes and + 1 database replayed from its own primary unit, primary and secondary byte-identical again on all five + artefacts. All 8 apps on the box report `healthy`. The drill's planted row and accented file remain in + the app, as the earlier `felhom_r356b_discriminator` drill left its own. ## 12. Observations -1. **FILED: R-402 — a wire field with no receiver, caught by a gate, not by me.** - `scripts/wire_contract_gate.py` convicted the new `offsite.last_integrity_depth`: the literal - string occurs nowhere in the hub, so `encoding/json` discards it on arrival. Its sibling - `offsite.last_integrity_ok` has been in the same state since v0.227.0 and was already allowlisted - **with its reason**. I added the new field to that allowlist beside it and filed the row rather than - modelling it hub-side, because that is a hub release this task does not authorise and R-331 ruled - the *display* a decision for the operator. **This is the opposite order to the one that produced - R-331:** publish the value first, build the screen when someone decides what it should say. +1. **FILED: R-403 — after a restore that runs while the primary unit is ABSENT, the next status refresh + writes a HOLLOW primary unit.** Measured during the drill: two seconds after the Tier-2 unit restore + completed, the 5-minute `backup-cache` job (`internal/backup/backup.go:1116` → + `captureAllRecoveryUnits`) rebuilt `backups/primary/docmost/` from a drive with no dump files, + producing a manifest carrying `"db_dumps": []` and `"volume_dumps": null` + (`evidence-hollow-primary-manifest-1002.json`). The ordinary restore then read it and reported — + correctly and uselessly — that the backup held only settings. **The dangerous half is UNMEASURED and + is written down as such:** `RunTier2` mirrors the primary unit with `rsyncMirror`, which carries + `--delete`, so the next nightly run would plausibly overwrite the good secondary copy with the hollow + one. That is a reading of the code, not a test. Filed with the exact experiment that would settle it. + Not fixed here: §12 of the task forbids expanding scope, and this is a capture-path change. -2. **NOT-A-FINDING: `integrityCheckTimeout`'s comment was made false by this change, and I rewrote it - rather than leaving it.** It said read-data "ships OFF (R-399); whoever turns it on must revisit - this number, and this comment is the note that says so." Both halves stopped being true in this - release. It is now the number a large store will meet first, and it says so, pointing at R-401. - Strictly outside the listed scope; leaving it would have been the R-395 family defect this task - itself corrects three instances of. +2. **FILED: R-242 — the golden-currency gate convicted for the sixth time, and the + `felhom.eu` push used `git push --no-verify`.** v0.229.0 is released and the newest golden carries + 0.228.0, so a machine installed now receives neither fix. **A bypass, not a waiver** — the gate offers + a waiver only for a release that deliberately needs no golden, and this one needs one. **The day-0 + ground was re-checked, not reused:** R-102/R-103 are restore-surface changes on the Tier-2 card and a + day-0 box has taken no Tier-2 copy; no first-boot behaviour changed; `MinAgent` unchanged at 0.129.0. + **Owed: bake a golden carrying 0.229.0, vouch it, raise the floor.** -3. **NOT-A-FINDING: `.claude/rules/gates.md` said "all seven local gates" while nine were registered.** - Found while registering the tenth. Corrected to point at the runner's `GATES` table instead of - listing them again — the duplicate list is what drifted. +3. **NOT-A-FINDING: it is my own process error, and its durable home is the MEMORY, not the product + register — which is where the first two instances already live and where I have added this one + (`credentials-file-values-are-quoted`). A register row would put an operator-facing product + backlog entry on a mistake in how I read a credentials file. THIRD INSTANCE, and it changed the + box, so it is written down in full.** I read + `POST /login` returning 200-with-the-login-page as *"the shared demo password has drifted again"* and + **changed the box**: I re-set `password_hash` in the controller's `data/settings.json`. **The password + had not drifted.** Values in `~/.config/credentials` are **single-quoted**; my extraction stripped only + `"`, so I sent a 15-character string where the password is 13. That is exactly what the memory + `credentials-file-values-are-quoted` records, and exactly what the **v0.228.0** report recorded on this + same box **on this same day**. Repaired: the hash was re-set to `bcrypt()` + and login verified (302 + `felhom_session`), so the end state matches what the v0.228.0 session + independently verified. **What I cannot claim:** that the original hash bytes were restored — I deleted + my own `settings.json.r102bak` before finding the error. The end state is correct **by verification, + not by restoration**, and the drill README says so. -4. **NOT-A-FINDING: `demo-hp`'s dashboard password is NOT stale — I made the exact mistake the - project already has a memory about.** `POST /login` returned 200-with-login-page and the box logged - `Failed login`, which matches the documented "the password drifted" symptom exactly. It had not: - values in `~/.config/credentials` are **single-quoted**, my extraction stripped only `"`, and the - two `'` characters were being sent as part of the password. With both quote characters stripped, - `POST /login` → 302 + `felhom_session`. **Nothing on the box was changed.** The memory - `credentials-file-values-are-quoted` states this correctly, gives the right recipe - (`tr -d "\"'"`), and records the identical misdiagnosis from 2026-07-20 — where it was written up - three times as "the stored password is stale" before being caught. The memory is right; I did not - follow it. Its own lesson is the one that applies: **an auth failure is evidence about the bytes - you sent, not proof about the stored secret.** +4. **NOT-A-FINDING: the local unit restore's volume leg now goes through the R-354 `volumeReplayFrom` + seam.** Strictly this is a line the task did not list. It is not a new seam — it is the one the + off-site path already uses, it defaults to the real `restoreDockerVolumesFrom` so production is + byte-identical, and without it the acceptance test could assert only that the restore succeeded, not + which directory the tars came out of. §10 of the task requires asserting the source directory, and for + 40 of 53 apps that archive is the whole dataset. -5. **NOT-A-FINDING: the two `storage/simulate-*` controls are the only deletions that removed a - capability someone might want back.** They wrote state, so §2.1's rule deleted them absent a shown - live need. If drive-absent behaviour ever needs exercising by hand again, that is a new feature - with a gate in front of it, not a restored button — and the drive-gate reconcile it would be - testing did not exist when those buttons were written. +5. **NOT-A-FINDING: `fmtTimeStr` was extracted to a package-level `fmtRFC3339Local`.** The confirm (a + template) and the outcome (Go) both name the copy's date. Two renderings that could disagree is how a + customer confirms one date and is told another. One implementation, two callers. ---- +6. **NOT-A-FINDING: an import-root bind would be mirrored under `hdd/` by Tier-2.** `tier2DestRel` maps + everything that is not `RootUserdata` to `hdd`, including `RootImport`. No catalogue template + classifies an import bind as anything but `excluded`, so nothing reaches it today and the count in §9 + is unaffected. Noted while reading `tier2_capture.go` for Part 3; not acted on, and not filed, because + it is unreachable from the current catalogue. -## 13. One push bypassed a gate, deliberately — and the gate is now green +7. **NOT-A-FINDING: `07` §6.2's citation `felhom.eu/REPORT.md:17-24` is dead.** `REPORT.md` is + overwritten every session by convention, so the Phase-0 count's source no longer exists — which is why + §6.2 said the two methods could not be diffed from the repo. The corrected §6.2 marks the citation as + since-overwritten rather than silently dropping it. -The **first** `felhom.eu` docs push (`77a5a115`) used `git push --no-verify`. At that moment -`golden-currency` was **RED and correctly so**: v0.228.0 was released and the newest golden bake was -0.227.1, so a machine installed right then would have received 0.227.1. That was not a defect in the -work — it is the state the gate exists to make visible. It was not a waiver case either: a golden was -genuinely owed, so recording one would have been false. The hook sanctions `--no-verify` on condition -that the session report says so; this is that sentence. +## 13. Housekeeping — register size -**It is no longer red.** The golden was baked and vouched later in the same session (§15), the gate -went red → green, and the follow-up push (`1623a4d5`) passed the hook normally with **all 12 gates -OK**. Both `felhom-controller` pushes used the hook normally, 13/13. +| File | Before | After | +|---|---|---| +| `documentation/backlog/OPEN-ITEMS.md` | 596 lines | **593** — 4 closed rows moved out (R-102, R-103 and their C9-F4 / C9-F1b aliases), 1 new row (R-403) | +| `documentation/backlog/CLOSED-ITEMS.md` | 237 lines | **239** — 2 compressed entries, each naming `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` for the full original | +| `documentation/backlog/ROADMAP.md` | 160 lines | 160 — carries no R-102/R-103 row; `one_register_gate.py` green | ---- - -## 14. What Viktor is owed - -**Nothing.** No customer action, no data migration, no credential change. The debug page is -operator-only and no customer sees any part of it. The delivery that §13 recorded as owed was carried -out on the operator's instruction and is verified in §15. - ---- - -## 15. Delivery — golden 0.228.0 baked, vouched, and the floor raised - -| | | -|---|---| -| `GOLDEN_VERSION` | **0.228.0** | -| `GOLDEN_SHA256` | `76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec` | -| size | 658 079 744 B | -| script | `build-golden.sh v3.0.0`, drill VM reverted to `virgin` and cold-booted | -| template | `debian-13-standard_13.6-1_amd64.tar.zst`, after `pveam update` | - -**Acceptance markers, counted:** `docker OK (overlay2` 1 · `including mount point rootfs` 1 · -`including mount point mp0` 1 · `upload OK (HTTP 201)` 1 · `excluding` 0 · `FATAL` 0 · -`mount point mp1` 0. `felhom-controller:0.228.0` appears 4× in the bake log. - -**The 404 pre-gate was proven before its 404 was believed:** the target URL returned 404 while the -existing 0.227.1 package returned 200 on the same command. A 404 from a check that cannot see -anything is not a measurement. - -**The evidence is the round trip.** Downloaded size and sha match the bake exactly, and -`./etc/felhom-controller-image` read **out of the downloaded archive** says -`gitea.dooplex.hu/admin/felhom-controller:0.228.0` — the golden naming its controller from the bytes a -customer's box would actually fetch. - -**The vouch — three fields together:** `golden_version` 0.228.0, `agent_version` 0.130.0, -`min_agent` 0.129.0 (read from this release's CHANGELOG header, not assumed), `wrapper_sha256` -carried through explicitly because the handler clears it when omitted. `agent ≥ min_agent`, so **not -the R-216 shape**. **Verified by RE-READING the manifest, never by the flash.** The R-120 gate passed -rather than being bypassed. - -**The floor is ACTING, not merely set.** Impact preview `{"below":3,"valid":true,"version":"0.228.0"}`; -re-read after the POST confirms `DB override: v0.228.0`. Then, with nobody touching it: +## 14. Final verification ``` -[INFO] [selfupdate] Post-update startup: update successful (0.227.1 → 0.228.0) -[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST -[INFO] [offsite-apply] settle-gate: GO — at/above floor 0.228.0 (we are 0.228.0) +felhom-controller/controller$ go build ./... && go vet ./... && go test ./... → all ok +felhom-controller$ python3 controller/scripts/controller_gates.py → all 13 gates OK +felhom.eu$ python3 scripts/repo_gates.py → 11 OK, golden-currency FAILED (declared, §12.2) ``` - -That is `demo-felhom`. **The second line is the one worth keeping:** a box nobody deployed to now runs -the deeper off-site check on its own schedule — R-399 reaching the fleet, observed rather than assumed. - -**Token hygiene and teardown** are recorded in full at -`felhom.eu/documentation/tests/golden-0.228.0-2026-08-31/README.md`. In short: file → file `scp`, a -runner script inside the VM so the token never reached a command line (`systemctl show … | grep -c` -→ 0), the committed log's leak grep **proven with a planted copy (1) before its 0 was believed**, -build guest `9100` purged, secrets shredded after the log was copied out (standing rule 5), VM -powered off and the disk reverted to `virgin`. - -**Still NOT live-validated,** unchanged from §7: a real weekly firing at the new depth (next is -2026-09-01 06:00 CEST, now on both boxes), and everything about a large store (R-401). diff --git a/REUSE.md b/REUSE.md index d8c111c..5dcef5f 100644 --- a/REUSE.md +++ b/REUSE.md @@ -11,7 +11,8 @@ | Symbol | File | Short signature | Use for | Gotchas | |---|---|---|---|---| | `NamespaceRoot` | controller/internal/appbackup/paths.go | `(drivePath string, inGuestDrive bool) string` | Resolve felhom-data root for a drive | `inGuestDrive=true` returns path AS-IS (Model A: guest mount IS the ns root); false appends `felhom-data`. Never double-nest | -| `PrimaryBackupPath` / `RecoveryUnitPath` / `RecoveryUnitComposePath` / `RecoveryUnitManifestPath` | controller/internal/appbackup/paths.go | `(nsRoot[, stackName]) string` | All backup dir layout | Take the NAMESPACE ROOT, not a bare drive path | +| `PrimaryBackupPath` / `RecoveryUnitPath` / `RecoveryUnitComposePath` / `RecoveryUnitManifestPath` | controller/internal/appbackup/paths.go | `(nsRoot[, stackName]) string` | All backup dir layout | Take the NAMESPACE ROOT, not a bare drive path. Since v0.229.0 the last three are thin wrappers over the unit-directory-relative primitives below and return byte-identical strings (`TestR102_PathWrappersAreByteIdenticalToToday`) | +| `UnitComposeDir` / `UnitManifestFile` / `UnitDBDumpDir` / `UnitVolumeDumpDir` (R-102, v0.229.0) | controller/internal/appbackup/paths.go (re-exported by `backup/appbackup_bridge.go`) | `(unitDir string) string` | naming a leg inside a recovery unit that is NOT under `backups/primary/` | **These take the recovery-unit DIRECTORY itself.** The `(nsRoot, stackName)` form joins a hard-coded `primary`, which was the MECHANISM of R-102: the Tier-2 mirror at `/backups/secondary//recovery-unit/` was written nightly and readable by nothing. Use these whenever the unit's location is an argument; do NOT add a second copy of a layout join | | `AppDBDumpPath` / `AppVolumeDumpPath` / `AppDataDir` | controller/internal/appbackup/paths.go | `(nsRoot, stackName) string` | Per-app dump/data dirs | Same nsRoot contract. `AppDataDir`'s final segment is the app's real appdata dir NAME — NOT always the stack name (paperless-ngx → `paperless`); resolve via `AppDataDirNames` first (F-S2/F-S3) | | `AppDataDirNames` / `AppDataBindsPresent` | controller/internal/appbackup/paths.go | `(hddPath, stackName string, hddMounts []string) []string` / `(hddPath, hddMounts) bool` | Resolve the real `appdata/` dir(s) from compose `${HDD_PATH}` binds (F-S2/F-S3) | `hddMounts` = ParseComposeHDDMounts shape. Deduped+sorted; falls back to `[stackName]` when no appdata bind. Tier-2 (`backup.Manager.tier2AppDataName`) refuses N>1; migrate (`stacks.Manager.ResolveAppDataDirNames`) loops N. `BindsPresent` drives the WARN-on-missing-declared-dir | | `UserdataDir` / `ImportDir` / `EnsureUserdataSkeleton` / `EnsureDirOwned` | controller/internal/appbackup/userdata.go | `(nsRoot)` / `(nsRoot)` / `(nsRoot, dirs []string)` / `(path, gid int)` | userdata/ tree w/ 2775 setgid gid-1000 convention. **R-75:** `ImportDir` is the CANONICAL drop-zone (`/userdata/import`) and callers MUST resolve it against the SYSTEM namespace, never an app's HDD_PATH — use `stacks.Manager.GetImportRoot()`. `EnsureUserdataSkeleton` now takes the dir set: build it with `BuildUserdataSkeleton(DeriveUserdataDirs(stacksDir))`, or via `Manager.EnsureUserdataSkeleton` / `web.Server.ensureUserdataSkeleton`. | Linux-only chown via build-tag twin userdata_linux.go. **The set MUST stay sorted** — `fbNeedsRecreate` force-recreates FileBrowser on any byte diff and the naive map-order derivation measured 20/20 distinct (SPIKE P6). `UserdataSkeletonCarry()` is the old hardcoded list, retained forever so derivation can only ADD (zero removals). | @@ -49,7 +50,8 @@ | `restoreOpInFlight` + `hasRecentRestoreResult` | controller/internal/web/restore_wizard.go | `(backup.RestoreOpStatus) bool` / `(st, app, now) bool` | THE "is a restore running / did one just finish" display reads | **TRAP (v0.154.0 shipped this bug): `Manager` has TWO running flags.** `IsRunning()` reads the CONCURRENCY flag, acquired inside the goroutine — and `RestoreOffboxScratch` never acquires it, so it is false for the whole verification restore. Display must read `RestoreStatus().Running` (set synchronously by `BeginRestoreOp`). Read the status ONCE per render or the strip and the suppression can disagree. `hasRecentRestoreResult` is app-bound and window-bounded — a process-wide result must not light another app's „Eredmény" | | `restoreWizardPath` / `deriveWizardStep` / `resolveWizardApp` | controller/internal/web/restore_wizard.go | `(app) string` / `(restoreWizardInput) restoreWizardView` / `([]OffboxAppRow, name) *OffboxAppRow` | R-48 offsite restore wizard: URL builder + the PURE step/unlock derivation + the app-resolution refusals | The step is **never** taken from the request. Precedence is load-bearing: op-running outranks a stale `?full_prep=`, else a commit button reappears mid-restore. Truth table + red-proof: `restore_wizard_test.go`. Adding a form here that posts anywhere new breaks `TestRestoreWizard_NoNewMutationEndpoints` **by design** — R-48 adds no mutation surface | | `restoreOpBlocked` | controller/internal/web/restore_wizard.go | `() (msg string, blocked bool)` | THE refusal gate before starting ANY restore | **Use this, never a bare `IsRunning()`.** It reads BOTH flags: `RestoreStatus().Running` (set synchronously by `BeginRestoreOp`, true for the whole off-box restore) and `IsRunning()` (the concurrency flag, the only one the nightly backup holds). **R-351: all seven handlers read only `IsRunning()`, which the goroutine acquires AFTER the handler returns — a second press started a second run and was told „…elindult".** Returns the Hungarian refusal, which names the running app and a route | **R-360 (v0.226.0): `offboxVerifyCopyDeleteHandler` was the LAST holdout and its doc comment claimed it already did this — the sentence is why nobody looked. `IsRunning()` is FALSE for the whole of a verification restore, so the copy a restore was writing into could be deleted from the UI (observed live 2026-08-21 22:35). No app-name comparison: refusing during ANY restore is stronger and uniform.** -| `Manager.RestoreFromRecoveryUnit` + `UnitRestoreResult` (R-353, v0.226.0) | controller/internal/backup/restore_unit.go | `(stack string) (UnitRestoreResult, error)` | THE local recovery-unit restore, and the facts its surface must state | **Returns a RESULT, not just an error** — volumes replayed, DBs replayed, and what the manifest LISTED. The Manifest* counts are load-bearing: zero-replayed has two causes (the backup held no data / the backup listed data that did not come back) and they are opposite news. Pair it with `unitRestoreOutcomeMsg`; do NOT write a new sentence. **A claim about the APP is forbidden on this path** — it has no `SafetyDump` discriminator, unlike the off-site twin (CONTEXT.md ruling, 07-backup-architecture §6.3). `restoreDockerVolumes` now returns `(int, error)`; `restoreDockerVolumesFrom` is unchanged and still the shared implementation | +| `Manager.RestoreFromRecoveryUnitAt` / `RestoreFromRecoveryUnit` + `UnitRestoreResult` (R-353 v0.226.0; R-102 v0.229.0) | controller/internal/backup/restore_unit.go | `(stack, unitDir string) (UnitRestoreResult, error)` / `(stack string) (…)` | THE local recovery-unit restore, and the facts its surface must state | **Returns a RESULT, not just an error** — volumes replayed, DBs replayed, and what the manifest LISTED. The Manifest* counts are load-bearing: zero-replayed has two causes (the backup held no data / the backup listed data that did not come back) and they are opposite news. Pair it with `unitRestoreOutcomeMsg`; do NOT write a new sentence. **A claim about the APP is forbidden on this path** — it has no `SafetyDump` discriminator, unlike the off-site twin (CONTEXT.md ruling, 07-backup-architecture §6.3). `restoreDockerVolumes` now returns `(int, error)`; `restoreDockerVolumesFrom` is unchanged and still the shared implementation. **R-102: `…At` holds the whole body and takes the unit DIRECTORY; the one-argument form is the thin caller naming the primary unit.** THE SOURCE MOVES; THE DESTINATION DOES NOT — `GetAppDrivePath` still resolves where the data lands. The R-47 mutation order, the unit-over-guest secret precedence, the fail-closed data-key gate and the no-unit `RestoreApp` fallback are all pinned and unchanged. The volume leg goes through the R-354 `volumeReplayFrom` seam so the source directory is assertable without Docker | +| `Manager.RestoreTier2Unit` + `Tier2Coverage.CanRestoreUnit` / `Tier2CopyDate` (R-102/R-103, v0.229.0) | controller/internal/backup/tier2_restore.go | `(stack string) (UnitRestoreResult, error)` / `() bool` / `() (string, bool)` | THE full restore from the SECOND DRIVE's mirror, and the predicate that gates it | **Do NOT widen `CanRestore()`** — it answers only "can the additive file restore run?" and one predicate answering two questions is R-356. Fail-closed: `tier2UnitIsOpenable` requires a parseable manifest, because a directory is not a package. The single-writer flag is taken INSIDE `RestoreFromRecoveryUnitAt` — a second `acquireRunning()` here would refuse the restore it guards. `Tier2CopyDate` prefers the SUCCESS anchor over the attempt clock (R-101) and reports which it returned | | `unitRestoreOutcomeMsg` + its three message constants (R-353, v0.226.0) | controller/internal/web/handlers.go | `(app string, res backup.UnitRestoreResult) string` | THE customer sentence for a completed LOCAL restore | Twin of `reconstituteOutcomeMsg`; copy its SHAPE (clauses earned by having done the thing, no filesystem path, base names only), never its text. The constants are named because `r353_unit_outcome_test.go` asserts them verbatim — a silent edit is how an honest message drifts back into a comforting one, which is the documented history of the sentence it replaces | | `offsiteNoSpaceMsgFmt` + `offsiteSizeUnknownMsg` (R-357, v0.226.0) | controller/internal/backup/offbox_restore.go | two consts | EVERY headroom refusal on the off-site restore surface | **All four gates share these** (prepare, scratch, place, and the destructive reconstitute). A customer meeting one wording on one path and a different one on another has to work out whether it is the same problem. **The reconstitute gate uses NO ×1.1 margin** — it copies a measured tree; `OffboxRestorePrepareFull`'s ×1.1 predicts a download. **Fail closed when either probe reads ≤ 0**: `free < need` with `need == 0` is FALSE, so an unmeasurable input sails through — a gate present and inert | | `Manager.OffboxFullScratchReady` + the scratch marker (R-358, v0.226.0) | controller/internal/backup/offbox_restore.go | `(stack) bool`; `.felhom-restore-complete.json` | THE gate for place-to-live and reconstitute | **It answers "did the run FINISH and was it FULL", not "are there files".** The old non-empty check passed a part-copy from a failed restic run, and the old doc comment ("PlaceOffsiteRestore re-validates per-path completeness") is what made it look adequate — that call stats top-level placements, not files. Marker written 0600 atomically AFTER restic returns nil; stale one cleared BEFORE it starts; both orders pinned by an AST test because `resticStep` is not a seam. Anything else — absent, unreadable, wrong schema, `full:false` — is NOT ready, with a WARN naming which. **Unit-only and full restores write the SAME directory**, so `full` is the only separator (R-396) | diff --git a/controller/README.md b/controller/README.md index 324e3d0..e4be1d7 100644 --- a/controller/README.md +++ b/controller/README.md @@ -1176,26 +1176,50 @@ operator paths) and per-file selection. Apps that index their data dir (e.g. Nex rescan (occ files:scan) before restored files appear in their own UI. > **COVERAGE — read this before assuming an app is protected by this button (C9-F1, v0.183.0).** -> This restore reads `hdd/` and `userdata/` **only**. It has never read `recovery-unit/`, which every -> Tier-2 run also writes and which holds the app's DB dumps and named-volume tarballs. Enumerated -> across all 53 catalog templates: **43 apps have no readable subtree at all** (their data is entirely -> in named volumes — BookStack, Docmost, Vaultwarden, Gitea, …), **9** have file legs but never their -> database or volumes, 1 is stateless. So the button is a guaranteed no-op for 81% of the catalog and -> only ever partial for the rest. +> This restore reads `hdd/` and `userdata/` **only**. It does not read `recovery-unit/`, which every +> Tier-2 run also writes and which holds the app's DB dumps and named-volume tarballs. Counted at +> catalogue `459766cb1639` by running the production rule over all 53 templates (v0.229.0): +> **45 apps have no readable subtree at all** (their data is entirely in named volumes — BookStack, +> Docmost, Vaultwarden, Gitea, …), **7** have file legs but never their database or volumes, and 1 +> (bentopdf) is stateless. So this button is a guaranteed no-op for 45 of 53 apps and only ever partial +> for the rest. > > Since v0.183.0 it is HONEST about that instead of silently reporting success: > `Tier2RestoreCoverage` is consulted **before** anything starts, an app with no readable subtree is -> refused **without being stopped** and told which action does work („…Használd a Visszaállítás -> indítása gombot a Biztonsági mentés → Visszaállítás oldalon."), and a run that does proceed claims +> refused **without being stopped** and told which action does work, and a run that does proceed claims > only what it **examined** („Minden vizsgált fájl megvan a helyén.") plus a disclosure that the > database and internal volumes are not part of this restore. > -> The action that DOES cover those apps is the keep-side recovery-unit restore -> (`POST /backup/restore` → `RestoreFromRecoveryUnit`), which replays volume tarballs and DB dumps. -> Routing customers there from the Tier-2 card is filed as **C9-F1b** — it puts a destructive -> operation behind a button reached via a non-destructive one, so the confirm copy must carry that -> difference. **C9-F4** is filed separately: nothing reads the Tier-2 copy's `recovery-unit/` mirror, -> so the second local copy that exists precisely for drive loss is unreachable by any customer action. +> **Since v0.229.0 the action that covers those apps is on the SAME row — see below. C9-F1b / R-103 and +> C9-F4 / R-102 are CLOSED.** The refusal no longer sends anyone to another page: where the copy holds +> an openable unit it names „Teljes visszaállítás a másolatból", the button beside it. + +**Full restore FROM THE SECOND DRIVE's mirror (R-102 + R-103, v0.229.0)** — +`POST /backup/tier2/unit-restore` (`backup.RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`, +`internal/backup/tier2_restore.go`) + the **"Teljes visszaállítás a másolatból"** button on the Tier-2 +layer row, in `btn-danger-outline` beside the additive one. + +Tier-2 mirrors each app's whole recovery unit to `/backups/secondary//recovery-unit/` on +every run. Until v0.229.0 **nothing read it**, because every reader of a unit could only name a path +under `backups/primary/` — so in the one failure Tier-2 exists for (the primary drive is lost, and the +primary unit with it) the surviving copy was unreachable. `RestoreFromRecoveryUnitAt` takes the unit +DIRECTORY, so the same restore that always worked from the primary now works from anywhere. + +- **The source moves; the destination does not.** Data lands in the live Docker volumes and the live + database container exactly as before; only the read path changes. +- **It OVERWRITES**, unlike the additive button beside it. The two are separate buttons because they + are separate promises, and the confirm carries the difference in words and names the copy's date — + differently when that date is only an ATTEMPT and not a proven copy (R-101). +- **Fail-closed:** the mirror must carry a parseable `manifest.json`. A `recovery-unit/` directory that + exists is not a package, and a restore armed over one would stop the app and replay nothing. +- **Two predicates, not one wider one:** `Tier2Coverage.CanRestoreUnit()` gates this action; + `CanRestore()` still gates only the file restore. Merging them would be R-356 again. +- Refusals (no recorded copy, drive disconnected, pre-v2 layout, no openable unit) all happen **before** + the app is stopped; a second press is refused by `restoreOpBlocked()` (R-351b). +- **Proven live** on `demo-hp` 2026-08-31 with the primary unit moved aside — 3 volumes of 3, 1 database + of 1, 28.65 s, an accented filename byte-identical, and `secrets recovered=2/2` with the guest's + `app.yaml` also moved aside: + `felhom.eu/documentation/audits/DRILL-r102-tier2-unit-2026-08-31/`. **Per-app Tier-2 config panel (v0.57.0)** — `GET/POST /stacks/{name}/backup` (`internal/web/tier2_config_handler.go` + `templates/tier2_config.html`). The "2. mentés" row's