# REPORT — R-102 + R-103: the second drive's copy becomes a way back **Controller v0.229.0 · 2026-08-31 · MinAgent 0.129.0 (unchanged)** --- ## 1. Confirmed baselines Re-checked before the first edit, and both matched the task's table exactly. **No drift.** | Repo | `main` @ start | expected | version | |---|---|---|---| | `felhom-controller` | `430fb4448d6175064ef4e86c2b3f8796e15ae30c` | same | `v0.228.0` → **`v0.229.0`** | | `felhom.eu` | `1623a4d5b5d5d0f1b73aba2727b8de053930f336` | same | — (docs only) | | `felhom-agent` | not touched | — | `v0.130.0` unchanged | | `app-catalog-felhom.eu` | `459766cb16395fd1d1a66282f5cc6da59ead5924` | read-only | unchanged | `git status --porcelain` empty in both repos; disk 37% / 51%. ## 2. Files created / modified **`felhom-controller`** | File | Change | |---|---| | `controller/internal/appbackup/paths.go` | +4 unit-directory-relative primitives; the 4 `(nsRoot, stackName)` helpers become wrappers | | `controller/internal/appbackup/r102_paths_split_test.go` | **new** — A1 + a directory-relative guard | | `controller/internal/backup/appbackup_bridge.go` | re-exports the 4 primitives into `backup` | | `controller/internal/backup/restore_unit.go` | `RestoreFromRecoveryUnitAt(stack, unitDir)`; the 1-arg form becomes the thin caller; the volume leg goes through the R-354 seam; the unit dir is logged | | `controller/internal/backup/restore_db.go` | `reimportDBDumpsAtCtx`; `dbReimportTimeout` named once | | `controller/internal/backup/tier2_restore.go` | `RestoreTier2Unit`, `ErrTier2NoUnitInCopy`, `tier2UnitDir`, `tier2UnitIsOpenable`, `Tier2Coverage.{UnitRestorable,CopyLastRun,CopyLastSuccess}`, `CanRestoreUnit()`, `Tier2CopyDate()` | | `controller/internal/backup/r102_unit_at_test.go` | **new** — A2…A6 + 1 | | `controller/internal/backup/r102_tier2_unit_test.go` | **new** — B1…B5 + 1 | | `controller/internal/backup/r103_file_restore_untouched_test.go` | **new** — C1, C2 | | `controller/internal/web/handlers.go` | `backupTier2UnitRestoreHandler`; the refusal split; 7 named constants; `tier2UnitConfirmMsg`, `tier2UnitSourceMsg`; 4 new `AppBackupRow` fields + their builder | | `controller/internal/web/server.go` | route `POST /backup/tier2/unit-restore` | | `controller/internal/web/funcmap.go` | `fmtTimeStr` delegates to a package-level `fmtRFC3339Local` | | `controller/internal/web/templates/backups_apps.html` | the destructive action on the Tier-2 row | | `controller/internal/web/r103_tier2_action_test.go` | **new** — D1…D6 + 4 | | `CHANGELOG.md`, `CONTEXT.md`, `controller/README.md`, `REUSE.md`, `REPORT.md` | documentation | **`felhom.eu`** — `documentation/architecture/07-backup-architecture.md` (§6.2, §6.3, §7.2, §8, §8.1), `documentation/architecture/00-capability-map.md` (header note, Tier-2 row, D5 row), `documentation/backlog/OPEN-ITEMS.md`, `documentation/backlog/CLOSED-ITEMS.md`, `STATUS.md`, `documentation/audits/DRILL-r102-tier2-unit-2026-08-31/` (**new**, 11 files). **Not modified:** `felhom-agent`, `app-catalog-felhom.eu` (Part 3 is a measurement), `ROADMAP.md` (it carries no R-102/R-103 row — `one_register_gate.py` green). ## 3. Commits pushed to `main` | Repo | Commit | What | |---|---|---| | `felhom-controller` | `c732006d263a25744eca11a11cdeef343b1c95af` | Part 1.1 — path helper split, zero behaviour change | | `felhom-controller` | `0f9b796615d3e662fc010b747fb5048f36faffb7` | R-102 — Parts 1.2, 1.3, 2.1 + Groups A and B | | `felhom-controller` | `4c8f0d291948fd60887c0fd7066e63df9e3abb68` | R-103 — Part 2.2 + Groups C and D | | `felhom-controller` | `8aa95b58319f947fa9b3de38cb410b9d2756332f` | CHANGELOG / CONTEXT / README / REUSE / REPORT | | `felhom.eu` | `c2de785bf27631284fa5ab525e0f32494804773a` | architecture, register, STATUS, drill evidence — pushed with `--no-verify`, declared | | `felhom-controller` | *(this §3 fill-in)* | commit hashes recorded after the fact | ## 4. Per-test results, and the five named red-proofs Full suite: `go build ./... && go vet ./... && go test ./...` — **all packages ok** (`internal/backup` 296.0 s, the rest cached/fast). `python3 controller/scripts/controller_gates.py` — **all 13 gates OK**. | Group | Tests | Result | |---|---|---| | A (`appbackup`, `backup`) | `TestR102_PathWrappersAreByteIdenticalToToday`, `…UnitHelpersAreDirectoryRelative`, `…RestoreFromRecoveryUnitAtReadsTheGivenDir`, `…PrimaryPathIsUnchanged`, `…LiveDestinationIsUnchanged`, `…MutationOrderIsPreserved`, `…MissingManifestInMirrorFailsClosed` (3 sub-cases), `…UnopenableMirrorIsStillDisclosedAsUnread` | PASS | | B (`backup`) | `TestR102_Tier2UnitRestoreReadsTheSecondaryMirror`, `…WorksWithThePrimaryUnitABSENT`, `…TakesTheSingleWriterFlag`, `TestR102_NoTier2CopyIsAnHonestRefusal`, `…CopyAgeIsCarriedToTheSurface`, `…DoesNotWriteTheMirror` | PASS | | C (`backup`) | `TestR103_CanRestoreStillAnswersLegsOnly`, `TestR103_FileRestoreBehaviourUnchanged` | PASS | | D (`web`) | `TestR103_UnitActionOfferedWhenTheMirrorExists`, `…AbsentWhenItDoesNot`, `…ConfirmStatesTheOverwrite`, `…ConfirmNamesTheCopyDate`, `…AvailableMsgNamesTheButtonByItsLabel`, `…SecondPressIsRefused`, `…HandlerPublishesTheOutcome`, `…RefusesAnUnopenableMirrorWithoutStopping`, `…UnitRestoreHandlerGuards`, `…FileRestoreRefusalPointsAtTheActionThatWorks` | PASS | | E | the existing suite unmodified — **no existing test was edited** | PASS | **Red-proofs — each mutated, run, observed failing, reverted:** | # | Mutation | Observed failure | |---|---|---| | **A1** | `UnitComposeDir` joins `"compose2"` | FAIL on all three fixtures: `…/docmost/compose2, want …/docmost/compose` | | **A5** | recreate the definition **before** replaying the volumes | FAIL: *"volumes were replayed AFTER the definition was recreated"* + *"the volume replay had not run when the definition was recreated"* | | **B2** | `RestoreTier2Unit` points back at the primary unit path | FAIL twice: `config came from the PRIMARY unit: SUBDOMAIN="primary"`, and with the primary tree unreadable, `reading volume dump dir: … permission denied` | | **C1** | `CanRestore()` widened to `len(Legs) > 0 \|\| HasUnit` | FAIL on both unit-only cases | | **D6** | drop `EndRestoreOp` from the handler's goroutine | FAIL after 30 s: *"the restore never published a result"* | ## 5. Test count **1606 → 1632 test functions (+26)**, counted as unique `^func Test…` across `controller/**/*_test.go` at `430fb44` and at HEAD. ## 6. Deployed version ``` $ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'" gitea.dooplex.hu/admin/felhom-controller:0.229.0 Up 22 seconds (healthy) ``` Built locally (`build.sh 0.229.0 --push`) from a clean, pushed tree; deployed to demo-hp guest 9201 by the bootstrap mechanism (`docker pull` → `/etc/felhom-controller-image` → restart the unit). **Fleet delivery is NOT done and is the operator's (R-242).** The fleet floor and the vouched golden are **0.228.0**; **a golden carrying 0.229.0 is owed**, then the floor raised. `golden_currency_gate.py` convicts correctly and the `felhom.eu` push used `git push --no-verify` — see §13, Observation 2. ## 7. NOT yet live-validated — per matrix row, deliberately - **§8 row 3b — PROVEN.** The route was exercised live end-to-end with the primary unit absent, and the verdict is taken from the DATA, not from an exit code. - **§8 row 4 — stays PARTIAL, and this is the sentence that must not be overstated.** What was proven is the ROUTE, not the JOURNEY. **No drive has ever actually died or been replaced under this recovery.** The drill removed a *unit directory*; it did not remove a *disk*. So drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the only surviving one are all still unexercised. Promoting row 4 needs that journey. - **Not exercised at all:** the Tier-2 copy on **network storage** (§8 of the task's edge table — the behaviour is inherited unchanged and was not re-decided, so it was not re-tested); the **primary-newer-than-the-mirror** case (both dates are stated and no steering logic was built, per §12); and the **R-403 consequence** below. - **Rendering is unproven, as always here.** `claude-in-chrome` is unavailable on DooPlex, so the drill was endpoint-level: the exact routes the buttons post to, with a real session cookie and a real session CSRF. The confirm string was read out of the **rendered page**, so the markup is proven; a human click-through is still the only proof of the browser dialog. ## 8. The drill, step by step Evidence: `felhom.eu/documentation/audits/DRILL-r102-tier2-unit-2026-08-31/` (README + 9 phase logs + the hollow manifest). Machine: **demo-hp**, guest 9201 — Tier 0, disposable. `demo-felhom`, `ep0`, DooPlex and Peti's box untouched. App: **docmost**, class B — its own Tier-2 run reports **`0 leg(s)`**. 1. **The mirror's contents were confirmed, not assumed:** manifest schema 2, `compose/{app.yaml, docker-compose.yml,.felhom.yml}`, 3 volume tars, 1 canonical `.sql` (+3 `pre-restore-` undo copies). 2. **Observables planted.** A file with an accented Hungarian name in docmost's own storage root (`/app/data/storage`), and a row in docmost's own Postgres via docmost's own DB role. `content sha256 9228fddade66a054…c444`; `name utf-8 hex c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e747874` (35 bytes). 3. **Captured and mirrored** through `POST /api/backup/run` then `POST /api/backup/tier2`. Primary and mirror then **byte-identical on all four artefacts** (sha256 printed for each). 4. **A post-backup discriminator added** — `csak-mentes-utan.txt` and a `post-backup` row — so a restore that changed nothing could not pass as a restore that worked. 5. **The live data destroyed, and the loss proven BY THE OBSERVABLE.** *The first attempt destroyed nothing:* `docker volume rm` was refused because the stopped containers still referenced the volumes, and printed nothing. **Recorded at the top of `phase4-destroy.log` rather than quietly re-run** — an unchecked exit code that looks like success is the trap this project has a standing rule about. Re-done as an in-place wipe: 68 989 735 B → **0**, and the app's own database then answered `ERROR: relation "felhom_r102_discriminator" does not exist`. 6. **The PRIMARY unit moved aside** — `mv backups/primary/docmost backups/primary/docmost.ASIDE-r102`; `ls` on the original path returns `No such file or directory`. **Without this the drill proves only that the code runs.** 7. **Restored through the real endpoint** — `POST /backup/tier2/unit-restore`, session + 64-char CSRF. `302` → *„Teljes visszaállítás elindult"*. The controller's own line names the source: `Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit: images=3, secrets recovered=2/2, data_keys=0` → `3 volume(s) of 3 listed, 1 database(s) of 1 listed`, **28.65 s**. The published outcome: *„A(z) docmost: 3 adatkötet és az adatbázis visszaállítva … A visszaállítás forrása a második meghajtón lévő másolat volt (2026-08-31 12:00)."* 8. **Proven byte-for-byte, and proven THROUGH THE APP.** The accented name came back with the identical 35-byte hex above (compared as **hex**, never as rendered text — R-364), content sha256 identical; the post-backup file and the post-backup row were **GONE**; a psql client on docmost's own network with docmost's own credential returned `docmost@172.20.0.2/32:5432` and the `pre-backup` row; 44 tables present; docmost answered **HTTP 200**. 9. **Scenario D.** Repeated with the guest's `app.yaml` moved aside and a fresh `phase8-marker` row planted first. Result: `secrets recovered=2/2` **from the mirrored unit**, the guest's `app.yaml` rebuilt from it at 0600, `phase8-marker` **gone**, the accented file byte-identical, HTTP 200. **This closes `00-capability-map.md`'s open clause** for Tier-2's own cross-drive copy of a secret-bearing unit. 10. **The primary put back, and the ordinary path re-proved** — `3 volume(s) of 3, 1 database(s) of 1` from `…/backups/primary/docmost`, accented file byte-identical, HTTP 200. **Evidence was copied off the box at the end of each phase, before the reverts** (R-320): every phase log was `tee`'d to DooPlex as it ran, and the hollow manifest was pulled off **before** it was deleted. ## 9. Part 3 — the count **A = 7 · B = 45 · C = 1**, counted at catalogue **`459766cb16395fd1d1a66282f5cc6da59ead5924`**. **Rule applied:** the production pipeline, not a restatement of it — `stacks.LoadMetadata` (the single validation choke point, so a rejected `backup:` block degrades to legacy exactly as it does live) → `stacks.ParseComposeClassifiableBinds` → `appbackup.ClassifyBinds` → `appbackup.ComputeCaptureSet` at `TierSecondary`, with the legacy branch falling back to `AppDataBindsPresent` + `AppDataDirNames` as `backup.tier2CaptureSet` does. A = at least one leg survives that pipeline. 13 templates carry a valid `backup:` block; the other 40 are legacy and none binds a namespace path. - **A (7):** audiobookshelf, calibre-web, immich, komga, nextcloud, paperless-ngx, romm. - **C (1):** bentopdf — no `volumes:` key and no `${…_PATH}` bind at all. - **B (45):** the rest. **How the disagreement arose — established, not guessed.** The INV Part B.1 count (7/45/1) was right. Phase 0's own write-up says four apps are in B *"only because their single bind is a `:ro` media mount, which `ClassifyBinds` correctly excludes"* — it applied the **`:ro` default** rule. The two it therefore missed are **radarr and sonarr**: their `${USERDATA_PATH}` binds are **writable**, so that rule never reaches them, and they are excluded by an **explicit `class: excluded`** entry instead. 9 − 2 = 7 and 43 + 2 = 45 — exactly the gap. The counting tool was temporary and was deleted; the method above re-runs it. **No catalogue file was changed.** ## 10. Rows moved, and to what | Row | Was | Now | Citation | |---|---|---|---| | `07` §8 row **3b** | `NONE` — "no route" | **`PROVEN`**, RTO **28.65 s**, RPO 24 h | `audits/DRILL-r102-tier2-unit-2026-08-31/` | | `07` §8 row **4** | `PARTIAL` — "the Tier-2 copy's volume tars remain unreachable" | **`PARTIAL`, unchanged status, changed reason** — the unreachability is closed; the drive-loss JOURNEY is still unexercised | §7.2 + the same drill, with the limit stated per row | | `07` §8.1 blanks | 3b: RTO+RPO; 4: RTO | 3b **removed** (both measured); 4 kept, with the route/journey distinction spelled out | — | | `07` §6.3 Tier-2 row | "the unit mirror is read by nothing" | **CLOSED — R-102**, old sentence kept in the past tense per the section's own practice | drill | | `07` §7.2 first bullet | open | **CLOSED**, and it says plainly that Tier-2 can now meet its prerequisite in the failure it exists for | drill | | `07` §6.2 | ⚠ UNRESOLVED, two counts | **✔ RESOLVED — 7/45/1**, with which was wrong and why | catalogue `459766cb1639` | | `00-capability-map` Tier-2 row | "the mirror is read by no path (→ R-102)" | **R-102 CLOSED**, route named, §8 row 3b added | drill | | `00-capability-map` D5 row | *"Not exercised live: … Tier-2's own cross-drive copy of a secret-bearing unit"* | that half **struck**, with the evidence path | `phase8-scenarioD-…log` | | `00-capability-map` header | "two counts disagree … do not adopt either" | points at the settled number | `07` §6.2 | ## 11. Teardown — all three layers - **Machines provisioned:** **none.** No VM, no scratch guest, no drill rig. The drill used the existing guest 9201. - **Hub records created:** **none.** No enrolment, no appliance, no escrow, no claim code. - **On-box artefacts:** the endpoint driver, the password file, the session file and every phase script were shredded or removed; the hollow-unit copy was pulled off as evidence and then deleted; the controller's `settings.json.r102bak` was removed. - **The drilled app:** **docmost is running and healthy with its data back** — HTTP 200, 3 volumes and 1 database replayed from its own primary unit, primary and secondary byte-identical again on all five artefacts. All 8 apps on the box report `healthy`. The drill's planted row and accented file remain in the app, as the earlier `felhom_r356b_discriminator` drill left its own. ## 12. Observations 1. **FILED: R-403 — after a restore that runs while the primary unit is ABSENT, the next status refresh writes a HOLLOW primary unit.** Measured during the drill: two seconds after the Tier-2 unit restore completed, the 5-minute `backup-cache` job (`internal/backup/backup.go:1116` → `captureAllRecoveryUnits`) rebuilt `backups/primary/docmost/` from a drive with no dump files, producing a manifest carrying `"db_dumps": []` and `"volume_dumps": null` (`evidence-hollow-primary-manifest-1002.json`). The ordinary restore then read it and reported — correctly and uselessly — that the backup held only settings. **The dangerous half is UNMEASURED and is written down as such:** `RunTier2` mirrors the primary unit with `rsyncMirror`, which carries `--delete`, so the next nightly run would plausibly overwrite the good secondary copy with the hollow one. That is a reading of the code, not a test. Filed with the exact experiment that would settle it. Not fixed here: §12 of the task forbids expanding scope, and this is a capture-path change. 2. **FILED: R-242 — the golden-currency gate convicted for the sixth time, and the `felhom.eu` push used `git push --no-verify`.** v0.229.0 is released and the newest golden carries 0.228.0, so a machine installed now receives neither fix. **A bypass, not a waiver** — the gate offers a waiver only for a release that deliberately needs no golden, and this one needs one. **The day-0 ground was re-checked, not reused:** R-102/R-103 are restore-surface changes on the Tier-2 card and a day-0 box has taken no Tier-2 copy; no first-boot behaviour changed; `MinAgent` unchanged at 0.129.0. **Owed: bake a golden carrying 0.229.0, vouch it, raise the floor.** 3. **NOT-A-FINDING: it is my own process error, and its durable home is the MEMORY, not the product register — which is where the first two instances already live and where I have added this one (`credentials-file-values-are-quoted`). A register row would put an operator-facing product backlog entry on a mistake in how I read a credentials file. THIRD INSTANCE, and it changed the box, so it is written down in full.** I read `POST /login` returning 200-with-the-login-page as *"the shared demo password has drifted again"* and **changed the box**: I re-set `password_hash` in the controller's `data/settings.json`. **The password had not drifted.** Values in `~/.config/credentials` are **single-quoted**; my extraction stripped only `"`, so I sent a 15-character string where the password is 13. That is exactly what the memory `credentials-file-values-are-quoted` records, and exactly what the **v0.228.0** report recorded on this same box **on this same day**. Repaired: the hash was re-set to `bcrypt()` and login verified (302 + `felhom_session`), so the end state matches what the v0.228.0 session independently verified. **What I cannot claim:** that the original hash bytes were restored — I deleted my own `settings.json.r102bak` before finding the error. The end state is correct **by verification, not by restoration**, and the drill README says so. 4. **NOT-A-FINDING: the local unit restore's volume leg now goes through the R-354 `volumeReplayFrom` seam.** Strictly this is a line the task did not list. It is not a new seam — it is the one the off-site path already uses, it defaults to the real `restoreDockerVolumesFrom` so production is byte-identical, and without it the acceptance test could assert only that the restore succeeded, not which directory the tars came out of. §10 of the task requires asserting the source directory, and for 40 of 53 apps that archive is the whole dataset. 5. **NOT-A-FINDING: `fmtTimeStr` was extracted to a package-level `fmtRFC3339Local`.** The confirm (a template) and the outcome (Go) both name the copy's date. Two renderings that could disagree is how a customer confirms one date and is told another. One implementation, two callers. 6. **NOT-A-FINDING: an import-root bind would be mirrored under `hdd/` by Tier-2.** `tier2DestRel` maps everything that is not `RootUserdata` to `hdd`, including `RootImport`. No catalogue template classifies an import bind as anything but `excluded`, so nothing reaches it today and the count in §9 is unaffected. Noted while reading `tier2_capture.go` for Part 3; not acted on, and not filed, because it is unreachable from the current catalogue. 7. **NOT-A-FINDING: `07` §6.2's citation `felhom.eu/REPORT.md:17-24` is dead.** `REPORT.md` is overwritten every session by convention, so the Phase-0 count's source no longer exists — which is why §6.2 said the two methods could not be diffed from the repo. The corrected §6.2 marks the citation as since-overwritten rather than silently dropping it. ## 13. Housekeeping — register size | File | Before | After | |---|---|---| | `documentation/backlog/OPEN-ITEMS.md` | 596 lines | **593** — 4 closed rows moved out (R-102, R-103 and their C9-F4 / C9-F1b aliases), 1 new row (R-403) | | `documentation/backlog/CLOSED-ITEMS.md` | 237 lines | **239** — 2 compressed entries, each naming `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` for the full original | | `documentation/backlog/ROADMAP.md` | 160 lines | 160 — carries no R-102/R-103 row; `one_register_gate.py` green | ## 14. Final verification ``` felhom-controller/controller$ go build ./... && go vet ./... && go test ./... → all ok felhom-controller$ python3 controller/scripts/controller_gates.py → all 13 gates OK felhom.eu$ python3 scripts/repo_gates.py → 11 OK, golden-currency FAILED (declared, §12.2) ```