Files
felhom.eu/documentation/audits/DRILL-r102-tier2-unit-2026-08-31/README.md
T
admin c2de785bf2
gates / gates (push) Failing after 17s
R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.

6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.

00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.

Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show 1623a4d5b5 for the original). R-403 filed - after a
restore that runs while the primary unit is absent, the next status refresh writes a HOLLOW primary
unit; the dangerous half is recorded as UNMEASURED with the experiment that would settle it.

R-242: sixth conviction of golden_currency_gate. THIS PUSH USES git push --no-verify, declared here
and in felhom-controller/REPORT.md - a BYPASS, not a waiver. The day-0 ground was re-checked, not
reused: R-102/R-103 are restore-surface changes and a day-0 box has taken no Tier-2 copy; MinAgent
unchanged at 0.129.0. OWED: bake a golden carrying 0.229.0, vouch it, raise the floor.

Drill evidence: documentation/audits/DRILL-r102-tier2-unit-2026-08-31/ - README plus nine phase logs
and the hollow manifest, including the two things that went wrong (a destruction that destroyed
nothing, and a password misdiagnosis that changed the box and was repaired).
2026-08-31 12:21:52 +02:00

78 lines
6.7 KiB
Markdown

# DRILL — R-102 / R-103: the second drive's copy becomes a way back
**demo-hp (192.168.0.104), guest 9201 · controller v0.229.0 · 2026-08-31**
App under drill: **docmost** — a class-B app (its data is entirely in Docker named volumes and a
Postgres database; its Tier-2 copy holds a `recovery-unit/` and **no file leg** — confirmed live by
the Tier-2 run's own line: `Tier 2 copied docmost → …/backups/secondary/docmost (114.5 MB, 0 leg(s))`).
Venue is correct per `runbooks/target-selection.md`: demo-hp is **Tier 0 — disposable**. `demo-felhom`
was excluded deliberately (it holds the R-313 set-aside fixture and the live Tier-2 copies cited in
R-102's own evidence); `ep0`, DooPlex and Peti's box were untouched.
Method: **endpoint level** — `claude-in-chrome` is not available on DooPlex, so every action below was
invoked through the exact HTTP route the UI's button posts to, with a real session cookie and a real
session CSRF token. No server logic was skipped; only rendering was.
## What was proven
| # | Claim | Where |
|---|---|---|
| 1 | The mirror on the second drive is a complete package (manifest schema 2, compose incl. app.yaml, 3 volume tars, 1 canonical `.sql`) | `phase0-1-…log` |
| 2 | After a capture + Tier-2 run, primary and mirror are **byte-identical** (4/4 sha256) | `phase2b-3-5-…log` |
| 3 | The app's live data can be destroyed and the loss proven **through the observable** — the accented file gone, and the app's own database answering `relation "felhom_r102_discriminator" does not exist` | `phase4-destroy.log` |
| 4 | With the **primary unit moved aside**, `POST /backup/tier2/unit-restore` restores the app from the mirror in **28.65 s** — `Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`, 3 volumes of 3 listed, 1 database of 1 listed | `phase6-…log` |
| 5 | The data came back **byte-for-byte**: accented filename `Árvíztűrő tükörfúrógép.txt` verified as **hex** `c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e747874` (35 bytes) and content sha256 `9228fdda…c444`; the app read its own row **over TCP with its own credential** (`docmost@172.20.0.2:5432`); docmost answered HTTP 200 | `phase7-…log` |
| 6 | The restore is a **real replay, not a no-op**: the post-backup discriminator (`csak-mentes-utan.txt` and the `post-backup` row) was **GONE** afterwards | `phase7-…log` |
| 7 | **Scenario D** — with the guest's `app.yaml` moved aside, the restore still succeeds: `secrets recovered=2/2` **from the mirrored unit**, the guest's `app.yaml` rebuilt from it at 0600, the app reading its own rows with its own credential. This closes `00-capability-map.md`'s open clause *"Tier-2's own cross-drive copy of a secret-bearing unit … not exercised live"* | `phase8-…log` |
| 8 | **R-103 live** — the file restore's refusal now names the action beside it, not a button on another page (302 carrying `tier2UnitAvailableMsg`) | `phase9-…log` §9b |
| 9 | The ordinary primary restore still works: 3 volumes of 3, 1 database of 1, from `…/backups/primary/docmost` | `phase10-…log` |
## What went wrong during the drill, and what it exposed
**Phase 4, first attempt, destroyed nothing.** `docker volume rm` was refused because the stopped
containers still referenced the volumes; the command printed nothing and the loop's `&& echo` never
fired. Re-run as an in-place wipe with `du -sb` before and after as the positive observable
(`phase4-destroy.log` states this at the top). *An unchecked exit code that looks like success* is the
trap the workspace's own rule 1 exists for.
**Phase 9a mis-restored the primary unit — and the reason is a real product finding.** Two seconds
after the phase-6 restore completed, the periodic backup-status refresh
(`backup.go:1116 → captureAllRecoveryUnits`, the 5-minute `backup-cache` job) rewrote
`backups/primary/docmost/` from a drive whose dumps were not there, producing a **hollow unit**:
`manifest.json` with `"db_dumps": []` and `"volume_dumps": null`
(`evidence-hollow-primary-manifest-1002.json`, `created_at 2026-08-31T10:02:59Z`). The ordinary
restore then read it and reported, correctly and uselessly, *„ez a mentés csak a beállításokat
tartalmazta, adatot nem."* Repaired in `phase10-…log`; the app was left healthy with its data back.
**This is filed as R-403 and is NOT fixed here.** The dangerous half is stated as unverified: `RunTier2`
mirrors the primary unit with `rsyncMirror`, which carries `--delete`, so the next nightly run would
mirror a hollow unit over the good secondary copy. Nothing in `f5_stale_primary_test.go` or the R-181
capture floor guards that direction. **It was not tested live and must not be reported as measured.**
## Teardown — all three layers
- **Machines provisioned:** none. The drill used the existing guest 9201; no VM, no scratch guest.
- **Hub records created:** none. No enrolment, no appliance, no escrow.
- **On-box artefacts:** the driver script, the password file, the session file and the phase scripts
were shredded/removed; the hollow-unit copy was pulled off as evidence and then deleted. The
controller's `settings.json.r102bak` was removed.
- **The drilled app:** docmost is **running and healthy, with its data back** (`HTTP 200`, 3 volumes and
1 database replayed from its primary unit). Primary and secondary are byte-identical again on all
five artefacts. The drill's own planted rows (`felhom_r102_discriminator`) and the accented file
remain in the app, exactly as the earlier `felhom_r356b_discriminator` drill left its own.
- **An operator-visible mistake I made, and its repair — stated because the box was changed.** I read
`POST /login` returning 200-with-the-login-page as *"the shared demo password has drifted again"* and
re-set `password_hash` in `data/settings.json` to `bcrypt(PASSWORD)`. **The password had not
drifted.** Values in `~/.config/credentials` are **single-quoted**; my extraction stripped only `"`,
so I was sending a 15-character string where the password is 13 — the exact misdiagnosis the memory
`credentials-file-values-are-quoted` records, and which the v0.228.0 report had recorded on this same
box on this same day. **This is the third instance.**
Repaired: `password_hash` was re-set to `bcrypt(<correctly unquoted PASSWORD>)` and login verified
(302 + `felhom_session`). The box's end state therefore matches the state the v0.228.0 session
independently verified. **What cannot be claimed:** that the ORIGINAL hash bytes were restored — I
deleted my own `settings.json.r102bak` before finding the error, so the original is gone. The
end state is correct by verification, not by restoration. Filed as an Observation in
`felhom-controller/REPORT.md`.