Files
felhom.eu/documentation/audits/DRILL-r102-tier2-unit-2026-08-31

DRILL — R-102 / R-103: the second drive's copy becomes a way back

demo-hp (192.168.0.104), guest 9201 · controller v0.229.0 · 2026-08-31

App under drill: docmost — a class-B app (its data is entirely in Docker named volumes and a Postgres database; its Tier-2 copy holds a recovery-unit/ and no file leg — confirmed live by the Tier-2 run's own line: Tier 2 copied docmost → …/backups/secondary/docmost (114.5 MB, 0 leg(s))).

Venue is correct per runbooks/target-selection.md: demo-hp is Tier 0 — disposable. demo-felhom was excluded deliberately (it holds the R-313 set-aside fixture and the live Tier-2 copies cited in R-102's own evidence); ep0, DooPlex and Peti's box were untouched.

Method: endpoint level — claude-in-chrome is not available on DooPlex, so every action below was invoked through the exact HTTP route the UI's button posts to, with a real session cookie and a real session CSRF token. No server logic was skipped; only rendering was.

What was proven

# Claim Where
1 The mirror on the second drive is a complete package (manifest schema 2, compose incl. app.yaml, 3 volume tars, 1 canonical .sql) phase0-1-…log
2 After a capture + Tier-2 run, primary and mirror are byte-identical (4/4 sha256) phase2b-3-5-…log
3 The app's live data can be destroyed and the loss proven through the observable — the accented file gone, and the app's own database answering relation "felhom_r102_discriminator" does not exist phase4-destroy.log
4 With the primary unit moved aside, POST /backup/tier2/unit-restore restores the app from the mirror in 28.65 s — Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit, 3 volumes of 3 listed, 1 database of 1 listed phase6-…log
5 The data came back byte-for-byte: accented filename Árvíztűrő tükörfúrógép.txt verified as hex c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e747874 (35 bytes) and content sha256 9228fdda…c444; the app read its own row over TCP with its own credential (docmost@172.20.0.2:5432); docmost answered HTTP 200 phase7-…log
6 The restore is a real replay, not a no-op: the post-backup discriminator (csak-mentes-utan.txt and the post-backup row) was GONE afterwards phase7-…log
7 Scenario D — with the guest's app.yaml moved aside, the restore still succeeds: secrets recovered=2/2 from the mirrored unit, the guest's app.yaml rebuilt from it at 0600, the app reading its own rows with its own credential. This closes 00-capability-map.md's open clause "Tier-2's own cross-drive copy of a secret-bearing unit … not exercised live" phase8-…log
8 R-103 live — the file restore's refusal now names the action beside it, not a button on another page (302 carrying tier2UnitAvailableMsg) phase9-…log §9b
9 The ordinary primary restore still works: 3 volumes of 3, 1 database of 1, from …/backups/primary/docmost phase10-…log

What went wrong during the drill, and what it exposed

Phase 4, first attempt, destroyed nothing. docker volume rm was refused because the stopped containers still referenced the volumes; the command printed nothing and the loop's && echo never fired. Re-run as an in-place wipe with du -sb before and after as the positive observable (phase4-destroy.log states this at the top). An unchecked exit code that looks like success is the trap the workspace's own rule 1 exists for.

Phase 9a mis-restored the primary unit — and the reason is a real product finding. Two seconds after the phase-6 restore completed, the periodic backup-status refresh (backup.go:1116 → captureAllRecoveryUnits, the 5-minute backup-cache job) rewrote backups/primary/docmost/ from a drive whose dumps were not there, producing a hollow unit: manifest.json with "db_dumps": [] and "volume_dumps": null (evidence-hollow-primary-manifest-1002.json, created_at 2026-08-31T10:02:59Z). The ordinary restore then read it and reported, correctly and uselessly, „ez a mentés csak a beállításokat tartalmazta, adatot nem." Repaired in phase10-…log; the app was left healthy with its data back.

This is filed as R-403 and is NOT fixed here. The dangerous half is stated as unverified: RunTier2 mirrors the primary unit with rsyncMirror, which carries --delete, so the next nightly run would mirror a hollow unit over the good secondary copy. Nothing in f5_stale_primary_test.go or the R-181 capture floor guards that direction. It was not tested live and must not be reported as measured.

Teardown — all three layers

  • Machines provisioned: none. The drill used the existing guest 9201; no VM, no scratch guest.

  • Hub records created: none. No enrolment, no appliance, no escrow.

  • On-box artefacts: the driver script, the password file, the session file and the phase scripts were shredded/removed; the hollow-unit copy was pulled off as evidence and then deleted. The controller's settings.json.r102bak was removed.

  • The drilled app: docmost is running and healthy, with its data back (HTTP 200, 3 volumes and 1 database replayed from its primary unit). Primary and secondary are byte-identical again on all five artefacts. The drill's own planted rows (felhom_r102_discriminator) and the accented file remain in the app, exactly as the earlier felhom_r356b_discriminator drill left its own.

  • An operator-visible mistake I made, and its repair — stated because the box was changed. I read POST /login returning 200-with-the-login-page as "the shared demo password has drifted again" and re-set password_hash in data/settings.json to bcrypt(PASSWORD). The password had not drifted. Values in ~/.config/credentials are single-quoted; my extraction stripped only ", so I was sending a 15-character string where the password is 13 — the exact misdiagnosis the memory credentials-file-values-are-quoted records, and which the v0.228.0 report had recorded on this same box on this same day. This is the third instance.

    Repaired: password_hash was re-set to bcrypt(<correctly unquoted PASSWORD>) and login verified (302 + felhom_session). The box's end state therefore matches the state the v0.228.0 session independently verified. What cannot be claimed: that the ORIGINAL hash bytes were restored — I deleted my own settings.json.r102bak before finding the error, so the original is gone. The end state is correct by verification, not by restoration. Filed as an Observation in felhom-controller/REPORT.md.