R-102 + R-103 CLOSED (controller v0.229.0) — architecture, register, STATUS, drill evidence
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
07-backup-architecture: 6.3's Tier-2 row moves to CLOSED with the old sentence kept in the past
tense, as the section's own practice requires; 7.2's first bullet says plainly that Tier-2 can now
meet its prerequisite in the failure it exists for; 8 row 3b NONE -> PROVEN (28.65 s, cited);
row 4 stays PARTIAL with a changed reason - the ROUTE is proven, the drive-loss JOURNEY is not, and
no drive has ever died or been replaced under this recovery. 8.1's blanks updated per row.
6.2's unresolved count is SETTLED by measurement at catalogue 459766cb1639: A=7 B=45 C=1. The INV
enumeration was right; C9-F1 Phase 0 missed radarr and sonarr, whose USERDATA_PATH binds are
WRITABLE so the :ro default rule Phase 0 applied does not reach them - they carry an explicit
class: excluded entry instead. Class C is bentopdf. No catalogue file was changed.
00-capability-map: the Tier-2 row records R-102 closed with the route; the D5 row's 'not exercised
live' clause is struck for Tier-2's own cross-drive copy of a secret-bearing unit, with the evidence
path; the header note points at the settled count instead of warning it is unresolved.
Register: R-102, R-103 and their C9-F4 / C9-F1b aliases closed and compressed into CLOSED-ITEMS
(596 -> 593 lines, each naming git show 1623a4d5b5 for the original). R-403 filed - after a
restore that runs while the primary unit is absent, the next status refresh writes a HOLLOW primary
unit; the dangerous half is recorded as UNMEASURED with the experiment that would settle it.
R-242: sixth conviction of golden_currency_gate. THIS PUSH USES git push --no-verify, declared here
and in felhom-controller/REPORT.md - a BYPASS, not a waiver. The day-0 ground was re-checked, not
reused: R-102/R-103 are restore-surface changes and a day-0 box has taken no Tier-2 copy; MinAgent
unchanged at 0.129.0. OWED: bake a golden carrying 0.229.0, vouch it, raise the floor.
Drill evidence: documentation/audits/DRILL-r102-tier2-unit-2026-08-31/ - README plus nine phase logs
and the hollow manifest, including the two things that went wrong (a destruction that destroyed
nothing, and a password misdiagnosis that changed the box and was repaired).
This commit is contained in:
@@ -94,6 +94,24 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
||||
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
||||
off-site permanently.
|
||||
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
|
||||
proven on `demo-hp`). Every night the box copied each app's whole recovery package onto the second
|
||||
drive — its settings, its database and its data. It did that for months. **Nothing could open those
|
||||
copies.** No button, no screen, no command. That mattered most in the one fault the second drive
|
||||
exists for: if the first drive dies, the package on it dies too, and the copy that survived could
|
||||
not be read. For **45** of the 53 apps that is everything they own.
|
||||
Now the same restore that always worked from the first drive can read the copy on the second one,
|
||||
and the button is on the app's own backup row. **Proved with the first drive's package taken away:**
|
||||
Docmost came back in 29 seconds — its database, all three data volumes, a file with a Hungarian
|
||||
accented name back byte for byte, and the app then read its own rows with its own password. Done
|
||||
again with the app's password file also taken away: the copy carried the passwords too (2 of 2).
|
||||
The screen that used to say „press that other button on another page" now offers the action itself.
|
||||
It says plainly that **this one overwrites** what is there — the gentle „Fájlok visszaállítása"
|
||||
beside it still only adds back missing files — and it names the date of the copy, so nobody puts
|
||||
last week over today by accident.
|
||||
- **We counted the apps this affects, and settled it.** Two of our own notes disagreed — 43 or 45.
|
||||
The answer is **45**, counted with the product's own rule against the live catalogue. The older
|
||||
count missed **radarr and sonarr**.
|
||||
- **The off-site restore now works for the other 40 apps** (R-356, controller 0.219.0, proven on
|
||||
`demo-hp`). It used to refuse before starting, tell the customer a running app „nincs telepítve",
|
||||
and send them to reinstall it "to the same place" — a place those 40 apps never offer, because they
|
||||
@@ -131,6 +149,16 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
|
||||
machine re-reads its whole store every week, however large it grows, and the first person to notice
|
||||
would be a customer whose upload is busy. The warning is there so that does not happen.
|
||||
- **After a restore from the second drive, the first drive's package is rewritten EMPTY** (R-403,
|
||||
found during the 0.229.0 drill, **not fixed**). Two seconds after the restore finished, the box's
|
||||
five-minute housekeeping rebuilt the first drive's package from a drive that had no data files on
|
||||
it, and wrote a package that lists nothing. The ordinary „Visszaállítás indítása" then read it and
|
||||
said, correctly and uselessly, that the backup held only settings. **What we did not test:** the
|
||||
nightly copy mirrors the first drive over the second one and deletes what is not there, so the
|
||||
next night could plausibly overwrite the good copy with the empty one. That is a reading of the
|
||||
code, not a measurement, and it is written down as unmeasured on purpose. **If you do nothing:**
|
||||
a customer who recovers from their second drive may find, the next morning, that the copy they
|
||||
recovered from has been replaced by an empty one. Settling it costs one test on a spare machine.
|
||||
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
||||
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
||||
is **under a year** away on the corrected measurement, not two.
|
||||
|
||||
@@ -117,15 +117,16 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
> about the **failure** and the row is about the **mechanism**; that is not a contradiction, and both
|
||||
> cite the same evidence.
|
||||
>
|
||||
> The matrix's blank RTO/RPO cells are deliberate: no number is estimated anywhere. Two counts of
|
||||
> Tier-2 app coverage disagree (9/43/1 vs 7/45/1) and are **both** recorded there, unresolved —
|
||||
> do not adopt either from this page.
|
||||
> The matrix's blank RTO/RPO cells are deliberate: no number is estimated anywhere. The two counts of
|
||||
> Tier-2 app coverage that used to disagree (9/43/1 vs 7/45/1) were **settled 2026-08-31 by measurement
|
||||
> at catalogue `459766cb1639`: A = 7 · B = 45 · C = 1** (`07-backup-architecture.md` §6.2, which also
|
||||
> records which prior count was wrong and why). Take the number from there, not from memory.
|
||||
|
||||
| Scenario | Components | Status | Evidence | Gap / roadmap |
|
||||
|---|---|---|---|---|
|
||||
| Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded **DEGRADED** rather than silently normal (installer Case A/B) | agent v0.113, host-install v1.22.0 | **PROVEN-LIVE** | `E2D-fresh-vm-2026-07-29` C1 (real 1.22.0 install, rc=0, `Day-0 provision SUCCESS`) + C2 (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort) | Case A (a second drive already present at install) has never fired naturally — only Case B has |
|
||||
| Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 **DB replay route → `07-backup-architecture.md` §8 row 3** |
|
||||
| Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 4, 5.** The matrix records that the copy's `recovery-unit/` mirror is read by no path (→ R-102) |
|
||||
| Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 3b, 4, 5.** **R-102 CLOSED 2026-08-31 (controller v0.229.0):** the copy's `recovery-unit/` mirror is now restorable — „Teljes visszaállítás a másolatból" / `POST /backup/tier2/unit-restore` — and was proven live on `demo-hp` with the primary unit moved aside (`audits/DRILL-r102-tier2-unit-2026-08-31/`, §8 row 3b, 28.65 s). R-103 closed with it: the refusal that used to name a button on another page now offers the action on the row itself |
|
||||
| Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Route + RTO → `07-backup-architecture.md` §8 rows 4, 10, 12, 15** (incl. the R-95 delete exposure and the R-104 stale-lock defect). **2026-08-04 (R-193/R-197, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`) — the row is NOT overclaiming and the guard is not the gap; the CADENCE is.** This row already recorded that a recreated data volume orphans the repo, and the spike confirms the mechanism at source: `WriteOffboxSecrets` (`offbox.go:392`) mints a fresh 256-bit repo password whenever `<DataDir>/offbox/repo_password` is absent, and **no automatic path ever consults the escrowed one** — `InjectOffboxPassword` has exactly one caller, a web form a human pastes into. What was NOT recorded is that this now fires on an **ordinary, planned, unattended guest rebuild**, on every box: measured on BOTH demo boxes 2026-08-03/04 by comparing `host_escrow.restic_pw_sha256` against `host_escrow_superseded.restic_pw_sha256` (demo-hp `8e03eddf…`→`8a9e33aa…`, demo-felhom `48741892…`→`c60c8bc7…`), orphaning **15 snapshots / 40.9 MB** and **36 snapshots / 1.14 GB** respectively. **demo-felhom is the important half:** it kept its DELIVERY (a stale staged secret restored the target in 76 s) and lost its REPOSITORY anyway, with **nothing marking the escrow stale for 13 h** — `escrow_stale` is wired to `ReissueCredentials`, the one path that does NOT change the repo password (**R-196**), and absent from the rebuild path that does. Status unchanged: the classify-and-move-aside guard remains PROVEN-LIVE and correct, and is predicted (not yet measured) to refuse the 2026-08-05 run on both boxes rather than start a silent fresh history |
|
||||
| Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live | controller v0.134/134.1/135 | **PROVEN-LIVE** (2026-07-20) | `CAMPAIGN-6D` accept legs (immich end-to-end from offsite alone) | **2026-07-19:** `audits/DIAG-immich-restore-2026-07-19.md` finds **no offsite path loads a DB dump** — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "**immich end-to-end from offsite alone**" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. **RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume** (`immich_postgres_data` is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" **overclaimed scope**: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → **PARTIAL**, scope-corrected. Evidence: `audits/DIAG-immich-restore-2026-07-19.md` (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). **2026-07-19, controller v0.148.0:** the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but **round 2 found it aborts against a running app** (`audits/DIAG-immich-restore-round2-2026-07-19.md`, H4: the replay races immich's own schema repair; `clip_index` recreated by the app 2 s before the dump's CREATE INDEX). **2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47)** — both restore paths now replay into a DB-ONLY window (`StartStackServices` brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. *(The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard `BackupStatus` fix. R-47 shipped in v0.153.0.)* **2026-07-20: the clean run HAPPENED** — endpoint-level supervised reconstitute of immich from snapshot `49e7cb46` (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no `already exists`, operation reported SUCCESS, immich's own DatabaseService logged `No schema drift detected` twice, 11 assets `active`, 4/4 containers healthy, 231 `public` indexes. **Operator confirmed the immich timeline renders correctly after the reconstitute** (screenshot held, 2026-07-20). Evidence: `felhom-controller/REPORT.md` §4b. **2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE.** The operator deleted the photos in immich own UI **and emptied the trash** (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. **`40 file(s) placed`** against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets `active`, `No schema drift detected`, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: `felhom-controller/REPORT.md` 4e **Route + RTO → `07-backup-architecture.md` §8 rows 3, 4** — the matrix also records that no offsite action unpacks the named-volume tars it captures (→ R-107) |
|
||||
| Manual `.fab` export/import: class-scoped capture, browser up/download, tunnel-proof chunking | controller v0.125/128/130/136 | **PROVEN-LIVE** | `CAMPAIGN-6D` P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking `CAMPAIGN-6B` P2 (100 MiB via real CF edge, 120 MiB→413) | Chunking proven at the real CF edge via `curl --resolve`; the **rendered browser file-picker** upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B *finding*; fix verified in 6D |
|
||||
@@ -150,7 +151,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Data migration between drives (all / per-app), crash-safe | controller | **PROVEN-LIVE** | `CAMPAIGN-6C` 4P-5 (scope=app round-trip, byte-identical); `storage-lifecycle-acceptance-2026-06-15` (two migrate-all runs via dashboard UI, sha256 byte-identical) | (Cited `CAMPAIGN-2` T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS |
|
||||
| NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated | controller v0.113–117, agent v0.81/84/85 | **PROVEN-LIVE** | `SPIKE-nas-verify-2026-07-11`, `SPIKE-nas-dsm-2026-07-11`, `CAMPAIGN-3-2026-07-11` (boot/reassert fixes) | |
|
||||
| **Network storage (NAS) is browse + bulk-media only — it may NOT host an app's data namespace** | controller v0.187.0 | **PROVEN-LIVE** (2026-07-30) | `audits/R108-network-app-namespace-2026-07-30.md`. Same-box before/after on demo-felhom through the exact endpoint the UI invokes (`POST /api/storage/migrate-app`, authenticated + CSRF): **pre-fix v0.186.0** the target was never examined — both a NAS-shaped and an unregistered path passed straight into `MigrateApp` and failed only on the app name (409); **post-fix v0.187.0** both are refused 400 with a Hungarian reason, while a real local drive still reaches `MigrateApp` (409) proving the guard is not over-broad. On demo-hp (the box with a REGISTERED share) the network-specific refusal fires. Non-effect verified in the registry: no `migrated_to`, nothing decommissioned, app `HDD_PATH` unchanged, no `backups/` on the share | **Closes R-108 and UNBLOCKS D5** (`07-backup-architecture.md` §7.3, §10.1). The share-root `:rslave` FileBrowser bind is deliberately UNCHANGED — load-bearing for automount wake, and unscopable — verified byte-identical by diffing demo-hp's generated compose before/after. Fails closed: `/mnt/felhom-drives` holds both kinds, so an unregistered path under it is un-classifiable and refused. **Not exercised:** the deploy POST and decommission-migrate refusals are unit-tested (non-effect, nil `stackMgr`) but were NOT live-fired — only migrate-app was. `.fab`-export-onto-NAS remains open (→ R-126) |
|
||||
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore, and Tier-2's own cross-drive copy of a secret-bearing unit (both unit-tested only). Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
||||
| **A Tier-1/Tier-2 app restore works from the DRIVE ALONE — the recovery unit carries the app's secrets (D5)** | controller v0.188.0 | **PROVEN-LIVE** (2026-07-30) | `audits/D5-drive-alone-restore-2026-07-30.md`, `07-backup-architecture.md` §7.4. On a scratch drill guest on felhom-pve, through the real endpoints (`POST /api/stacks/{app}/deploy` → `POST /api/backup/run` → `POST /backup/restore`): AdventureLog (`SECRET_KEY` **data_key** + `DB_PASSWORD`) restored with the guest's `app.yaml` **moved aside** → `secrets recovered=2/2`, **27.6 s**, `Restore-from-unit completed`. **The observable is the DATA, not the exit code:** the app itself then read the seeded customer row **over TCP with its own credential** (`connected_as=adventurelog over_TCP=True`), 51 Django tables intact, and the discriminator held — the pre-backup row returned while a row added AFTER the backup was gone, so the volume tar was genuinely restored. The unit held **no `.sql` dump**, so the DB came back from the volume tar, which is exactly the case a regenerated password would have broken. **The withheld half is proven too:** Grafana's `type: password` admin login was live in its container and `ENC:` in the guest, yet appeared in **0 files** anywhere under the backup namespace, and the unit's app.yaml header names it as withheld | **Ruling (operator, 2026-07-30): `type: secret` travels, `type: password` NEVER does, minus the `nonPortableSecrets` code register (`vaultwarden/ADMIN_TOKEN`).** Plaintext on the drive, like the data — defensible **only because** the internet-reachable class is withheld; the two are coupled and must not be relaxed independently. **The brief's own proposal (data_key-only) was tested and rejected:** the flag is unreliable (4 encryption keys the catalog itself labels as such are unflagged → **R-127**) and a DB password is not resettable in practice (`POSTGRES_PASSWORD` is ignored once PGDATA is non-empty). Precedence: **the unit wins** over the guest, because the unit's secrets match the data being restored. The fail-closed data-key gate is UNCHANGED. **Not exercised live:** the withheld-class O4 regeneration on restore (unit-tested only). ~~Tier-2's own cross-drive copy of a secret-bearing unit~~ — **EXERCISED LIVE 2026-08-31 (controller v0.229.0):** docmost restored from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit` with the guest's `app.yaml` **moved aside** AND the primary unit moved aside, `secrets recovered=2/2` (`APP_SECRET`, `DB_PASSWORD`) taken from the MIRRORED unit's `compose/app.yaml`; the guest's `app.yaml` was rebuilt from it at 0600 and the app then read its own rows over TCP with its own credential. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/phase8-scenarioD-guest-appyaml-aside.log`. Venue was a **fixture-class scratch guest**, correct per `runbooks/target-selection.md` since D5's claim is about restore CODE, not the install path |
|
||||
| **A restore SAYS what it returned, and refuses what it cannot do** — the four restore-surface truth defects from the 2026-08-21 drill | controller **v0.226.0** (R-353, R-357, R-358, R-360, R-396) | **PROVEN-LIVE (2026-08-30) for three of the four; R-357 is IMPLEMENTED only** | `audits/evidence-r353-r360-live-2026-08-30/live-validation.txt`, controller `CHANGELOG.md` v0.226.0 + `REPORT.md`. Driven on `demo-hp` through the endpoints the UI invokes (no browser on DooPlex; the residual is client-side rendering). **R-353:** the sentence read off the customer's own wizard page — `A(z) opengist: 1 adatkötet visszaállítva — az alkalmazás újraindult.` with real counts (1 volume of 1 listed, 0 databases of 0 listed, and correctly no database clause). **R-360:** in the exact flag state that produced the bug (display flag true, concurrency flag false) the delete was refused and **a planted canary file survived**. **R-358/R-396:** a `mode=unit` restore wrote `{"schema":1,…,"full":false}` at mode 0600 with no `.tmp` left, and the gate logged `scratch holds a UNIT-ONLY restore … place-to-live stays closed` | **WHAT IS AND IS NOT CLAIMED, split deliberately.** **R-357 (the destructive restore's free-space gate) is IMPLEMENTED, NOT PROVEN-LIVE** — filling a real filesystem is a drill step, not a build step, so it rests on seam tests (`SetOffboxFreeFn`, `SetOffboxSizer`, and the new `SetOffboxLatestSnapshotFn`) whose central assertion is that `StopStack` was never called. **R-353's Scenario B — the "backup held only settings" sentence — was NOT reproduced live either**, and the reason is stated rather than glossed: no app on `demo-hp` still has a data-less unit (the drill's opengist has been recaptured and now lists one volume dump), and falsifying a manifest to produce it is the hand-set-state shortcut this project forbids. That branch is unit-proven only. **This row is about the MESSAGE and the REFUSALS, not the recovery mechanism** — `07-backup-architecture.md` §8 row 3 keeps its PROVEN status because the restore always did return what the unit held; what it could not do was say so |
|
||||
| **The off-site store is VERIFIED on a cadence — something checks that the customer's backups are still readable** | controller **v0.228.0** (R-359, R-397, R-399) | **PROVEN-LIVE (2026-08-30, re-proven at FULL DEPTH 2026-08-31) for the check, the notifier and the hazard control; the SCHEDULED FIRING is IMPLEMENTED only** | `documentation/tests/r359-integrity-2026-08-30/`. Driven on `demo-hp` through the endpoint the debug button invokes. A throwaway repo was built, checked healthy (**negative control first**), then one pack corrupted; the live store was checked read-only in **35.0 s**; and the notifier fired end to end — `Event pushed: backup_integrity_ok (info)`. The hazard control was observed live: a second check fired while the first held the single-writer flag returned `skipped:true, duration_ms:0` — **it never ran restic at all** | **⚠ WHAT AN `ok` MEANS — CHANGED 2026-08-31 (R-399, controller v0.228.0): the check now RE-READS THE DATA.** The default is `--read-data-subset=100%`, so an `ok` means every stored byte was downloaded and re-hashed, not merely that the catalogue hangs together. **The reason is measured, and it is why the default must not be turned back down to save four seconds:** a pack corrupted WITHOUT a size change made a structure check return `no errors were found`, exit 0, while every `--read-data*` form caught it. Cost curve on 134.3 MB: structure 35.0 s · 10% 35.9 s · 50% 37.3 s · 100% 39.2 s — **and those do NOT extrapolate**, which is why v0.228.0 ships a slow-check WARN (R-401) rather than a rotation schedule. `off` returns a box to structure depth. **PROVEN-LIVE at the new depth 2026-08-31 on `demo-hp`**, endpoint-level, with the restic argv observed from the guest: default → `… check --read-data-subset=100%`, 38.7 s; `off` → `… check`, 34.7 s. **The weekly firing at the new depth is IMPLEMENTED only** — the job is confirmed REGISTERED on BOTH demo boxes (`Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST`), which is not the same claim. `demo-felhom` reached 0.228.0 by SELF-UPDATE on the 2026-08-31 floor raise and re-registered the job itself, so the depth change is on the fleet and not only on the box that was deployed to by hand. **This is a readability check and NOT a restore-test** — R-87 remains open and the two are routinely conflated because their register rows are adjacent |
|
||||
| USB drive enrollment + unplug detection + recommission | controller, agent | **PROVEN-LIVE** | `storage-lifecycle-acceptance-2026-06-15` E4 (yanked-while-running → agent auto-rebind) + E10 (re-enroll, data intact); `CAMPAIGN-4`/`6A` (3 USB re-establish across device-letter reshuffle) | (Cited `RUNBOOK-usb` could NOT complete a wizard enrollment; `CAMPAIGN-2` legs were auth-hollow.) Fresh-USB **wizard enrollment** specifically still unproven |
|
||||
|
||||
@@ -308,17 +308,33 @@ entry wins; else a `:ro` reader is *excluded*; else a writable bind is *mandator
|
||||
| **Tier-2** (`mandatory + optional`) | **7** — the 4 above plus audiobookshelf, komga, romm |
|
||||
| legacy resolver path (apps with no block) | **0** — no non-block template binds a namespace path |
|
||||
|
||||
> ### ⚠️ UNRESOLVED — two counts of the same thing disagree
|
||||
> ### ✔ RESOLVED 2026-08-31 (controller v0.229.0) — **A = 7 · B = 45 · C = 1**
|
||||
>
|
||||
> | source | count |
|
||||
> |---|---|
|
||||
> | C9-F1 Phase 0, as shipped | **A = 9 / B = 43 / C = 1** — `felhom.eu/REPORT.md:17-24`, restated `backlog/OPEN-ITEMS.md:31` |
|
||||
> | INV Part B.1, independent enumeration at catalog `4252121` | **A = 7 / B = 45 / C = 1** |
|
||||
> | source | count | verdict |
|
||||
> |---|---|---|
|
||||
> | C9-F1 Phase 0, as shipped | **A = 9 / B = 43 / C = 1** — `felhom.eu/REPORT.md:17-24` (since overwritten), restated `backlog/OPEN-ITEMS.md:31` | **WRONG by two apps** |
|
||||
> | INV Part B.1, independent enumeration at catalog `4252121` | **A = 7 / B = 45 / C = 1** | **CORRECT** |
|
||||
> | Measured at catalogue `459766cb16395fd1d1a66282f5cc6da59ead5924`, 2026-08-31 | **A = 7 / B = 45 / C = 1** | adopted |
|
||||
>
|
||||
> Both use the same definition ("templates whose Tier-2 copy can hold a readable file leg"). The
|
||||
> difference is two apps and **neither number is adopted here**. The Phase-0 enumeration is described
|
||||
> in prose but the script is not committed, so the two methods cannot be diffed from the repo.
|
||||
> **This must be resolved before either figure is used to size anything.**
|
||||
> **The method, so it can be re-run rather than re-argued.** A throwaway `main` inside the controller
|
||||
> module drove the PRODUCTION rule over all 53 template directories — `stacks.LoadMetadata` (the single
|
||||
> validation choke point, so a rejected `backup:` block degrades to legacy exactly as it does live) →
|
||||
> `stacks.ParseComposeClassifiableBinds` → `appbackup.ClassifyBinds` → `appbackup.ComputeCaptureSet` at
|
||||
> `TierSecondary`, with the legacy branch falling back to `AppDataBindsPresent` + `AppDataDirNames` as
|
||||
> `backup.tier2CaptureSet` does. **A = at least one leg survives that pipeline.** 13 templates carry a
|
||||
> valid `backup:` block; 40 are legacy and none of them binds a namespace path, so all 40 are B or C.
|
||||
>
|
||||
> **A (7):** audiobookshelf, calibre-web, immich, komga, nextcloud, paperless-ngx, romm.
|
||||
> **C (1):** bentopdf — it declares no `volumes:` key and no `${…_PATH}` bind at all.
|
||||
> **B (45):** everything else.
|
||||
>
|
||||
> **How the earlier disagreement arose, established rather than guessed.** Phase 0's own write-up
|
||||
> (controller `CHANGELOG.md`, v0.183.0) says four apps — plex, jellyfin, emby, navidrome — are in B
|
||||
> "only because their single bind is a `:ro` media mount, which `ClassifyBinds` correctly excludes". It
|
||||
> applied the `:ro` **default** rule. The two apps it therefore missed are **radarr and sonarr**: their
|
||||
> `${USERDATA_PATH}/media/*` and `${USERDATA_PATH}/downloads` binds are **writable**, so the `:ro` rule
|
||||
> does not reach them, and they are excluded by an **explicit** `class: excluded` entry instead. 9 − 2 =
|
||||
> 7 and 43 + 2 = 45, which is exactly the gap. No catalogue file was changed; this is a measurement.
|
||||
|
||||
**[FACT] The class-B consequence is real regardless of which count is right.** For an app whose data
|
||||
lives entirely in named volumes, the Tier-2 copy holds a full `recovery-unit/` and **no readable
|
||||
@@ -333,7 +349,7 @@ refuses **before** stopping the app and names the action that works.
|
||||
| tier | captured | read back by that tier's restore | gap |
|
||||
|---|---|---|---|
|
||||
| Tier-1 | unit incl. volume tars + DB dumps | all of it | none |
|
||||
| Tier-2 | unit mirror **+** file legs | `hdd/` and `userdata/` **only** (`tier2_restore.go:101-104`) | **the unit mirror is read by nothing** — `RecoveryUnitPath` resolves to `backups/primary/` (`appbackup/paths.go:46-48`) → **R-102** |
|
||||
| Tier-2 | unit mirror **+** file legs | the file restore reads `hdd/` and `userdata/`; **since controller v0.218.0's Tier-3 sibling and now v0.229.0, a SECOND action reads the unit mirror itself** (`RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`, `tier2_restore.go`) | **CLOSED — R-102** |
|
||||
| Tier-3 | unit (incl. volume tars) + mandatory legs | files + DB replay **+ the named-volume tars, replayed from the scratch unit** (`offbox_reconstitute.go` `volReplay`, controller **v0.218.0**); the unit itself is still **skipped** on the way to live (`offbox_reconstitute.go:284-289`; placed only if the live unit is absent, `offbox_restore.go:352-356`) | **CLOSED — R-107** |
|
||||
|
||||
**[FACT] 2026-08-22 — the Tier-3 row above was corrected; the Tier-2 row was NOT.** Until controller
|
||||
@@ -344,13 +360,25 @@ proven live on `demo-hp`. The old sentence is kept here, in the past tense, beca
|
||||
erases what was believed leaves the next reader no way to tell a fixed gap from one that was never
|
||||
noticed.
|
||||
|
||||
**R-102 — the Tier-2 half — is NOT closed and nothing in this correction touches it.** The Tier-2 row
|
||||
above stands exactly as written: the secondary unit mirror is still read by nothing. Do not read
|
||||
"R-107 closed" as covering both; they were always two register rows, and only one of them moved.
|
||||
**[FACT] 2026-08-31 — the Tier-2 half is now closed too (R-102, controller v0.229.0), and the old
|
||||
sentence is kept here in the past tense for the reason the paragraph above gives.** Until v0.229.0 this
|
||||
table said of Tier-2: *"the unit mirror is read by nothing — `RecoveryUnitPath` resolves to
|
||||
`backups/primary/` (`appbackup/paths.go:46-48`)"*, and that was true from the day Tier-2 shipped until
|
||||
2026-08-31. The mechanism was a hard-coded `primary` segment: every reader of a recovery unit could
|
||||
only name a path under it. `appbackup` now also exposes four **unit-directory-relative** primitives,
|
||||
`Manager.RestoreFromRecoveryUnitAt(stack, unitDir)` holds the restore body, and `RestoreTier2Unit`
|
||||
points it at `<dest>/backups/secondary/<app>/recovery-unit/`. Proven live on `demo-hp` **with the
|
||||
primary unit moved aside** — `audits/DRILL-r102-tier2-unit-2026-08-31/`.
|
||||
|
||||
**[FACT]** Tier-2's gap is the sharper one because of *when* it bites: Tier-2 exists for the case
|
||||
**Two things did NOT change, and both are load-bearing.** THE SOURCE MOVED; THE DESTINATION DID NOT —
|
||||
data still lands in the live named volumes and the live database container. And the FILE restore's
|
||||
reach is unchanged: `CanRestore()` still answers only "are there file legs?", and
|
||||
`tier2UnitNotCoveredMsg` is still appended where that restore runs, so a clean file result never reads
|
||||
as a clean bill of health for the database.
|
||||
|
||||
**[FACT]** Tier-2's gap was the sharper one because of *when* it bit: Tier-2 exists for the case
|
||||
where the primary drive is lost — and in exactly that case the primary unit is gone while this
|
||||
mirror survives on the second drive, unreachable by any customer action.
|
||||
mirror survives on the second drive. Until v0.229.0 it was unreachable by any customer action.
|
||||
|
||||
**[FACT] 2026-08-23 — taking the undo copy used to DESTROY the app's own database backup (R-361).**
|
||||
`writeSafetyDump` called `DumpOne` into the app's own unit directory and renamed the result to
|
||||
@@ -605,10 +633,20 @@ anywhere under the backup namespace. Full record: `audits/D5-drive-alone-restore
|
||||
|
||||
**[FACT]** Two instances, both current:
|
||||
|
||||
- **Tier-2 vs primary-drive loss.** Tier-2's stated purpose is surviving the loss of the primary
|
||||
drive. In that failure the primary recovery unit is gone; the surviving mirror on the second drive
|
||||
is `backups/secondary/<app>/recovery-unit/`, which **no code path reads** (§6.3). For the 45-or-43
|
||||
class-B apps the restore is a guaranteed no-op in exactly its designed scenario. → **R-102**
|
||||
- **~~Tier-2 vs primary-drive loss~~ — CLOSED 2026-08-31, controller v0.229.0 (R-102).** Tier-2's
|
||||
stated purpose is surviving the loss of the primary drive. It was true until v0.229.0 that in that
|
||||
failure the primary recovery unit is gone while the surviving mirror on the second drive —
|
||||
`backups/secondary/<app>/recovery-unit/` — was read by **no code path** (§6.3), so for the **45**
|
||||
class-B apps (§6.2, count settled the same day) the restore was a guaranteed no-op in exactly its
|
||||
designed scenario.
|
||||
**Tier-2 can now meet its prerequisite in the failure it exists for**, and that is stated plainly
|
||||
because it is the whole point: the restore was run on `demo-hp` **with the primary unit moved aside**
|
||||
and returned 3 volumes of 3 and 1 database of 1 in 28.65 s, byte-for-byte, with the app then reading
|
||||
its own row over TCP with its own credential — and again with the guest's `app.yaml` also moved
|
||||
aside, `secrets recovered=2/2` from the mirrored unit. Evidence:
|
||||
`audits/DRILL-r102-tier2-unit-2026-08-31/`.
|
||||
**What the drill did NOT cover, stated per §8 below:** the full drive-loss journey — a genuinely
|
||||
absent or replaced physical drive — was not run. Only the recovery-unit half was.
|
||||
- **Tier-3 vs guest loss — HALF of this closed.** Tier-3 holds the volume tars and the DB dump. It
|
||||
was true until controller v0.218.0 that the tars were unpacked only by the Tier-1 path; since
|
||||
v0.218.0 the reconstitution replays them itself (`volReplay`) → **R-107 CLOSED 2026-08-22**. What
|
||||
@@ -824,8 +862,8 @@ crosses the line — **R-158**.
|
||||
| 2 | **An app's data directory is destroyed** | the guest, the other tiers | same route | **customer** | **46 s** (43 files) | 24 h | **PROVEN** | CAMPAIGN-9 A3 — loss verified by a positive observable (doc download 200 → 404), then 16/16 documents usable |
|
||||
| 3 | **An app's DB and named volumes are lost** | the guest, the unit | „Visszaállítás indítása" — Tier-1 unit restore (the **only** path that unpacks volume tars) | **customer** | **18.25 s** (path execution) · **27.6 s** (D5 drill, guest `app.yaml` absent) | 24 h | **PROVEN** (content recovery proven 2026-07-30) | CAMPAIGN-9 A2 proved the path executes; **D5 v0.188.0 closed the content gap** — after a restore with the guest's `app.yaml` moved aside, the app read the seeded row **over TCP with its own credential**, the pre-backup row returned and a post-backup row was gone (so the tar was really restored). No `.sql` dump in the unit ⇒ the DB came back from the volume tar. **2026-08-30 — this row KEEPS its PROVEN status and the reason is worth stating: R-353 was a defect in the MESSAGE, not in the mechanism.** The restore really did return what the unit held, every time; what it could not do was say so, because the count was discarded one call deep. Controller v0.226.0 fixed the sentence and changed nothing about the recovery path. A status that measures whether data comes back must not move because a status line was wrong |
|
||||
| 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery |
|
||||
| 3b | *same, for a class-B app via Tier-2* | — | **no route** — Tier-2 never reads the unit mirror | — | | | **NONE** | §6.3; **R-102** |
|
||||
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 or 9 of 53 apps); Tier-3 reconstitute for files + DB **+ the named-volume tars since controller v0.218.0** (`volReplay`); **the Tier-2 copy's volume tars remain unreachable** | **customer** (both) | | 24 h | **PARTIAL** | §7.2; **R-102** (open); **R-107 CLOSED 2026-08-22, v0.218.0** |
|
||||
| 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** |
|
||||
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore |
|
||||
| 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go:359-393` |
|
||||
| 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session |
|
||||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||||
@@ -845,8 +883,7 @@ Per the rule that a blank is a finding, here they are:
|
||||
|
||||
| row | blank | why |
|
||||
|---|---|---|
|
||||
| 3b | RTO, RPO | no route exists to time |
|
||||
| 4 | RTO | no drive-loss recovery has ever been timed |
|
||||
| 4 | RTO | no drive-loss recovery has ever been timed. **The Tier-2 unit restore inside it now is** — 28.65 s, row 3b, 2026-08-31 — but that is the route, not the journey: no drive has been removed or replaced under a recovery |
|
||||
| 5 | RTO | never timed; the rebuild is a normal Tier-2 run |
|
||||
| 6 | — | RTO present, but it is a **restore into a scratch guest on the same host**; a restore *to a different host* has never been timed |
|
||||
| 8 | RTO | a host has never been rebuilt as itself (INV Part D1) |
|
||||
|
||||
@@ -0,0 +1,77 @@
|
||||
# DRILL — R-102 / R-103: the second drive's copy becomes a way back
|
||||
|
||||
**demo-hp (192.168.0.104), guest 9201 · controller v0.229.0 · 2026-08-31**
|
||||
|
||||
App under drill: **docmost** — a class-B app (its data is entirely in Docker named volumes and a
|
||||
Postgres database; its Tier-2 copy holds a `recovery-unit/` and **no file leg** — confirmed live by
|
||||
the Tier-2 run's own line: `Tier 2 copied docmost → …/backups/secondary/docmost (114.5 MB, 0 leg(s))`).
|
||||
|
||||
Venue is correct per `runbooks/target-selection.md`: demo-hp is **Tier 0 — disposable**. `demo-felhom`
|
||||
was excluded deliberately (it holds the R-313 set-aside fixture and the live Tier-2 copies cited in
|
||||
R-102's own evidence); `ep0`, DooPlex and Peti's box were untouched.
|
||||
|
||||
Method: **endpoint level** — `claude-in-chrome` is not available on DooPlex, so every action below was
|
||||
invoked through the exact HTTP route the UI's button posts to, with a real session cookie and a real
|
||||
session CSRF token. No server logic was skipped; only rendering was.
|
||||
|
||||
## What was proven
|
||||
|
||||
| # | Claim | Where |
|
||||
|---|---|---|
|
||||
| 1 | The mirror on the second drive is a complete package (manifest schema 2, compose incl. app.yaml, 3 volume tars, 1 canonical `.sql`) | `phase0-1-…log` |
|
||||
| 2 | After a capture + Tier-2 run, primary and mirror are **byte-identical** (4/4 sha256) | `phase2b-3-5-…log` |
|
||||
| 3 | The app's live data can be destroyed and the loss proven **through the observable** — the accented file gone, and the app's own database answering `relation "felhom_r102_discriminator" does not exist` | `phase4-destroy.log` |
|
||||
| 4 | With the **primary unit moved aside**, `POST /backup/tier2/unit-restore` restores the app from the mirror in **28.65 s** — `Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`, 3 volumes of 3 listed, 1 database of 1 listed | `phase6-…log` |
|
||||
| 5 | The data came back **byte-for-byte**: accented filename `Árvíztűrő tükörfúrógép.txt` verified as **hex** `c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e747874` (35 bytes) and content sha256 `9228fdda…c444`; the app read its own row **over TCP with its own credential** (`docmost@172.20.0.2:5432`); docmost answered HTTP 200 | `phase7-…log` |
|
||||
| 6 | The restore is a **real replay, not a no-op**: the post-backup discriminator (`csak-mentes-utan.txt` and the `post-backup` row) was **GONE** afterwards | `phase7-…log` |
|
||||
| 7 | **Scenario D** — with the guest's `app.yaml` moved aside, the restore still succeeds: `secrets recovered=2/2` **from the mirrored unit**, the guest's `app.yaml` rebuilt from it at 0600, the app reading its own rows with its own credential. This closes `00-capability-map.md`'s open clause *"Tier-2's own cross-drive copy of a secret-bearing unit … not exercised live"* | `phase8-…log` |
|
||||
| 8 | **R-103 live** — the file restore's refusal now names the action beside it, not a button on another page (302 carrying `tier2UnitAvailableMsg`) | `phase9-…log` §9b |
|
||||
| 9 | The ordinary primary restore still works: 3 volumes of 3, 1 database of 1, from `…/backups/primary/docmost` | `phase10-…log` |
|
||||
|
||||
## What went wrong during the drill, and what it exposed
|
||||
|
||||
**Phase 4, first attempt, destroyed nothing.** `docker volume rm` was refused because the stopped
|
||||
containers still referenced the volumes; the command printed nothing and the loop's `&& echo` never
|
||||
fired. Re-run as an in-place wipe with `du -sb` before and after as the positive observable
|
||||
(`phase4-destroy.log` states this at the top). *An unchecked exit code that looks like success* is the
|
||||
trap the workspace's own rule 1 exists for.
|
||||
|
||||
**Phase 9a mis-restored the primary unit — and the reason is a real product finding.** Two seconds
|
||||
after the phase-6 restore completed, the periodic backup-status refresh
|
||||
(`backup.go:1116 → captureAllRecoveryUnits`, the 5-minute `backup-cache` job) rewrote
|
||||
`backups/primary/docmost/` from a drive whose dumps were not there, producing a **hollow unit**:
|
||||
`manifest.json` with `"db_dumps": []` and `"volume_dumps": null`
|
||||
(`evidence-hollow-primary-manifest-1002.json`, `created_at 2026-08-31T10:02:59Z`). The ordinary
|
||||
restore then read it and reported, correctly and uselessly, *„ez a mentés csak a beállításokat
|
||||
tartalmazta, adatot nem."* Repaired in `phase10-…log`; the app was left healthy with its data back.
|
||||
|
||||
**This is filed as R-403 and is NOT fixed here.** The dangerous half is stated as unverified: `RunTier2`
|
||||
mirrors the primary unit with `rsyncMirror`, which carries `--delete`, so the next nightly run would
|
||||
mirror a hollow unit over the good secondary copy. Nothing in `f5_stale_primary_test.go` or the R-181
|
||||
capture floor guards that direction. **It was not tested live and must not be reported as measured.**
|
||||
|
||||
## Teardown — all three layers
|
||||
|
||||
- **Machines provisioned:** none. The drill used the existing guest 9201; no VM, no scratch guest.
|
||||
- **Hub records created:** none. No enrolment, no appliance, no escrow.
|
||||
- **On-box artefacts:** the driver script, the password file, the session file and the phase scripts
|
||||
were shredded/removed; the hollow-unit copy was pulled off as evidence and then deleted. The
|
||||
controller's `settings.json.r102bak` was removed.
|
||||
- **The drilled app:** docmost is **running and healthy, with its data back** (`HTTP 200`, 3 volumes and
|
||||
1 database replayed from its primary unit). Primary and secondary are byte-identical again on all
|
||||
five artefacts. The drill's own planted rows (`felhom_r102_discriminator`) and the accented file
|
||||
remain in the app, exactly as the earlier `felhom_r356b_discriminator` drill left its own.
|
||||
- **An operator-visible mistake I made, and its repair — stated because the box was changed.** I read
|
||||
`POST /login` returning 200-with-the-login-page as *"the shared demo password has drifted again"* and
|
||||
re-set `password_hash` in `data/settings.json` to `bcrypt(PASSWORD)`. **The password had not
|
||||
drifted.** Values in `~/.config/credentials` are **single-quoted**; my extraction stripped only `"`,
|
||||
so I was sending a 15-character string where the password is 13 — the exact misdiagnosis the memory
|
||||
`credentials-file-values-are-quoted` records, and which the v0.228.0 report had recorded on this same
|
||||
box on this same day. **This is the third instance.**
|
||||
|
||||
Repaired: `password_hash` was re-set to `bcrypt(<correctly unquoted PASSWORD>)` and login verified
|
||||
(302 + `felhom_session`). The box's end state therefore matches the state the v0.228.0 session
|
||||
independently verified. **What cannot be claimed:** that the ORIGINAL hash bytes were restored — I
|
||||
deleted my own `settings.json.r102bak` before finding the error, so the original is gone. The
|
||||
end state is correct by verification, not by restoration. Filed as an Observation in
|
||||
`felhom-controller/REPORT.md`.
|
||||
+36
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"schema_version": 2,
|
||||
"app_name": "docmost",
|
||||
"display_name": "Docmost",
|
||||
"controller_version": "0.229.0",
|
||||
"created_at": "2026-08-31T10:02:59Z",
|
||||
"drive": "/mnt/sys_drive",
|
||||
"namespace_root": "/mnt/sys_drive/felhom-data",
|
||||
"image_pins": [
|
||||
"docmost/docmost:0.95.0",
|
||||
"postgres:16-alpine",
|
||||
"redis:7-alpine"
|
||||
],
|
||||
"secret_env_vars": [
|
||||
"APP_SECRET",
|
||||
"DB_PASSWORD"
|
||||
],
|
||||
"data_key_env_vars": null,
|
||||
"secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore",
|
||||
"config_files": [
|
||||
"docker-compose.yml",
|
||||
".felhom.yml",
|
||||
"app.yaml"
|
||||
],
|
||||
"db_dumps": [],
|
||||
"volume_dumps": null,
|
||||
"checksums": {
|
||||
".felhom.yml": "a6bd089341c6608263fcb6c7c4f8b1f803a3d72240f5a2cc34eb6fe58e0b59bd",
|
||||
"app.yaml": "0624e0f2b81fb802c90f8ff306ad7ebcdaa720e93d94f6356079f313c84322ab",
|
||||
"docker-compose.yml": "3920e17042a2f6103abd28bcf641cc22f6d6c1850384bea1d0896826ec482496"
|
||||
},
|
||||
"portable_secret_env_vars": [
|
||||
"APP_SECRET",
|
||||
"DB_PASSWORD"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,57 @@
|
||||
############ PHASE 0 — pre-state ############
|
||||
UTC 2026-08-31T09:49:01Z
|
||||
--- controller image
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.229.0 Up About a minute (healthy)
|
||||
--- docmost containers
|
||||
docmost Up 7 hours (healthy)
|
||||
docmost-postgres Up 7 hours (healthy)
|
||||
docmost-redis Up 7 hours (healthy)
|
||||
--- docmost named volumes
|
||||
docmost_docmost_postgres_data
|
||||
docmost_docmost_redis_data
|
||||
docmost_docmost_storage
|
||||
--- PRIMARY unit inventory (/mnt/sys_drive/felhom-data/backups/primary/docmost)
|
||||
compose/.felhom.yml
|
||||
compose/app.yaml
|
||||
compose/docker-compose.yml
|
||||
db-dumps/docmost-postgres.sql
|
||||
db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||
db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
|
||||
db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
|
||||
manifest.json
|
||||
volume-dumps/docmost_docmost_postgres_data.tar
|
||||
volume-dumps/docmost_docmost_redis_data.tar
|
||||
volume-dumps/docmost_docmost_storage.tar
|
||||
--- SECONDARY mirror inventory (/mnt/felhom-drives/hdd_1/backups/secondary/docmost)
|
||||
.felhom-tier2-layout
|
||||
recovery-unit/compose/.felhom.yml
|
||||
recovery-unit/compose/app.yaml
|
||||
recovery-unit/compose/docker-compose.yml
|
||||
recovery-unit/db-dumps/docmost-postgres.sql
|
||||
recovery-unit/db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||
recovery-unit/db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
|
||||
recovery-unit/db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
|
||||
recovery-unit/manifest.json
|
||||
recovery-unit/volume-dumps/docmost_docmost_postgres_data.tar
|
||||
recovery-unit/volume-dumps/docmost_docmost_redis_data.tar
|
||||
recovery-unit/volume-dumps/docmost_docmost_storage.tar
|
||||
--- SECONDARY mirror sha256 (the copy the restore will read)
|
||||
84d04a7f93f437a747d66cbbdda9a4deaf858001f4bb941ad00bd911109fd769 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_postgres_data.tar
|
||||
40e861b5e2ff33aaaf0a58c4808e821a9e131562f150102e8c1d54aebd4e36f6 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_redis_data.tar
|
||||
12c696bee0ff4d46f05beff54c0be719e325b46e806558a05e183a8846bb6304 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_storage.tar
|
||||
f3d8da33a3d16ee9a1a5fd3464285765f93fd2050cfe6fa920d8733c379c2793 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql
|
||||
0fd7b2ebfc42e735ea6abeee53178153454303f5c717a973492a0cc89e143557 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/manifest.json
|
||||
|
||||
############ PHASE 1 — plant the observables ############
|
||||
--- 1a. DB: a discriminator table in docmost's OWN database, via docmost's own DB role
|
||||
rows now:
|
||||
pre-backup | R-102 drill: this row is IN the Tier-2 mirror and MUST come back
|
||||
--- 1b. FILE: an accented Hungarian filename in docmost's OWN storage root (/app/data/storage)
|
||||
filename utf8 hex : c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e747874
|
||||
filename bytes : 35
|
||||
content sha256 : 9228fddade66a05449b77afaf66645b1c956fa5d208c9d1a972169b17623c444
|
||||
path exists : True
|
||||
--- storage dir listing (byte-safe)
|
||||
\303\201rv\303\255zt\305\261r\305\221\ t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt
|
||||
--- storage dir names as hex
|
||||
c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e7478740a <- Árvíztűrő tükörfúrógép.txt
|
||||
+43
@@ -0,0 +1,43 @@
|
||||
############ PHASE 10 — repair: put the REAL primary unit back, then restore the app ############
|
||||
UTC 2026-08-31T10:07:09Z
|
||||
WHAT WENT WRONG IN 9a, stated plainly: the primary unit directory had been RE-CREATED by a
|
||||
capture that ran 2 s after the phase-6 restore (manifest created_at 2026-08-31T10:02:59Z,
|
||||
controller_version 0.229.0, db_dumps: [], volume_dumps: null). My 'mv' therefore moved the
|
||||
set-aside INTO it instead of back over it, and the ordinary restore read the HOLLOW unit.
|
||||
|
||||
--- 10a. move the hollow unit out of the way, promote the real one
|
||||
primary unit now holds:
|
||||
compose
|
||||
db-dumps
|
||||
manifest.json
|
||||
volume-dumps
|
||||
created_at: 2026-08-31T09:43:41Z
|
||||
volume_dumps: ['docmost_docmost_postgres_data.tar', 'docmost_docmost_redis_data.tar', 'docmost_docmost_storage.tar']
|
||||
db_dumps: ['docmost-postgres.sql']
|
||||
total 116720
|
||||
drwxr-xr-x 2 root root 4096 Aug 31 09:50 .
|
||||
drwxr-xr-x 5 root root 4096 Aug 31 09:43 ..
|
||||
-rw-r--r-- 1 root root 70135296 Aug 31 09:50 docmost_docmost_postgres_data.tar
|
||||
-rw-r--r-- 1 root root 49370624 Aug 31 09:50 docmost_docmost_redis_data.tar
|
||||
-rw-r--r-- 1 root root 2560 Aug 31 09:49 docmost_docmost_storage.tar
|
||||
|
||||
--- 10b. the ORDINARY primary restore, through the real endpoint
|
||||
302 https://127.0.0.1:443/backups/restore?flash=Vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
|
||||
{"ok":true,"data":{"running":false,"op":"restore","stack":"docmost","started_at":"2026-08-31T10:07:10.071526585Z","last":{"op":"restore","stack":"docmost","ok":true,"message":"A(z) docmost: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult.","finished_at":"2026-08-31T10:07:39.399730776Z"},"last_recent":true}}
|
||||
|
||||
2026/08/31 10:07:10 restore_unit.go:313: [INFO] [backup] Restoring docmost from recovery unit /mnt/sys_drive/felhom-data/backups/primary/docmost: images=3, secrets recovered=2/2, data_keys=0
|
||||
2026/08/31 10:07:39 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: docmost — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
|
||||
|
||||
--- 10c. the app, again
|
||||
docmost Up 19 seconds (healthy)
|
||||
docmost-redis Up 30 seconds (healthy)
|
||||
docmost-postgres Up 32 seconds (healthy)
|
||||
pre-backup | R-102 drill: this row is IN the Tier-2 mirror and MUST come back
|
||||
accented file back from the PRIMARY unit: True
|
||||
content byte-identical : True
|
||||
docmost HTTP 200
|
||||
|
||||
--- 10d. did a capture immediately rewrite the primary unit again?
|
||||
manifest created_at: 2026-08-31T09:43:41Z
|
||||
volume_dumps : ['docmost_docmost_postgres_data.tar', 'docmost_docmost_redis_data.tar', 'docmost_docmost_storage.tar']
|
||||
db_dumps : ['docmost-postgres.sql']
|
||||
@@ -0,0 +1,5 @@
|
||||
############ PHASE 2 — capture, then mirror ############
|
||||
UTC 2026-08-31T09:49:38Z
|
||||
--- POST /api/backup/run (the nightly Tier-1: DB dump + volume tars + recovery unit)
|
||||
200
|
||||
waiting for the run to finish...
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
############ PHASE 2b — the MIRROR now carries the planted state ############
|
||||
UTC 2026-08-31T10:01:02Z
|
||||
total 116720
|
||||
drwxr-xr-x 2 root root 4096 Aug 31 09:50 .
|
||||
drwxr-xr-x 5 root root 4096 Aug 31 09:43 ..
|
||||
-rw-r--r-- 1 root root 70135296 Aug 31 09:50 docmost_docmost_postgres_data.tar
|
||||
-rw-r--r-- 1 root root 49370624 Aug 31 09:50 docmost_docmost_redis_data.tar
|
||||
-rw-r--r-- 1 root root 2560 Aug 31 09:49 docmost_docmost_storage.tar
|
||||
f46a2fc3aa9a7ae2502d83b1c6ef27e503102f5ba71a0c6559246d9674c8e3b1 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_postgres_data.tar
|
||||
a8df17c444e41f54762e122ce1be998315c969015a1580bb8abdc7211cfa1a73 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_redis_data.tar
|
||||
88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_storage.tar
|
||||
9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql
|
||||
--- accented file inside the MIRROR's storage tar:
|
||||
drwxr-xr-x 1000/1000 0 2026-08-31 09:49 ./
|
||||
-rw-r--r-- root/root 65 2026-08-31 09:49 ./\303\201rv\303\255zt\305\261r\305\221 t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt
|
||||
--- pre-backup row inside the MIRROR's .sql (count):
|
||||
1
|
||||
--- PRIMARY vs MIRROR: are the three tars and the .sql byte-identical?
|
||||
IDENTICAL volume-dumps/docmost_docmost_postgres_data.tar f46a2fc3aa9a7ae2502d83b1c6ef27e503102f5ba71a0c6559246d9674c8e3b1
|
||||
IDENTICAL volume-dumps/docmost_docmost_redis_data.tar a8df17c444e41f54762e122ce1be998315c969015a1580bb8abdc7211cfa1a73
|
||||
IDENTICAL volume-dumps/docmost_docmost_storage.tar 88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d
|
||||
IDENTICAL db-dumps/docmost-postgres.sql 9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28
|
||||
|
||||
############ PHASE 3 — the POST-backup discriminator (must be GONE after the restore) ############
|
||||
post-backup file sha256: ef0ae720c3d2649a1340e66644273f22b3c3d1fadea7e6fc3eeb6156120c9996
|
||||
rows now:
|
||||
post-backup
|
||||
pre-backup
|
||||
storage now:
|
||||
csak-mentes-utan.txt
|
||||
\303\201rv\303\255zt\305\261r\305\221\ t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt
|
||||
|
||||
############ PHASE 4 — DESTROY the app's live data ############
|
||||
--- volumes remaining (positive observable of destruction):
|
||||
docmost_docmost_postgres_data
|
||||
docmost_docmost_redis_data
|
||||
docmost_docmost_storage
|
||||
--- PROOF OF LOSS through the observable, not the filesystem:
|
||||
the app's discriminator table is now: 2
|
||||
the accented file is now: csak-mentes-utan.txt
|
||||
\303\201rv\303\255zt\305\261r\305\221\ t\303\274k\303\266rf\303\272r\303\263g\303\251p.txt
|
||||
/var/lib/docker/volumes/docmost_docmost_storage
|
||||
|
||||
############ PHASE 5 — move the PRIMARY recovery unit ASIDE ############
|
||||
--- primary unit present?
|
||||
ls: cannot access '/mnt/sys_drive/felhom-data/backups/primary/docmost': No such file or directory
|
||||
--- what remains under backups/primary/:
|
||||
bookstack
|
||||
docmost.ASIDE-r102
|
||||
kimai
|
||||
opengist
|
||||
paperless
|
||||
privatebin
|
||||
--- the mirror is untouched:
|
||||
88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/volume-dumps/docmost_docmost_storage.tar
|
||||
9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28 /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit/db-dumps/docmost-postgres.sql
|
||||
@@ -0,0 +1,32 @@
|
||||
############ PHASE 4 (redo) — DESTROY the app's live data, for real ############
|
||||
UTC 2026-08-31T10:01:46Z
|
||||
NOTE: the first attempt used 'docker volume rm', which docker REFUSED because the stopped
|
||||
containers still referenced the volumes. It printed nothing and destroyed nothing.
|
||||
An unchecked exit code that looks like success is exactly the trap this drill tests for.
|
||||
|
||||
docmost Exited (1) 42 seconds ago
|
||||
docmost-redis Exited (0) 42 seconds ago
|
||||
docmost-postgres Exited (0) 29 seconds ago
|
||||
--- sizes BEFORE the wipe
|
||||
68989735 /var/lib/docker/volumes/docmost_docmost_postgres_data/_data
|
||||
128 /var/lib/docker/volumes/docmost_docmost_storage/_data
|
||||
49419992 /var/lib/docker/volumes/docmost_docmost_redis_data/_data
|
||||
wiped docmost_docmost_postgres_data -> entries remaining: 0
|
||||
wiped docmost_docmost_redis_data -> entries remaining: 0
|
||||
wiped docmost_docmost_storage -> entries remaining: 0
|
||||
--- sizes AFTER the wipe
|
||||
0 /var/lib/docker/volumes/docmost_docmost_postgres_data/_data
|
||||
0 /var/lib/docker/volumes/docmost_docmost_storage/_data
|
||||
0 /var/lib/docker/volumes/docmost_docmost_redis_data/_data
|
||||
|
||||
--- PROOF OF LOSS through the OBSERVABLE, not the filesystem
|
||||
1) the accented file:
|
||||
(empty listing above = the file is gone)
|
||||
2) the app's own database — start the DB container and ask it for the rows:
|
||||
ERROR: relation "felhom_r102_discriminator" does not exist
|
||||
LINE 1: SELECT count(*) FROM felhom_r102_discriminator;
|
||||
^
|
||||
^ an error or an empty database IS the proof: the rows cannot be read any more.
|
||||
|
||||
--- the PRIMARY unit is still aside:
|
||||
ls: cannot access '/mnt/sys_drive/felhom-data/backups/primary/docmost': No such file or directory
|
||||
+35
@@ -0,0 +1,35 @@
|
||||
############ PHASE 6 — the NEW Tier-2 unit restore, through the real endpoint ############
|
||||
UTC 2026-08-31T10:02:28Z
|
||||
Endpoint: POST /backup/tier2/unit-restore (the exact route the row's button posts to)
|
||||
Method: endpoint-level (no browser on DooPlex) — session cookie + session CSRF, as the UI sends.
|
||||
|
||||
--- the surface's own answer BEFORE the press: is the action offered for docmost?
|
||||
action="/backup/tier2/unit-restore"
|
||||
<button type="submit" class="btn btn-xs btn-danger-outline" data-confirm="Ez a művelet FELÜLÍRJA az alkalmazás jelenlegi adatait – az adatbázisát és a belső köteteit is – a második meghajtón lévő másolattal. Ami a másolat óta keletkezett, elveszik. A másolat kelte: 2026-08-31 12:00. A mellette lévő „Fájlok visszaállítása” ezzel szemben csak a hiányzó fájlokat pótolja, és semmit nem ír felül. Az alkalmazás a művelet idejére leáll.">Teljes visszaállítás a másolatból</button>
|
||||
--
|
||||
|
||||
--- POST
|
||||
302 https://127.0.0.1:443/backups/apps?flash=Teljes+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
|
||||
|
||||
--- waiting for the async restore to publish its result...
|
||||
{"ok":true,"data":{"running":false,"op":"tier2-unit-restore","stack":"docmost","started_at":"2026-08-31T10:02:28.860491563Z","last":{"op":"tier2-unit-restore","stack":"docmost","ok":true,"message":"A(z) docmost: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-08-31 12:00).","finished_at":"2026-08-31T10:02:57.511163864Z"},"last_recent":true}}
|
||||
|
||||
--- controller log for the restore
|
||||
2026/08/31 09:52:59 backup.go:1077: [INFO] [backup] Found 13 DB dump files across drives
|
||||
2026/08/31 09:57:59 backup.go:1077: [INFO] [backup] Found 13 DB dump files across drives
|
||||
2026/08/31 10:00:07 tier2.go:446: [INFO] [backup] Tier 2 run complete: 8 app(s) processed (incl. volume-only — F6)
|
||||
2026/08/31 10:02:28 handlers.go:1842: [WARN] [web] Tier-2 UNIT restore requested (async, OVERWRITES live data): stack=docmost from 172.18.0.4:33772
|
||||
2026/08/31 10:02:28 tier2_restore.go:180: [WARN] [backup] Tier-2 UNIT restore for docmost from the secondary mirror /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit — this OVERWRITES live app data
|
||||
2026/08/31 10:02:28 restore_unit.go:313: [INFO] [backup] Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit: images=3, secrets recovered=2/2, data_keys=0
|
||||
2026/08/31 10:02:29 restore.go:148: [INFO] [backup] Restoring Docker volume docmost_docmost_postgres_data for docmost
|
||||
2026/08/31 10:02:30 restore.go:177: [DEBUG] [backup] Volume docmost_docmost_postgres_data restored successfully
|
||||
2026/08/31 10:02:30 restore.go:148: [INFO] [backup] Restoring Docker volume docmost_docmost_redis_data for docmost
|
||||
2026/08/31 10:02:30 restore.go:177: [DEBUG] [backup] Volume docmost_docmost_redis_data restored successfully
|
||||
2026/08/31 10:02:30 restore.go:148: [INFO] [backup] Restoring Docker volume docmost_docmost_storage for docmost
|
||||
2026/08/31 10:02:31 restore.go:177: [DEBUG] [backup] Volume docmost_docmost_storage restored successfully
|
||||
2026/08/31 10:02:31 restore.go:182: [INFO] [backup] Restored 3 Docker volume(s) for docmost
|
||||
2026/08/31 10:02:31 restore_db.go:77: [INFO] [backup] Restore docmost: replaying DB dump into docmost-postgres (postgres)
|
||||
2026/08/31 10:02:33 dbdump.go:799: [INFO] [backup] Imported DB dump docmost-postgres.sql into docmost-postgres (postgres)
|
||||
2026/08/31 10:02:33 restore_db.go:87: [INFO] [backup] Restore docmost: replayed 1 DB dump(s)
|
||||
2026/08/31 10:02:57 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: docmost — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
|
||||
2026/08/31 10:02:57 handlers.go:1852: [INFO] [web] Tier-2 unit restore completed (async): stack=docmost in 28.650568445s (volumes 3/3, dbs 1/1)
|
||||
+28
@@ -0,0 +1,28 @@
|
||||
############ PHASE 7 — prove the data came back, THROUGH THE APP ############
|
||||
UTC 2026-08-31T10:03:32Z
|
||||
--- docmost stack state
|
||||
docmost Up 48 seconds (healthy)
|
||||
docmost-redis Up 58 seconds (healthy)
|
||||
docmost-postgres Up About a minute (healthy)
|
||||
|
||||
--- 7a. the accented file: name bytes as HEX (R-364 — never judged as rendered text)
|
||||
entries in the app's storage root: 1
|
||||
name_hex : c3817276c3ad7a74c5b172c5912074c3bc6bc3b67266c3ba72c3b367c3a9702e747874
|
||||
name_render: Árvíztűrő tükörfúrógép.txt
|
||||
sha256 : 9228fddade66a05449b77afaf66645b1c956fa5d208c9d1a972169b17623c444
|
||||
|
||||
PLANTED accented name present, byte-for-byte : True
|
||||
content byte-identical to the planted bytes : True
|
||||
POST-backup file 'csak-mentes-utan.txt' GONE : True
|
||||
|
||||
--- 7b. the DATABASE, read by a client on the app's OWN network with the app's OWN credential
|
||||
docmost@172.20.0.2/32:5432
|
||||
pre-backup | R-102 drill: this row is IN the Tier-2 mirror and MUST come back
|
||||
|
||||
^ current_user + inet_server_addr proves it is the APP's role over TCP, not a local socket.
|
||||
|
||||
--- 7c. the app's own tables survived the replay (a count, not a claim about the app)
|
||||
44
|
||||
|
||||
--- 7d. docmost answers over HTTP (its own interface is up)
|
||||
docmost HTTP 200
|
||||
+58
@@ -0,0 +1,58 @@
|
||||
############ PHASE 8 — SCENARIO D: the guest's app.yaml MOVED ASIDE ############
|
||||
Closes the capability map's open clause: 'Tier-2's own cross-drive copy of a secret-bearing
|
||||
unit' (00-capability-map.md, the D5 row) was unit-tested only.
|
||||
UTC 2026-08-31T10:04:22Z
|
||||
|
||||
--- 8a. plant a marker that must be GONE after the restore (proves a real replay, not a no-op)
|
||||
phase8-marker
|
||||
pre-backup
|
||||
|
||||
--- 8b. move the GUEST's app.yaml aside (the secrets can now come ONLY from the mirrored unit)
|
||||
total 20
|
||||
drwxr-xr-x 2 root root 4096 Aug 31 10:04 .
|
||||
drwxr-xr-x 58 root root 4096 Aug 21 16:00 ..
|
||||
-rw-r--r-- 1 root root 2332 Aug 31 10:02 .felhom.yml
|
||||
-rw------- 1 root root 498 Aug 31 10:02 app.yaml.ASIDE-r102
|
||||
-rw-r--r-- 1 root root 3105 Aug 31 10:02 docker-compose.yml
|
||||
|
||||
--- 8c. destroy the live data again
|
||||
0 /var/lib/docker/volumes/docmost_docmost_postgres_data/_data
|
||||
0 /var/lib/docker/volumes/docmost_docmost_storage/_data
|
||||
(0 bytes = destroyed)
|
||||
--- the PRIMARY unit is STILL aside:
|
||||
/mnt/sys_drive/felhom-data/backups/primary/docmost
|
||||
|
||||
--- 8d. restore from the MIRROR, through the real endpoint
|
||||
302 https://127.0.0.1:443/backups/apps?flash=Teljes+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
|
||||
{"ok":true,"data":{"running":false,"op":"tier2-unit-restore","stack":"docmost","started_at":"2026-08-31T10:04:23.364841193Z","last":{"op":"tier2-unit-restore","stack":"docmost","ok":true,"message":"A(z) docmost: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-08-31 12:00).","finished_at":"2026-08-31T10:04:51.894947224Z"},"last_recent":true}}
|
||||
|
||||
--- 8e. SECRETS RECOVERED (the line that closes the clause)
|
||||
2026/08/31 10:02:28 restore_unit.go:313: [INFO] [backup] Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit: images=3, secrets recovered=2/2, data_keys=0
|
||||
2026/08/31 10:02:57 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: docmost — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
|
||||
2026/08/31 10:04:23 restore_unit.go:313: [INFO] [backup] Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit: images=3, secrets recovered=2/2, data_keys=0
|
||||
2026/08/31 10:04:51 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: docmost — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
|
||||
|
||||
--- 8f. the guest's app.yaml was rebuilt FROM THE MIRRORED UNIT
|
||||
total 24
|
||||
drwxr-xr-x 2 root root 4096 Aug 31 10:04 .
|
||||
drwxr-xr-x 58 root root 4096 Aug 21 16:00 ..
|
||||
-rw-r--r-- 1 root root 2332 Aug 31 10:04 .felhom.yml
|
||||
-rw------- 1 root root 498 Aug 31 10:04 app.yaml
|
||||
-rw------- 1 root root 498 Aug 31 10:02 app.yaml.ASIDE-r102
|
||||
-rw-r--r-- 1 root root 3105 Aug 31 10:04 docker-compose.yml
|
||||
(app.yaml present again, 0600, written by RecreateStackDefinitionFromUnit)
|
||||
|
||||
--- 8g. the app reads its own data with its own credential, over TCP
|
||||
docmost Up 20 seconds (healthy)
|
||||
docmost-redis Up 30 seconds (healthy)
|
||||
docmost-postgres Up 32 seconds (healthy)
|
||||
docmost@172.20.0.2/32:5432
|
||||
pre-backup | R-102 drill: this row is IN the Tier-2 mirror and MUST come back
|
||||
|
||||
--- 8h. the accented file, again as HEX
|
||||
entries: ['Árvíztűrő tükörfúrógép.txt']
|
||||
accented name byte-for-byte: True
|
||||
content byte-identical : True
|
||||
|
||||
--- 8i. docmost's own HTTP interface
|
||||
docmost HTTP 200
|
||||
+42
@@ -0,0 +1,42 @@
|
||||
############ PHASE 9 — put it back, and prove the ORDINARY path still works ############
|
||||
UTC 2026-08-31T10:05:32Z
|
||||
--- 9a. restore the PRIMARY unit and remove the app.yaml set-aside
|
||||
primary unit restored: compose docmost.ASIDE-r102 manifest.json
|
||||
app.yaml set-aside removed (the live app.yaml is the one the restore rebuilt)
|
||||
bookstack
|
||||
docmost
|
||||
kimai
|
||||
opengist
|
||||
paperless
|
||||
privatebin
|
||||
|
||||
--- 9b. R-103 live: the FILE restore now points at the action beside it, not another page
|
||||
302 https://127.0.0.1:443/backups/apps?flash_error=Ennek+az+alkalmaz%C3%A1snak+az+adatai+nem+f%C3%A1jlokban%2C+hanem+az+alkalmaz%C3%A1s+saj%C3%A1t+adatb%C3%A1zis%C3%A1ban+%C3%A9s+k%C3%B6teteiben+vannak+%E2%80%94+az+alkalmaz%C3%A1s+nem+%C3%A1llt+le.+Ezeket+a+mellette+l%C3%A9v%C5%91+%E2%80%9ETeljes+vissza%C3%A1ll%C3%ADt%C3%A1s+a+m%C3%A1solatb%C3%B3l%E2%80%9D+gombbal+tudod+visszahozni+ugyanerr%C5%91l+a+m%C3%A1solatr%C3%B3l.+Figyelem%3A+az+a+m%C5%B1velet+FEL%C3%9CL%C3%8DRJA+a+jelenlegi+adatokat%2C+m%C3%ADg+ez+a+gomb+csak+a+hi%C3%A1nyz%C3%B3+f%C3%A1jlokat+p%C3%B3tolja.
|
||||
|
||||
--- 9c. R-102/R-103 live: an app whose Tier-2 copy has NO unit is refused (Scenario G shape)
|
||||
(calibre-web's copy is on the SSD, state-only — checking what the surface offers per app)
|
||||
8
|
||||
^ apps offering the destructive action
|
||||
|
||||
--- 9d. destroy again, then the ORDINARY PRIMARY restore (POST /backup/restore)
|
||||
destroyed: 0 /var/lib/docker/volumes/docmost_docmost_postgres_data/_data
|
||||
302 https://127.0.0.1:443/backups/restore?flash=Vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
|
||||
{"ok":true,"data":{"running":false,"op":"restore","stack":"docmost","started_at":"2026-08-31T10:05:33.081614501Z","last":{"op":"restore","stack":"docmost","ok":true,"message":"A(z) docmost: a beállítások visszaálltak — az alkalmazás újraindult. FIGYELEM: ez a mentés csak a beállításokat tartalmazta, adatot nem. Az alkalmazás adatai NEM álltak vissza ebből a mentésből.","finished_at":"2026-08-31T10:05:57.630738384Z"},"last_recent":true}}
|
||||
|
||||
--- 9e. which unit did the ORDINARY restore read?
|
||||
2026/08/31 10:02:28 restore_unit.go:313: [INFO] [backup] Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit: images=3, secrets recovered=2/2, data_keys=0
|
||||
2026/08/31 10:02:57 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: docmost — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
|
||||
2026/08/31 10:04:23 restore_unit.go:313: [INFO] [backup] Restoring docmost from recovery unit /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit: images=3, secrets recovered=2/2, data_keys=0
|
||||
2026/08/31 10:04:51 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: docmost — 3 volume(s) of 3 listed, 1 database(s) of 1 listed
|
||||
2026/08/31 10:05:33 restore_unit.go:313: [INFO] [backup] Restoring docmost from recovery unit /mnt/sys_drive/felhom-data/backups/primary/docmost: images=3, secrets recovered=2/2, data_keys=0
|
||||
2026/08/31 10:05:57 restore_unit.go:391: [INFO] [backup] Restore-from-unit completed: docmost — 0 volume(s) of 0 listed, 0 database(s) of 0 listed
|
||||
|
||||
--- 9f. final state
|
||||
docmost Up 14 seconds (healthy)
|
||||
docmost-postgres Up 24 seconds (healthy)
|
||||
docmost-redis Up 24 seconds (healthy)
|
||||
ERROR: relation "felhom_r102_discriminator" does not exist
|
||||
LINE 1: select id||chr(32)||chr(124)||chr(32)||note from felhom_r102...
|
||||
^
|
||||
accented file back from the PRIMARY unit: False
|
||||
docmost HTTP 200
|
||||
@@ -235,3 +235,5 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
||||
such a sentence in place; do not delete it.
|
||||
| **R-399** | **How deep should the off-site integrity check go — Viktor ruled full depth.** Shipped in controller **v0.228.0**, 2026-08-31. Evidence: `felhom-controller/REPORT.md` (v0.228.0) — restic argv observed from the guest at both depths on `demo-hp`. **Reasoning kept:** *the structure check does not detect a size-preserving pack corruption — measured 2026-08-30, plain `restic check` reported `no errors were found` and exited 0 over a damaged pack that every read-data form caught. That is the reason for the default and it is what should stop anyone turning it back down to save four seconds.* *An empty value means "not configured", therefore the default; `off` is the off token, because a setting with no off switch is not a setting.* *A malformed value falls back to the DEFAULT, never to structure — falling back to structure would silently remove the protection on a typo, which is R-357's shape.* **Superseded by R-401** for anything about a large store. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user