R-403 CLOSED (controller v0.230.0), R-404 filed as a decision for Viktor
gates / gates (push) Failing after 17s
gates / gates (push) Failing after 17s
07-backup-architecture gains section 8.2, placed beside row 5 on purpose: the derived-copy rebuild rule is UNCHANGED and section 8.2 names the single exception, so a future reader who finds RunTier2 skipping a leg does not fix it back. It carries the measurement (120 082 104 B -> 7 036 B on the shipped v0.229.0), the four-case table, why hollow is a manifest question and not a size question, why the data legs are deliberately not guarded, and why the capture job is not guarded either. 00-capability-map: the Tier-2 row's status does NOT move, stated explicitly rather than left ambiguous. R-403 removes a way the route could be DESTROYED between uses; it does not change what the route can be relied on for. Register: R-403 CLOSED and compressed into CLOSED-ITEMS (594 -> 593 open lines). R-404 FILED as a DECISION and deliberately not acted on - should a documents-only push be subject to the golden-currency gate, now that it has been correctly bypassed six times? Both sides stated, plus what happens if Viktor does nothing. The gate was NOT changed. R-242: seventh conviction, and the FIRST where the day-0 ground does not apply - R-403 is a defect in the nightly Tier-2 copy, which a day-0 box starts running on its first night. This push uses git push --no-verify, declared here and in felhom-controller/REPORT.md. A golden carrying 0.230.0 is owed and is more urgent than the previous six. STATUS: the R-403 item moves out of 'Broken' into what works, in plain words; the delivery item now says a golden is owed and that the fleet carries the defect; R-404 goes into the decide section with its do-nothing outcome. Drill evidence: nine phase logs, including the two things that went wrong (a repair whose rsync was not installed in the guest and silently did nothing, and a session that expired mid-run so a POST did nothing).
This commit is contained in:
@@ -13,21 +13,33 @@ delivered; both demo machines are on it and nothing is waiting on you about this
|
|||||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||||
nothing.*
|
nothing.*
|
||||||
|
|
||||||
1. **Nothing about delivery — the golden train is current.** Golden **0.229.0** was baked, published,
|
1. **A golden carrying 0.230.0 is owed.** Golden **0.229.0** is vouched and the fleet floor is 0.229.0
|
||||||
round-trip verified, vouched, and the fleet floor raised to 0.229.0 on 2026-08-31 (the second bake
|
— **and 0.229.0 is the build that deletes a good copy** (R-403, measured today). Both demo machines
|
||||||
that day; 0.228.0 was the first). Both demo machines run it; **`demo-felhom` got there by itself**
|
and any machine installed right now carry that defect. `demo-hp` has been updated to 0.230.0 by
|
||||||
and re-registered its jobs without anyone touching it. A machine installed today receives 0.229.0
|
hand; `demo-felhom` has not. **If you do nothing:** the fix stays on one machine and a newly
|
||||||
and everything shipped today, **including the restore from the second drive**. Evidence:
|
installed box gets the defect. Baking and vouching 0.230.0 and raising the floor closes it — the
|
||||||
`documentation/tests/golden-0.229.0-2026-08-31/`.
|
same three-field change as this morning.
|
||||||
|
|
||||||
2. **Nothing else about this release.** Everything in 0.229.0 is a fix to code that ships in the
|
2. **Nothing else about this release.** Everything in 0.229.0 is a fix to code that ships in the
|
||||||
controller image; no customer action, no data migration, no credential change.
|
controller image; no customer action, no data migration, no credential change.
|
||||||
|
|
||||||
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
3. **Whether a documents-only push should still be checked for a missing golden** (R-404). We have now
|
||||||
|
skipped that check **six times**, each time for a written reason: it runs on every push to the
|
||||||
|
website/documentation repository, including pushes that change nothing a machine installs.
|
||||||
|
**A guard we correctly skip six times is teaching us to skip it.**
|
||||||
|
**The case for narrowing it:** a documents-only push cannot be the one that finishes a release, so
|
||||||
|
only checking pushes that touch real code would fire on exactly the risky ones and end the habit.
|
||||||
|
**The case against:** the check was earned — a release went out while machines were still being
|
||||||
|
installed with the previous one, three times in three days — and narrowing a guard is how the thing
|
||||||
|
it was built for comes back.
|
||||||
|
**If you do nothing:** nothing breaks, the skipping stays routine, and the count keeps rising.
|
||||||
|
I have NOT changed it; this is yours to decide and mine to build.
|
||||||
|
|
||||||
|
4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||||
|
|
||||||
4. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||||
as it misled one by an hour.
|
as it misled one by an hour.
|
||||||
|
|
||||||
@@ -95,6 +107,20 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
|||||||
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
longer pastes raw database text at you (it was 615 bytes once, including rows out of your own
|
||||||
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
database), and the undo copies no longer pile up forever — three per app, and they were being copied
|
||||||
off-site permanently.
|
off-site permanently.
|
||||||
|
- **An empty package can no longer wipe out a good one** (R-403, controller 0.230.0). Yesterday we
|
||||||
|
wrote this down as *suspected* and said plainly it had not been tested. **We tested it first, and it
|
||||||
|
was real.** On a demo machine, on yesterday's build: an app's copy on the second drive went from
|
||||||
|
**120 MB — four database backups and three data archives — to 7 KB, nothing left**, in a single
|
||||||
|
nightly run, and the run reported success. The cause was that the nightly job only asked *does the
|
||||||
|
folder exist* before copying over it, and an empty package is a folder that exists.
|
||||||
|
Now the nightly job refuses to replace a **complete** package with an **empty** one. It keeps what it
|
||||||
|
has, says so on the app's own backup page, and carries on with everything else. Proven on the same
|
||||||
|
machine, in the same state: **all seven files still there, byte for byte.**
|
||||||
|
Two more things came with it. The page no longer calls that copy fresh when the run did not refresh
|
||||||
|
it — it names the real date of the package instead. And after a restore from the second drive, the
|
||||||
|
first drive's package is filled back in immediately, so the empty state that started all this cannot
|
||||||
|
happen again. A copy that legitimately gets smaller still gets smaller; only the one dangerous case
|
||||||
|
is fenced.
|
||||||
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
|
- **The copy on the second drive can now bring an app back** (R-102 + R-103, controller 0.229.0,
|
||||||
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
|
proven on `demo-hp`, **and delivered** — golden 0.229.0 is vouched and the fleet floor is raised, so
|
||||||
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
|
a machine installed today has it). Every night the box copied each app's whole recovery package onto the second
|
||||||
@@ -151,16 +177,6 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
|||||||
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
|
takes longer than five minutes writes a warning naming this item. **If you do nothing:** every
|
||||||
machine re-reads its whole store every week, however large it grows, and the first person to notice
|
machine re-reads its whole store every week, however large it grows, and the first person to notice
|
||||||
would be a customer whose upload is busy. The warning is there so that does not happen.
|
would be a customer whose upload is busy. The warning is there so that does not happen.
|
||||||
- **After a restore from the second drive, the first drive's package is rewritten EMPTY** (R-403,
|
|
||||||
found during the 0.229.0 drill, **not fixed**). Two seconds after the restore finished, the box's
|
|
||||||
five-minute housekeeping rebuilt the first drive's package from a drive that had no data files on
|
|
||||||
it, and wrote a package that lists nothing. The ordinary „Visszaállítás indítása" then read it and
|
|
||||||
said, correctly and uselessly, that the backup held only settings. **What we did not test:** the
|
|
||||||
nightly copy mirrors the first drive over the second one and deletes what is not there, so the
|
|
||||||
next night could plausibly overwrite the good copy with the empty one. That is a reading of the
|
|
||||||
code, not a measurement, and it is written down as unmeasured on purpose. **If you do nothing:**
|
|
||||||
a customer who recovers from their second drive may find, the next morning, that the copy they
|
|
||||||
recovered from has been replaced by an empty one. Settling it costs one test on a spare machine.
|
|
||||||
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
- **We ask the off-site box a question about once a second** (R-336) — ~85,000 a day for a box we
|
||||||
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
write to weekly. The leak that made this dangerous is fixed (R-344); the volume is not. The ceiling
|
||||||
is **under a year** away on the corrected measurement, not two.
|
is **under a year** away on the corrected measurement, not two.
|
||||||
|
|||||||
@@ -126,7 +126,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
|||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded **DEGRADED** rather than silently normal (installer Case A/B) | agent v0.113, host-install v1.22.0 | **PROVEN-LIVE** | `E2D-fresh-vm-2026-07-29` C1 (real 1.22.0 install, rc=0, `Day-0 provision SUCCESS`) + C2 (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort) | Case A (a second drive already present at install) has never fired naturally — only Case B has |
|
| Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded **DEGRADED** rather than silently normal (installer Case A/B) | agent v0.113, host-install v1.22.0 | **PROVEN-LIVE** | `E2D-fresh-vm-2026-07-29` C1 (real 1.22.0 install, rc=0, `Day-0 provision SUCCESS`) + C2 (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort) | Case A (a second drive already present at install) has never fired naturally — only Case B has |
|
||||||
| Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 **DB replay route → `07-backup-architecture.md` §8 row 3** |
|
| Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 **DB replay route → `07-backup-architecture.md` §8 row 3** |
|
||||||
| Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 3b, 4, 5.** **R-102 CLOSED 2026-08-31 (controller v0.229.0):** the copy's `recovery-unit/` mirror is now restorable — „Teljes visszaállítás a másolatból" / `POST /backup/tier2/unit-restore` — and was proven live on `demo-hp` with the primary unit moved aside (`audits/DRILL-r102-tier2-unit-2026-08-31/`, §8 row 3b, 28.65 s). R-103 closed with it: the refusal that used to name a button on another page now offers the action on the row itself |
|
| Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 3b, 4, 5.** **R-403 (controller v0.230.0, 2026-08-31) changes NO status on this row and that is stated rather than left ambiguous:** the nightly copy now refuses to replace a complete unit package with an empty one (`07` §8.2), which removes a way the route could be DESTROYED between uses — it does not change what the route can be relied on for, and every leg it covers is the same one it covered yesterday. Proven live on demo-hp: the same state that deleted 120 082 104 B on v0.229.0 preserved all 7 files on v0.230.0. **R-102 CLOSED 2026-08-31 (controller v0.229.0):** the copy's `recovery-unit/` mirror is now restorable — „Teljes visszaállítás a másolatból" / `POST /backup/tier2/unit-restore` — and was proven live on `demo-hp` with the primary unit moved aside (`audits/DRILL-r102-tier2-unit-2026-08-31/`, §8 row 3b, 28.65 s). R-103 closed with it: the refusal that used to name a button on another page now offers the action on the row itself |
|
||||||
| Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Route + RTO → `07-backup-architecture.md` §8 rows 4, 10, 12, 15** (incl. the R-95 delete exposure and the R-104 stale-lock defect). **2026-08-04 (R-193/R-197, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`) — the row is NOT overclaiming and the guard is not the gap; the CADENCE is.** This row already recorded that a recreated data volume orphans the repo, and the spike confirms the mechanism at source: `WriteOffboxSecrets` (`offbox.go:392`) mints a fresh 256-bit repo password whenever `<DataDir>/offbox/repo_password` is absent, and **no automatic path ever consults the escrowed one** — `InjectOffboxPassword` has exactly one caller, a web form a human pastes into. What was NOT recorded is that this now fires on an **ordinary, planned, unattended guest rebuild**, on every box: measured on BOTH demo boxes 2026-08-03/04 by comparing `host_escrow.restic_pw_sha256` against `host_escrow_superseded.restic_pw_sha256` (demo-hp `8e03eddf…`→`8a9e33aa…`, demo-felhom `48741892…`→`c60c8bc7…`), orphaning **15 snapshots / 40.9 MB** and **36 snapshots / 1.14 GB** respectively. **demo-felhom is the important half:** it kept its DELIVERY (a stale staged secret restored the target in 76 s) and lost its REPOSITORY anyway, with **nothing marking the escrow stale for 13 h** — `escrow_stale` is wired to `ReissueCredentials`, the one path that does NOT change the repo password (**R-196**), and absent from the rebuild path that does. Status unchanged: the classify-and-move-aside guard remains PROVEN-LIVE and correct, and is predicted (not yet measured) to refuse the 2026-08-05 run on both boxes rather than start a silent fresh history |
|
| Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Route + RTO → `07-backup-architecture.md` §8 rows 4, 10, 12, 15** (incl. the R-95 delete exposure and the R-104 stale-lock defect). **2026-08-04 (R-193/R-197, `audits/SPIKE-offsite-credential-recovery-2026-08-04.md`) — the row is NOT overclaiming and the guard is not the gap; the CADENCE is.** This row already recorded that a recreated data volume orphans the repo, and the spike confirms the mechanism at source: `WriteOffboxSecrets` (`offbox.go:392`) mints a fresh 256-bit repo password whenever `<DataDir>/offbox/repo_password` is absent, and **no automatic path ever consults the escrowed one** — `InjectOffboxPassword` has exactly one caller, a web form a human pastes into. What was NOT recorded is that this now fires on an **ordinary, planned, unattended guest rebuild**, on every box: measured on BOTH demo boxes 2026-08-03/04 by comparing `host_escrow.restic_pw_sha256` against `host_escrow_superseded.restic_pw_sha256` (demo-hp `8e03eddf…`→`8a9e33aa…`, demo-felhom `48741892…`→`c60c8bc7…`), orphaning **15 snapshots / 40.9 MB** and **36 snapshots / 1.14 GB** respectively. **demo-felhom is the important half:** it kept its DELIVERY (a stale staged secret restored the target in 76 s) and lost its REPOSITORY anyway, with **nothing marking the escrow stale for 13 h** — `escrow_stale` is wired to `ReissueCredentials`, the one path that does NOT change the repo password (**R-196**), and absent from the rebuild path that does. Status unchanged: the classify-and-move-aside guard remains PROVEN-LIVE and correct, and is predicted (not yet measured) to refuse the 2026-08-05 run on both boxes rather than start a silent fresh history |
|
||||||
| Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live | controller v0.134/134.1/135 | **PROVEN-LIVE** (2026-07-20) | `CAMPAIGN-6D` accept legs (immich end-to-end from offsite alone) | **2026-07-19:** `audits/DIAG-immich-restore-2026-07-19.md` finds **no offsite path loads a DB dump** — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "**immich end-to-end from offsite alone**" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. **RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume** (`immich_postgres_data` is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" **overclaimed scope**: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → **PARTIAL**, scope-corrected. Evidence: `audits/DIAG-immich-restore-2026-07-19.md` (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). **2026-07-19, controller v0.148.0:** the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but **round 2 found it aborts against a running app** (`audits/DIAG-immich-restore-round2-2026-07-19.md`, H4: the replay races immich's own schema repair; `clip_index` recreated by the app 2 s before the dump's CREATE INDEX). **2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47)** — both restore paths now replay into a DB-ONLY window (`StartStackServices` brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. *(The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard `BackupStatus` fix. R-47 shipped in v0.153.0.)* **2026-07-20: the clean run HAPPENED** — endpoint-level supervised reconstitute of immich from snapshot `49e7cb46` (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no `already exists`, operation reported SUCCESS, immich's own DatabaseService logged `No schema drift detected` twice, 11 assets `active`, 4/4 containers healthy, 231 `public` indexes. **Operator confirmed the immich timeline renders correctly after the reconstitute** (screenshot held, 2026-07-20). Evidence: `felhom-controller/REPORT.md` §4b. **2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE.** The operator deleted the photos in immich own UI **and emptied the trash** (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. **`40 file(s) placed`** against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets `active`, `No schema drift detected`, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: `felhom-controller/REPORT.md` 4e **Route + RTO → `07-backup-architecture.md` §8 rows 3, 4** — the matrix also records that no offsite action unpacks the named-volume tars it captures (→ R-107) |
|
| Offsite restore: local-preferred scratch, unit-only default, full two-step, missing-only place-to-live | controller v0.134/134.1/135 | **PROVEN-LIVE** (2026-07-20) | `CAMPAIGN-6D` accept legs (immich end-to-end from offsite alone) | **2026-07-19:** `audits/DIAG-immich-restore-2026-07-19.md` finds **no offsite path loads a DB dump** — all three buttons are file-only (R-43). The mechanics in this row's title are each proven; the phrase "**immich end-to-end from offsite alone**" is what is contested, since a DB-indexed app cannot be reconstituted by any offsite action. **RULED 2026-07-19 (Viktor): 6D's destruction hit the FILE TREE ONLY — the database survived in its named volume** (`immich_postgres_data` is a named volume in both the v2 and v3 template eras), so "immich end-to-end from offsite alone" **overclaimed scope**: the file half was proven, the DB half was never destroyed and therefore never restored. Row downgraded PROVEN-LIVE → **PARTIAL**, scope-corrected. Evidence: `audits/DIAG-immich-restore-2026-07-19.md` (no offsite path could replay a DB at all) + the P-FAB destructive re-import (the proven-replay evidence, on the LOCAL path). **2026-07-19, controller v0.148.0:** the DB half now exists in code (R-43 + R-44) and its replay reached a live box — but **round 2 found it aborts against a running app** (`audits/DIAG-immich-restore-round2-2026-07-19.md`, H4: the replay races immich's own schema repair; `clip_index` recreated by the app 2 s before the dump's CREATE INDEX). **2026-07-20, controller v0.153.0: H4 IS CLOSED (R-47)** — both restore paths now replay into a DB-ONLY window (`StartStackServices` brings up the database service alone; the app starts only after the replay exits 0), with a fail-closed refusal when a dump has no identifiable DB service. *(The earlier note here said "closes in v0.149" — that was wrong: v0.149.0 was the F3 dashboard `BackupStatus` fix. R-47 shipped in v0.153.0.)* **2026-07-20: the clean run HAPPENED** — endpoint-level supervised reconstitute of immich from snapshot `49e7cb46` (the very snapshot that aborted in round 2): stop → DB-service-only start → replay rc-0 → full start, no `already exists`, operation reported SUCCESS, immich's own DatabaseService logged `No schema drift detected` twice, 11 assets `active`, 4/4 containers healthy, 231 `public` indexes. **Operator confirmed the immich timeline renders correctly after the reconstitute** (screenshot held, 2026-07-20). Evidence: `felhom-controller/REPORT.md` §4b. **2026-07-20, LATER THE SAME DAY — the destructive drill RAN and the row now earns PROVEN-LIVE.** The operator deleted the photos in immich own UI **and emptied the trash** (the step whose absence makes a drill prove nothing — the round-1 lesson), then restored through the customer-facing UI. **`40 file(s) placed`** against the 6 of the earlier non-destructive run — the files were really gone and really came back — plus 1 DB dump replayed rc-0, 11 assets `active`, `No schema drift detected`, timeline confirmed by the operator. This is the destroy-then-recover proof the 6D downgrade asked for, and it was taken through the customer own buttons, not endpoint shortcuts. Evidence: `felhom-controller/REPORT.md` 4e **Route + RTO → `07-backup-architecture.md` §8 rows 3, 4** — the matrix also records that no offsite action unpacks the named-volume tars it captures (→ R-107) |
|
||||||
| Manual `.fab` export/import: class-scoped capture, browser up/download, tunnel-proof chunking | controller v0.125/128/130/136 | **PROVEN-LIVE** | `CAMPAIGN-6D` P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking `CAMPAIGN-6B` P2 (100 MiB via real CF edge, 120 MiB→413) | Chunking proven at the real CF edge via `curl --resolve`; the **rendered browser file-picker** upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B *finding*; fix verified in 6D |
|
| Manual `.fab` export/import: class-scoped capture, browser up/download, tunnel-proof chunking | controller v0.125/128/130/136 | **PROVEN-LIVE** | `CAMPAIGN-6D` P-FAB / Accept #1 (1.7 GB full circle, byte-identical, app boots); chunking `CAMPAIGN-6B` P2 (100 MiB via real CF edge, 120 MiB→413) | Chunking proven at the real CF edge via `curl --resolve`; the **rendered browser file-picker** upload leg is still Viktor's open full-circle test (6C ran it NOT-RUN). C6B-F1 was the 6B *finding*; fix verified in 6D |
|
||||||
|
|||||||
@@ -864,7 +864,7 @@ crosses the line — **R-158**.
|
|||||||
| 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery |
|
| 3c | *same, with the GUEST GONE (secrets unavailable)* | the drive | Tier-1 unit restore — **the unit carries the portable secrets** | **customer** | **27.6 s** | 24 h | **PROVEN** | D5, §7.4. Before v0.188.0 this row was **NONE**: the data-key gate refused and the customer's own copy of their own data was not a recovery |
|
||||||
| 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** |
|
| 3b | *same, for a class-B app via Tier-2* | the Tier-2 copy on the second drive | „Teljes visszaállítás a másolatból" — **Tier-2 unit restore** (`POST /backup/tier2/unit-restore` → `RestoreTier2Unit` → `RestoreFromRecoveryUnitAt`), controller **v0.229.0** | **customer** | **28.65 s** (3 volumes, 1 database, 114.5 MB unit) | 24 h | **PROVEN** (2026-08-31) | `audits/DRILL-r102-tier2-unit-2026-08-31/`. docmost — a class-B app whose Tier-2 run reports **0 leg(s)** — restored **with the primary unit moved aside**: 3 volumes of 3 and 1 database of 1, from `/mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit`. **The observable is the DATA:** an accented Hungarian filename returned byte-for-byte (verified as hex, R-364) and the app read its own row **over TCP with its own credential**; the post-backup discriminator was **gone**, so the tar was genuinely replayed. Repeated with the guest's `app.yaml` also aside → `secrets recovered=2/2` from the mirrored unit. **R-102 CLOSED** |
|
||||||
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore |
|
| 4 | **Primary drive dies** | Tier-2 copy on the second drive; Tier-3 offsite; the guest | Tier-2 for **file legs** (7 of 53 apps) **and, since controller v0.229.0, for the unit mirror — its volume tars and DB dump, i.e. the whole of what the other 45 own**; Tier-3 reconstitute for files + DB **+ the named-volume tars since v0.218.0** (`volReplay`) | **customer** (all) | | 24 h | **PARTIAL** | §7.2. **Both unreachability gaps are now closed — R-107 (v0.218.0) and R-102 (v0.229.0).** This row stays **PARTIAL** deliberately: what is proven is the ROUTE (row 3b, live, primary unit absent), not the JOURNEY. **No drive has ever actually died or been replaced under this recovery** — the drill removed a unit directory, not a disk, so drive re-attachment by `durable_id`, the agent's enrolment of a replacement, and a Tier-2 copy read from a drive that is the ONLY surviving one are all still unexercised. Promoting this row to PROVEN needs that journey, not another unit restore |
|
||||||
| 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go:359-393` |
|
| 5 | **Secondary drive dies** | everything the customer uses | none needed — Tier-2 is a derived copy, rebuilt on the next run (`07` §8 migration rule: *"Migration = rebuild, not preserve"*) | automatic | | 24 h | **PROVEN** (by construction) | tier2 v2 layout marker + rebuild, `internal/backup/tier2.go`. **The derived-copy rule is UNCHANGED by R-403 (controller v0.230.0) and the single exception is stated in §8.2 below — read it before "fixing" a skip you find in the code** |
|
||||||
| 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session |
|
| 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session |
|
||||||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||||||
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
||||||
@@ -877,6 +877,47 @@ crosses the line — **R-158**.
|
|||||||
| 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix |
|
| 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix |
|
||||||
| 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** |
|
| 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** |
|
||||||
|
|
||||||
|
### 8.2 The ONE exception to row 5's derived-copy rule (R-403, controller v0.230.0)
|
||||||
|
|
||||||
|
**[DESIGN] Row 5 stands: the secondary IS a derived copy and it IS rebuilt on the next run.** Nothing
|
||||||
|
below weakens that, and a future reader who finds `RunTier2` skipping a leg and does not find this
|
||||||
|
section will "fix" it back — which is why it is here and not only in a register row.
|
||||||
|
|
||||||
|
**[FACT] What was measured, on the shipped v0.229.0, on `demo-hp` 2026-08-31.** An app's Tier-2 copy
|
||||||
|
went from **120 082 104 B** (4 database dumps + 3 named-volume tars) to **7 036 B** (none of either)
|
||||||
|
in one nightly run, and the run recorded itself a success. The mechanism was three individually
|
||||||
|
correct lines: `RunTier2` guarded the unit leg with `os.Stat` alone — *does the folder exist* —
|
||||||
|
`rsyncMirror` is `rsync -a --delete`, and nothing between them compared source to destination. **An
|
||||||
|
empty recovery unit is a folder that exists.** Evidence:
|
||||||
|
`audits/DRILL-r403-tier2-delete-2026-08-31/`.
|
||||||
|
|
||||||
|
**Why it bites harder since 2026-08-30:** R-102 made that mirror a LIVE recovery route (§6.3, §8 row
|
||||||
|
3b). Deleting it used to cost a copy nobody could open; it now costs the route itself.
|
||||||
|
|
||||||
|
**[DESIGN] The exception, stated exactly.** The unit leg — and ONLY the unit leg — is skipped when the
|
||||||
|
SOURCE unit carries no data and the DESTINATION unit does. Everything else is unchanged:
|
||||||
|
|
||||||
|
| source unit | destination unit | behaviour |
|
||||||
|
|---|---|---|
|
||||||
|
| complete | complete | mirror, with `--delete`, as before |
|
||||||
|
| complete | hollow or absent | mirror (the normal first copy) |
|
||||||
|
| hollow | hollow | mirror — both sides agree, nothing is at risk |
|
||||||
|
| **hollow** | **complete** | **skip the unit leg, preserve the destination, warn, record for the surface** |
|
||||||
|
|
||||||
|
**"Hollow" is a MANIFEST question, never a size question** — the manifest lists no database dump and
|
||||||
|
no volume tar; absent or unparseable counts as hollow, fail closed. A unit with a fat compose capture
|
||||||
|
and no dumps is the dangerous shape; a 360-byte unit belonging to a tiny app is healthy.
|
||||||
|
|
||||||
|
**The data legs are NOT guarded and must not be.** A classified app's copy legitimately shrinks as
|
||||||
|
`export` drops out of its class set (`tier2.go` header), and fencing that would be calling this row's
|
||||||
|
own decision a defect.
|
||||||
|
|
||||||
|
**[DESIGN] And the CAUSE is closed at the other end.** The hollow primary was written by the 5-minute
|
||||||
|
capture job **two seconds** after a Tier-2 unit restore. Since v0.230.0 `RestoreTier2Unit` refills an
|
||||||
|
absent or hollow primary unit from the mirror **inside the call, before returning**, so no capture can
|
||||||
|
observe the hollow state. **The capture itself is deliberately not guarded:** a capture that describes
|
||||||
|
an empty drive as empty is correct, and guarding it would make the manifest lie.
|
||||||
|
|
||||||
### 8.1 The blank cells, listed explicitly
|
### 8.1 The blank cells, listed explicitly
|
||||||
|
|
||||||
Per the rule that a blank is a finding, here they are:
|
Per the rule that a blank is a finding, here they are:
|
||||||
|
|||||||
+18
@@ -0,0 +1,18 @@
|
|||||||
|
=== THE WARN, verbatim ===
|
||||||
|
2026/08/31 12:14:24 tier2.go:383: [WARN] [backup] Tier 2 docmost: unit leg SKIPPED — the recovery unit on the source drive lists no database dumps and no volume tars, while the existing copy at /mnt/felhom-drives/hdd_1/backups/secondary/docmost/recovery-unit does. The copy was PRESERVED rather than replaced with an empty one (R-403). The other legs continue.
|
||||||
|
2026/08/31 12:14:24 tier2.go:424: [INFO] [backup] Tier 2 copied docmost → /mnt/felhom-drives/hdd_1/backups/secondary/docmost (14.9 KB, 0 leg(s), 0s) [unit leg SKIPPED — existing package preserved, R-403]
|
||||||
|
|
||||||
|
=== the run summary line ===
|
||||||
|
2026/08/31 12:14:24 tier2.go:424: [INFO] [backup] Tier 2 copied docmost → /mnt/felhom-drives/hdd_1/backups/secondary/docmost (14.9 KB, 0 leg(s), 0s) [unit leg SKIPPED — existing package preserved, R-403]
|
||||||
|
|
||||||
|
=== AFTER the run ===
|
||||||
|
db-dumps: 4 volume-dumps: 3 size: 120082104
|
||||||
|
9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28 db-dumps/docmost-postgres.sql
|
||||||
|
73917ba6bc3072dfc7b5be9c6df4f8361da7e987230f5d56f7b62f397fe15ef1 db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||||
|
4c134c2ced74df26f49ef1910694cbbd25f2598549bb4b5bb145aac054935949 db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
|
||||||
|
13e5a864701966d9e4053b5bb7dd800cca3d77ebb07f4fd2f32c86389422af25 db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
|
||||||
|
f46a2fc3aa9a7ae2502d83b1c6ef27e503102f5ba71a0c6559246d9674c8e3b1 volume-dumps/docmost_docmost_postgres_data.tar
|
||||||
|
a8df17c444e41f54762e122ce1be998315c969015a1580bb8abdc7211cfa1a73 volume-dumps/docmost_docmost_redis_data.tar
|
||||||
|
88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d volume-dumps/docmost_docmost_storage.tar
|
||||||
|
|
||||||
|
VERDICT: PRESERVED — all 7 files still there (v0.229.0 left 0)
|
||||||
+23
@@ -0,0 +1,23 @@
|
|||||||
|
######## SCENARIO B LIVE — the same state, on the FIXED build ########
|
||||||
|
UTC 2026-08-31T12:04:11Z
|
||||||
|
controller under test: gitea.dooplex.hu/admin/felhom-controller:0.230.0
|
||||||
|
|
||||||
|
--- recreate the hollow primary (hand-made this time; the R-102 path was proven in phase 1b)
|
||||||
|
hollow primary manifest written: db_dumps=[] volume_dumps=null
|
||||||
|
primary tree:
|
||||||
|
compose/.felhom.yml
|
||||||
|
compose/app.yaml
|
||||||
|
compose/docker-compose.yml
|
||||||
|
manifest.json
|
||||||
|
|
||||||
|
=== BEFORE the run ===
|
||||||
|
secondary db-dumps: 4 volume-dumps: 3 size: 120082104
|
||||||
|
9f676376f759733f5b62e590e4a2b31dddd66ff49990df3394332b790a092a28 db-dumps/docmost-postgres.sql
|
||||||
|
73917ba6bc3072dfc7b5be9c6df4f8361da7e987230f5d56f7b62f397fe15ef1 db-dumps/pre-restore-20260822T162347Z-docmost-postgres.sql
|
||||||
|
4c134c2ced74df26f49ef1910694cbbd25f2598549bb4b5bb145aac054935949 db-dumps/pre-restore-20260822T162708Z-docmost-postgres.sql
|
||||||
|
13e5a864701966d9e4053b5bb7dd800cca3d77ebb07f4fd2f32c86389422af25 db-dumps/pre-restore-20260822T215432Z-docmost-postgres.sql
|
||||||
|
f46a2fc3aa9a7ae2502d83b1c6ef27e503102f5ba71a0c6559246d9674c8e3b1 volume-dumps/docmost_docmost_postgres_data.tar
|
||||||
|
a8df17c444e41f54762e122ce1be998315c969015a1580bb8abdc7211cfa1a73 volume-dumps/docmost_docmost_redis_data.tar
|
||||||
|
88f21f491d0766aa7a1fc9eba5866e5fffd7a72fa640c55f7bccf575f2ba751d volume-dumps/docmost_docmost_storage.tar
|
||||||
|
|
||||||
|
=== POST /api/backup/tier2 ===
|
||||||
+10
@@ -0,0 +1,10 @@
|
|||||||
|
rows found: 8
|
||||||
|
app notice FIGYELEM skipped? package date in the confirm
|
||||||
|
bookstack False False False 2026-08-31 14:03
|
||||||
|
calibre-web False False False 2026-08-31 14:03
|
||||||
|
docmost True True True 2026-08-31 11:43
|
||||||
|
kimai False False False 2026-08-31 14:03
|
||||||
|
opengist False False False 2026-08-31 14:03
|
||||||
|
paperless-ngx False False False 2026-08-31 14:03
|
||||||
|
privatebin False False False 2026-08-31 14:03
|
||||||
|
romm False False False 2026-08-31 14:03
|
||||||
+19
@@ -0,0 +1,19 @@
|
|||||||
|
######## SCENARIO E LIVE — the primary is filled back in ########
|
||||||
|
UTC 2026-08-31T12:24:38Z
|
||||||
|
--- BEFORE: the primary is hollow (this is the state a Tier-2 unit restore leaves on 0.229.0)
|
||||||
|
created_at: 2026-08-31T12:08:49Z db_dumps: [] volume_dumps: None
|
||||||
|
compose/.felhom.yml
|
||||||
|
compose/app.yaml
|
||||||
|
compose/docker-compose.yml
|
||||||
|
manifest.json
|
||||||
|
|
||||||
|
--- the R-102 restore, through the real endpoint
|
||||||
|
302 https://127.0.0.1:443/backups/apps?flash=Teljes+vissza%C3%A1ll%C3%ADt%C3%A1s+elindult+%E2%80%94+az+%C3%A1llapot+itt+friss%C3%BCl.
|
||||||
|
{"ok":true,"data":{"running":false,"op":"tier2-unit-restore","stack":"docmost","started_at":"2026-08-31T12:24:38.178701187Z","last":{"op":"tier2-unit-restore","stack":"docmost","ok":true,"message":"A(z) docmost: 3 adatkötet és az adatbázis visszaállítva — az alkalmazás újraindult. A visszaállítás forrása a második meghajtón lévő másolat volt (2026-08-31 14:23).","finished_at":"2026-08-31T12:25:07.772790659Z"},"last_recent":true}}
|
||||||
|
|
||||||
|
--- IMMEDIATELY after the call returned (no waiting): is the primary a real package?
|
||||||
|
created_at: 2026-08-31T09:43:41Z db_dumps: ['docmost-postgres.sql'] volume_dumps: ['docmost_docmost_postgres_data.tar', 'docmost_docmost_redis_data.tar', 'docmost_docmost_storage.tar']
|
||||||
|
volume tars on the app drive: 3 db dumps: 4
|
||||||
|
|
||||||
|
--- the rehydrate log line, verbatim
|
||||||
|
2026/08/31 12:25:07 tier2_restore.go:253: [INFO] [backup] docmost: primary unit refilled from the secondary mirror (R-403) — 3 volume tar(s), 4 database dump(s) now on the app's own drive
|
||||||
+45
@@ -0,0 +1,45 @@
|
|||||||
|
######## TEARDOWN / final state ########
|
||||||
|
UTC 2026-08-31T12:33:57Z
|
||||||
|
--- a normal Tier-2 run now that BOTH sides are complete (the unit leg must NOT be skipped)
|
||||||
|
200
|
||||||
|
2026/08/31 12:33:57 tier2.go:424: [INFO] [backup] Tier 2 copied docmost → /mnt/felhom-drives/hdd_1/backups/secondary/docmost (114.5 MB, 0 leg(s), 0s)
|
||||||
|
unit-leg skips in this run: 0
|
||||||
|
|
||||||
|
--- both copies
|
||||||
|
PRIMARY : 3 tars, 4 dumps, 120082104 bytes
|
||||||
|
SECONDARY : 3 tars, 4 dumps, 120082104 bytes
|
||||||
|
IDENTICAL volume-dumps/docmost_docmost_postgres_data.tar
|
||||||
|
IDENTICAL db-dumps/docmost-postgres.sql
|
||||||
|
|
||||||
|
--- the surface is back to normal (no stale notice anywhere)
|
||||||
|
adatcsomagja occurrences: 0
|
||||||
|
|
||||||
|
--- remove the drill safety net and the driver
|
||||||
|
safekeeping removed
|
||||||
|
.bash_history
|
||||||
|
.bashrc
|
||||||
|
.docker
|
||||||
|
.profile
|
||||||
|
.ssh
|
||||||
|
settings.json.drill-backup
|
||||||
|
|
||||||
|
--- every app healthy?
|
||||||
|
docmost Up 9 minutes (healthy)
|
||||||
|
docmost-redis Up 9 minutes (healthy)
|
||||||
|
docmost-postgres Up 9 minutes (healthy)
|
||||||
|
romm Up 3 hours (healthy)
|
||||||
|
romm-db Up 3 hours (healthy)
|
||||||
|
romm-redis Up 3 hours (healthy)
|
||||||
|
privatebin Up 3 hours (healthy)
|
||||||
|
paperless-webserver Up 3 hours (healthy)
|
||||||
|
paperless-postgres Up 3 hours (healthy)
|
||||||
|
paperless-redis Up 3 hours (healthy)
|
||||||
|
opengist Up 3 hours (healthy)
|
||||||
|
kimai Up 3 hours (healthy)
|
||||||
|
kimai-db Up 3 hours (healthy)
|
||||||
|
calibre-web Up 3 hours (healthy)
|
||||||
|
bookstack Up 3 hours (healthy)
|
||||||
|
bookstack-db Up 3 hours (healthy)
|
||||||
|
filebrowser Up 9 days (healthy)
|
||||||
|
cloudflared Up 9 days
|
||||||
|
traefik Up 9 days
|
||||||
@@ -237,3 +237,4 @@ Compressed here to title, shipping version, evidence, and the sentences that sta
|
|||||||
| **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
|
| **R-400** | **A third of the debug page posted to endpoints that did not exist — and three of the seven fetched on page LOAD.** Shipped in controller **v0.228.0**, 2026-08-31. 24 referenced / 17 dispatched became 18 / 18. `backup/crossdrive` implemented (proven live: real Tier-2 copies for three apps); `backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status` and both `storage/simulate-*` deleted with their panels and JavaScript. **Reasoning kept:** *implement or delete FIRST, register the gate SECOND — a registered-but-failing gate refuses every push.* *Keep `handleDebugAPI`'s exact-match switch with its `NotFound` default; a prefix match would have made the defect invisible instead of merely silent.* *A panel left behind renders nothing forever, which is how this class hides.* *A debug control that simulates or mutates storage state is deleted unless a live need can be shown — that is where drives get unenrolled and data gets stranded.* Enforced by `controller/scripts/debug_route_gate.py`, both directions, red-proofed. Original text: `git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
| **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
| **R-102** (was **C9-F4**) | **Tier-2 wrote a full `recovery-unit/` mirror on every run and no code path read it** - `RecoveryUnitPath` joined a hard-coded `backups/primary/`, so in the one failure Tier-2 exists for the surviving copy was unopenable. Shipped in controller **v0.229.0**: four unit-directory-relative path primitives in `appbackup`, `RestoreFromRecoveryUnitAt(stack, unitDir)`, `RestoreTier2Unit`. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *THE SOURCE MOVES; THE DESTINATION DOES NOT* - `unitDir` changes only where a unit is READ from; data still lands in the live volumes and the live database container, resolved by `GetAppDrivePath` exactly as the capture is, because a restore that also relocated an app's data would be a migration wearing a restore's label. And: *a directory that exists is not a package* - the Tier-2 route refuses fail-closed unless the mirror carries a parseable manifest. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE with the primary unit moved aside** (`07` §8 row 3b -> PROVEN, 28.65 s; row 4 stays PARTIAL - the drive-loss JOURNEY is still unexercised) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
| **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
| **R-103** (was **C9-F1b**) | **The Tier-2 no-coverage refusal named the working action but did not route to it** - it sent the customer to a button on another page for data that R-102 made restorable on the page they were already looking at. Shipped in controller **v0.229.0**: `POST /backup/tier2/unit-restore` and „Teljes visszaállítás a másolatból” on the Tier-2 row. Evidence: `audits/DRILL-r102-tier2-unit-2026-08-31/`. **Reasoning kept:** *a destructive operation reached from a non-destructive surface must carry the difference in the CONFIRM, not in the label* - the two actions stay two buttons because they are two promises, and the confirm names the copy's date, differently when that date is only an attempt clock (R-101). And: *two questions, two predicates* - `CanRestore()` was NOT widened to cover the unit; one predicate answering two questions is R-356, which refused 40 running apps for months. And: `tier2UnitNotCoveredMsg` was NOT deleted, because it is appended where the FILE restore ran and is still exactly true of it. | **CLOSED 2026-08-31 - controller v0.229.0, PROVEN-LIVE** (the refusal now carries `tier2UnitAvailableMsg`, verified at the endpoint) | full text: `git show 1623a4d5b5d5:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
| **R-403** | **A poorer copy deleted a richer one: an EMPTY recovery unit on the primary drive was mirrored over a COMPLETE copy on the second drive, with `--delete`.** Shipped in controller **v0.230.0**. **MEASURED before it was fixed** — on the shipped v0.229.0, on demo-hp: 120 082 104 B (4 database dumps + 3 volume tars) -> 7 036 B (none of either) in one nightly run, recorded as a success. Evidence: `audits/DRILL-r403-tier2-delete-2026-08-31/`. **Reasoning kept:** *hollowness is a MANIFEST question, never a size question* - a unit with a fat compose capture and no dumps is the dangerous shape and a 360-byte unit for a tiny app is healthy; absent or unparseable manifest counts as hollow, fail closed. *The guard fences ONE shape and not shrinking* - `07` §8 row 5's derived-copy rebuild is a DESIGN DECISION, `--delete` stays, the data legs are untouched, and only source-hollow-over-destination-complete is refused (§8.2 records the exception beside the rule so nobody 'fixes' it back). *The rehydrate happens INSIDE the restore* - the hollow manifest was written two seconds later by the 5-minute capture job, so any follow-up job races it; and *the capture is deliberately NOT guarded*, because a capture describing an empty drive as empty is correct and guarding it would make the manifest lie. *A warning that fires on everything costs the same as the comforting lie it replaces* - the first draft flagged 'package older than the run', which is true of every healthy app, and four healthy apps on the box would have been warned. | **CLOSED 2026-08-31 - controller v0.230.0, PROVEN-LIVE both ways** (the loss reproduced on v0.229.0, then the same state preserved on v0.230.0 with all 7 files sha256-identical) | full text: `git show 66156c619fd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||||
|
|||||||
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user