2a56f557d048b7b76d387d67cb7ab51636647468
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0f9b796615 |
R-102: the recovery unit on the second drive becomes a way back
Tier-2 mirrors each app's whole recovery unit to <dest>/backups/secondary/<app>/recovery-unit/ on every run and has done for months. Nothing read it. In the one failure Tier-2 exists for - the primary drive is lost, and the primary unit with it - the surviving copy could not be opened by any action in the product (07-backup-architecture 6.3, 7.2). Part 1.2: RestoreFromRecoveryUnitAt(stack, unitDir) holds the whole body; RestoreFromRecoveryUnit is the thin caller naming the primary unit. ONE implementation, two callers. The SOURCE moves; the DESTINATION does not - live Docker volumes, the live database container, the guest's definition, all unchanged. The R-47 mutation order, the secret reconciliation with unit-over-guest precedence, the fail-closed data-key gate and the no-unit fallback with CountsUnknown are untouched. reimportDBDumpsAtCtx is the bounded-context twin of reimportDBDumpsFrom; the 35-minute bound is now named once so the two paths cannot drift. The R-354 volume-replay seam is reused rather than a second one invented, which is what lets the acceptance test assert the volume leg's source directory. Part 1.3: RestoreTier2Unit resolves the recorded copy, refuses fail-closed unless the mirror carries a parseable manifest - a directory is not a package - and delegates. The single-writer flag is taken inside RestoreFromRecoveryUnitAt, not beside it. Part 2.1: Tier2Coverage gains UnitRestorable and the copy's dates. CanRestore() is NOT widened; it still answers only 'can the file restore run?'. One predicate answering two questions is R-356, which refused 40 running apps for months. Tests: A2-A6 and B1-B5, plus two non-regression guards. The Tier-2 fixtures build their mirror with the production RunTier2, so the claim is 'the copy Tier-2 writes is the copy this restore reads'. Red-proofs: A5 (swap volumes/recreate -> fails on the order), B2 (point the reader back at the primary -> fails with the mirror never reaching the redeploy, and with permission denied once the primary tree is unreadable). |
||
|
|
c0c8fe67bf |
An unknown drawn as a zero: the defect v0.226.0's own fix introduced
gates / gates (push) Failing after 13s
Writing the REPORT's observation "the no-unit fallback already reports a zero result, which is honest" exposed that the sentence was FALSE. A zero UnitRestoreResult is Scenario B's shape. So RestoreFromRecoveryUnit's fallback to RestoreApp -- which returns only an error, and whose signature is deliberately out of scope -- would have printed "ez a mentes csak a beallitasokat tartalmazta, adatot nem" over a restore that may have replayed the app's entire dataset. That is an unknown drawn as a zero: the exact R-88 failure direction this whole change exists to remove, re-introduced by the change. UnitRestoreResult now carries CountsUnknown, the fallback sets it, and there is a fourth sentence claiming only what is known -- the restore ran, the app is back, and we cannot say what came back. RestoreApp's signature is untouched. Pinned by TestUnitRestoreOutcome_NoUnitFallbackSaysUnknownNotEmpty. The A5 seam test was corrected too: its fixture has no recovery unit, so it exercises exactly this path and had been asserting the wrong sentence -- it now asserts the unknown, which is what pins the fallback to it. IT WAS THE observations GATE REFUSING THE PUSH THAT FORCED THE RE-READ. A gate written to stop findings dying in an overwritten REPORT.md caught a live defect instead. Also files R-397 (NotifyIntegrityOK/Failed are dead code AND the monitoring page advertises a weekly integrity check that does not exist) and R-398 (resticStep is not a seam, which is why R-358's ordering needed an AST test) rather than leaving them in a file that is overwritten every session. REPORT.md is the full run record: baselines re-confirmed, per-test results, the five red-proofs with their observed output, the live validation with verbatim Hungarian messages, what was NOT validated and why, teardown across three layers, and the register 165 -> 167 -> 161. Green gate clean: 28 packages, rc 0. All 12 controller gates OK. |
||
|
|
b8af72764d |
R-353/R-357/R-358/R-360: the restore tells the truth (v0.226.0)
gates / gates (push) Successful in 11s
Four defects on the restore surface, all proven on demo-hp during the 2026-08-21 backup-truth drill, all still in shipped code. They share one acceptance idea: a restore surface must state what it actually did, and must refuse what it cannot do. VERSION NOTE. The task specifying this targeted v0.224.0 against baseline |
||
|
|
4ed938cce4 |
D5: an app restore works from the drive alone (v0.188.0)
The recovery unit on the customer's drive now carries the PORTABLE secret class, so Tier-1/Tier-2 restore no longer depends on the whole-guest tier. A customer needs the drive and nothing else. Part 0's rulings overturned the brief's recommendation, on evidence: - the data_key flag is untrustworthy (4+ encryption keys the catalog itself labels as such are unflagged) -> R-127 - a DB password is not resettable in practice: POSTGRES_PASSWORD is ignored once PGDATA is non-empty, so a regenerated value leaves the app unable to authenticate against its own restored rows while the dump replay still reports success (proven on a throwaway postgres:16-alpine) Ruling (operator): type:secret travels, type:password never does, minus the nonPortableSecrets code register. Plaintext -- withholding the internet- reachable class is what licenses that, and the two are coupled. Precedence: the UNIT WINS over the guest -- the unit's secrets were captured in the same run as the dumps beside them, so they match the data being restored. The fail-closed data-key gate is unchanged. Secret values are never logged; the manifest records NAMES only. |
||
|
|
78ff991f1c |
v0.153.0 — R-47: the DB replay no longer races the app, on BOTH restore paths
Closes R-47. No new agent coupling — MinAgent stays 0.90.0. The replay needs a running DB container, so both restore paths started the WHOLE stack first, giving the application a window to rebuild the very schema objects the dump was about to create. Measured live on 2026-07-19 (H4, DIAG-immich-restore-round2): immich-server rebuilt clip_index two seconds before the dump's CREATE INDEX, the replay aborted "already exists" under ON_ERROR_STOP=1, and immich reported schema drift. The data survived only because pg_dump emits COPY before CREATE INDEX. Both paths now open a DB-ONLY window: only the stack's database service(s) come up, the dump is replayed with the app still down, and the full start runs only after the replay exits 0. Fail-closed: a dump with no identifiable DB service refuses BEFORE the first mutation. Every exit from the window still does a best-effort full start, so a failed restore never leaves a box with a database and no application. New: appbackup.DBServiceNames (yaml.v3 services-map parse — never a line scan; immich's top-level volume keys are the decoy) sharing dbTypeForImage with DiscoverDatabases; stacks.Manager.StartStackServices (refuses an empty list — argument-less `up -d` is a full start); RedeployFromEnv split into PersistUnitRedeployConfig + its unchanged tail. StackDataProvider's RecreateStackFromUnit becomes RecreateStackDefinitionFromUnit — the hidden `up -d` inside the old name is what carried the defect on the local path. 19 new tests (ordering plus state-at-replay-time, zero-mutation fail-closed effects, replay-failure bring-up, parser decoys, empty-list refusal); three companion red-proofs run and reverted. 23/23 packages green. Not yet live-validated: STOP-1 supervised reconstitute, golden 0.153.0. |
||
|
|
a52851e79e |
fix(backup): O4 — generate a replacement for unrecoverable resettable secrets on restore
The proceed-path for a missing RESETTABLE secret redeployed the app with the secret blank (compose "Defaulting to a blank string" → exit 1, live-hit in the 2026-07-04 drill Phase 5). Now the restore generates a fresh credential instead: - stacks.Manager.GenerateSecretForField: replacement value from the field's catalog generate spec via the deploy flow's generateValue (no logic copied); refuses data-keys (defense-in-depth), spec-less and non-secret fields. - backup.Manager.SetSecretGenerator seam (wired in main.go), consulted in RestoreFromRecoveryUnit AFTER the untouched fail-closed gate, for missing names NOT in DataKeyEnvVars. The generated value rides fullEnv into RecreateStackFromUnit → RedeployFromEnv → SaveAppConfig, so it persists encrypted in the guest app.yaml and round-trips on the next backup/restore (no second write path). reconcileRestoreSecrets stays pure and untouched. - WARNs now discriminate: "generated replacement for X (credential was reset)" vs "X unrecoverable and has no generator — app may fail to start". Values are never logged (asserted in test). - Residual case (documented, not pretended away): if a restored volume tar carries the OLD internal credential hash, the app may still fail auth until a manual in-DB reset — generation fully fixes only the fresh-init case. Companion red-proof: pre-fix behaviour (generation skipped) fails TestRestoreGeneratesMissingResettableSecret on the non-empty DB_PASSWORD assertion (verified, reverted). Data-key gate proven unreachable by generation in TestRestoreGenerationNeverReachesDataKeys. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6 |
||
|
|
0b9450e356 |
F17: per-app restore now replays the captured .sql DB dump (.sql wins)
The restore paths (RestoreFromRecoveryUnit + the RestoreApp fallback) repopulated Docker volume tars but NEVER replayed the captured <stack>-<dbtype>.sql dump, so DB-resident data (e.g. rows in a DB whose data dir is a bind mount) did not come back — the romm marker round-trip in the audit lost the row. New appbackup.ImportDump (read-side counterpart to DumpOne) replays a .sql/.sql.gz into the running DB using the live container's OWN discovered credentials (no env threading; reuses DiscoveredDB + getMariaDBPassword). backup.reimportDBDumps orchestrates it AFTER volume restore + stack bring-up, so the logical dump WINS over any volume-tar copy of the DB (operator-chosen precedence). pg_dump --clean --if-exists and mariadb-dump (default --add-drop-table) make replay idempotent; psql ON_ERROR_STOP=1 surfaces real import errors. Also: volume-restore per-volume failures and DB-import failures now SURFACE (the restore returns an error) instead of a swallowed WARN, so a failed data restore cannot read as success. Tests (restore_db_test.go, injectable discover/import seams): imports when dump+DB present, failure surfaces, no-dump skips discovery, dump-but-no-matching-DB is a non-fatal skip. Live DB round-trip to be validated post-deploy. |
||
|
|
7863e62f29 |
v0.54.0: Phase 2b — restore-from-recovery-unit + fail-closed data-key gate
Restore recreates an app from its on-drive unit + the guest's own secrets, regenerating nothing. reconcileRestoreSecrets (pure, unit-tested) merges the unit's non-secret env with secrets recovered from the live app.yaml and FAILS CLOSED if a data-encrypting key is unrecoverable (refuse — a PBS whole-guest restore is needed — rather than regenerate and corrupt). Resettable secrets missing → warn + proceed. - backup: RestoreFromRecoveryUnit (manifest -> recover secrets -> gate -> restore volumes -> recreate definition + redeploy w/ re-pull); falls back to volume-only. - seams: RecoverStackSecrets/RecreateStackFromUnit (adapter +encKey), stacks.RedeployFromEnv. Wired into /backup/restore. - tests: gate (refuse/proceed/verbatim) + data_key parsing. Gate + reconcile + data_key parsing unit-tested; capture live-validated (v0.53.1). Full readable-data e2e vs AdventureLog needs the auth-gated dashboard restore — pending. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |