Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
9.2 KiB
REPORT — TASK C1: restore-path fixes (F1 dead restore UI + F3 volume dumps + O4 blank-secret redeploy)
Date: 2026-07-05 · Class: implementation (code + tests + deploy + live validation)
Baseline: main @ a39ab65 (v0.98.3, matched the task's confirmed baseline exactly; not behind origin)
Shipped: v0.99.0, live on guest 9201 (gitea.dooplex.hu/admin/felhom-controller:0.99.0, healthy)
Scope honoured: F1 + F3 + O4 only. F2 (one-click in-place class-C restore) untouched — remains open for TASK C2. No F9 bypass; every restore/backup action went through the real endpoints/UI.
Commits (all on main)
| Hash | What |
|---|---|
f413f95 |
F1 — GET /api/backup/snapshots (backup.ListRestorePoints + api route + guards + tests) |
73378a8 |
F3 — runVolumeDumps wired into the backup run (+ seam + tests) |
a52851e |
O4 — resettable-secret generation on restore (GenerateSecretForField + SetSecretGenerator + tests) |
88362da |
docs — CHANGELOG v0.99.0, CONTEXT, REUSE, controller/README |
felhom.eu b5d92a8 |
docs — backup-architecture.md (volume dumps real, snapshots endpoint, O4 + residual case) |
Files touched (code): internal/api/router.go, internal/backup/{restore_points.go(new), backup.go, restore_unit.go}, internal/stacks/deploy.go, cmd/controller/main.go. Tests (all new files): internal/api/backup_snapshots_test.go, internal/backup/{restore_points_test.go, volume_dumps_test.go, restore_secrets_gen_test.go}, internal/stacks/deploy_secretgen_test.go.
What was fixed
- F1: the restore panel fetched
GET /api/backup/snapshots— a restic-era route that no longer existed — so the dropdown never populated and "Visszaállítás indítása" never enabled. Now backed byListRestorePoints: at most ONE honest entry (the current recovery unit;time= newest artifact mtime,short_id:"helyi",tier:1always — Tier-2 copies are NOT restorable viaPOST /backup/restoreand are never listed),drive_labelfrom the registry. Guards: traversal/empty → 400 (validStackParam, same semantics as web'svalidStackName; api↔web import cycle prevents sharing the symbol — recorded in REUSE §6), unknown stack → 404, no unit →ok:true, data:[]. No template change needed. - F3:
DumpAppVolumesSafehad no caller →volume-dumps/never produced. NewrunVolumeDumpsinrunDBDumpsInternal, BEFOREcaptureAllRecoveryUnits(manifests enumerate fresh tars). Gate order is load-bearing: protected + has-volumes checks precede the Safe call (it stops the stack before its own check). Disconnected/decommissioned drives skip DB-loop-style. Failures land in the summary (FAIL <app> volumes:) and fail the run. Zero-DB early return removed. Seam:dumpVolumesSafefunc field. - O4: a missing resettable secret no longer redeploys blank: generated from the catalog field's
generatespec (stacks.GenerateSecretForField— refusesdata_key/spec-less/non-secret fields — via thebackup.SetSecretGeneratorseam), persisted encrypted through the existingRecreateStackFromUnit→RedeployFromEnv→SaveAppConfigpath (no second write path; round-trips on later backups).reconcileRestoreSecrets(the frozen fail-closed gate) untouched. WARNs discriminate generated-vs-ungenerable; values never logged.
Tests & companion red-proofs
go build ./... && go vet ./... && go test ./... green after every commit. Top-level test functions: 272 → 286 (+14). All three companions executed and reverted:
| Companion mutation | Expected red | Outcome |
|---|---|---|
F1: ListRestorePoints hollowed to always-[] |
unit-on-disk tests fail | FAILED as expected ("want exactly 1 restore point" family), reverted, green |
F3: volume gate neutered (len(vols)==0 check removed) |
gating tests fail | FAILED as expected (dump invoked for [nextcloud rallly]; _VolumelessNeverStopped saw dumped=1), reverted, green |
| O4: generation skipped (pre-fix behaviour) | TestRestoreGeneratesMissingResettableSecret fails |
FAILED as expected (DB_PASSWORD = "", want the generated replacement), reverted, green |
Extra non-hollow assertions: the generated secret VALUE is asserted absent from the logs (secrets safety); TestRestoreGenerationNeverReachesDataKeys proves the fatal gate fires before generation ever runs; the api tests dispatch through Router.ServeHTTP (the F1 bug WAS a missing route).
Deploy + live validation (§13, on felhom-pve / guest 9201 only)
Build: ./build.sh 0.99.0 --push on 180 (explicit git pull first). Deploy: 9201 bootstrap mechanism. Verify: gitea.dooplex.hu/admin/felhom-controller:0.99.0 Up (healthy), startup logs clean.
- Snapshots endpoint:
?stack=nextcloud→{"ok":true,"data":[{"time":"2026-07-05T09:58:34Z","short_id":"helyi","tier":1,"drive_label":"Tárhely (felhom-usb)"}]};?stack=sparkyfitness→ same shape,drive_label:""(SSD fallback);?stack=../../etc→ 400;?stack=ghostapp→ 404. - UI end-to-end (claude-in-chrome — closes the drill's Phase-7 PARTIAL for restore):
/backups→ Alkalmazás=Nextcloud → snapshot dropdown populated ("2026-07-05 vasárnap 12:01 (helyi) — 1. szint, Tárhely (felhom-usb)") → confirm checkbox → button enabled → click → flash "nextcloud visszaállítva (helyi)." Screenshot-verified; sidebar shows 0.99.0. - Drill Phase-2 rerun (the F3 red now green): sentinel
DRILL-VOL-c4c5e110…(sha256560cb2a7…) planted insidenextcloud_nextcloud_html→POST /api/backup/run→ 5 volume dumps produced (log: "App-data backup completed: 4 databases (810.2 KB total), 5 volume dump(s)");nextcloud_nextcloud_html.tarcontains the sentinel; manifest enumerates all 3 nextcloud tars → sentinel deleted from the live volume → UI restore → sentinel back, sha256 identical (560cb2a7…). Restore log: 3 volumes restored + DB dump replayed + completed; nextcloud healthy. - Gates live: protected/infra stacks (felhom-controller, traefik, cloudflared, filebrowser) never stopped by the run (still "Up 29 hours" afterwards). Deviation from the task's expectation: the spec assumed rallly is volume-less — on this box every deployed app has named volumes (rallly's postgres uses
rallly_rallly_postgres_data), so all five were legitimately dumped and no volume-less app existed to observe live; that negative is covered byTestRunVolumeDumps_VolumelessNeverStopped(a no-seam variant driving the realDumpAppVolumesSafepath against a fake provider). - O4 spot-check (drill Phase-5 rerun):
DB_PASSWORDremoved from sparkyfitness app.yaml (backed up) → restore → loggenerated replacement for [DB_PASSWORD] — the credential was reset…,recovered=3/4, data_keys=2; the value persisted asENC:in app.yaml and appears in no log. The documented residual case then hit live: the (now-existing, F3-produced) postgres volume tar was restored carrying the OLD credential hash, the server auth'd with the NEW one → server unhealthy → compose dependency failure → the restore surfaced the error honestly (flash_error, F17 no-swallow) instead of claiming success. Healed by restoring the original app.yaml + endpoint restore →recovered=4/4, completed, all healthy.
Cleanup
Volume sentinel removed from the live volume; a final backup run refreshed all tars (sentinel confirmed absent from nextcloud_nextcloud_html.tar); no *.drillbak remains; sparkyfitness app.yaml back to original. Final state: all app containers healthy, /backups 200, no controller ERRORs. Guest 9201 only; Peti's box untouched.
Observations (not acted on — candidates for follow-up TASKs)
- O4 residual amplified by F3 (MEDIUM — product decision): for apps whose DB lives in a named volume, restoring the volume tar resurrects the OLD credential hash; combined with a generated NEW password the app fails auth (hit live, above). Options: skip restoring a DB-data volume when the SQL dump will be replayed anyway (the dump is authoritative per F17), or in-DB credential reset after replay.
internal/backup/restore_unit.go(volume restore vs dump replay ordering). - compose stderr truncation (LOW):
composeExeckeeps the FIRST ~500 chars, so the restoreflash_errorshowed only volume-label warnings, not the decisive "dependency failed: container … is unhealthy" tail. Consider keeping the LAST 500 chars.internal/stacks/manager.go. volume "X" already exists but was not created by Docker Composewarnings (LOW/cosmetic):restoreDockerVolumesrecreates volumes without compose labels; compose then warns on every subsequent up.- Demo box unauth/empty-CSRF state (pre-existing, still open): the dashboard accepts POSTs with an empty
_csrfand no session, andfelhom.demo-felhom.euis internet-reachable without auth — needs a separate check whether this is intentional demo setup or a gap (drill §2.2 flagged it; unchanged by this task). - Nightly downtime note (by design, worth customer-facing wording): volume-bearing apps now stop briefly during the nightly backup (stop-first is the locked torn-tar policy) — nextcloud was down ~16 s in the live run; the backups page already carries the "az alkalmazások rövid időre leállnak" hint for whole-guest backup, but the app-data run has no equivalent note.