Night 2026-10-06: R-366 design + hub CHANGELOG (escrow kept on a backup-key change)
gates / gates (push) Successful in 2m35s
gates / gates (push) Successful in 2m35s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -26,3 +26,4 @@ no reboot.
|
||||
| R-822 | **design written** (`design-R-822.md`): pick — accept the residual and close (operator's word); option B filed as **R-895** (P2) | — | (this batch) |
|
||||
| R-895 | **OPENED** — the hub's clean-up-window check trusts the box's own snapshot counts (needs a design + a read-only measurement) | 5 | (this batch) |
|
||||
| (07 §6.5 area) | stale sentence fixed: „until that ships nothing prunes" — the window runs live since 2026-10-05 | 2 | (this batch) |
|
||||
| R-366 | **hub fix on main** (escrow retained when the backup key changes; red-proved) + **design** (`design-R-366.md`, pick B); the old wrong alarm was already gone (R-727) | 45 | (this batch) |
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
# R-366 — a reinstall orphans the whole-guest off-site archives — design proposal + first slice built (night 2026-10-06)
|
||||
|
||||
Baselines read: felhom.eu `8e2dc204` (hub v0.140.0), felhom-agent `74b5eae`. Architecture: `07-backup-architecture.md`
|
||||
§5 (encryption, „PBS … per-customer `encryption-key`") and §6.1 (off-site tier: server-side prune `keep-last 2`);
|
||||
`06-offsite-connectivity.md` §3.5 (key custody: the escrow wraps the PBS key K under the recovery code R).
|
||||
|
||||
## 1. The problem, and what is still true today
|
||||
- **Measured 2026-08-21** (hub event 3016): after demo-hp's reinstall the restore test failed on a pre-reinstall archive
|
||||
with `wrong key … manifest's key 3f:4f:65:c0… does not match provided key dd:d1:d8:53…`. Read as „a restore test
|
||||
failed", not as „every whole-guest archive from before the reinstall is unreadable here".
|
||||
- **The wrong verdict is already gone.** Agent v0.138.0 (R-727, decision 51) skips any archive written with another key
|
||||
before it picks a restore-test candidate (`felhom-agent/internal/backup/runner.go:360-374`). The skip is one INFO
|
||||
line per archive (`runner.go:684-698`). So the hub no longer sees a false failure — and now sees **nothing**.
|
||||
- **The August archives no longer exist (inferred, not measured — ep0 is fenced tonight).** The off-site tier is pruned
|
||||
on ep0 with `keep-last 2` per group (`07` §6.1 table); a reinstalled box writes the same group `ct/9201` in the same
|
||||
namespace (measured for R-727 on 2026-09-30), and demo-hp has written weekly copies since 2026-08-21. So the old
|
||||
archives were pruned in early September. **Question 1 of the row („are they recoverable?") is moot for demo-hp.**
|
||||
- **The real remaining gap, found tonight (read in source, then pinned by a test):** the hub keeps an old escrow only
|
||||
when its *restic* password differs (`hub/internal/store/store.go` `SaveHostEscrow`, the rule before this fix). A
|
||||
reinstall mints a **new** K (`felhom-agent/configs/felhom-pbs-apply:99`, `--encryption-key autogen`). Since R-241 a
|
||||
rebuilt box keeps its restic password (memory `retained-key-is-operator-only`; not re-measured tonight), so the escrow
|
||||
PUT carries the same restic sha with a new key fingerprint — and the row holding the **old K was overwritten**. That
|
||||
made every pre-reinstall archive unopenable **for good**, not only for the new box. Red test observed:
|
||||
`r366/red-key-change.txt` („a new backup key with the same restic password must supersede (retain the old K)").
|
||||
|
||||
## 2. What the code does today (after the slice below)
|
||||
- Agent: foreign-key archives skipped from restore-test candidacy, logged by name (`runner.go:360-374, 684-698`).
|
||||
- Hub, on branch `night-r366` (`70b07fdf`, not on main): `SaveHostEscrow` retains the current row when the restic sha
|
||||
**or** the key fingerprint changes (`backupKeyChanged`: case and space ignored; an empty side is unknown, never a
|
||||
change). The audit event `escrow_superseded` and the log line now name both causes. The repo-key-changed alarm
|
||||
(`maybeEmitRepoKeyChanged`) still fires only on a restic change — a K-only change does not raise it.
|
||||
- Nobody is told that „the previous install's whole-guest archives exist and this box cannot open them" in the ~14 days
|
||||
before ep0's prune removes them.
|
||||
|
||||
## 3. Options
|
||||
**A. Retain the old K (built tonight, slice 1) and stop there.** One condition and two tests in the hub. Cost: one more
|
||||
retained row per reinstall (opaque, R-wrapped — the hub learns nothing). Wrong case: a producer that sends a different
|
||||
fingerprint *format* for the same key would add a retained row per ceremony; case/space are normalised, and today one
|
||||
producer (the agent's escrow PUT) sends it. Restoring from the old K stays operator-only (R-304).
|
||||
|
||||
**B. A + tell the operator.** The agent counts foreign-key archives per tier in the host report (a new additive field);
|
||||
the hub shows one line on the host page and raises one `info` operator event per new key: „N whole-guest archives on
|
||||
felhom-pbs were written with key X by a previous install; this box cannot open them; the hub retains key X: yes/no;
|
||||
ep0 prunes them after two new copies." Cost: agent + hub, an additive report field (wire contract gate), one event
|
||||
type (allow-list). Wrong case: noise on a returning customer every reinstall — once per key, so bounded.
|
||||
|
||||
**C. Key continuity: a reinstall re-uses the escrowed K instead of `autogen`.** The new box would read its own history.
|
||||
Cost: the reinstall then needs R at install time (the customer's code) — the DR consume path, made a normal path. Changes
|
||||
a promise (what a reinstall does with old backups) and custody → the operator's.
|
||||
|
||||
## 4. The pick — PROPOSAL
|
||||
**B, built in two slices; slice 1 (A) is done and waits on its branch.** A closes the only *irreversible* part (the old
|
||||
key destroyed). B removes the silence without changing any promise. C is a product decision and stays a question.
|
||||
|
||||
## 5. First slice and its proof
|
||||
- **Built:** hub `70b07fdf` on `night-r366` — `TestSaveHostEscrow_R366_NewBackupKeySameResticPasswordRetainsOldKey`
|
||||
(red observed on the old rule) and `TestSaveHostEscrow_R366_SameKeyOrUnknownFingerprintDoesNotSupersede`
|
||||
(idempotence: same key in another case, or an empty fingerprint, retains nothing). Full hub `go test ./...` green.
|
||||
felhom.eu `repo_gates.py --fast` could not judge the worktree (sibling clones absent next to it — „controller clone
|
||||
not found at /mnt/5_hdd/felhom.eu/wt/…"); it must run on `main` after the cherry-pick.
|
||||
- **Live proof (next day, not tonight):** on scratch box 9202 only — re-run the escrow PUT with the same restic sha and
|
||||
a new fingerprint through the agent's ceremony (or a reinstall of 9202), then read the host page's „N superseded
|
||||
escrow blobs retained" (positive observable) and, as the control from another channel, the hub log line
|
||||
`superseded an escrow with a different passphrase or backup key (R-366)`.
|
||||
- Slice 2 (B): red test first — a host report with one foreign-key archive on felhom-pbs must produce exactly one
|
||||
operator event naming the key, and a second identical report none.
|
||||
|
||||
## 6. Questions for the operator
|
||||
1. **Deploy slice 1 (the hub keeps the old backup key after a reinstall)?** My pick: yes, in tonight's or tomorrow's hub
|
||||
release. If you do nothing: the next reinstall of a box destroys the only key to its earlier whole-guest backups.
|
||||
2. **Should a reinstall keep reading its old whole-guest backups (option C — it would need the household's recovery code
|
||||
during the reinstall)?** If you do nothing: a reinstalled box starts a new off-site history; the old one is opened
|
||||
only by the operator, by hand, and only until ep0 prunes it (about two weeks).
|
||||
@@ -0,0 +1,7 @@
|
||||
## GREEN after fix, branch night-r366 70b07fdf
|
||||
--- PASS: TestSaveHostEscrow_RetainsIdentityBlob (0.04s)
|
||||
--- PASS: TestSaveHostEscrow_SupersedesWithoutIdentityBlob (0.03s)
|
||||
--- PASS: TestSaveHostEscrow_RetainsSuperseded (0.03s)
|
||||
--- PASS: TestSaveHostEscrow_R366_NewBackupKeySameResticPasswordRetainsOldKey (0.03s)
|
||||
--- PASS: TestSaveHostEscrow_R366_SameKeyOrUnknownFingerprintDoesNotSupersede (0.03s)
|
||||
ok gitea.dooplex.hu/admin/felhom-hub/internal/store 0.169s
|
||||
@@ -0,0 +1,6 @@
|
||||
## RED: pre-fix rule (restic sha only), hub 8e2dc204
|
||||
--- FAIL: TestSaveHostEscrow_R366_NewBackupKeySameResticPasswordRetainsOldKey (0.03s)
|
||||
r366_escrow_key_change_test.go:26: a new backup key with the same restic password must supersede (retain the old K)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/store 0.068s
|
||||
FAIL
|
||||
@@ -154,7 +154,7 @@ stopping line that lies.
|
||||
| **R-105** | Backup & restore | P2 | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-06 night: TRACED in source — not a fault in a running path.** `hosts.dr_record_json` has no writer and no reader; `host_escrow.directive_json` is filled only by the by-hand `-directive` flag (the wizard's fixed argv carries none, so every wizard escrow stores `{}`), and its only routes (`/re-enroll`, `/restore-directive`) have no client. The built DR path reads the recipe, tenantsync and the escrow blob. Design with a pick (retire both and correct `05` §9/§11, `06` §3.5 — needs the operator's word): `audits/night-burndown-2026-10-06/design-R-105.md`. The `drives` third was not re-measured (the hub-DB read was refused). | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC |
|
||||
| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator |
|
||||
| **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC |
|
||||
| **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
|
||||
| **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** **2026-10-06 night: the wrong verdict was already gone** (agent v0.138.0, R-727 skips another key's archives); demo-hp's August archives are pruned by ep0's keep-last 2 (inferred, ep0 fenced). **NEW, the real gap:** the hub retained an escrow only on a restic-password change, so a reinstall's new backup key K overwrote the only copy of the old one — **fixed on felhom.eu main (hub, unreleased; red-proved `audits/night-burndown-2026-10-06/r366/`)**. Design `audits/night-burndown-2026-10-06/design-R-366.md` (pick B: retain + tell the operator; C, key continuity, is the operator's). Not closed: the operator signal (archives made with another key) is the design's second slice. | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC |
|
||||
| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. | — | Read back the 2026-10-07 night (local tier) and the ~2026-10-08 night (both tiers on demo-hp); measure one press (Part D) | CC |
|
||||
| **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. | **OPEN — filed 2026-10-06** | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC |
|
||||
| **R-894** | Backup & restore | P3 | **After an agent restart, an UNREADABLE off-site storage makes the off-site tier look DUE, so the box asks for a copy that cannot be made.** MEASURED 2026-10-05 on demo-hp, read 2026-10-06 (`audits/readback-2026-10-07/C/`): the last off-site copy was 2026-10-01 20:15Z (ep0's own listing, verify `ok`), so the 7-day tier was NOT due; the agent had restarted at 04:57 local; at 06:25 `GET …/storage/felhom-pbs/content` answered 500 *Can't connect to 10.77.0.1:8007*; `newestArchiveOn` returned `unknown` and fell back to the in-memory record (`internal/localapi/server.go`), which a restart empties (`internal/backup/store.go` — memory only, R-348); so the tier read DUE, the controller requested it, and vzdump failed (*could not activate storage 'felhom-pbs'*). By design the controller stops the apps before it asks (`07` §6.4); whether it did that night is NOT KNOWN (the controller's log was lost to a later restart). The hub got `whole_guest_backup_failed` (error); its operator mail then failed (fixed in the hub this session: a failed operator mail is retried). The code's rule is deliberate: an unreadable storage must not suppress a backup. | **OPEN — filed 2026-10-06** **2026-10-06 night: fixed on agent main `74b5eae` (the row's first option — the newest success per tier on disk, read only when the storage cannot be read; `07` §6.1); ships with agent v0.150.0, closes when delivered.** | a design: keep the newest success per tier on disk (as `RestoreTestState` does) so the fallback is the last known copy, not "never"; or report an unreachable storage so no app is stopped for it | Decide the fallback; build it in the agent with a test that restarts the agent and then cannot read the storage; measure once | CC |
|
||||
|
||||
@@ -1,3 +1,7 @@
|
||||
## Unreleased (2026-10-06 night) — a reinstall's new backup key no longer overwrites the old one in the escrow (R-366)
|
||||
|
||||
- hub: the current escrow row is kept (retained, operator-only) when the whole-guest backup key's fingerprint changes, not only when the restic password changes. A reinstall mints a new PBS key while R-241 keeps the restic password, so the old rule overwrote the only copy of the old key and every pre-reinstall whole-guest archive became unopenable. An empty fingerprint on either side is unknown, not a change. Tests `TestSaveHostEscrow_R366_*` (red-proved: `documentation/audits/night-burndown-2026-10-06/r366/`).
|
||||
|
||||
## v0.140.0 — a failed operator mail is sent again; a Docker set is approved only after the engine showed it reports a memory kill (decision 157) (2026-10-06)
|
||||
|
||||
**Operator action on deploy: none.** Until agent v0.150.0 reaches the ring-0 boxes, no Docker set can be approved (none is pending: 29.8.2 is approved since 2026-10-04).
|
||||
|
||||
Reference in New Issue
Block a user