RECON: trace the offsite DR chain link by link — it does not join up (R-198..R-201)
gates / gates (push) Successful in 7s

Read-only recon of the escrow -> recovery chain, from a dead node to an open
repository. No production code, no build, no version bump.

Headline: the hub's superseded-escrow retention does NOT retain the offsite
repository password. host_escrow_superseded has no identity_blob column and
demoteCurrentEscrowTx copies only the K-escrow blob, so what survives a
supersession is the PBS datastore key, not the restic repo password. The next
escrow ceremony -- which the system tells a rebuilt box's customer to run --
destroys the last copy. Both demo boxes crossed that line on 2026-08-04.

Also established:
- the hub's blob-serving endpoints (re-enroll / restore-directive) have zero
  callers anywhere: agent, hub UI, scripts, runbooks (R-199)
- POST /backup/offbox/inject-password is routed and handled but no template
  contains the form (R-200)
- nothing in the recovery path has ever been exercised; the one live
  round-trip proof (2026-06-10) predates the ResticRepoPassword field (R-201)
- a fail-closed mint refusal IS implementable: the report ACK already carries
  escrow{identity_blob_present, restic_pw_sha256} and the controller discards
  it whenever no offbox target exists

Corrections: yesterday's spike annotated (candidate (b) overturned in part --
unattended recovery is impossible, customer-present is not); capability-map
retention claim struck through and replaced with what the code does.

Deliverable: documentation/audits/RECON-offsite-dr-chain-2026-08-04.md
Register: new R-198..R-201; R-193 and R-192 updated; STATUS.md refreshed.
This commit is contained in:
2026-08-04 12:16:04 +02:00
parent d26f49ad68
commit 3f2b7bc023
6 changed files with 726 additions and 310 deletions
+50
View File
@@ -35,6 +35,25 @@ Proven end to end on real hardware.
**The next scheduled off-site run is tomorrow at 04:15, and it will refuse rather than quietly start
a new history** — the machine already knows how to recognise this and asks before resetting.
**There is a decision here for you — see "Waiting on you".** *(R-193, R-192, and two new items)*
- **And the safety net we believed was under all of this is not there.** We traced the whole path today —
from a dead machine to reopened backups — and it does not join up. The important part is not that
several steps are manual; it is this: **the central hub keeps the old sealed key when a machine
re-seals, and we found it keeps the wrong one.** It keeps the key for the whole-machine backups and
**not** the key for the off-site file backups — the one this entire problem is about. So the moment
a rebuilt machine is asked to make a new recovery code, which is exactly what it asks the customer to
do, **the last copy of the old key is gone.** Both demo machines crossed that line yesterday morning.
**This means the 51 orphaned backups were beyond reach even if you had kept the recovery codes**
which is a relief in one narrow sense and much worse in every other. It also means the customer is
currently told, in Hungarian on their own screen, that their old backups "may later be restorable
with the matching recovery code". That sentence is not true today. It is a small fix — one missing
column — and it must land before anything else here. *(R-198)*
- **Nothing in the recovery path has ever been performed.** Making a recovery code is proven, live,
with a real customer. **Using one is not.** No sealed package has ever been handed back to a machine
(the hub has the buttons for it; nothing anywhere presses them), no recovered key has ever been put
back, no old backup store has ever been reopened, and no file has ever been restored from a recovered
key. There is a form missing, a step missing, and a connection missing between two parts of the
system. We designed the exercise that would prove it end to end — see "Waiting on you".
*(R-199, R-200, R-201)*
- *(fixed 4 Aug)* ~~The weekly off-site backup reports FAILED although it worked.~~ It uploaded fine
and then tripped on a tidy-up step it is deliberately not allowed to perform. The machine no longer
asks — tidying up is the endpoint's job, and **that was checked first**: the endpoint has been doing
@@ -106,6 +125,32 @@ Proven end to end on real hardware.
nothing else, and it is the piece whose absence let a gigabyte of history go quiet for thirteen
hours. What we are asking you to choose is (a). *(R-193; the full reasoning is in the findings
document)*
**Correction to (b), from today:** "old backups are kept" is true, but "openable later with the
recovery code" is **not** — see R-198 above. Option (b) is only honest once that is fixed.
- **Your ruling about the recovery screen has been priced, and it can be built.** You said: a freshly
installed machine that finds a sealed package waiting should say so loudly, offer the customer a box
to type their recovery code into, and show what would come back before doing anything. **All of it is
buildable, and one part is already free** — the hub is *already* telling every machine, on every
check-in, that a sealed package exists and which key it covers, and the machine currently throws that
message away. Showing a preview is also cheap: listing what is in an off-site store reads it without
writing to it, so the customer can see how many backups, from when, and for which apps before
committing. The real work is one new connection: the unsealing has to happen in the part that runs on
the Proxmox host, because the customer-facing part cannot do it — the same crossing your recovery-code
ceremony already makes in the other direction. **One thing for you to weigh, which we deliberately did
not decide:** a screen that takes a recovery code and then shows what is in a backup store is reachable
by anyone with the household's dashboard password, and the preview reveals backup dates and app names.
*(R-193)*
- **Should we run the proof?** Nothing about recovery has ever actually been done, so we designed the
exercise: on the spare demo machine, make a recovery code and **keep it**, put a marked file into an
app, back it up off-site, wipe the machine, reinstall, recover with the saved code, and check the
marked file comes back byte for byte. About half a day, of which under an hour needs you. The failure
we are watching for is precise: **if the recovered machine reports one backup instead of the ones we
put there, it started a fresh history and the proof failed** — "the store opened" is not good enough.
It costs that machine's contents if it goes wrong, which is acceptable there. **Do R-198 first**, or
the exercise walks a path that is missing a piece. *(R-201)*
- **The orphaned backups still on the storage box.** About 1.2 GB across the two demo machines, sitting
in set-aside stores nobody can open and nothing prunes, using the 50 GB allowance. We did not touch
them. **Delete, or leave?** *(R-193)*
- **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a
session log; nothing suggests anyone else saw it. *(R-132)*
- **One small question, not urgent.** The automatic check cannot see which version you have told
@@ -115,6 +160,11 @@ Proven end to end on real hardware.
## Changed since last update
- **2026-08-04 (latest)** — Traced the whole recovery path end to end and it does not join up. The hub
keeps the wrong sealed key when a machine re-seals, so a rebuilt machine's old off-site key is lost the
moment the customer is asked for a new recovery code. Nothing in the recovery path has ever been
performed. Your ruling about the recovery screen is priced and buildable. The proof exercise is
designed and awaiting your go-ahead. *(R-198, R-199, R-200, R-201, R-193)*
- **2026-08-04 (later)** — Established what a machine rebuild actually destroys: not the storage
password we had been re-issuing, but the key that encrypts the backups. Both demo machines lost
access to their off-site history, one of them silently. A decision is now waiting on you. The daily