REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
gates / gates (push) Successful in 18s

Opens with Part 1's answer because everything reads differently after it: a customer's own account can
reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY.

The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather
than cited - the control write to the account home succeeded and was cleaned up, the write into
/.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes.

storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence -
which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's
remedy can ever be product-driven.

STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two
minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today -
the live firing was required to prove delivery and nothing was deleted.

Five of my own mistakes are named, including the one that matters most: my first escalation-only test
was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved
and the latch was never consulted. That is why red-proofs are run.
This commit is contained in:
2026-09-01 14:33:16 +02:00
parent 65c82c4aa0
commit 17d92e71a1
3 changed files with 266 additions and 15 deletions
+21 -15
View File
@@ -18,7 +18,7 @@ not an evening's work.**
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **One thing is waiting on you — item 4, and it takes ten minutes.** Otherwise: Both problems the overnight test found are fixed and proven on
1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on
the real machines:
- the background job that could delete a live restore's lock now waits its turn — and the check
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
@@ -37,24 +37,30 @@ nothing.*
two register lines in the hub (already live). No customer action, no data migration, no
credential change.
4. **The copy that holds the customers' documents and photos can still be deleted by the box that
made it** (R-95 — first on the list since July, and this is the first time it has reached this
page). I studied it today and did not change anything. **One thing is yours and it takes ten
minutes.** Our notes say a safety net is switched on: the storage provider takes a snapshot every
night and keeps seven. **I could not find a single one.** I looked from both machines, using the
credentials they already have, and checked that my method could see other things and could
correctly fail to see a made-up name. Either the snapshots are not being taken, or they are
invisible to the machines — and if they are invisible, they are also useless to them: getting one
back would be you, in the provider's control panel. **I did not log in to check, because your own
notes say that question is yours.** **If you do nothing:** the register keeps saying the net is
armed, and nobody knows whether it is. **Please open the Storage Box panel and tell me whether
snapshots exist.** The answer changes which fix is worth building.
4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on
the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in
the control panel: seven of them, one a day. **And they are better than we thought** — I tried to
write into the snapshot area from a customer's machine and the storage **refused**, while the same
write to its normal folder worked. So a machine that wipes its own backup **cannot touch the
snapshots of it**. The worst case is losing about a day, then copying the rest back file by file.
**That is much smaller than what the notes have said since July.** I have corrected the notes.
**What I still need from you:** the machines can see the snapshot *door* but not what is inside —
only the main account can. So getting data back is you, in a browser, for now. **If you read one
snapshot's name off the panel and send it to me, one command settles whether the machines can
reach them directly** — and if they can, recovery becomes something the product does by itself.
**If you do nothing:** it stays a manual job for you, which is workable but slow.
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a
test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it
once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`.
**No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one
mail.
6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
as it misled one by an hour.