REPORT + CONTEXT + STATUS: the record corrected, and the thing worth building shipped
gates / gates (push) Successful in 18s
gates / gates (push) Successful in 18s
Opens with Part 1's answer because everything reads differently after it: a customer's own account can reach the snapshot DOOR and is REFUSED writes to it, but sees the tree EMPTY. The write-refusal is the load-bearing sentence of the whole R-95 re-scope and it is now PROVEN rather than cited - the control write to the account home succeeded and was cleaned up, the write into /.zfs/snapshot returned `dest open ...: Failure`, and nothing was left behind. Identical on both boxes. storage-box-pool-1 IS u629488, so the emptiness is per-sub-account filtering rather than absence - which means recovery is an operator act in a browser today (R-432), and that decides whether R-95's remedy can ever be product-driven. STATUS carries two items for Viktor in plain words: read one snapshot name off the panel (two minutes, and it may make recovery product-reachable), and IGNORE the alarm mail he received today - the live firing was required to prove delivery and nothing was deleted. Five of my own mistakes are named, including the one that matters most: my first escalation-only test was HOLLOW and its red-proof PASSED. It re-swept the same report, so the baseline had already moved and the latch was never consulted. That is why red-proofs are run.
This commit is contained in:
@@ -18,7 +18,7 @@ not an evening's work.**
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **One thing is waiting on you — item 4, and it takes ten minutes.** Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
1. **Two things are waiting on you — item 4 (a snapshot name, two minutes) and item 5 (ignore one alarm mail).** Otherwise: Both problems the overnight test found are fixed and proven on
|
||||
the real machines:
|
||||
- the background job that could delete a live restore's lock now waits its turn — and the check
|
||||
that finds the next one like it is a test, not a comment, so it cannot come back quietly;
|
||||
@@ -37,24 +37,30 @@ nothing.*
|
||||
two register lines in the hub (already live). No customer action, no data migration, no
|
||||
credential change.
|
||||
|
||||
4. **The copy that holds the customers' documents and photos can still be deleted by the box that
|
||||
made it** (R-95 — first on the list since July, and this is the first time it has reached this
|
||||
page). I studied it today and did not change anything. **One thing is yours and it takes ten
|
||||
minutes.** Our notes say a safety net is switched on: the storage provider takes a snapshot every
|
||||
night and keeps seven. **I could not find a single one.** I looked from both machines, using the
|
||||
credentials they already have, and checked that my method could see other things and could
|
||||
correctly fail to see a made-up name. Either the snapshots are not being taken, or they are
|
||||
invisible to the machines — and if they are invisible, they are also useless to them: getting one
|
||||
back would be you, in the provider's control panel. **I did not log in to check, because your own
|
||||
notes say that question is yours.** **If you do nothing:** the register keeps saying the net is
|
||||
armed, and nobody knows whether it is. **Please open the Storage Box panel and tell me whether
|
||||
snapshots exist.** The answer changes which fix is worth building.
|
||||
4. **Good news, and I was wrong yesterday.** I told you I could not find the nightly safety net on
|
||||
the storage. **I was looking for the wrong name.** The snapshots are there, exactly as you saw in
|
||||
the control panel: seven of them, one a day. **And they are better than we thought** — I tried to
|
||||
write into the snapshot area from a customer's machine and the storage **refused**, while the same
|
||||
write to its normal folder worked. So a machine that wipes its own backup **cannot touch the
|
||||
snapshots of it**. The worst case is losing about a day, then copying the rest back file by file.
|
||||
**That is much smaller than what the notes have said since July.** I have corrected the notes.
|
||||
**What I still need from you:** the machines can see the snapshot *door* but not what is inside —
|
||||
only the main account can. So getting data back is you, in a browser, for now. **If you read one
|
||||
snapshot's name off the panel and send it to me, one command settles whether the machines can
|
||||
reach them directly** — and if they can, recovery becomes something the product does by itself.
|
||||
**If you do nothing:** it stays a manual job for you, which is workable but slow.
|
||||
|
||||
5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
5. **You will have received an alarm email from me today about `demo-hp` losing 65 backups. It is a
|
||||
test and nothing is wrong.** I built the new "someone deleted the backups" alarm and had to fire it
|
||||
once for real to prove it reaches you. Subject: `[Felhom] 🔴 demo-hp: offsite_snapshots_dropped`.
|
||||
**No backups were deleted.** **If you do nothing:** nothing — but please do not act on that one
|
||||
mail.
|
||||
|
||||
6. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
|
||||
|
||||
6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
7. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is
|
||||
wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session,
|
||||
as it misled one by an hour.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user