chaos night Phase 2: the off-site restore cannot be done, and why - two instruments agree
gates / gates (push) Successful in 22s
gates / gates (push) Successful in 22s
restic with the box's own credentials: 'Fatal: wrong password or no key found', exit status 1. The product's own surface: orphaned true, snapshots 0, status error. The repository is orphaned because this box is a REBUILD - its restic password was minted fresh, so the existing snapshots cannot be opened. The product surfaced that honestly as a true alarm in round 1. So Phase 2's 'restore one DB-backed app from off-site' has nothing to restore from. That is a fact about the fixture, not a product failure. Recorded alongside: ep0 holds two intact whole-guest snapshots for this box, including tonight's 21:59:54Z copy - round 6's off-site leg, the one that ran by itself after I killed the local leg. Its file index is four times the size of the afternoon copy. So the data did leave the house. Stated as a limit, not glossed: that is a LISTING, not a verification. A PBS verify would prove restorability and writes state, so it was not run - ep0 is read-only for evidence tonight. Tenth instrument slip recorded: my first restic probe printed 'exit status 0' beneath a fatal error, because the zero belonged to the head at the end of the pipe. restic 0.14 has sftp.command, not sftp.args - established by asking 'restic options' rather than assuming. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+41
@@ -0,0 +1,41 @@
|
||||
# ep0: the whole-guest off-site copies for THIS box - contents, 2026-09-17T00:27:45Z, READ ONLY
|
||||
# Taken because "the data left the house" is a claim, and a claim needs a look inside.
|
||||
|
||||
group /mnt/pbs-datastore/ns/tester-1/ct/9201
|
||||
owner felhom@pbs!tester-1
|
||||
|
||||
2026-09-16T17:27:32Z catalog.pcat1.didx 4496
|
||||
client.log.blob 1196
|
||||
index.json.blob 688
|
||||
pct.conf.blob 401
|
||||
root.pxar.didx 61056
|
||||
|
||||
2026-09-16T21:59:54Z catalog.pcat1.didx 5216 <- TONIGHT's copy
|
||||
client.log.blob 1397
|
||||
index.json.blob 687
|
||||
pct.conf.blob 402
|
||||
root.pxar.didx 246616
|
||||
|
||||
Both snapshots carry a full set: a manifest (index.json.blob), a file index (root.pxar.didx), a
|
||||
catalog, the guest config and the client log. Neither is a stub or a half-written directory.
|
||||
|
||||
## WHY THE SECOND ONE MATTERS TONIGHT
|
||||
21:59:54Z is round 6's OFF-SITE leg - the one that started BY ITSELF after I killed the local leg
|
||||
that could never have fit. Its file index is 246616 bytes against 61056 for the afternoon copy, i.e.
|
||||
roughly four times the indexed content, which is consistent with a guest that had by then been
|
||||
seeded with twelve apps. So the whole-guest copy of tonight's box is on ep0, and the data really did
|
||||
leave the house on the night the local tier could not hold it.
|
||||
|
||||
## THE HONEST LIMIT - presence is still not success
|
||||
This is a LISTING, not a verification. It proves the files exist and are shaped like a real backup.
|
||||
It does not prove every chunk is readable. A PBS verify job WOULD prove that, and it is deliberately
|
||||
NOT run: it writes verify state into the datastore, and ep0 is read-only for evidence tonight.
|
||||
So the correct claim is: the off-site copy is PRESENT and well-formed, and its restorability has
|
||||
not been tested this session.
|
||||
|
||||
## AND THIS IS A DIFFERENT STORE FROM THE ORPHANED ONE
|
||||
These PBS snapshots (whole-guest vzdump to ep0) are not the same thing as the restic app-backup
|
||||
repository on the Storage Box, which is orphaned and holds zero readable snapshots. Tonight the box
|
||||
had a working off-site path for the WHOLE GUEST and a broken one for PER-APP restores, at the same
|
||||
time. Any sentence that says "off-site backup works" or "off-site backup is broken" without naming
|
||||
which of the two is meant would be wrong in one direction or the other.
|
||||
+42
@@ -0,0 +1,42 @@
|
||||
# PHASE 2: the off-site restore CANNOT be performed on this box - and two independent
|
||||
# instruments agree on why. 2026-09-17T00:26Z
|
||||
|
||||
## What the brief asked for
|
||||
"restore one DB-backed app from off-site onto scratch 9202."
|
||||
|
||||
## Instrument 1 - restic itself, using the box's own credentials, READ ONLY
|
||||
repository sftp:u629488-sub4@u629488-sub4.your-storagebox.de:/home/felhom-repo (port 23)
|
||||
credentials the box's own ssh_key, known_hosts and repo_password, exactly as the product uses them
|
||||
result Fatal: wrong password or no key found
|
||||
exit status 1 (restic's OWN status, captured to a file first - see the slip below)
|
||||
|
||||
## Instrument 2 - the product's own status surface, which is what a customer would see
|
||||
GET /backup/offbox/status ->
|
||||
{"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z",
|
||||
"orphaned":true, ... ,"repo_size_human":"","snapshots":0,"status":"error"}
|
||||
|
||||
The two agree: the repository is ORPHANED and holds ZERO snapshots this box can read.
|
||||
There is nothing to restore from, so the step cannot be performed. This is a fact about the
|
||||
fixture, not a failure of the product.
|
||||
|
||||
## WHY the repository is orphaned, and why that is CORRECT behaviour
|
||||
This box is a REBUILD for an existing customer. On a rebuild the restic repository password is
|
||||
minted fresh, so the snapshots already sitting in the remote store were written under a password
|
||||
that no longer exists. The product did not hide this: `offbox_repo_orphaned` fired as a true alarm
|
||||
in round 1 at 21:09 and was counted as true in the truth table. It is a known and documented shape.
|
||||
|
||||
## THE ONE THING IN THAT JSON WORTH A SECOND LOOK
|
||||
`"status":"error"` sits beside `"last_error":""`, with `last_run` set and a 1m45s duration.
|
||||
That is this project's "a timestamp records an ATTEMPT, not a RESULT" shape. A run that produced
|
||||
ZERO snapshots still recorded a last_run and an empty error string. Whether that misleads anyone
|
||||
depends entirely on whether a customer-facing surface renders last_run without consulting status -
|
||||
which is being checked separately. It is NOT filed as a defect on the JSON alone.
|
||||
|
||||
## MY TENTH INSTRUMENT SLIP, and it is the oldest trap in this project
|
||||
My first probe printed:
|
||||
Fatal: option sftp.args is not known
|
||||
--- exit status above: 0 ---
|
||||
The zero was the exit status of `head` at the end of my pipeline, not restic's. A reassuring 0
|
||||
printed directly beneath a fatal error. restic 0.14.0 has `sftp.command`, not `sftp.args` -
|
||||
confirmed by asking `restic options` rather than assuming. The re-run writes restic's output to a
|
||||
file and reads `$?` immediately, so the status belongs to the command it is printed next to.
|
||||
Reference in New Issue
Block a user