chaos night Phase 2: the off-site restore cannot be done, and why - two instruments agree
gates / gates (push) Successful in 22s

restic with the box's own credentials: 'Fatal: wrong password or no key found',
exit status 1. The product's own surface: orphaned true, snapshots 0, status
error. The repository is orphaned because this box is a REBUILD - its restic
password was minted fresh, so the existing snapshots cannot be opened. The
product surfaced that honestly as a true alarm in round 1.

So Phase 2's 'restore one DB-backed app from off-site' has nothing to restore
from. That is a fact about the fixture, not a product failure.

Recorded alongside: ep0 holds two intact whole-guest snapshots for this box,
including tonight's 21:59:54Z copy - round 6's off-site leg, the one that ran
by itself after I killed the local leg. Its file index is four times the size
of the afternoon copy. So the data did leave the house.

Stated as a limit, not glossed: that is a LISTING, not a verification. A PBS
verify would prove restorability and writes state, so it was not run - ep0 is
read-only for evidence tonight.

Tenth instrument slip recorded: my first restic probe printed 'exit status 0'
beneath a fatal error, because the zero belonged to the head at the end of the
pipe. restic 0.14 has sftp.command, not sftp.args - established by asking
'restic options' rather than assuming.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-17 02:28:52 +02:00
parent be99cf7c74
commit 7c8a299cd1
2 changed files with 83 additions and 0 deletions
@@ -0,0 +1,41 @@
# ep0: the whole-guest off-site copies for THIS box - contents, 2026-09-17T00:27:45Z, READ ONLY
# Taken because "the data left the house" is a claim, and a claim needs a look inside.
group /mnt/pbs-datastore/ns/tester-1/ct/9201
owner felhom@pbs!tester-1
2026-09-16T17:27:32Z catalog.pcat1.didx 4496
client.log.blob 1196
index.json.blob 688
pct.conf.blob 401
root.pxar.didx 61056
2026-09-16T21:59:54Z catalog.pcat1.didx 5216 <- TONIGHT's copy
client.log.blob 1397
index.json.blob 687
pct.conf.blob 402
root.pxar.didx 246616
Both snapshots carry a full set: a manifest (index.json.blob), a file index (root.pxar.didx), a
catalog, the guest config and the client log. Neither is a stub or a half-written directory.
## WHY THE SECOND ONE MATTERS TONIGHT
21:59:54Z is round 6's OFF-SITE leg - the one that started BY ITSELF after I killed the local leg
that could never have fit. Its file index is 246616 bytes against 61056 for the afternoon copy, i.e.
roughly four times the indexed content, which is consistent with a guest that had by then been
seeded with twelve apps. So the whole-guest copy of tonight's box is on ep0, and the data really did
leave the house on the night the local tier could not hold it.
## THE HONEST LIMIT - presence is still not success
This is a LISTING, not a verification. It proves the files exist and are shaped like a real backup.
It does not prove every chunk is readable. A PBS verify job WOULD prove that, and it is deliberately
NOT run: it writes verify state into the datastore, and ep0 is read-only for evidence tonight.
So the correct claim is: the off-site copy is PRESENT and well-formed, and its restorability has
not been tested this session.
## AND THIS IS A DIFFERENT STORE FROM THE ORPHANED ONE
These PBS snapshots (whole-guest vzdump to ep0) are not the same thing as the restic app-backup
repository on the Storage Box, which is orphaned and holds zero readable snapshots. Tonight the box
had a working off-site path for the WHOLE GUEST and a broken one for PER-APP restores, at the same
time. Any sentence that says "off-site backup works" or "off-site backup is broken" without naming
which of the two is meant would be wrong in one direction or the other.
@@ -0,0 +1,42 @@
# PHASE 2: the off-site restore CANNOT be performed on this box - and two independent
# instruments agree on why. 2026-09-17T00:26Z
## What the brief asked for
"restore one DB-backed app from off-site onto scratch 9202."
## Instrument 1 - restic itself, using the box's own credentials, READ ONLY
repository sftp:u629488-sub4@u629488-sub4.your-storagebox.de:/home/felhom-repo (port 23)
credentials the box's own ssh_key, known_hosts and repo_password, exactly as the product uses them
result Fatal: wrong password or no key found
exit status 1 (restic's OWN status, captured to a file first - see the slip below)
## Instrument 2 - the product's own status surface, which is what a customer would see
GET /backup/offbox/status ->
{"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z",
"orphaned":true, ... ,"repo_size_human":"","snapshots":0,"status":"error"}
The two agree: the repository is ORPHANED and holds ZERO snapshots this box can read.
There is nothing to restore from, so the step cannot be performed. This is a fact about the
fixture, not a failure of the product.
## WHY the repository is orphaned, and why that is CORRECT behaviour
This box is a REBUILD for an existing customer. On a rebuild the restic repository password is
minted fresh, so the snapshots already sitting in the remote store were written under a password
that no longer exists. The product did not hide this: `offbox_repo_orphaned` fired as a true alarm
in round 1 at 21:09 and was counted as true in the truth table. It is a known and documented shape.
## THE ONE THING IN THAT JSON WORTH A SECOND LOOK
`"status":"error"` sits beside `"last_error":""`, with `last_run` set and a 1m45s duration.
That is this project's "a timestamp records an ATTEMPT, not a RESULT" shape. A run that produced
ZERO snapshots still recorded a last_run and an empty error string. Whether that misleads anyone
depends entirely on whether a customer-facing surface renders last_run without consulting status -
which is being checked separately. It is NOT filed as a defect on the JSON alone.
## MY TENTH INSTRUMENT SLIP, and it is the oldest trap in this project
My first probe printed:
Fatal: option sftp.args is not known
--- exit status above: 0 ---
The zero was the exit status of `head` at the end of my pipeline, not restic's. A reassuring 0
printed directly beneath a fatal error. restic 0.14.0 has `sftp.command`, not `sftp.args` -
confirmed by asking `restic options` rather than assuming. The re-run writes restic's output to a
file and reads `$?` immediately, so the status belongs to the command it is printed next to.