diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/ep0-snapshots-verified-present.txt b/documentation/audits/evidence-chaos-night-2026-09-17/ep0-snapshots-verified-present.txt new file mode 100644 index 00000000..923c48ff --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/ep0-snapshots-verified-present.txt @@ -0,0 +1,41 @@ +# ep0: the whole-guest off-site copies for THIS box - contents, 2026-09-17T00:27:45Z, READ ONLY +# Taken because "the data left the house" is a claim, and a claim needs a look inside. + +group /mnt/pbs-datastore/ns/tester-1/ct/9201 +owner felhom@pbs!tester-1 + + 2026-09-16T17:27:32Z catalog.pcat1.didx 4496 + client.log.blob 1196 + index.json.blob 688 + pct.conf.blob 401 + root.pxar.didx 61056 + + 2026-09-16T21:59:54Z catalog.pcat1.didx 5216 <- TONIGHT's copy + client.log.blob 1397 + index.json.blob 687 + pct.conf.blob 402 + root.pxar.didx 246616 + +Both snapshots carry a full set: a manifest (index.json.blob), a file index (root.pxar.didx), a +catalog, the guest config and the client log. Neither is a stub or a half-written directory. + +## WHY THE SECOND ONE MATTERS TONIGHT +21:59:54Z is round 6's OFF-SITE leg - the one that started BY ITSELF after I killed the local leg +that could never have fit. Its file index is 246616 bytes against 61056 for the afternoon copy, i.e. +roughly four times the indexed content, which is consistent with a guest that had by then been +seeded with twelve apps. So the whole-guest copy of tonight's box is on ep0, and the data really did +leave the house on the night the local tier could not hold it. + +## THE HONEST LIMIT - presence is still not success +This is a LISTING, not a verification. It proves the files exist and are shaped like a real backup. +It does not prove every chunk is readable. A PBS verify job WOULD prove that, and it is deliberately +NOT run: it writes verify state into the datastore, and ep0 is read-only for evidence tonight. +So the correct claim is: the off-site copy is PRESENT and well-formed, and its restorability has +not been tested this session. + +## AND THIS IS A DIFFERENT STORE FROM THE ORPHANED ONE +These PBS snapshots (whole-guest vzdump to ep0) are not the same thing as the restic app-backup +repository on the Storage Box, which is orphaned and holds zero readable snapshots. Tonight the box +had a working off-site path for the WHOLE GUEST and a broken one for PER-APP restores, at the same +time. Any sentence that says "off-site backup works" or "off-site backup is broken" without naming +which of the two is meant would be wrong in one direction or the other. diff --git a/documentation/audits/evidence-chaos-night-2026-09-17/phase2-offsite-restore-blocked.txt b/documentation/audits/evidence-chaos-night-2026-09-17/phase2-offsite-restore-blocked.txt new file mode 100644 index 00000000..2e600a53 --- /dev/null +++ b/documentation/audits/evidence-chaos-night-2026-09-17/phase2-offsite-restore-blocked.txt @@ -0,0 +1,42 @@ +# PHASE 2: the off-site restore CANNOT be performed on this box - and two independent +# instruments agree on why. 2026-09-17T00:26Z + +## What the brief asked for +"restore one DB-backed app from off-site onto scratch 9202." + +## Instrument 1 - restic itself, using the box's own credentials, READ ONLY + repository sftp:u629488-sub4@u629488-sub4.your-storagebox.de:/home/felhom-repo (port 23) + credentials the box's own ssh_key, known_hosts and repo_password, exactly as the product uses them + result Fatal: wrong password or no key found + exit status 1 (restic's OWN status, captured to a file first - see the slip below) + +## Instrument 2 - the product's own status surface, which is what a customer would see + GET /backup/offbox/status -> + {"last_duration":"1m45s","last_error":"","last_run":"2026-09-16T21:09:00Z", + "orphaned":true, ... ,"repo_size_human":"","snapshots":0,"status":"error"} + +The two agree: the repository is ORPHANED and holds ZERO snapshots this box can read. +There is nothing to restore from, so the step cannot be performed. This is a fact about the +fixture, not a failure of the product. + +## WHY the repository is orphaned, and why that is CORRECT behaviour +This box is a REBUILD for an existing customer. On a rebuild the restic repository password is +minted fresh, so the snapshots already sitting in the remote store were written under a password +that no longer exists. The product did not hide this: `offbox_repo_orphaned` fired as a true alarm +in round 1 at 21:09 and was counted as true in the truth table. It is a known and documented shape. + +## THE ONE THING IN THAT JSON WORTH A SECOND LOOK +`"status":"error"` sits beside `"last_error":""`, with `last_run` set and a 1m45s duration. +That is this project's "a timestamp records an ATTEMPT, not a RESULT" shape. A run that produced +ZERO snapshots still recorded a last_run and an empty error string. Whether that misleads anyone +depends entirely on whether a customer-facing surface renders last_run without consulting status - +which is being checked separately. It is NOT filed as a defect on the JSON alone. + +## MY TENTH INSTRUMENT SLIP, and it is the oldest trap in this project +My first probe printed: + Fatal: option sftp.args is not known + --- exit status above: 0 --- +The zero was the exit status of `head` at the end of my pipeline, not restic's. A reassuring 0 +printed directly beneath a fatal error. restic 0.14.0 has `sftp.command`, not `sftp.args` - +confirmed by asking `restic options` rather than assuming. The re-run writes restic's output to a +file and reads `$?` immediately, so the status belongs to the command it is printed next to.