chaos night: Phase 2 and the interventions section written up
gates / gates (push) Successful in 22s
gates / gates (push) Successful in 22s
Phase 2: all twelve front doors 200 and all 26 containers healthy (traefik and cloudflared carry no healthcheck, so a bare Up is correct). The off-site restore could not be done and two independent instruments agree why - restic itself and the product's own status surface both say the repository is orphaned with zero readable snapshots, because this box is a rebuild whose restic password was minted fresh. The product surfaced that honestly within seconds. What DID leave the house: the whole-guest copy on ep0, two intact snapshots including tonight's 21:59:54Z one. Stated as a limit: that is a listing, not a verification, and a PBS verify writes state so it was not run. Household loop: 204 probes, 7 flagged, only 2 real events - both single-sample outages during the two abrupt stops. Three were my own classifier counting a 301 as a failure, corrected in the log's own words. Ten of twelve rounds left no mark, which is the 2-minute sampling rate and not proof of nothing. Catalog bump verified reverted. Mail delivery proof recorded: raised and delivered are two different claims and only one had evidence before tonight. Interventions: 1 of 4. Both pre-declared presses unused - both prompt claims they insured against turned out true, and the F-14 path was measured live for the first time. Phase 0's seeding repairs listed separately because that damage was mine; harness acts excluded with the reason stated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -541,18 +541,90 @@ cuts, resets, full disks, severed networks and a drive pulled out of a running m
|
||||
which nothing was done produced nothing — no alarm, no restart, no drift. The alarm feed is not
|
||||
simply noisy.
|
||||
|
||||
### Rounds 8-12 — *(written up individually above)*
|
||||
|
||||
PENDING
|
||||
|
||||
|
||||
## Phase 2 — the morning after
|
||||
|
||||
PENDING
|
||||
**Every app answered through its own front door.** All twelve real names, measured at 00:19:16Z —
|
||||
`cloud`, `inventory`, `media`, `paperless`, `paste`, `photos`, `recipes`, `share`, `status`,
|
||||
`travel`, `vault`, `wiki` — **200 on the public path, every one**. And healthy is not inferred from
|
||||
a door: all **26 containers** report `Up … (healthy)`, except `traefik` and `cloudflared`, which
|
||||
carry no healthcheck and show a bare `Up`. The four showing seven minutes are the drive-backed apps
|
||||
restarted after round 11.
|
||||
|
||||
**The off-site restore could not be done, and two independent instruments agree why.** The brief asked
|
||||
for one DB-backed app restored from off-site onto scratch 9202. It has nothing to restore from:
|
||||
|
||||
| instrument | answer |
|
||||
|---|---|
|
||||
| `restic`, with the box's own key, password file and `known_hosts` | `Fatal: wrong password or no key found` — exit status **1**, read from restic itself rather than from the end of a pipeline |
|
||||
| the product's own status surface, which is what a customer sees | `{"orphaned":true, … ,"snapshots":0,"status":"error"}` |
|
||||
|
||||
The repository is orphaned **because this box is a rebuild for an existing customer**: on a rebuild
|
||||
the restic password is minted fresh, so snapshots written under the old one can never be opened
|
||||
again. That is a known, documented shape, and **the product surfaced it honestly** — the true
|
||||
`offbox_repo_orphaned` alarm of round 1, mailed to the operator within two seconds of the run.
|
||||
|
||||
**What did leave the house.** The *whole-guest* off-site copy is a different store and it worked: ep0
|
||||
holds two intact snapshots for this box, including tonight's at **21:59:54Z** — round 6's off-site
|
||||
leg, the one that started by itself after I killed the local leg. Its file index is roughly four
|
||||
times the afternoon copy's, consistent with a guest by then carrying twelve apps. **Stated as a
|
||||
limit:** that is a *listing*, not a verification. A PBS verify job would prove restorability and it
|
||||
**writes** verify state, so it was not run — ep0 is read-only for evidence tonight.
|
||||
|
||||
**The household loop.** 204 probes from 21:06:30Z to 00:18:10Z. Seven lines flagged as failures, of
|
||||
which **only two are real events** — `wiki` during round 2's power cut and `cloud` during round 10's
|
||||
hard reset, each a single sample, each healed before the next probe. The other three were **my own
|
||||
classifier** counting a 301 redirect as a dashboard failure; the log carries that correction in its
|
||||
own words at 21:11:25Z and the wrong lines were left in place so the correction stays visible. Ten of
|
||||
the twelve rounds left **no mark at all** in this log, including the twenty minutes with the drive
|
||||
pulled — the loop samples each name every two minutes, so that silence is the instrument's sampling
|
||||
rate and **not** evidence the household saw nothing.
|
||||
|
||||
**The catalog bump: verified reverted, not remembered.** Clean tree, local exactly level with
|
||||
`origin/main` (0 ahead, 0 behind), the nextcloud template still on its original `redis:7-alpine` pin
|
||||
and `catalog_since: "2026-07-18"`, newest commit 2026-09-15. The bump was prepared, **refused by the
|
||||
catalog's own gates** (image-resolvable and volume-persistence both INCONCLUSIVE — its own canary
|
||||
failed, so the verdict was UNDETERMINED, never a pass) and reverted before any push.
|
||||
|
||||
**And one delivery proof the truth table could not give.** The mailbox shows every alarm **arrived**
|
||||
at the operator, not merely that it was stored — round 11's whole set, round 6's tier-naming backup
|
||||
failure, round 1's orphan warning, the OOM, the deploy failure and my own thin-pool `storage_fill_critical`.
|
||||
In a project where 91 events once sat in a database having e-mailed nobody, *raised* and *delivered*
|
||||
are two different claims, and only one of them had evidence before tonight.
|
||||
|
||||
## Interventions — counted, with the reason for each verdict
|
||||
|
||||
PENDING
|
||||
**One.** At 21:59:45Z in round 6 I killed the **local leg** of the whole-guest backup. The arithmetic
|
||||
that forced it: a ~29 GB source being written into a 14 GB root filesystem at ~16 MB/s, i.e. under
|
||||
four minutes to a full `/` on the nested host, mid-round. What it cost: a leg that could never have
|
||||
succeeded. What happened next without me: **the off-site leg started by itself from the same snapshot
|
||||
and succeeded in about eight and a half minutes.** Filed as **R-548**.
|
||||
|
||||
My first note on it said „I stopped the backup". That was wrong and is corrected in the evidence: I
|
||||
stopped **a leg** of it, and the box completed the other one unaided.
|
||||
|
||||
**Both pre-declared presses went unused.** O1, the „Send self-bind link" button, was not needed —
|
||||
the automatic mail was already waiting (18:17:46Z) and the box bound with **zero** operator presses.
|
||||
O2, „Re-issue PBS credentials", was not needed either — the acknowledged-delete path re-issued them
|
||||
by itself (`pbsdr_auto_reissue`, 20:19Z). **Both prompt claims they were insurance against turned out
|
||||
to be true**, and the F-14 path was measured live for the first time.
|
||||
|
||||
**Counted separately, because it is not a round result:** Phase 0's seeding repairs. I filled the LVM
|
||||
thin pool to 100 % by firing twelve deploys at once, then repaired the damage — a guest restart to
|
||||
clear an `emergency_ro` remount, dropping corrupt image layers, a remove-with-data and one consistent
|
||||
re-deploy after a re-seed minted fresh database passwords over initialised volumes, and a rewritten
|
||||
`APP_KEY`. **The damage was mine, not the product's**, the product's behaviour throughout was correct,
|
||||
and every repair went through the product's own endpoints rather than by hand-running compose. Listed
|
||||
in full in `evidence-chaos-night-2026-09-17/interventions.txt` so the distinction is visible rather
|
||||
than convenient.
|
||||
|
||||
**Not counted, and why:** acts on my **own instruments** — moving the dashboard password after a
|
||||
power cut cleared `/tmp`, rewriting the injector and the runner, re-creating the disk guard as a real
|
||||
unit. Counting those would flatter the night in one direction and pad the stop-rule count in the
|
||||
other. **Round 4 is the clearest case of declining to intervene:** the tunnel was left dead on
|
||||
purpose — „NOT restarting it by hand — whether it returns by itself IS the measurement". It
|
||||
returned by itself in 97 s.
|
||||
|
||||
**Standing against the stop rule: 1 of 4.** The night ran its full twelve rounds.
|
||||
|
||||
## Teardown — three layers, stated
|
||||
|
||||
|
||||
Reference in New Issue
Block a user