CAMPAIGN-11: hygiene, what-did-not-run, venue end state, and the session report
gates / gates (push) Successful in 8s

The recovery codes are shredded with the plant->find->shred->fail-to-find
control the brief asks for, and THE CONTROL PAID FOR ITSELF ON ITS FIRST RUN:
it found the Phase 0 code in ~/.config/credentials as R_CAMPAIGN_11 — a copy
this session did not create and would never have looked for. Without it, a
'codes shredded' claim would have been false. That key was removed from the
shared file with a verified diff (every other line identical, nine keys intact)
and HUB_PW re-tested at hub:200.

Consequence stated plainly rather than left to be discovered:
/home/felhom-repo.orphaned-20260805 (12 535 KB, the three Phase 0 sentinels) is
now permanently unopenable — which is what the set-aside screen promises, and
teardown removes it anyway.

Venue left WORKING and said so: ONLINE, 4 containers healthy, backup target not
degraded, off-site on 2 snapshots. Two things a future session needs: the raw
/mnt/{adatok,mentes} mounts are deliberately left unmounted (R-220's
workaround), and the appliance root credential was shredded — re-fetch it from
the hub.

REPORT-campaign11-phase24.md rather than REPORT.md, per the repo's
parallel-session rule.

No product code changed. No version bumped.
This commit is contained in:
2026-08-06 04:32:56 +02:00
parent df6081e60b
commit 9c1d05d360
3 changed files with 248 additions and 2 deletions
+124
View File
@@ -0,0 +1,124 @@
# REPORT — CAMPAIGN 11, Phases 2 and 4 (2026-08-05 22:38 → 08-06 04:35, unattended)
> Written as a `REPORT-<topic>.md` sibling rather than into `REPORT.md`, per this repo's
> parallel-session rule (`CLAUDE.md:82-87`) — the brief states another session commits here.
## 1. The venue's state at the end — **WORKING**
`c11-36d660` **ONLINE**, agent `0.125.0`, 1/1 guests, reporting on schedule (last seen 04:29:54).
All four containers healthy (`calibre-web`, `filebrowser`, `felhom-controller:0.201.0`, `traefik`).
Backup target `{"degraded":false,"label":"mentes","target":"felhom-backup"}`. Off-site: **2 snapshots**,
`last_status: ok`, `last_success 2026-08-06T02:15:24Z`, `escrow_state: escrowed`.
Two things a future session must know:
- **The raw `/mnt/adatok` and `/mnt/mentes` mounts are deliberately left unmounted** — R-220's
workaround, without which no app can be deployed on a rebuilt box.
- **The appliance's vaulted root credential was shredded with the codes.** Re-fetch it from the hub
(`POST /hosts/c11-36d660/reveal-recovery-credential`) — the designed path.
## 2. §4.1 and §4.2
**§4.1 — MEASURED, twice, and the brief's own plan for taking it was wrong.**
The box renders `GetFloor()` = **`0.200.0`** (`/settings`), and a **cold-started** controller logs
`settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)` against the same line reading
`floor still unknown after 1m30s` while the hold was in force. Both `ResolveManagedFloor` hold branches
serve `Floor=""` (pinned by `managed_floor_test.go:94`), so a non-empty floor proves an ACK carried
one. The hub's HELD lines ran every 15 min to 22:12:06 then stopped, with a liveness control.
**The correction:** `SetFloor`'s line is `u.dbg(...)`, gated on `logging.level=debug` and written to
the logger — it can **never** reach the debug ring, so the restart the brief prescribed would have
produced nothing. Caught by running a level census on the ring first.
**§4.2 — still NOT measured, deliberately.** The fix is present (`needsOffsiteCredential` now retires
on the *target*, not the key), but the venue **has** a target, so the box correctly does not declare.
The state that exercises it needs a rebuild. **Phase 4 did supply its negative control:** zero
`needs_credential` and zero `offsiteheal` in five hours.
## 3. Every fault
| | injected? | result |
|---|---|---|
| F1 wrong code ×3 | yes | refused ×3, **no lockout**, **nothing written** (mtimes frozen), I4 clean — but **M4, not M1** → R-226 |
| F2 no code | yes | **PASS** — screen states nobody can replace it, offers set-aside; empty POST → „Add meg a helyreállítási kódot." in 0.027 s |
| F3 hub unreachable | yes (302→exit 7) | **FAIL** — M4 for a correct code in **0.0556 s** → R-224 |
| F4 agent stopped | yes (:8443 gone) | **FAIL** — M4 in **0.0299 s**; gate answered `source=version` → R-224 |
| F5 store unreachable | yes (TCP-OPEN→CLOSED) | **PASS** — M3, R-217's false claims absent |
| F6 „Most nem" | yes | **PASS** — full page silenced, entry point survives, unlock still works |
| F7 set aside + change of mind | yes | set-aside **PASS** (12 535 KB untouched); afterwards **FAIL** → R-228 |
| F8 restart mid-unlock | yes (StartedAt moved) | no half-state ✅, raw English `Bad Gateway` ❌ → R-227 |
| F9 never-had-off-site | precondition not staged | **PARTIAL PASS** — R-215's gate proven live |
| F10 missing mandatory path | **NOT INJECTED** | harness — three attempts, all self-healed |
| F11 offline a window | yes | **PASS** — stale + recovered, an operator mail each way |
## 4. The four messages, as they rendered
- **M1** never appeared — unreachable on this box (R-226).
- **M2** never appeared — the capability gate passed on a cached version (R-224).
- **M3** „A kulcs visszakerült, de a mentések listáját most nem sikerült beolvasni…" — F5, correct.
- **M4** „Ez a kód nem nyitja meg azt a csomagot, amit most őrzünk…" — F1, F3 **and** F4. Three
different causes, one sentence.
## 5. Invariants
**I1, I5, I7 held. I4 held on the product** (breached by the harness — see §7). **I3 breached twice**
(R-227's raw `Bad Gateway`; R-220's refusal naming an impossible action). **I6 breached twice**
(R-224, R-225). **I2 recorded as untested**, because F10 could not be injected.
## 6. Phase 4
**All five daily jobs fired exactly once, on time.** The 04:15 off-site run produced
`snapshot_count` **1 → 2** unprompted. **Nothing on the must-not list fired.**
`tier2-backup`'s 118 ms was suspected of being a silent no-op and **DISPROVED** (818.5 KB verified on
the backup drive). **A correction to my own pre-registration:** `backup_run_digest` is a *test
filename*, not an event type — the real one is `backup_run_failures`, correctly silent on a clean
night. Two absences were answered rather than assumed: the restore-test's silence was pre-registered
as correct; the whole-guest tier's is **explicitly unresolved**, because routine local-api calls are
invisible at INFO (a five-hour search returns 0 on a box that demonstrably served them).
## 7. Harness faults, separated from the product's
Five, all mine: (1) `source ~/.config/credentials` **echoed two demo-box recovery codes** into the
transcript; (2) its values are quoted, so a bare `cut -d=` yields the wrong secret; (3) the recovery
page carries no `<meta>` CSRF — my length check caught it; (4) **two failed reachability controls**
(`nc` absent; the container has no IPv6 route) reported as failed controls, not results; (5) a fixed
temp filename in my guest runner let three collectors delete each other's script, and a waiter keyed
on `date +%H -ge 4` fired at 23:xx — both caught because the evidence contradicted the claim.
## 8. Suspicions investigated and disproved
The floor being still held (**disproved** — measured served); an I5 disagreement over store size
(**disproved** — rounding); `tier2-backup` no-opping (**disproved** — the copy is real).
## 9. New findings
**R-224** misattributed unlock failures · **R-225** `0 snapshots · 0 GB` on an unread store ·
**R-226** M1 unreachable after re-escrow · **R-227** raw `Bad Gateway` · **R-228** the set-aside
history is invisible. **The highest register ID had NOT moved** — it was R-223 on arrival and R-223
when I minted, re-checked immediately before writing.
## 10. Documents
`documentation/audits/CAMPAIGN-11-recovery-journey-2026-08-05.md` — the campaign document, all four
phases, in Campaign 10's shape. It records as **still owed**: the three-layer teardown, §4.2's
positive half, F9's literal precondition, F10's real injection, the `STALE → DOWN` arm, and **a
re-walk** — the capability map's recovery row stays **FAIL** until one passes.
## 11. Recovery codes — shredded, with the control
Plant → find → shred → fail to find. **The control paid for itself immediately**: it found the Phase 0
code in `~/.config/credentials` as `R_CAMPAIGN_11`, a copy this session did not create. Without it,
"codes shredded" would have been **false**. That key was removed with a verified diff and `HUB_PW`
re-tested (`hub:200`). **`/home/felhom-repo.orphaned-20260805` (12 535 KB) is now permanently
unopenable** — as the set-aside screen promises, and teardown removes it anyway.
## 12. Teardown — OWED
Nothing removed. Three layers named in the campaign document §11, plus the off-site side: the
sub-account now holds **two** repositories, and `demo-felhom`/`demo-hp` namespaces on ep0 must not be
touched.
## 13. What did not run
F10 as specified, F9's literal precondition, §4.2's positive half, `STALE → DOWN`, the whole-guest
tier's due-ness, the retained package's read path (unbuilt), and any re-walk. **No product code
changed; no version bumped.** CI green by run ID for every push (**179184**).
@@ -554,8 +554,17 @@ accumulated before; the Phase 0 census found **none** outstanding, and c11 must
## 12. Hygiene
- **The recovery codes** live in `~/.config/campaign11/` on DooPlex, `0600`, and are reported in §13
of the session report with their shred and its positive control.
- **The recovery codes are SHREDDED**, with the plant→find→shred→fail-to-find control the brief asks
for. **The control paid for itself on its first run**: it found the Phase 0 code in
`~/.config/credentials` as **`R_CAMPAIGN_11`** — a copy this session did not create and would never
have looked for, which would have made a "codes shredded" claim **false**. That key was removed from
the shared file carefully (backup → exact-match removal → `diff` proving every other line identical →
`HUB_PW` re-verified at `hub:200` → backup shredded). Everything else was `shred -u`'d on both hosts
and the absence re-swept; only the planted controls remained, and they were shredded too.
**⚠ Consequence, stated plainly: `/home/felhom-repo.orphaned-20260805` (12 535 KB, the three Phase 0
sentinels) is now permanently unopenable.** That is what the set-aside screen promises will happen,
teardown removes the repository anyway, and R-222 means no read path existed for it regardless — but
the door is now shut for good.
- **No secret is written into any committed file.** Every hash quoted here is a sha256 prefix; the
codes' contents appear nowhere.
- **`git add -A` was never used** — every commit staged explicit paths, and `git status --porcelain`
@@ -567,3 +576,34 @@ accumulated before; the Phase 0 census found **none** outstanding, and c11 must
rule 1.
- **No version was bumped** in any repo.
---
## 13. What did not run, and why
| | why |
|---|---|
| **F10 as specified** | **Not injectable.** Three attempts, each with a control: the app, then the controller's monitor, then the run itself recreate the mandatory directory within ~1 s. The state does not exist on a deployed app of this kind. **Harness, not product** |
| **F9's literal precondition** | A box that never had off-site backups needs a **rebuild**, which the brief forbids before Phase 4. The assertion that failed in Phase 1 (the page not consulting its predicate) was tested instead, and passes |
| **§4.2's positive half** | Needs shape (a) — hub blob present, key placed, **no target**. The venue has a target, so the box correctly does not declare. Also needs a rebuild |
| **`STALE → DOWN` escalation** | F11 was ended once `stale` and `recovered` had both fired, because Phase 4 needed the venue back. The >1 h arm is untested |
| **The whole-guest tier's due-ness** | **Not resolvable with existing instruments** — routine local-api calls and the hub's deadline monitor are both invisible at INFO. Recorded as an open question, not scored |
| **The retained package's read path** | Does not exist (R-199's inventory is unbuilt). R-222 was re-confirmed live rather than re-tested |
| **A re-walk of the journey** | **Deliberately out of scope.** These faults are not a re-walk, and the capability map's recovery row stays **FAIL** until one passes |
| **Any fix** | Brief rule 1 — file and continue. **No product code was changed and no version bumped** |
---
## 14. Venue state at the end — **WORKING**
| | |
|---|---|
| Hub | `c11-36d660` **ONLINE**, agent `0.125.0`, 1/1 guests, reporting on schedule (last seen 04:29:54) |
| Containers | `calibre-web` · `filebrowser` · `felhom-controller:0.201.0` · `traefik` — all healthy |
| Drives | both enrolled; backup target `{"degraded":false,"label":"mentes","target":"felhom-backup"}` |
| Off-site | **2 snapshots**, `last_status: ok`, `last_success 2026-08-06T02:15:24Z`, 51 206 B, `escrow_state: escrowed` |
| Set aside | `/home/felhom-repo.orphaned-20260805` — 12 535 KB, intact, and now **permanently unopenable** (§12) |
| Left in place | the raw `/mnt/adatok` and `/mnt/mentes` mounts remain **unmounted** — R-220's workaround, without which no app can be deployed on a rebuilt box. The stable `/mnt/felhom-drives/*` mounts are what everything uses |
| Access | the appliance's vaulted root credential was **shredded with the codes**, so a future session must re-fetch it from the hub (`POST /hosts/c11-36d660/reveal-recovery-credential`) — the designed path |
**Nothing is left broken, and nothing is left running that should not be.**
@@ -833,3 +833,85 @@ self-heal, which is R-218's negative control.
was true it is said so: the restore-test's silence was predicted in advance, and the whole-guest
tier's silence is left explicitly unresolved rather than counted as a pass.
---
# Hygiene — the recovery codes, shredded with a positive control
The brief's warning was earned: *"a sweep pointed at a path that did not exist inside the guest and
its zero hits meant nothing. Prove the sweep works before trusting it."* So the sweep was proved
first, and **it immediately found a copy this session did not know about.**
**The sweep searches by CONTENT, never by filename**, and the pattern is passed to `grep -f` from a
file so the code never appears on a command line or in a process argument.
### 1. Plant → 2. Find (the positive control)
A copy of each code was planted at a known extra path. The sweep over DooPlex (`~/.config`,
`/tmp/claude-1000`, `/tmp`) and demo-hp (`/root`, `/tmp`) returned:
```
phase0: ~/.config/credentials ← ⚠ NOT KNOWN TO THIS SESSION
~/.config/campaign11/R_C11_phase0.txt
<scratchpad>/PLANTED_phase0.txt ← the planted control, found ✅
<scratchpad>/R_C11_phase0.strip
demo-hp:/root/R_C11_phase0.strip
phase3: ~/.config/campaign11/R_C11_phase3.txt · PLANTED_phase3.txt · .strip · demo-hp:/root/…
```
> **The control paid for itself on its first run.** `~/.config/credentials` — the *shared* credential
> store — held the Phase 0 recovery code as **`R_CAMPAIGN_11`** (line 10). Nothing in this session put
> it there. Without the planted-copy control there would have been no reason to sweep at all, and a
> "codes shredded" claim would have been **false**.
>
> That file is the same one whose values are echoed by a failed `source` (§harness trap 1). A recovery
> code living there is the two hazards composed.
### 3. Shred
**DooPlex** — all six files in `~/.config/campaign11/` (`R_C11_phase0`, `R_C11_phase3`,
`dashboard_pw`, `managed_root_pw`, `retrieval_passphrase`, `root_pw`), then the directory itself;
the scratchpad's `.strip` files, `appliance_pw.txt`, `dashpw.txt`, `reveal.json`. All with `shred -u`.
**`~/.config/credentials`** — `R_CAMPAIGN_11` removed. Done carefully because that file also holds
`HUB_PW`, on which this session depended: a backup was taken, the key removed by exact match, and
then **verified** — `diff` of every other line reports **identical**, the nine remaining keys are
unchanged (`HETZNER_API PASSWORD TS_KEY HUB_PW ISO_S3_* R_DEMO-FELHOM R_DEMO-HP`), and `HUB_PW` still
authenticates (`hub:200`). The backup was shredded afterwards.
**demo-hp** — the four `.strip` files, `.c11pw`, `.c11dashpw`, `.c11sess`, `f8_out.txt` shredded; the
helper scripts removed.
### 4. Fail to find
Re-swept with the planted copies as the pattern — **the only remaining matches were the planted files
themselves**, on both hosts. demo-hp returned nothing at all. The planted copies were then shredded
and their absence verified by path, along with `~/.config/campaign11/`.
**The captured HTML pages are covered by that sweep**, not merely assumed clean: step 4 searched
`/tmp/claude-1000` recursively, which contains every `*.html` capture taken this session, and none
matched.
**Not swept, and why:** the guest was not searched for the two real codes, because putting the pattern
there to search for it would be the leak. **F1 already covers the mechanism** — a sweep of the guest's
data dir, `/var/log`, `/tmp` and `docker logs` for the *wrong* code returned zero product hits with a
planted canary passing first, and the same code-handling path served every unlock.
### ⚠ One consequence, stated plainly
**Shredding the Phase 0 code makes `/home/felhom-repo.orphaned-20260805` permanently unopenable.**
That is 12 535 KB holding the three Phase 0 sentinels. It is exactly what the set-aside screen tells a
customer will happen, teardown will remove the repository anyway, and R-222 means no read path exists
for it regardless — but it is a door that is now closed for good, and it should be closed knowingly
rather than discovered later.
### Other hygiene
- **No secret is in any committed file.** Every hash quoted is a sha256 prefix.
- **`git add -A` never used**; `git status --porcelain` checked before each commit; no foreign file
was swept (a parallel session shares this clone).
- **The pre-push hook is armed** (`core.hooksPath=.githooks`) and ran green before every push;
**no `--no-verify`**.
- **CI confirmed green by run ID** for all pushes: runs **179184**, each matched to its `head_sha`.
- **No product code changed. No version bumped.**