Files
felhom.eu/REPORT-campaign11-phase24.md
T
admin 9c1d05d360
gates / gates (push) Successful in 8s
CAMPAIGN-11: hygiene, what-did-not-run, venue end state, and the session report
The recovery codes are shredded with the plant->find->shred->fail-to-find
control the brief asks for, and THE CONTROL PAID FOR ITSELF ON ITS FIRST RUN:
it found the Phase 0 code in ~/.config/credentials as R_CAMPAIGN_11 — a copy
this session did not create and would never have looked for. Without it, a
'codes shredded' claim would have been false. That key was removed from the
shared file with a verified diff (every other line identical, nine keys intact)
and HUB_PW re-tested at hub:200.

Consequence stated plainly rather than left to be discovered:
/home/felhom-repo.orphaned-20260805 (12 535 KB, the three Phase 0 sentinels) is
now permanently unopenable — which is what the set-aside screen promises, and
teardown removes it anyway.

Venue left WORKING and said so: ONLINE, 4 containers healthy, backup target not
degraded, off-site on 2 snapshots. Two things a future session needs: the raw
/mnt/{adatok,mentes} mounts are deliberately left unmounted (R-220's
workaround), and the appliance root credential was shredded — re-fetch it from
the hub.

REPORT-campaign11-phase24.md rather than REPORT.md, per the repo's
parallel-session rule.

No product code changed. No version bumped.
2026-08-06 04:32:56 +02:00

125 lines
7.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — CAMPAIGN 11, Phases 2 and 4 (2026-08-05 22:38 → 08-06 04:35, unattended)
> Written as a `REPORT-<topic>.md` sibling rather than into `REPORT.md`, per this repo's
> parallel-session rule (`CLAUDE.md:82-87`) — the brief states another session commits here.
## 1. The venue's state at the end — **WORKING**
`c11-36d660` **ONLINE**, agent `0.125.0`, 1/1 guests, reporting on schedule (last seen 04:29:54).
All four containers healthy (`calibre-web`, `filebrowser`, `felhom-controller:0.201.0`, `traefik`).
Backup target `{"degraded":false,"label":"mentes","target":"felhom-backup"}`. Off-site: **2 snapshots**,
`last_status: ok`, `last_success 2026-08-06T02:15:24Z`, `escrow_state: escrowed`.
Two things a future session must know:
- **The raw `/mnt/adatok` and `/mnt/mentes` mounts are deliberately left unmounted** — R-220's
workaround, without which no app can be deployed on a rebuilt box.
- **The appliance's vaulted root credential was shredded with the codes.** Re-fetch it from the hub
(`POST /hosts/c11-36d660/reveal-recovery-credential`) — the designed path.
## 2. §4.1 and §4.2
**§4.1 — MEASURED, twice, and the brief's own plan for taking it was wrong.**
The box renders `GetFloor()` = **`0.200.0`** (`/settings`), and a **cold-started** controller logs
`settle-gate: GO — at/above floor 0.200.0 (we are 0.201.0)` against the same line reading
`floor still unknown after 1m30s` while the hold was in force. Both `ResolveManagedFloor` hold branches
serve `Floor=""` (pinned by `managed_floor_test.go:94`), so a non-empty floor proves an ACK carried
one. The hub's HELD lines ran every 15 min to 22:12:06 then stopped, with a liveness control.
**The correction:** `SetFloor`'s line is `u.dbg(...)`, gated on `logging.level=debug` and written to
the logger — it can **never** reach the debug ring, so the restart the brief prescribed would have
produced nothing. Caught by running a level census on the ring first.
**§4.2 — still NOT measured, deliberately.** The fix is present (`needsOffsiteCredential` now retires
on the *target*, not the key), but the venue **has** a target, so the box correctly does not declare.
The state that exercises it needs a rebuild. **Phase 4 did supply its negative control:** zero
`needs_credential` and zero `offsiteheal` in five hours.
## 3. Every fault
| | injected? | result |
|---|---|---|
| F1 wrong code ×3 | yes | refused ×3, **no lockout**, **nothing written** (mtimes frozen), I4 clean — but **M4, not M1** → R-226 |
| F2 no code | yes | **PASS** — screen states nobody can replace it, offers set-aside; empty POST → „Add meg a helyreállítási kódot." in 0.027 s |
| F3 hub unreachable | yes (302→exit 7) | **FAIL** — M4 for a correct code in **0.0556 s** → R-224 |
| F4 agent stopped | yes (:8443 gone) | **FAIL** — M4 in **0.0299 s**; gate answered `source=version` → R-224 |
| F5 store unreachable | yes (TCP-OPEN→CLOSED) | **PASS** — M3, R-217's false claims absent |
| F6 „Most nem" | yes | **PASS** — full page silenced, entry point survives, unlock still works |
| F7 set aside + change of mind | yes | set-aside **PASS** (12 535 KB untouched); afterwards **FAIL** → R-228 |
| F8 restart mid-unlock | yes (StartedAt moved) | no half-state ✅, raw English `Bad Gateway` ❌ → R-227 |
| F9 never-had-off-site | precondition not staged | **PARTIAL PASS** — R-215's gate proven live |
| F10 missing mandatory path | **NOT INJECTED** | harness — three attempts, all self-healed |
| F11 offline a window | yes | **PASS** — stale + recovered, an operator mail each way |
## 4. The four messages, as they rendered
- **M1** never appeared — unreachable on this box (R-226).
- **M2** never appeared — the capability gate passed on a cached version (R-224).
- **M3** „A kulcs visszakerült, de a mentések listáját most nem sikerült beolvasni…" — F5, correct.
- **M4** „Ez a kód nem nyitja meg azt a csomagot, amit most őrzünk…" — F1, F3 **and** F4. Three
different causes, one sentence.
## 5. Invariants
**I1, I5, I7 held. I4 held on the product** (breached by the harness — see §7). **I3 breached twice**
(R-227's raw `Bad Gateway`; R-220's refusal naming an impossible action). **I6 breached twice**
(R-224, R-225). **I2 recorded as untested**, because F10 could not be injected.
## 6. Phase 4
**All five daily jobs fired exactly once, on time.** The 04:15 off-site run produced
`snapshot_count` **1 → 2** unprompted. **Nothing on the must-not list fired.**
`tier2-backup`'s 118 ms was suspected of being a silent no-op and **DISPROVED** (818.5 KB verified on
the backup drive). **A correction to my own pre-registration:** `backup_run_digest` is a *test
filename*, not an event type — the real one is `backup_run_failures`, correctly silent on a clean
night. Two absences were answered rather than assumed: the restore-test's silence was pre-registered
as correct; the whole-guest tier's is **explicitly unresolved**, because routine local-api calls are
invisible at INFO (a five-hour search returns 0 on a box that demonstrably served them).
## 7. Harness faults, separated from the product's
Five, all mine: (1) `source ~/.config/credentials` **echoed two demo-box recovery codes** into the
transcript; (2) its values are quoted, so a bare `cut -d=` yields the wrong secret; (3) the recovery
page carries no `<meta>` CSRF — my length check caught it; (4) **two failed reachability controls**
(`nc` absent; the container has no IPv6 route) reported as failed controls, not results; (5) a fixed
temp filename in my guest runner let three collectors delete each other's script, and a waiter keyed
on `date +%H -ge 4` fired at 23:xx — both caught because the evidence contradicted the claim.
## 8. Suspicions investigated and disproved
The floor being still held (**disproved** — measured served); an I5 disagreement over store size
(**disproved** — rounding); `tier2-backup` no-opping (**disproved** — the copy is real).
## 9. New findings
**R-224** misattributed unlock failures · **R-225** `0 snapshots · 0 GB` on an unread store ·
**R-226** M1 unreachable after re-escrow · **R-227** raw `Bad Gateway` · **R-228** the set-aside
history is invisible. **The highest register ID had NOT moved** — it was R-223 on arrival and R-223
when I minted, re-checked immediately before writing.
## 10. Documents
`documentation/audits/CAMPAIGN-11-recovery-journey-2026-08-05.md` — the campaign document, all four
phases, in Campaign 10's shape. It records as **still owed**: the three-layer teardown, §4.2's
positive half, F9's literal precondition, F10's real injection, the `STALE → DOWN` arm, and **a
re-walk** — the capability map's recovery row stays **FAIL** until one passes.
## 11. Recovery codes — shredded, with the control
Plant → find → shred → fail to find. **The control paid for itself immediately**: it found the Phase 0
code in `~/.config/credentials` as `R_CAMPAIGN_11`, a copy this session did not create. Without it,
"codes shredded" would have been **false**. That key was removed with a verified diff and `HUB_PW`
re-tested (`hub:200`). **`/home/felhom-repo.orphaned-20260805` (12 535 KB) is now permanently
unopenable** — as the set-aside screen promises, and teardown removes it anyway.
## 12. Teardown — OWED
Nothing removed. Three layers named in the campaign document §11, plus the off-site side: the
sub-account now holds **two** repositories, and `demo-felhom`/`demo-hp` namespaces on ep0 must not be
touched.
## 13. What did not run
F10 as specified, F9's literal precondition, §4.2's positive half, `STALE → DOWN`, the whole-guest
tier's due-ness, the retained package's read path (unbuilt), and any re-walk. **No product code
changed; no version bumped.** CI green by run ID for every push (**179184**).