Files
felhom.eu/REPORT-r241-spike-golden-teardown-2026-08-07.md
T
admin 094e93e828
gates / gates (push) Successful in 13s
finalwalk teardown complete; R-244 filed; session report
All five layers gone, each verified with a positive control that must
survive and does:
  VM 324 + 4 disks   -> absent (VM 300 drill-r50 remains)
  hub: 13 tables at 0, incl. BOTH escrow tables (demo-felhom/demo-hp/peti remain)
  Storage Box u629488-sub4 -> gone (sub1/2/3 remain)
  ep0 PBS ns finalwalk     -> gone (demo-felhom, demo-hp remain)
  WireGuard 10.77.0.5      -> gone from the LIVE wg show on ep0, not just
                              the hub DB (.2/.3/.4/.250 remain)
14.06 GiB reclaimed against 15 G measured before deletion.

R shredded with a planted-copy control: plant -> search finds both ->
shred -> the same search finds 0. The zero was not believed until the
instrument was proven.

R-244 (NEW): a FULL census after the cascade logged COMPLETE full teardown
found 61 rows still matching finalwalk. Four sources are deliberate
provenance; the fifth, app_log_issues (29 rows), is NOT covered by the
residue purge - and it is systematic: c11 40, rewalk 20, part4 24 still
present from the 2026-08-06 teardown, whose ledger recorded zero
occurrences. That claim used a narrower query than a census and does not
hold; the correction is recorded in both the prior ledger and the register
rather than the measurement quietly redone.

No secret material is involved. The table is a fleet-wide aggregate: 12 of
the 29 rows are finalwalk-only orphans, 17 are shared with LIVE customers
and must be de-referenced, not deleted - very likely why the leg was never
written. Not fixed; a cascade change needs its own red-proof.

Lesson, and it is the reusable part: a per-table absence query is not a
census.
2026-08-07 09:59:16 +02:00

195 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — R-241 spike · golden 0.205.0 · finalwalk teardown (2026-08-07)
*A sibling report: `REPORT.md` is overwritten per-session and a parallel session shares this clone.*
**Three parts, in the order the venue's perishability required.** The spike read VM 324 before
anything else touched the fleet; the bake and the vouch followed; the teardown went last.
---
## 1. THE SESSION'S ANSWER — R-241 is a MINTING defect, not a screen-predicate defect
**This reverses the fix.** The recovery screen was telling the truth: there genuinely was nothing
recoverable under the key the box held, because **the box minted that key itself, over the top of a
sealed package it already knew the hub was holding for it.** Mending the predicate would have papered
over a box quietly making its own backup history unopenable.
**Three measurements, taken from the venue this session, not copied from last night's journal:**
1. **`WriteOffboxSecrets` (`offbox.go:411`) mints on ONE input — does the file exist.** No settings
read at all, while its two neighbours in the same file, `OffsiteRecoveryOffer()` (`:1412`) and
`needsOffsiteCredential()` (`:1377`), both consult `GetHubEscrowIdentityPresent()`. **The same fact
is available on three paths and used on two.**
2. **That flag was not merely available — it was the precondition of the chain that reached the
minting.** The 5-minute retry job logs only when `RetryIfDeclared` fires, which requires the
declaration, which requires the flag. The venue logged
`credential retry: … (the box still declares a need; retrying)` at **02:48:03Z** and five times
after — **thirty minutes and six ticks before the mint at 03:18:06Z.**
3. **The box computed the right answer and threw it away.** At **03:28:03Z** — thirty-five minutes
before the customer looked — `EscrowAutoConfirmer.Reconcile` (`escrow_confirm.go:154`) logged
`the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local
9b4a9a9dcec7…) … staying pending`. It is recomputed on every report cycle, **never persisted,
never surfaced.**
**And the hub explicitly disclaims doing this**`offsiteheal`'s package doc: *"it never runs, or asks
for, an escrow ceremony… **credential automatic, key customer-present** — the ruling this session
implements and must not quietly widen."* The repository key is created **in the seam between two sides
that each honoured their contract**, by a helper doing exactly what its doc comment says. That is why
it survived review: every individual comment is accurate.
**Corroboration from both sides, measured independently here:** the two key hashes on the venue
(`30ef574f…` current, `9b4a9a9d…` in `repo_password.selfheal-aside`, mtime 03:18:06) reproduce the
journal's figures exactly; and the hub's `host_escrow` row for `finalwalk-ed05d6` seals
`restic_pw_sha256 = 30ef574f…` with a 572-byte `identity_blob` and **no superseded row** — so the hash
the box logged as *"hub hash"* is confirmed from the database, not inferred.
**Q2 — shape (b) is structurally unreachable**, not merely unfired: `markOrphaned()` has one producer,
`ensureOffboxRepo()`, reached only from inside a run, and every run passes `if !m.offboxEscrowed()`
(`offbox.go:743`) **first**. The box was `pending` and could not stop being pending, because the
auto-confirm flips only on a hash match. *Positive control, because an absent log line is not
evidence:* the scheduler was alive throughout (241 `agent-channel-health`, 120 `stack-scan`, 48
`offsite-credential-retry`, 16 `hub-report`), and `offbox-backup` is a `sched.Daily` leg whose slot
fell **before** the destruction — the absence is explained, not just observed.
**Q7 — the „Helyreállítási kód létrehozása" button.** It does **not** destroy the data: `SaveHostEscrow`
demotes the current row into `host_escrow_superseded` **copy-before-delete, `identity_blob` included**
(R-198). But the read path for a retained package is **unbuilt** (R-199), so it converts a
one-screen-away self-service recovery into one needing an operator and tooling that does not exist.
**And it re-enables the screen while invalidating the code that screen accepts** — a worse trap than a
plain dead end, because it looks like progress. *(Reasoned from code and the hub schema; I did not
press it — that is a state change and would have destroyed the evidence.)*
**Two findings nobody asked for**, both filed: **R-243** (a box in this state silently stops backing up
and **no alarm fires** — three individually-correct exclusions leave one state unobserved) and the trap
in the obvious fix (`ResetOrphanedRepo` clears the orphan flag without a ceremony, so a hash-mismatch
discriminator alone would re-offer the screen forever to a customer who declined the old data).
**Output:** `audits/SPIKE-r241-recovery-offer-2026-08-07.md` — question/method/measurement/ruling, the
Q4 seven-state table, ranked options, and **four operator decisions stated and left unanswered.**
**No product code was written, in any part of this session.**
---
## 2. GOLDEN 0.205.0 — baked, verified, VOUCHED (R-239 CLOSED)
| | |
|---|---|
| version | **0.205.0** · sha256 `8f49b2e8ccbc86a49df821fee9fb00c07293758811d3d0f0512dd0cf5fd54ee8` |
| size | 656,937,561 bytes (uncompressed 2,003,343,360) |
| MinAgent | 0.127.0 |
**Round-trip verified rather than trusted:** the published bytes were fetched back (HTTP 200, byte
count matches), re-hashed **independently of the baker** (matches), `zstd -t`'d clean, and
`./etc/felhom-controller-image` was read **out of the downloaded archive**
`felhom-controller:0.205.0`. That last step is the one that matters, because `GOLDEN_VERSION` is
derived from the tag argument and could have been right over stale content. A **third witness**: the
hub's own dropdown lists `0.205.0` with `data-sha="8f49b2e8…"`.
All acceptance markers pass (`docker OK (overlay2` ×1, `including mount point` ×2 = rootfs + mp0,
`upload OK (HTTP 201)` ×1, and 0 each for `excluding`/`FATAL`/`ERROR:`/`WARN:`); unit
`Result=success`, `ExecMainStatus=0`. Bake VM torn down: CT 9100 purged, secrets shredded, qemu
observed gone via `ps -eo comm`, `drill.qcow2` reverted to `virgin`.
**Secret hygiene:** token copied file→file; the invocation lives in an in-VM runner that reads the
token itself, so it never reached a command line (`systemctl show … | grep -c -F <token>`**0**);
literal-value leak grep on the **committed** log → **0**, and **the instrument was proven first**
token appended to a throwaway copy → 1 hit → copy shredded → the 0 is a measurement.
**Vouched, with the operator's approval.** In the event it was a **ONE-field change, not three** — read
from live `hub_settings` before and after, not assumed:
| field | before | after |
|---|---|---|
| `artifact_golden_version` | 0.203.0 | **0.205.0** |
| `artifact_golden_sha256` | `3039c6ff…` | **`8f49b2e8…`** |
| `artifact_agent_version` | 0.127.0 | 0.127.0 — **unchanged** |
| `artifact_min_agent` | 0.127.0 | 0.127.0 — **unchanged** |
**A trap worth naming:** `wrapper_sha256` is read from the form and **cleared when omitted**. A
headless POST that forgets it silently drops the PBS-DR wrapper hash. It was carried through
explicitly and verified present afterwards. **The R-120 gate passed exactly** — the newest controller
the fleet reports is 0.205.0 (demo-hp), so a **0.204.0 golden would have been REFUSED**. Vouching is
reversible; a bake never deletes an older golden's package.
**§4.1's systemic half is recorded, NOT built** → **R-242**, with three proposed shapes and a stated
earliest-catch (a `repo_gates.py` comparison of the manifest's `golden_version` against the newest
released controller — it fires on the push that creates the gap, before any box is installed).
---
## 3. TEARDOWN — `finalwalk`, all five layers
Enumerated first and cross-checked against the hub's **own** delete-preview, which agreed in every
field. Matched on **identity**, never on size.
The cascade refuses a live host, so VM 324 was stopped at 07:23:56Z (guarded on `qm config 324`
reading `name: finalwalk-appliance``demo-hp` also carries a guest 9201) and aged past the 30 m
`stale_threshold`, read from the deployed config. `POST /configs/finalwalk/delete` with all six gates
→ the hub's leg-by-leg log shows host deleted, off-site sub-account 285071 deprovisioned, PBS
namespace deprovisioned (`existed=true` — the positive observable), claim reset, residue purged
(**124 rows, matching the preview exactly**), `COMPLETE … full teardown`. Then `qm destroy 324
--purge`.
**Every layer verified absent with a positive control that must survive, and does:** VM 324 gone (VM
300 `drill-r50` remains) · all 13 hub tables at 0 including **both** escrow tables (demo-felhom 50,823
/ demo-hp 7,640 / peti 1,827 rows remain) · Storage Box `u629488-sub4` gone (sub1/2/3 remain) · `ep0`
namespace `finalwalk` gone (demo-felhom, demo-hp remain) · WireGuard `10.77.0.5` gone **from the live
`wg show` on ep0, not merely from the hub DB** (`.2/.3/.4/.250` remain). **14.06 GiB reclaimed**,
against 15 G measured before deletion.
**R shredded with a planted-copy control** — plant → search finds both → shred → same search finds 0.
The zero was not believed until the instrument was proven.
### ⚠ And a correction I am reporting rather than quietly fixing
A **full census** (every table, every column) after the "COMPLETE" cascade found **61 rows still
matching `finalwalk`**. Four sources are deliberate (`events`, `notification_log`, `host_deletions`,
`customer_resets` — the cascade's header says provenance outlives every tier). **The fifth is a gap:**
`app_log_issues`, 29 rows, not covered by the residue purge — and **systematic**, with `c11` 40,
`rewalk` 20 and `part4` 24 still present from the 2026-08-06 teardown, **whose ledger recorded "0
occurrences"**. That claim used a narrower query than a census and does not hold. Both the prior ledger
and the register now carry the correction.
**No secret material is involved.** The table is a fleet-wide aggregate; 12 of the 29 rows are
`finalwalk`-only orphans and **17 are shared with live customers and must be de-referenced, not
deleted** — very likely why the leg was never written. Filed as **R-244**, not fixed: a cascade change
needs its own red-proof. **The reusable lesson: a per-table absence query is not a census.**
---
## Register
| ID | Movement |
|---|---|
| **R-241** | **RULED** — minting defect. Diagnosed, **not fixed**; four operator decisions owed |
| **R-239** | **CLOSED** — golden 0.205.0 baked, verified and vouched |
| **R-242** | **NEW** — a release is not delivered until a golden carries it; recorded, **not built** |
| **R-243** | **NEW** — the R-241 state silently stops off-site backups with no alarm |
| **R-244** | **NEW**`app_log_issues` survives the delete cascade, across all four torn-down venues |
**Highest register ID moved R-241 → R-244.**
## Verification
| | |
|---|---|
| commits | `71c43f87c240` (spike) · `08b75e602e01` (bake evidence) · `db578cd44d3e` (R-239 closed, map + STATUS) · this one |
| CI | run **233** `71c43f87c240` success · run **234** `08b75e602e01` success — matched by `head_sha`, pulled not assumed |
| `--no-verify` | **not used.** The pre-push hook ran `repo_gates.py --fast` on every push and reported `gates OK` |
| gates | `python3 scripts/repo_gates.py --fast` → all six OK before each commit |
## What did not run, and why
- **No product code, deliberately** — Part 1 was a question, and the answer changes what the fix
should be. Beginning a candidate before the operator rules on §"THE OPERATOR'S DECISION" would
prejudge it.
- **Q4 row 7 and Q7's post-button behaviour are reasoned from code, not measured.** Both need a state
change on the venue; either would have destroyed the evidence for everything else. They need a fresh
fixture and their own session — and **row 7 shapes the fix**, so it matters.
- **R-242, R-243 and R-244 are filed and not built** — R-242 because the task scoped it record-only;
the other two because they surfaced inside a spike and an operation, and each needs its own
red-proof.
- **The exact wall-clock at which `hub_escrow_identity_present` first became true is not measured** —
`SetHubEscrowIdentityPresent` writes only on change and logs nothing. The bound that matters is
established by control flow: true at or before **02:48:03Z**.