3a539ea530
gates / gates (push) Successful in 11s
The first vouch of agent 0.125.0 was REFUSED by R-120's gate: the artifacts form saves as a unit, and golden 0.192.0 was older than the newest controller the fleet reports. The golden had been stale since before controller 0.193.0, so the Day-0 manifest had been effectively unvouchable for days and nobody had cause to notice. Correcting STATUS.md and OPEN-ITEMS, which said "one dropdown". Golden 0.201.0 baked in the drill VM, round-trip verified from Gitea (e730d7cab343eb35...f007654, 658165766 B), teardown clean, 0 token hits in the saved log. Manifest now agent=0.125.0 golden=0.201.0 min_agent=0.125.0, read back after the save. The held floor: the HOLD is established positively (hub log + the box's own "floor still unknown"). Its RELEASE is not — no HELD line and no dashboard reason are both absences, the served path logs nothing by design, and the box-side positive needs a restart the venue cannot take before Phase 2. Recorded as deduction plus a pinned test, not a live positive.
1205 lines
63 KiB
Markdown
1205 lines
63 KiB
Markdown
# CAMPAIGN 11 — step-by-step journal
|
||
|
||
Every observable, in the order it was taken. Times are CEST (demo-hp's clock) unless marked UTC.
|
||
|
||
---
|
||
|
||
## Pre-flight — baselines re-read on arrival (2026-08-05)
|
||
|
||
| What | Value | How read |
|
||
|---|---|---|
|
||
| `felhom-controller` `main` | **v0.200.0** @ `a315d623b857` | `git rev-parse HEAD` == `origin/main`, tree clean |
|
||
| `felhom-agent` `main` | **v0.125.0** @ `0404f60e6a7b` | same |
|
||
| `felhom.eu` `main` | hub **v0.96.0** @ `f45b1f6761ff` | same |
|
||
| hub, LIVE | `gitea.dooplex.hu/admin/felhom-hub:0.96.0` | `kubectl -n felhom-system get deploy hub -o jsonpath=…` |
|
||
| ISO, published | `felhom-installer-1.26.1-pve9.2-1.iso` · 1 705 322 496 B · sha `f3cc86d5f0ec…59a6` | `curl -I` on `iso.felhom.eu` + published `.sha256` |
|
||
| ISO, local copy on demo-hp | **sha `f3cc86d5f0ec…59a6`, 1 705 322 496 B — identical** | `sha256sum` on demo-hp |
|
||
| Highest register ID | **R-213** | fresh `grep -rhoE 'R-[0-9]{1,3}' --include=*.md` over all four repos |
|
||
|
||
**All three cited commits match the brief exactly.** Two brief assumptions were stale and reading
|
||
fresh caught both:
|
||
|
||
1. The brief says installer **1.25.0**; the published artifact is **1.26.1** (since 2026-07-31).
|
||
Same class Campaign 10 caught, same file.
|
||
2. The brief's §11 says *"Five scratch customers have accumulated"*. **There are none.** The hub holds
|
||
exactly five customer records and every one is legitimate: `david` (a real prospect),
|
||
`demo-felhom`, `demo-hp`, `drill-r50` (the fixture, BLOCKED), `peti-felhom` (the external tester).
|
||
`c10-soak`, `sess-c`, `sess-d`, `sess-f` are all gone.
|
||
|
||
### The finding that decided the campaign's shape, found before anything was built
|
||
|
||
Day-0 artifact manifest, read from the hub's `/configuration`:
|
||
|
||
| | Vouched | Newest that exists | `main` |
|
||
|---|---|---|---|
|
||
| golden (bakes the controller) | **0.192.0** | **0.192.0** (no newer golden exists) | controller **0.200.0** |
|
||
| agent | **0.120.0** | **0.125.0** (published in Gitea) | **0.125.0** |
|
||
| global controller floor | **0.156.0** | — | — |
|
||
|
||
The controller's auto-update target is **always the floor, never latest**
|
||
(`selfupdate/updater.go:64`). So a box installed from today's published ISO lands on controller
|
||
**0.192.0 and stays there** — which is before R-203's fix (0.197.0), before the credential self-heal
|
||
(0.199.0), and **before the recovery screen this campaign exists to test (0.200.0)**.
|
||
|
||
Both demo boxes run agent **0.125.0** while the manifest vouches **0.120.0** — R-121's drift family,
|
||
live, on the day-0 channel rather than on a box.
|
||
|
||
**Operator ruling, 2026-08-05:** install stock, record what a customer gets today, then raise the
|
||
**per-customer** floor (`configs.go:1229`) to 0.200.0 and let the box converge. Per-customer means the
|
||
other four customers are untouched.
|
||
|
||
### Two further operator rulings taken before building
|
||
|
||
- **ep0 + the Hetzner Storage Box are written to, additively.** Off-site hard-requires the DR tier
|
||
(`hub/internal/web/configs.go:1298`), and the DR tier provisions a PBS namespace + token on **ep0**
|
||
via tenantsync plus a WireGuard peer; enabling off-site provisions a Storage Box sub-account. All
|
||
three are creates on the ordinary customer path — nothing existing is modified or deleted — and
|
||
teardown removes the campaign's own. Campaign 10 avoided this only by running with Tier 3 OFF, which
|
||
is not available to a campaign about off-site recovery. **The brief's I7 wording ("ep0 read-only")
|
||
was therefore relaxed deliberately, by ruling, and is recorded here rather than quietly widened.**
|
||
- **Recovery codes** travel file→file into a `0600` file, are moved by the operator into their own
|
||
store, and are then shredded with a planted-copy positive control.
|
||
|
||
---
|
||
|
||
## Venue
|
||
|
||
| | |
|
||
|---|---|
|
||
| Host | `demo-hp` (HP t740), Tier 0, the designated drill host. Reached **by SSH key, first try** — R-129 stands |
|
||
| VM | **321 `c11-appliance`** — q35/OVMF (`pre-enrolled-keys=0`), 4 cores, 8 GB, `cpu=host` |
|
||
| Disks | `scsi0` 200 G system · `scsi1` 50 G · `scsi2` 50 G, qcow2 |
|
||
| Storage | **`c11-scratch`**, `dir` at **`/mnt/nvme-1tb` — the mount ROOT** (a subdirectory fails the agent's `exactMount` check; Campaign 10 §1) |
|
||
| Box | `c11.felhom.eu` @ **192.168.0.105/24**, gw/DNS 192.168.0.1, on demo-hp's `vmbr0` |
|
||
| Hub customer | **`c11`** "Campaign 11", domain `c11.felhom.eu`, DR tier ON, **off-site initially OFF** |
|
||
| Host id | **`c11-36d660`** |
|
||
| Untouched | `drill-r50` (VM 300, stopped), guest 9201 on both demo boxes, `local-lvm`, `felhom-backup`, every other hub customer |
|
||
|
||
### Harness constraint — the appliance is NOT reachable from DooPlex
|
||
|
||
DooPlex is `192.168.0.180/24` on `eno1` and **demo-hp is a different site with an overlapping
|
||
`192.168.0.0/24`** — `ssh demo-hp` resolves to Tailscale `100.76.96.79`, and DooPlex's ARP for
|
||
`192.168.0.105` is `INCOMPLETE` while demo-hp pings it in 0.4 ms. Every HTTP call to the appliance
|
||
therefore goes **DooPlex → (Tailscale) → demo-hp → vmbr0 → appliance**, through a SOCKS proxy
|
||
(`ssh -N -D 11080 demo-hp`). No server logic is skipped — this is a transport, not a shortcut.
|
||
**Harness, not product.**
|
||
|
||
---
|
||
|
||
## Phase 0
|
||
|
||
### Step 1 — install from the published ISO, through the real day-0 path
|
||
|
||
Driven blind through the **Terminal UI** (`qm monitor screendump` → PNG → read visually; `qm sendkey`
|
||
for input). Campaign 10's three traps, all handled:
|
||
|
||
- **GRUB's default is the graphical entry** and the graphical entry is undrivable (monitor mouse
|
||
events do not move the guest cursor). The menu appears ~6 s in with a 15 s countdown; the first
|
||
attempt sent `down` from a *separate* round trip and missed the window, booting graphical (a grey
|
||
screen with a cursor). Reset and sent `down`+`ret` **inside the same remote command** — the console
|
||
then read *"A Felhom telepítése indul szöveges módban…"*.
|
||
- **The guest keymap defaulted to Hungarian while `sendkey` emits US scancodes.** Changed to
|
||
**U.S. English** before typing anything. **Positive control:** the administrator email was typed
|
||
through the identical code path and rendered **`c11@felhom.eu`** on screen — the `@` is `shift-2` on
|
||
a US layout and `AltGr+V` on a Hungarian one, so a correct `@` proves the mapping for the 24-char
|
||
root password that is masked and cannot be read back. *(Deviation from a real Hungarian install,
|
||
recorded: the installed console keymap is US. The customer journey is the web dashboard, so this is
|
||
cosmetic — but it is a deviation and it is not buried.)*
|
||
- **`--boot` set in its own `qm set` after the disks existed**, and verified from `qm config`
|
||
(`boot: order=scsi0`) before the first boot. `Automatically reboot` was **unchecked** so the reboot
|
||
into the freshly installed system was deliberate rather than a race with the ISO.
|
||
|
||
Only `scsi0` was attached during the install, so the installer offered exactly one target
|
||
(`/dev/sda (QEMU HARDDISK) (200.00 GiB)`) and could not reach a data drive. The two 50 G drives were
|
||
attached afterwards.
|
||
|
||
Summary screen, verbatim: `ext4` · `/dev/sda` · `Europe/Budapest` · `U.S. English` ·
|
||
`c11@felhom.eu` · `nic0` · `c11.felhom.eu` · `192.168.0.105/24` · `192.168.0.1` · `192.168.0.1`.
|
||
|
||
**Day-0 fired on first boot**, unaided. The console showed, in Hungarian:
|
||
|
||
```
|
||
Felhom — a doboz készen áll, és a párosításra vár.
|
||
Párosító kód: RPM-YNF
|
||
Nyisd meg az e-mailben kapott linket, és add meg ezt a kódot és a jelszavadat.
|
||
Ez a képernyő magától frissül — nincs teendő a doboznál, és nyugodtan itt hagyhatod bekapcsolva.
|
||
```
|
||
|
||
and the hub's **Unclaimed appliances** table carried it within a minute: uuid
|
||
`27f35841-0609-410d-af52-76d5c0dbf801` (matching the VM's smbios uuid), pairing code `RPM-YNF`,
|
||
MAC `bc:24:11:c8:4e:ce`, hardware `Standard PC (Q35 + ICH9, 2009)` / `AMD Ryzen Embedded V1756B` /
|
||
7.7 GB, and **three SSH host keys**.
|
||
|
||
Bound to `c11` through the real endpoint (`POST /appliances/20/bind`, HTTP 303) → the row reads
|
||
`bound → Campaign 11`. Host **`c11-36d660`** enrolled ~3 minutes later: `ONLINE`, agent **0.120.0**,
|
||
`PBS wrapper matches vouched`, desired generation 2.
|
||
|
||
**Day-0 provision SUCCESS, 14:12:37 CEST** (`journalctl -u felhom-bootstrap`), 7 minutes after the
|
||
bind:
|
||
|
||
```
|
||
[OK] controller: Up 19 seconds (healthy) (after ~0s)
|
||
[INFO] controller image: gitea.dooplex.hu/admin/felhom-controller:0.192.0
|
||
[OK] Day-0 provision SUCCESS — vmid=9201 host_id=c11-36d660 customer=c11
|
||
golden=local:backup/vzdump-lxc-9100-2026_08_05-14_10_02.tar.zst
|
||
[INFO] root@pam was rotated + vaulted at step 4b — retrieve at hub → host page
|
||
felhom-bootstrap: host-install SUCCESS — writing done-flag, disabling unit, scrubbing secrets
|
||
```
|
||
|
||
`/etc/felhom/.bootstrap-done` written 14:12, unit `inactive (dead)` and disabled, 1 min 35 s CPU.
|
||
**The installer's root password stopped working exactly as Campaign 10's A2 recorded** — the first
|
||
SSH attempt with it was refused, and the hub's vaulted managed credential
|
||
(`POST /hosts/c11-36d660/reveal-recovery-credential`, `set_at 2026-08-05T12:09:34Z`,
|
||
`username root@pam`) authenticated first try. That refusal is the **positive observable** that the
|
||
day-0 credential handover completed.
|
||
|
||
The controller landed on **0.192.0** — the golden's baked version, exactly as predicted from the
|
||
manifest.
|
||
|
||
### Step 1b — the per-customer floor, and the convergence it drives
|
||
|
||
`POST /customers/c11/floor` with `min_controller_version=0.200.0` (HTTP 303). The customers table then
|
||
reads `c11 … 0.192.0 … v0.200.0 (override)` while **every other customer still reads `v0.156.0`** —
|
||
the blast radius is zero, measured rather than asserted.
|
||
|
||
**The box converged in under two minutes**, unaided:
|
||
|
||
```
|
||
gitea.dooplex.hu/admin/felhom-controller:0.200.0 Up 9 seconds (healthy)
|
||
```
|
||
|
||
So the recovery screen IS reachable on a fresh box — but only behind an operator floor override,
|
||
because no golden carrying it exists.
|
||
|
||
### FINDING (candidate R-214) — the physical console never stops asking to be paired
|
||
|
||
Thirty minutes after `Day-0 provision SUCCESS`, with the host `ONLINE` in the hub and the customer
|
||
guest running, the appliance's physical console still read:
|
||
|
||
```
|
||
Felhom — a doboz készen áll, és a párosításra vár.
|
||
Párosító kód: RPM-YNF
|
||
...
|
||
Ez a képernyő magától frissül — nincs teendő a doboznál, és nyugodtan itt hagyhatod bekapcsolva.
|
||
```
|
||
|
||
**Established from source, not inferred.** `print_pairing_banner` is called from exactly one place —
|
||
the pairing wait loop (`scripts/iso/felhom-bootstrap.sh:456`), every 10 cycles (5 min). A census of
|
||
`/dev/console` writers across the entire day-0 path returns **two**, both in `felhom-bootstrap.sh`:
|
||
the banner (`:76`) and the R-59 network-refusal screen (`:252`). **`scripts/felhom-host-install.sh`
|
||
writes to the console not at all.** So after the bind lands, nothing repaints, and the banner is the
|
||
terminal state of the screen.
|
||
|
||
Two things are wrong at once, and the second is the one that makes it a defect rather than a cosmetic
|
||
gap: the screen shows a **stale pairing code** for a box that is already bound, and the banner
|
||
**asserts of itself** that it refreshes. That is `CLAUDE.md`'s "a comment asserting an invariant needs
|
||
a test pinning it" family, in customer-facing Hungarian copy on the one surface a customer looks at
|
||
when they are already unsure.
|
||
|
||
Mitigating, and stated so the severity is not overclaimed: the same banner says
|
||
*"nincs teendő a doboznál"* — the customer's real route is the e-mailed link, not the box.
|
||
|
||
### Step 2 — the claim, and the constraint it exposed
|
||
|
||
The dashboard is served by traefik on 443 in the guest (`192.168.0.207`), not by the controller
|
||
directly; `Host: felhom.c11.felhom.eu` is mandatory. `GET /` → 302 → **`/claim`**, titled
|
||
*"A szerver beállítása — Campaign 11"*, asking for the *beállító kód* and a new password (min. 12
|
||
characters).
|
||
|
||
**The claim code is e-mail-only by design and the campaign has to live with it.** The hub stores a
|
||
hash (`store.RotateClaimCode`, `store.go:1667`) and the resend path states it outright:
|
||
*"the hub stores a hash — no plaintext code exists to display"* (`configs.go:806`). It went first to
|
||
`c11@felhom.eu`, a mailbox that does not exist.
|
||
|
||
**This is not a Phase-0 inconvenience — it decides whether Phase 1 can be claimed as unaided.** Phase 1
|
||
and Phase 3 each need a further claim, and the controller's `--print-reset-code` escape hatch is a
|
||
command line inside the guest. **Operator ruling, 2026-08-05:** point c11's registered address at the
|
||
operator's own mailbox and have them relay each code — which is precisely what a real customer does,
|
||
and what R-119 already records as *"the single human step"*.
|
||
|
||
`POST /configs/c11/edit` with `email=<the operator's address>` (HTTP 303), then
|
||
`POST /configs/c11/claim-resend` (HTTP 303) → *"Nyitott — kód kiküldve just now, generation 2"*.
|
||
|
||
Claim driven through the real customer form (`POST /claim`, code + a 20-char password, HTTP 302 → `/`),
|
||
landing on **`/launcher`** ("Indítópult"), controller **0.200.0**. The box is claimed and the
|
||
household password is customer-owned.
|
||
|
||
### F9, taken here rather than in Phase 2 — and it FAILS
|
||
|
||
F9's stated right answer for a box that never had off-site backups is **"no screen, no entry point,
|
||
nothing"**. `c11` is exactly that box: installed today, off-site never configured. **Two of the three
|
||
hold. The third does not.**
|
||
|
||
| | Expected | Observed |
|
||
|---|---|---|
|
||
| full-page interruption | none | **none** — `GET /` → `/launcher`, not `/recovery` ✅ |
|
||
| entry point in the backups area | none | **none** — `/backups/remote` reads *"Még nincs beállítva távoli mentési cél."* and contains **zero** `/recovery` links ✅ |
|
||
| the page itself | should not exist for this box | **`GET /recovery` → HTTP 200, and it asserts the machine was reinstalled and that sealed backups exist** ❌ |
|
||
|
||
Verbatim, from the rendered page on a box with no backups of any kind:
|
||
|
||
> *„Ezt a gépet újratelepítették. A korábbi, **házon kívüli mentéseid megvannak** — a Felhom központi
|
||
> rendszere őriz hozzájuk egy lezárt csomagot."*
|
||
|
||
**Ground truth, measured, not inferred.** Hub SQLite snapshot taken WITH its `-wal` and `-shm`
|
||
(`PRAGMA integrity_check` → `ok`; freshness by positive observable — newest `host_reports.received_at`
|
||
`2026-08-05 12:55:05` against `datetime('now')` `12:55:39`, **34 s old**):
|
||
|
||
```sql
|
||
SELECT host_id, length(identity_blob) FROM host_escrow;
|
||
demo-felhom-8363b5|572
|
||
demo-hp-bb76ea|572
|
||
-- c11-36d660: NO ROW AT ALL
|
||
```
|
||
|
||
So `identity_blob_present` is false in every ACK c11 has ever received, `GetHubEscrowIdentityPresent()`
|
||
is false, and **`OffsiteRecoveryOffer()` correctly returns false. The predicate is right. The page
|
||
never asks it.**
|
||
|
||
**The mechanism, from source.**
|
||
|
||
- `recoveryPageHandler` (`internal/web/recovery_handlers.go:65-67`) is three lines and calls
|
||
`renderRecovery` with **no gate**.
|
||
- `renderRecovery` sets `data["Offer"] = s.recoveryOffer()` (`:76`) and **`recovery.html` never reads
|
||
it** — `grep -c '\.Offer'` over the template returns **0**. The template's only top-level branch is
|
||
`{{if .Unlocked}}`, so every non-unlocked render prints the headline above.
|
||
- **Its POST sibling IS gated**: `recoveryUnlockHandler:141` opens with
|
||
`if !s.recoveryOffer() { redirect to /backups/remote }`. The action is protected; the page is not.
|
||
- **And the other template gets it right**, in the same release, for the same claim:
|
||
`backups_remote.html:11` wraps the entry point in `{{if .RecoveryOffer}}` — which is why no entry
|
||
point appeared. One gated, one not, same sentence.
|
||
|
||
**Why the green suite missed it.** `TestRecovery_B_DoesNotAppearForAnyoneElse/never had off-site
|
||
backups (no hub package)` (`recovery_test.go:200-208`) asserts **`f.s.recoveryOffer()`** and issues no
|
||
`GET /recovery`. It pins the mechanism; the consequence is unpinned. That is `CLAUDE.md`'s own rule —
|
||
*prefer a test that asserts the consequence over one that asserts the mechanism* — and the seam-wiring
|
||
rule's inverse case: usually the gate is missing, here the gate exists and the render does not consult
|
||
it.
|
||
|
||
**Severity, stated without inflation.** There is no link to `/recovery` on such a box, so a customer
|
||
reaches it only by typing the URL or following a stale bookmark. But what they are then told is false
|
||
in two ways at once — *your machine was reinstalled* and *your off-site backups exist* — on the one
|
||
screen whose entire purpose is to be believed about backups. It is I6: an absence reported as a fact.
|
||
→ **R-215**
|
||
|
||
### Ground state for teardown — ep0, captured BEFORE anything else
|
||
|
||
The DR tier provisioned itself on ep0 the moment the customer was saved with `dr_tier=on` and the WG
|
||
peer registered. Recorded so teardown removes exactly the campaign's own and nothing else:
|
||
|
||
```
|
||
namespaces: c11 demo-felhom demo-hp ← c11 is the campaign's
|
||
tokens: felhom@pbs!c11 (+ the two demo) ← c11 is the campaign's
|
||
acl: felhom@pbs → /datastore/felhom-offsite/c11 DatastoreBackup
|
||
felhom@pbs!c11 → /datastore/felhom-offsite/c11 DatastoreBackup
|
||
df /mnt/pbs-datastore: 98G total, 9.9G used, 11%
|
||
```
|
||
|
||
### Step 2b — the drives, through the real endpoints
|
||
|
||
Both 50 G drives appeared as raw candidates (`GET /api/disks/candidates` → `initialize: [/dev/sdb, /dev/sdc]`)
|
||
and were enrolled through the customer flow (`POST /api/storage/init`, polled on
|
||
`/api/storage/init/status` to `phase: done` — Campaign 10's F6 detached-job path):
|
||
|
||
| Drive | Device | Stable path | Role |
|
||
|---|---|---|---|
|
||
| `adatok` | `/dev/sdb` | `/mnt/felhom-drives/adatok` | app data, `set_default` |
|
||
| `mentes` | `/dev/sdc` | `/mnt/felhom-drives/mentes` | **backup target** |
|
||
|
||
The backup-target offer behaved exactly as Campaign 10 recorded — offered, accepted through
|
||
`POST /api/storage/backup-target/assign`, `restart_required: true`, and the state flipped only after
|
||
the agent restart it asked for:
|
||
|
||
```
|
||
before: {"degraded":true,"target":"local","message":"A rendszermentés jelenleg ugyanazon a lemezen van…"}
|
||
after: {"degraded":false,"known":true,"label":"Mentes","target":"felhom-backup"}
|
||
```
|
||
|
||
That healthy read is this campaign's **I5/I6 baseline**.
|
||
|
||
### Step 3 — the app and the three sentinels
|
||
|
||
**`calibre-web`** — the drill's own fixture app, and the only sensible choice: it declares
|
||
`userdata: media/books class: mandatory` (`app-catalog-felhom.eu/templates/calibre-web/.felhom.yml:21-23`),
|
||
it is single-container, and its file leg is a bind rather than a named volume (the 2026-08-04 drill
|
||
established that a named-volume-only app has **no restorable file leg** — the off-site tier tars it and
|
||
the customer restore flow never unpacks it).
|
||
|
||
Deployed through the real API (`POST /api/stacks/calibre-web/deploy`, HTTP 202) with
|
||
`HDD_PATH=/mnt/felhom-drives/adatok` — a real enrolled drive, which is the normal customer shape and
|
||
deliberately **not** R-203's `/mnt/sys_drive` case. Reached `running` / `Up About a minute (healthy)`.
|
||
|
||
Three sentinels, written into the mandatory path, hashed after `sync`:
|
||
|
||
| # | File | Bytes | sha256 |
|
||
|---|---|---|---|
|
||
| A | `C11-SENTINEL-A.txt` | 54 | `beb9175d52325bf583039beb6db1ecc80ea0781720e262c83230ed46a166df27` |
|
||
| B | `C11-őrszem-ékezetes-árvíztűrő.txt` | 73 | `7c8cb0ad040856440cc96bb1ad8f336d04f458138d85e6904ee759e690743a87` |
|
||
| C | `C11-SENTINEL-C-12MB.bin` | 12 582 912 | `e012e76f3502de7c608d5b35f67b89c2404edad49eea92a1fc35a1bf02cbb9c0` |
|
||
|
||
*(Method, stated: the sentinels were written on the box, not through the dashboard. Phase 1's
|
||
no-command-line rule governs Phase 1; fixture construction is not part of what it measures.)*
|
||
|
||
### Step 4 — the off-site tier, and what enabling it actually did
|
||
|
||
`POST /configs/c11/edit` with `offsite_enabled / shared / 50 GB / bx11`. The hub provisioned a Storage
|
||
Box sub-account and the box picked up the target:
|
||
|
||
```
|
||
[offsite] shared provisioned for c11 (subaccount 284166, user u629488-sub4)
|
||
[offsite] host-key scan u629488-sub4.your-storagebox.de failed (attempt 1/6 … 5/6):
|
||
lookup … no such host ← DNS for a just-created sub-account
|
||
[INFO] Customer config updated: c11 ← succeeded on attempt 6
|
||
```
|
||
|
||
**Worth recording rather than treating as noise:** first-provision blocks on DNS propagation for the
|
||
new sub-account hostname and took **over 60 s of backoff** (2+4+8+16+30). The retry ladder is doing
|
||
exactly its job — but the HTTP request outlived a 120 s client timeout while the hub completed the work
|
||
anyway, so a client that gives up here would be looking at a *succeeded* provision it thinks failed.
|
||
No R minted: the hub's own copy already warns *"offsite provisioning can take up to a minute"*.
|
||
|
||
Target on the box: `u629488-sub4@u629488-sub4.your-storagebox.de:/home/felhom-repo`, quota 50 GB.
|
||
`calibre-web` toggled on for off-site (`POST /backup/offbox/toggle`).
|
||
|
||
### Step 5 — the escrow ceremony, and R
|
||
|
||
Preflight (`GET /api/escrow/preflight`) — **six of six green**, including the one that matters:
|
||
|
||
```
|
||
pbs_storage_id ok (felhom-pbs) · dr_tier ok (DR tier applied) · age_binary ok (/usr/bin/age)
|
||
hub_upload ok · staged_secret ok (staged secret present) · sudo_grant ok
|
||
```
|
||
|
||
Ceremony driven through the customer wizard's own endpoints (`/api/escrow/start` with a password
|
||
re-auth → `/api/escrow/status` → `/api/escrow/claim`):
|
||
|
||
```
|
||
phase: done · restic_pw_sealed: TRUE · uploaded: true · entropy_bits: 129.24
|
||
key_fingerprint: 3c:f7:76:82:e3:0c:85:d7:…:1d:04:43:b8
|
||
claimable: true → claimed: true, claim_expires_in_sec 585 → 0
|
||
```
|
||
|
||
**`restic_pw_sealed: true` is the field this whole campaign rests on** — the off-site repository
|
||
password is inside the sealed bundle.
|
||
|
||
**R was captured file→file** into a `0600` file and never rendered. Its shape was verified without
|
||
printing it: **ten hyphen-separated alphabetic words, 85 characters**, consistent with the ~129-bit EFF
|
||
phrase the ceremony reports.
|
||
|
||
Hub-side, immediately after (SQLite snapshot with `-wal`, freshness 9 s):
|
||
|
||
```
|
||
host_escrow(c11-36d660): blob 383 B · identity_blob 572 B · stale_at NULL
|
||
restic_pw_sha256 626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303
|
||
created_at 2026-08-05T13:11:11Z
|
||
```
|
||
|
||
The box then read *„A helyreállítási kód letétbe helyezve."* and `escrow_state: escrowed`.
|
||
|
||
### Step 5b — the off-site backup, and the sentinels listed BY NAME
|
||
|
||
`POST /backup/offbox/run` → `Fut…` → **`✓ Rendben`**, 1 m 6 s.
|
||
|
||
**The gate is not the green tick.** `restic snapshots` + `restic ls -l latest`, run against the
|
||
repository with its own credentials:
|
||
|
||
```
|
||
snapshot f3d9cd67d539c00359f0454ea7a782f45beac8406c55275bee5a6886afa8d791 (short f3d9cd67)
|
||
time 2026-08-05T13:13:32.326542714Z
|
||
hostname c11 tags [felhom-offbox, calibre-web]
|
||
paths /mnt/felhom-drives/adatok/backups/primary/calibre-web
|
||
/mnt/felhom-drives/adatok/userdata/media/books
|
||
|
||
-rw-r--r-- 1000 1000 54 …/userdata/media/books/C11-SENTINEL-A.txt
|
||
-rw-r--r-- 1000 1000 12582912 …/userdata/media/books/C11-SENTINEL-C-12MB.bin
|
||
-rw-r--r-- 1000 1000 73 …/userdata/media/books/C11-őrszem-ékezetes-árvíztűrő.txt
|
||
-rw-r--r-- 1000 1000 413696 …/userdata/media/books/metadata.db
|
||
+ the recovery unit: compose/{.felhom.yml,app.yaml,docker-compose.yml}, manifest.json,
|
||
volume-dumps/calibre-web_calibre_web_config.tar (389 632 B)
|
||
```
|
||
|
||
**All three sentinels are in the snapshot, by name, at the right sizes — and the Hungarian filename
|
||
survived the whole chain into restic byte-intact.** The mojibake gate the brief warned about does not
|
||
fire here.
|
||
|
||
**Independent live confirmation of R-203's fix, on a shape the drill never tested.** The capture set's
|
||
path is `/mnt/felhom-drives/adatok/userdata/media/books` and the app's live bind is the *same*
|
||
directory (`docker inspect`: `bind /mnt/felhom-drives/adatok/userdata/media/books -> /calibre-library`).
|
||
On a non-system drive `NamespaceRoot` correctly omits the `felhom-data` segment for both, so the two
|
||
roots agree. *(The import leg still resolves under `/mnt/sys_drive/felhom-data/userdata/import/calibre`
|
||
— a different drive AND the `felhom-data` root — which is correct for an `excluded`-class path but is
|
||
the same asymmetry R-203 was filed about. It is harmless here because nothing mandatory lives there.)*
|
||
|
||
### Step 6 — the pre-destruction truth
|
||
|
||
Box's own state (`settings.json`, secrets stripped):
|
||
|
||
```
|
||
offbox: enabled true · escrow_state "escrowed"
|
||
last_run/last_success 2026-08-05T13:14:26Z · last_status "ok" · last_duration 1m6s
|
||
snapshot_count 1 · repo_size_bytes 12 611 522 ("12.0 MB") · quota_gb 50
|
||
hub_escrow_identity_present: true
|
||
```
|
||
|
||
Hub's state: `identity_blob` **572 B**, `restic_pw_sha256` `626e4246…c303`, `stale_at` NULL,
|
||
Key Escrow **present**, DR Recipe **present**.
|
||
|
||
### I5 at the Phase 0 → Phase 1 boundary — investigated, and DISPROVED
|
||
|
||
At 15:20 CEST the hub's operator view read **`Snapshots 0 · Repo Size 0 MB · Integrity Unknown`**
|
||
while the box held 1 snapshot and 12 611 522 B. That is exactly the shape I5 exists to catch, and it
|
||
would have been easy to file.
|
||
|
||
**It was cadence, not disagreement.** c11's legacy controller report last landed at 13:11:43 UTC —
|
||
**three minutes before** the backup finished at 13:14:26. Rather than judge from a single stale read,
|
||
a poll ran until either convergence or a full interval elapsed. It converged at **13:23:10 UTC**, 8 m
|
||
44 s after the backup:
|
||
|
||
```
|
||
{"enabled":true,"escrow_state":"escrowed","last_run":"2026-08-05T13:14:26Z","last_status":"ok",
|
||
"last_success":"2026-08-05T13:14:26Z","snapshot_count":1,"repo_size_bytes":12611522,"quota_gb":50}
|
||
```
|
||
|
||
Byte-for-byte what the box's own `settings.json` holds. **I5 PASS.** No `R-n` minted.
|
||
|
||
The general lesson, and the reason this is written down rather than dropped: **on this fleet the hub's
|
||
view lags the box's by up to one report interval, so "the two views disagree" is only a finding after
|
||
an interval has passed.** A verdict taken from one read at the wrong instant is Campaign 10 §5.3's
|
||
mistake — reading the two halves at different instants of a transition.
|
||
|
||
---
|
||
|
||
## Phase 1 — the clean journey, unaided
|
||
|
||
**Method note, stated before the result.** There is no browser on DooPlex, so every customer action is
|
||
driven at the endpoint the UI itself invokes (`felhom-controller/CLAUDE.md`'s standard method), through
|
||
the SOCKS transport described in the venue section. **The rule Phase 1 actually tests — no command
|
||
line INSIDE THE GUEST — is enforced literally**: after the destruction, nothing reaches guest 9201
|
||
except HTTP requests a browser could have made. Actions on the appliance *host* (the Proxmox layer) are
|
||
the customer's installer, not the guest.
|
||
|
||
**No rollback archive was taken, deliberately.** `qm snapshot 321` refused — *"snapshot feature is not
|
||
available"*, because the OVMF `efidisk0` is `raw` and a raw disk blocks qcow2 snapshots. Rather than
|
||
convert a disk mid-campaign, the run proceeded without one: the venue is disposable and its loss is
|
||
authorised, and **the only failure that matters — the off-site repository failing to open — is not
|
||
protected by a local archive anyway.**
|
||
|
||
### Step 1 — the destruction, 2026-08-05T13:35:19Z
|
||
|
||
Recorded immediately before, so the loss is measured rather than assumed:
|
||
|
||
```
|
||
rootfs local-lvm:vm-9201-disk-0 32G · mp0 local-lvm:vm-9201-disk-1 70G (/var/lib/felhom)
|
||
containers: felhom-controller 0.200.0 · calibre-web · filebrowser · traefik
|
||
/mnt/felhom-drives/adatok/userdata/media/books: all three sentinels + metadata.db (13M total)
|
||
/mnt/felhom-drives/adatok/backups/primary/calibre-web: the Tier-1 recovery unit
|
||
```
|
||
|
||
Then:
|
||
|
||
```
|
||
pct stop 9201 ; pct destroy 9201 --purge
|
||
Logical volume "vm-9201-disk-0" successfully removed.
|
||
Logical volume "vm-9201-disk-1" successfully removed.
|
||
purging CT 9201 from related configurations..
|
||
pct list → (empty)
|
||
rm -rf /mnt/felhom-drives/{adatok,mentes}/* ← a hardware loss takes the drives too
|
||
du -sh → 4.0K / 4.0K
|
||
```
|
||
|
||
**The guest, its 102 GB of volumes, both enrolled drives and the entire local backup tier are gone.
|
||
The customer's three sentinels now exist in exactly one place on Earth: the off-site restic repository
|
||
at `u629488-sub4.your-storagebox.de:/home/felhom-repo`.**
|
||
|
||
This is a deliberately *harder* destruction than the 2026-08-04 night drill, which wiped only the
|
||
controller's data volume. The brief asks for "the guest and its data volume, as a hardware loss would",
|
||
and being attended rather than unattended makes the guest-destroy + reprovision chain safe to walk.
|
||
**The host identity `c11-36d660` survives** — which is the point: the hub's sealed package is keyed to
|
||
the HOST, so destroying the host as well would make recovery impossible by construction rather than by
|
||
defect.
|
||
|
||
### Step 2 — reinstall through the same day-0 path, started 13:36:07Z
|
||
|
||
The hub's own Setup Command, `felhom-host-install.sh` **v1.25.0** fetched live from
|
||
`https://felhom.eu/scripts/` — the documented path that *"does what the ISO's package would have
|
||
done"*. The retrieval passphrase went in via `--passphrase-file` (a `0600` file), never on argv.
|
||
|
||
Completed **13:38:06Z** — `Day-0 provision SUCCESS — vmid=9201 host_id=c11-36d660 customer=c11`,
|
||
guest back at **192.168.0.227**, controller **0.192.0** healthy. **`root@pam` rotated + re-vaulted
|
||
again.** Then, unprompted, the per-customer floor did its job:
|
||
|
||
```
|
||
13:35:19Z destroy
|
||
13:36:07Z reinstall start
|
||
13:38:06Z controller 0.192.0 healthy ← the GOLDEN's version
|
||
13:38:42Z controller 0.200.0 healthy ← floor-driven self-update, 36 s later
|
||
```
|
||
|
||
**The assertion that keeps recovery possible — PASSED.** Read from the hub with the `-wal`, after the
|
||
destruction:
|
||
|
||
```
|
||
host_escrow(c11-36d660): identity_blob 572 B · restic_pw_sha256 626e424670248db3… (UNCHANGED)
|
||
created/updated 2026-08-05 13:11:11 (UNCHANGED) · stale_at NULL
|
||
host_escrow_superseded: rows 3 and 4 only — the two demo boxes. NOTHING re-escrowed itself.
|
||
```
|
||
|
||
**And the rebuilt box already knew.** Its fresh `settings.json`, before any human touched it:
|
||
`hub_escrow_identity_present: true`, no `offbox` key, no `claimed` key — precisely shape (a) of
|
||
`OffsiteRecoveryOffer()`.
|
||
|
||
### Steps 3–4 — claim, and the first thing the customer sees
|
||
|
||
The rebuild left the box unclaimed, so the hub rotated a **reset code (generation 3)** to the
|
||
registered address; the operator relayed it, as R-119 records is the one human step.
|
||
|
||
```
|
||
13:42:41.101Z POST /claim (code + a new password) → 302 → /
|
||
GET / → 200 https://192.168.0.227/recovery ← NOT /launcher
|
||
```
|
||
|
||
> **The recovery screen appears without being sought.** The customer's very first landing after
|
||
> claiming their rebuilt machine is the screen that tells them their backups survived. ✅
|
||
|
||
### Step 5 — the screen, read as a customer
|
||
|
||
It says what is waiting, when it was sealed, and that nobody can replace a lost code:
|
||
|
||
> *„Ezt a gépet újratelepítették. A korábbi, **házon kívüli mentéseid megvannak** — a Felhom központi
|
||
> rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-05T13:11:11Z** zártunk le."*
|
||
>
|
||
> *„**A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az
|
||
> üzemeltető. Ez szándékos… Ha a kód elveszett, a korábbi mentések nem nyithatók meg többé."*
|
||
>
|
||
> *„Ebben a lépésben **semmit nem állítunk vissza és semmi nem változik**."*
|
||
|
||
All three of step 5's questions answered, in plain Hungarian. The set-aside option was correctly
|
||
**not** offered, with its reason stated rather than the button merely hidden: *„ez a lehetőség akkor
|
||
válik elérhetővé, ha a gép már újra kapcsolódott a házon kívüli tárhelyhez"* — the `CanSetAside`
|
||
gate (`ResetOrphanedRepo` refuses unless orphaned) working, and explaining itself.
|
||
|
||
**One copy defect on this screen:** the seal date renders as the raw RFC3339 UTC string
|
||
**`2026-08-05T13:11:11Z`** to a Hungarian household. `recoverySealedAt()` is documented as *"the human
|
||
date the hub says the sealed package was created"*; it is not humanised anywhere. Minor beside what
|
||
follows, and recorded rather than dropped.
|
||
|
||
### Step 6 — the recovery code was entered, and REFUSED. **Phase 1 FAILS HERE.**
|
||
|
||
```
|
||
13:43:30.800Z POST /recovery/unlock (the code exactly as issued)
|
||
13:43:30.953Z HTTP 200 in 0.134 s
|
||
```
|
||
|
||
> *„A megadott helyreállítási kódot nem fogadtuk el. Ellenőrizd, hogy mind a tíz szót pontosan,
|
||
> szóközökkel elválasztva írtad be — a kis- és nagybetűk nem számítanak."*
|
||
|
||
**0.134 s is the tell.** `age`'s scrypt is deliberately slow; a real unseal cannot finish that fast.
|
||
The controller's own log names the truth:
|
||
|
||
```
|
||
[WARN] [web] recovery: unlock failed: agentapi: POST /escrow/recover-offsite-password: HTTP 404
|
||
```
|
||
|
||
**The agent does not have the endpoint.** `POST /escrow/recover-offsite-password` ships in agent
|
||
**v0.125.0** (*"the agent opens the sealed bundle and returns one field"*, 2026-08-04, R-199 links
|
||
7–8). This box runs the **vouched day-0 agent, 0.120.0**.
|
||
|
||
**So the machine told the customer their recovery code was wrong, when the code was perfect and the
|
||
machine simply could not ask.** That is the single worst outcome the brief names — an unrelated
|
||
failure reported as a wrong code — and it happened in the CLEAN journey, not under an injected fault.
|
||
|
||
→ **R-216**
|
||
|
||
### The counterfactual — measured, so R-216 is a fact and not an inference
|
||
|
||
Agent **0.125.0** installed on the campaign box **only** (published artifact, sha
|
||
`f7d8339b53d92a6c45be7eaf189469a041b6b00b758a64511c0479beae7016b3`, verified on the box against the
|
||
hub's own vouched value; **no global manifest change, so no other customer moved**). The *identical*
|
||
code was then re-entered:
|
||
|
||
| | agent 0.120.0 | agent 0.125.0 |
|
||
|---|---|---|
|
||
| verdict | *„A megadott helyreállítási kódot nem fogadtuk el"* | **„A mentéseid elérhetők — A kulcs visszakerült"** |
|
||
| elapsed | **0.134 s** | **1.113 s** — the scrypt cost, visible |
|
||
| outcome | — | `recovery: the offsite repository key was recovered and placed (outcome=installed)` |
|
||
|
||
Nothing else changed. **The code was always correct; the machine could not ask, and blamed the customer.**
|
||
|
||
### R-217 — the same screen then asserted three things it could not know
|
||
|
||
The successful unlock rendered the honest error *and* a contradiction beneath it:
|
||
|
||
> *„A kulcs visszakerült. A gép még most kapcsolódik újra a házon kívüli tárhelyhez…"* **and**
|
||
> *„A tároló megnyílt, és van benne tartalom, de nem tudtuk alkalmazásokhoz rendelni."*
|
||
|
||
The store did **not** open, its content is **unknown**, and attribution is **not** the problem.
|
||
`recovery_handlers.go:184` builds `backup.OffsiteInventory{}` on the failure path — `Empty` defaults to
|
||
**false** — so the template's `InvUntagged` (`!Empty && len(Apps)==0`) fires. The type's own doc
|
||
comment names precisely this hazard:
|
||
|
||
> *"Empty is true when the repository opened cleanly and holds no snapshots… it is named rather than
|
||
> inferred from `len(Apps)==0`, **which is also what a failed read looks like**."*
|
||
|
||
The field built to prevent it was defaulted past. → **R-217**
|
||
|
||
### R-218 — succeeding at recovery is what stops the machine asking for what it still needs
|
||
|
||
**The store never opened.** 47 minutes after the unlock, against a screen promising *"néhány perc"*.
|
||
Not slow — **deadlocked**, and each half is independently sufficient.
|
||
|
||
**Measured, from the hub's own report stream:**
|
||
|
||
```
|
||
13:38:06 (offsite absent)
|
||
13:38:46 (offsite absent)
|
||
13:39:54 {"state":"needs_credential", …} ← the box asking. v0.199.0 working.
|
||
13:42:43 {"state":"needs_credential", …} ← second report = the debounce satisfied
|
||
13:47:03 hub: "offsiteheal: re-staged the stored one-time offsite secret for customer c11
|
||
(declared needs_credential across 2 reports) — the box re-consumes on its next cycle"
|
||
13:47:35 THE CUSTOMER'S RECOVERY UNLOCK PLACES THE REPOSITORY PASSWORD
|
||
13:53:42 (offsite absent) ← the box has stopped asking
|
||
14:08:42 (offsite absent)
|
||
14:23:42 (offsite absent)
|
||
```
|
||
|
||
**Half 1 — the box stops asking the moment recovery succeeds.** `needsOffsiteCredential`
|
||
(`backup/offbox.go`) short-circuits on the repository password:
|
||
|
||
```go
|
||
if _, ok := m.OffboxRepoPasswordHash(); ok {
|
||
return false // we still hold our repository password: not a fresh data area
|
||
}
|
||
```
|
||
|
||
Placing that password is the recovery screen's **entire job**. Doing it makes
|
||
`OffboxReportStatus()` return nil, the `offsite` object vanishes from the report, and
|
||
`offsiteheal`'s reconciler — which acts on that one declaration and nothing else
|
||
(`reconciler.go:39-52`) — can never fire again. The box's state now:
|
||
|
||
```
|
||
/data/offbox/ repo_password 64 B (13:47) ← the recovered key, present
|
||
no known_hosts, no ssh_key, no applied_marker
|
||
settings.json offbox key ABSENT ← no coordinates, no transport credential
|
||
hub_escrow_identity_present true
|
||
```
|
||
|
||
**A correct key, and nowhere to use it.**
|
||
|
||
**Half 2 — the re-stage nudges nothing.** `Reconciler.heal` re-arms the secret, logs *"the box
|
||
re-consumes on its next cycle"*, emits an event, and returns. A census of the whole `offsiteheal`
|
||
package for `poke|bumpIntent|SetHostDesired|generation` returns **zero hits** — no poke, no
|
||
generation bump, nothing that would make the box re-pull. And the controller's apply-bridge retries
|
||
only *"on next config refresh/restart"* (its own 13:38:51 line). Last `Config downloaded for customer
|
||
c11`: **13:38:00** — **nine minutes BEFORE the re-stage**, and none since.
|
||
|
||
So the hub holds an **unconsumed** credential (`one_time_secrets`: one c11 row, created
|
||
`2026-08-05 13:07:03`, `consumed_at` NULL after the re-arm) that the box has no reason to come and
|
||
collect, while the box no longer says it needs one.
|
||
|
||
> **The remedy fires correctly, 32 seconds before the customer succeeds — and the customer's success
|
||
> is what closes the door on it.** Two shipped features, each correct alone, cancel each other on the
|
||
> exact path they were both built for.
|
||
|
||
→ **R-218**. This is the direct cause of Phase 1 not reaching an open store.
|
||
|
||
### The intervention (deviation 3, operator-authorised) — and what it proved
|
||
|
||
`systemctl restart felhom-controller-bootstrap.service` at **14:31:33Z** — the apply-bridge's own
|
||
documented retry trigger, not a code change. It worked in **17 seconds**:
|
||
|
||
```
|
||
14:31:46 [offsite-apply] settle-gate: GO — at/above floor 0.200.0
|
||
14:31:50 [offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
|
||
```
|
||
|
||
**Which confirms R-218 exactly**: nothing was wrong with the credential, the target or the key — the
|
||
only thing missing was anything at all to trigger a retry.
|
||
|
||
**And the recovered key survived the re-apply, byte-exact — three independent sources agreeing:**
|
||
|
||
| source | sha256 |
|
||
|---|---|
|
||
| hub `restic_pw_sha256`, recorded BEFORE the destruction | `626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303` |
|
||
| the file on the rebuilt box | `626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303` |
|
||
| mtime of that file | **13:47** — the unlock. Untouched by the 14:31 re-apply |
|
||
|
||
The re-apply wrote `known_hosts`, `ssh_key` and `applied_marker` at 14:31 and **left
|
||
`repo_password` alone**, exactly as `WriteOffboxSecrets` documents. R-193's orphaning mechanism did
|
||
**not** fire.
|
||
|
||
### R-219 — the listing the screen promises can never render on the shape it exists for
|
||
|
||
The screen's headline promise is explicit:
|
||
|
||
> *„Ha megadod a kódot, **feloldjuk a mentéseid zárolását és megmutatjuk, mi van bennük** — melyik
|
||
> alkalmazás, mikorról, mekkora."*
|
||
|
||
**On shape (a) — the pristine rebuilt box, the shape `OffsiteRecoveryOffer`'s own header calls the
|
||
customer's actual state — that listing cannot render. Ever.** By construction:
|
||
|
||
1. `ApplyOffsiteTarget` → `WriteOffboxSecrets` **auto-generates the repository password whenever none
|
||
is present** (`offbox.go:400-403`). So **a target cannot exist without a repository password.**
|
||
2. Shape (a) is defined by *no repository password*. Contrapositive of (1): **no password ⇒ no target.**
|
||
3. `OffsiteInventoryList` needs a target → returns `ErrNoOffsiteTarget` → the handler takes the
|
||
failure branch every time.
|
||
|
||
Measured live at 13:47:35: the unlock succeeded and the customer got
|
||
*„a mentéseid listája néhány perc múlva jelenik meg"* — plus R-217's false *„A tároló megnyílt, és van
|
||
benne tartalom"*. **Not once did the promised listing appear.**
|
||
|
||
**And there is no second chance.** After the unlock places the password, `recoveryOffer()` evaluates
|
||
`OffboxRepoPasswordHash()` → ok, then `OffboxOrphaned()` → `t != nil && …` → **false** with no target.
|
||
So the offer retires itself, and `recoveryUnlockHandler`'s opening guard redirects. Measured, after
|
||
the re-apply:
|
||
|
||
```
|
||
POST /recovery/unlock → 302 /backups/remote ← the listing is permanently unreachable
|
||
GET / → 200 /launcher ← the screen has stepped aside
|
||
```
|
||
|
||
The customer's single opportunity to see what is in their backups is the unlock response itself — and
|
||
on shape (a) that response is structurally guaranteed not to contain it. → **R-219**
|
||
|
||
*(The screen stepping aside once the job is done is CORRECT and is recorded as a pass: `GET /` went
|
||
back to `/launcher`. The defect is that the one thing it promised was never delivered before it went.)*
|
||
|
||
### Step 8 blocked — the restore page is empty exactly when it is needed
|
||
|
||
`/backups/restore`, immediately after a successful recovery on a rebuilt box:
|
||
|
||
```
|
||
„Nincs távoli mentésre jelölt alkalmazás — a kijelölés a … oldalon történik."
|
||
„Nincs ellenőrző másolat a meghajtón."
|
||
„Nincs telepített alkalmazás."
|
||
```
|
||
|
||
The post-unlock screen's own button — *„Tovább a visszaállításhoz"* — lands here. **The restore flow is
|
||
keyed on locally installed, off-site-toggled apps, and a rebuilt box by definition has none**; the
|
||
snapshot in the repository (`calibre-web`, with its recovery unit) is invisible to it. This is R-213's
|
||
territory ("putting files back is not built"), but sharper than *not built*: the entry point exists,
|
||
is reachable, and is empty precisely on the path the recovery screen hands the customer to.
|
||
|
||
### R-220 — after a rebuild the customer's own drives cannot be re-enrolled, and the refusal names an impossible action
|
||
|
||
Redeploying the app was the next step, so the app needs its data drive. The deploy refused, in
|
||
Hungarian, through the real endpoint:
|
||
|
||
> *„A megadott tárhely nem azonosítható regisztrált meghajtóként, ezért alkalmazás
|
||
> adatkönyvtáraként nem használható. **Válasszon a listából csatlakoztatott meghajtót.**"*
|
||
|
||
**The list is empty.** `GET /api/disks/candidates` → `initialize: []`, `attach: []`, while the
|
||
customer's two drives sit mounted and visible in `/api/disks`. The wizard at `/storage/attach`
|
||
renders *„Betöltés…"* over nothing.
|
||
|
||
**The seam, named:** `felhom-agent/internal/storage/claim.go:84-85`
|
||
|
||
```go
|
||
if n.mountpoint != "" && !underFelhomDrives(n.mountpoint) {
|
||
return false, "device is mounted at " + n.mountpoint + " (" + n.name + ")"
|
||
}
|
||
```
|
||
|
||
A device mounted anywhere outside `/mnt/felhom-drives` is *claimed* and never offered as a candidate.
|
||
After the rebuild, `/dev/sdb` was mounted at **both** `/mnt/felhom-drives/adatok` (the bind the guest
|
||
sees) **and `/mnt/adatok`** — the raw enrol mount that `disks.go:606` leaves behind **by design**,
|
||
which `eject` also deliberately preserves. The controller's registry, however, lives in `settings.json`
|
||
and went with the wipe.
|
||
|
||
> **The drive is invisible to the re-enrolment wizard because of the mount that enrolling it created.**
|
||
|
||
**RED-PROOF, both directions.** Unmounting only `/mnt/adatok` and `/mnt/mentes` — changing nothing
|
||
else — flipped the scan immediately:
|
||
|
||
```
|
||
before: initialize: [] attach: []
|
||
after : initialize: [/dev/sdb, /dev/sdc] attach: [(/dev/sdb, ext4), (/dev/sdc, ext4)]
|
||
```
|
||
|
||
Both drives then re-enrolled through the real endpoint with their data intact
|
||
(`POST /api/storage/attach` → `{"registered":true,"where":"/mnt/felhom-drives/adatok"}`), and the
|
||
deploy that had returned **409** returned **202**.
|
||
|
||
**Why this is a defect and not an edge case.** (a) It is on the rebuild path, which is the whole
|
||
subject of the recovery feature. (b) It breaches **I3** — the refusal names an action ("choose a
|
||
connected drive from the list") that the customer cannot perform, because the list it points at is
|
||
empty for a reason unrelated to anything they did. (c) **Campaign 10 reached this exact state by hand
|
||
and recorded it as its own harness error** (§4a: *"hand-driving format + assign first left the drives
|
||
mounted but unregistered — invisible to both the raw-candidate scan and `/disks`"*). **Here the
|
||
product's own rebuild path arrives there.** What was a harness mistake a week ago is now a product
|
||
state. → **R-220**
|
||
|
||
*(Deviation 4, recorded: the unmount was an operator action on the appliance host. It doubles as the
|
||
red-proof, and everything after it went through the customer's real endpoints.)*
|
||
|
||
### Step 8 — the sentinels, restored and compared
|
||
|
||
`calibre-web` redeployed (70.1 s) onto the re-enrolled drive, toggled for off-site, and restored
|
||
through the customer's own two-step full-restore flow:
|
||
|
||
```
|
||
step 1 (size gate) POST /backup/offbox/restore mode=full → 302 …&full_size=12.8+MB
|
||
step 2 (confirm) POST /backup/offbox/restore mode=full&confirm=1
|
||
14:43:27 → 14:43:43 [offbox] restored calibre-web (f3d9cd67, full=true)
|
||
→ /mnt/felhom-drives/adatok/backups/offsite-restore/calibre-web
|
||
```
|
||
|
||
**16 seconds**, out of snapshot `f3d9cd67` — the one taken **before** the wipe. Non-destructive: the
|
||
live data was untouched, exactly as the message says (*„A meglévő adatok változatlanok."*).
|
||
|
||
| # | file | expected sha256 | restored sha256 | |
|
||
|---|---|---|---|---|
|
||
| A | `C11-SENTINEL-A.txt` | `beb9175d…6df27` | `beb9175d52325bf583039beb6db1ecc80ea0781720e262c83230ed46a166df27` | **✅** |
|
||
| B | `C11-őrszem-ékezetes-árvíztűrő.txt` | `7c8cb0ad…43a87` | `7c8cb0ad040856440cc96bb1ad8f336d04f458138d85e6904ee759e690743a87` | **✅** |
|
||
| C | `C11-SENTINEL-C-12MB.bin` | `e012e76f…bb9c0` | `e012e76f3502de7c608d5b35f67b89c2404edad49eea92a1fc35a1bf02cbb9c0` | **✅** |
|
||
|
||
**All three byte-identical**, sizes intact (54 / 73 / 12 582 912 B), and **the Hungarian filename
|
||
survived the entire round trip** — disk → restic → SFTP → Storage Box → restic restore → disk —
|
||
without a single mangled character. The mojibake gate never fired.
|
||
|
||
---
|
||
|
||
## Phase 1 — VERDICT: **FAIL**
|
||
|
||
The pass condition is *"a sentinel's sha256 is byte-identical **AND** no step needed a command line
|
||
inside the guest."* **The brief asks which of the two failed, and it matters enormously:**
|
||
|
||
| half | result |
|
||
|---|---|
|
||
| **the data** | **PASS** — all three sentinels byte-identical, including a 12 MB binary and a non-ASCII Hungarian name |
|
||
| **the journey** | **FAIL** — four operator interventions, none available to a customer |
|
||
|
||
**The backup promise holds. The recovery journey does not.**
|
||
|
||
### The four interventions, each one a customer dead end
|
||
|
||
| # | Where it stopped | What was needed | Finding |
|
||
|---|---|---|---|
|
||
| 1 | the correct recovery code was refused as wrong | install agent 0.125.0 | **R-216** |
|
||
| 2 | the store never opened; the box stopped asking for its credential | restart the controller | **R-218** |
|
||
| 3 | the promised listing never rendered, and became unreachable | *(none — structurally impossible)* | **R-219** |
|
||
| 4 | the app could not be redeployed; its drives were unenrollable | unmount the raw mounts on the host | **R-220** |
|
||
|
||
### The RTO the business gets to quote
|
||
|
||
**Undefined — the unaided journey does not complete.** The attended elapsed time, stated so it is not
|
||
mistaken for the customer number:
|
||
|
||
```
|
||
13:42:41 login (claim submitted)
|
||
13:43:30 recovery code entered → REFUSED (R-216)
|
||
13:47:35 recovery code accepted (after intervention 1) +4m 54s
|
||
14:31:50 off-site target applied (after intervention 2) +49m 09s
|
||
14:42:42 app redeployed (after intervention 4) +59m 61s
|
||
14:43:43 sentinels restored + verified +61m 02s
|
||
```
|
||
|
||
**61 minutes with an operator, three of whose four blockers required root on the appliance.** The only
|
||
segment that reflects the product working is the last one: **16 seconds to pull 12.8 MB back out of the
|
||
off-site repository once everything was in place.**
|
||
|
||
For comparison, the 2026-08-04 night drill measured the same restore leg on a smaller fixture and also
|
||
needed four undocumented manual steps (R-204). **Three of those four were closed on 2026-08-05.
|
||
This run found four new ones.**
|
||
|
||
---
|
||
|
||
## Phase 3 — the retention proof (R-198)
|
||
|
||
### Step 1 — the pre-supersession record
|
||
|
||
Hub SQLite with its `-wal`, freshness **2 s** (`MAX(received_at) 14:47:19` vs `now 14:47:21`):
|
||
|
||
```
|
||
host_escrow (c11-36d660): blob 383 B · identity_blob 572 B · stale_at NULL
|
||
restic_pw_sha256 626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303
|
||
created 2026-08-05T13:11:11Z
|
||
host_escrow_superseded: id 3 demo-hp-bb76ea 383 identity_blob NULL 2026-08-04 07:15:36
|
||
id 4 demo-felhom-… 383 identity_blob NULL 2026-08-04 07:20:08
|
||
```
|
||
|
||
**Two superseded rows, neither for c11, and both carrying `identity_blob` NULL** — the pre-v0.93.0
|
||
shape the RECON documented. This is the "before" the whole phase turns on.
|
||
|
||
### Step 2 — destroy and reinstall, WITHOUT recovering (14:47:43Z)
|
||
|
||
Same destruction as Phase 1 (guest purged, both drives wiped to `4.0K`), same day-0 path, and **the
|
||
recovery screen deliberately not used** so the box mints a fresh key.
|
||
|
||
### R-216, part 4 — the day-0 path REVERTS the fix on every rebuild
|
||
|
||
The reinstall put the broken agent back, and said so:
|
||
|
||
```
|
||
[INFO] agent (existing): felhom-agent 0.125.0
|
||
[INFO] manifest: agent v0.120.0 (sha a7763d31b55b5ce7…), golden v0.192.0
|
||
[OK] installed /usr/local/bin/felhom-agent (felhom-agent 0.120.0)
|
||
```
|
||
|
||
**Not a skip — a deliberate downgrade to the vouched version.** The first reinstall had logged
|
||
`[SKIP] agent v0.120.0 already installed`; this one found 0.125.0, compared it against the manifest,
|
||
and installed 0.120.0 **over** it.
|
||
|
||
> **So an operator who fixes a box by hand has it re-broken by the next rebuild — which is precisely
|
||
> the event that makes the recovery feature necessary in the first place.** The vouched-artifact
|
||
> channel does not merely default to a stale agent; it actively restores it. Until the manifest vouches
|
||
> ≥ 0.125.0, no box can hold the fix across the only event that needs it.
|
||
|
||
### Steps 2b–2c — the box mints a fresh key, exactly as R-193 predicts
|
||
|
||
The rebuilt box came up **claimed-less with no `offbox` directory at all** — no repository password —
|
||
and `hub_escrow_identity_present: true`. It declared its need on schedule:
|
||
|
||
```
|
||
14:50:43 {"state":"needs_credential", …}
|
||
14:55:42 {"state":"needs_credential", …} ← debounceReportsDefault = 2 satisfied
|
||
14:57:03 hub: offsiteheal re-staged the stored one-time offsite secret for c11
|
||
```
|
||
|
||
The reconciler's own constants explain the timing: `debounceReportsDefault = 2`,
|
||
`tickIntervalDefault = 5 * time.Minute` (`offsiteheal/reconciler.go:70,74`). **The hub side works.**
|
||
|
||
**And R-218 reproduced exactly.** The box again did not consume the re-staged credential on its own;
|
||
the same lever was applied (`systemctl restart felhom-controller-bootstrap.service`, 14:57:18Z) and the
|
||
tier configured **17 seconds later** — the identical 17 s as in Phase 1:
|
||
|
||
```
|
||
14:57:31 [offsite-apply] settle-gate: GO
|
||
14:57:35 [offsite-apply] offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow)
|
||
```
|
||
|
||
**Second occurrence, same signature, same fix.** R-218 is not a one-off.
|
||
|
||
**The fresh key, which is the whole premise of Phase 3:**
|
||
|
||
| | sha256 of the repository password |
|
||
|---|---|
|
||
| Phase 0/1 (recovered, sealed under the FIRST code) | `626e424670248db3a25b0a48aa8a114da9061529823ce9cea7d95d9e0ee3c303` |
|
||
| **after this rebuild (newly minted)** | **`e11a6c542b73477ad02635673286f22f56ba8e1d5d9706efb4b0a0ff94473d57`** |
|
||
|
||
Different keys. The Phase 0 history (snapshot `f3d9cd67`, holding all three sentinels) is now encrypted
|
||
under a password this box no longer has — **R-193's orphaning mechanism, live, on demand.** That is
|
||
precisely the state whose recoverability the retention is supposed to preserve.
|
||
|
||
### A third instance of the "nothing triggers a refresh" family
|
||
|
||
The escrow preflight refused immediately after the tier configured:
|
||
|
||
```
|
||
{"id":"pbs_storage_id","ok":false,"detail":"escrow.pbs_storage_id not configured"} ← ok:false overall
|
||
(the other five green: dr_tier, age_binary, hub_upload, staged_secret, sudo_grant)
|
||
```
|
||
|
||
The ceremony cannot start until the controller's config carries `escrow.pbs_storage_id`, which arrives
|
||
on a config refresh — the same channel R-218 showed nothing reliably triggers. Recorded as an
|
||
observation on the same family rather than as a separate finding; it resolved on the next refresh.
|
||
|
||
### R-221 — a rebuilt box cannot run the escrow ceremony at all, and the cause is a lifetime mismatch
|
||
|
||
With the tier configured, the ceremony's preflight refused:
|
||
|
||
```
|
||
{"id":"pbs_storage_id","ok":false,"detail":"escrow.pbs_storage_id not configured"} ← overall ok:false
|
||
(the other five green)
|
||
```
|
||
|
||
`escrow.pbs_storage_id` is **agent-side** and is seeded by the pbsdr bridge —
|
||
`finishConverged → seedEscrowStorageID` (`felhom-agent/internal/pbsdr/manager.go:427-438`), whose own
|
||
comment says *"pbsdr convergence that just seeded escrow.pbs_storage_id flips this row green with no
|
||
restart."* Measured on the box: **`grep -c escrow /etc/felhom-agent/agent.json` → 0.**
|
||
|
||
**The two things have different lifetimes, and that is the whole defect.**
|
||
|
||
| | lives | survives a guest rebuild? |
|
||
|---|---|---|
|
||
| the convergence marker | `/var/lib/felhom-agent/pbsdr/marker.json` — on the **HOST** | **YES** (`applied_at 2026-08-05T12:25:07Z`, from Phase 0) |
|
||
| the seeded key | `/etc/felhom-agent/agent.json` — **rewritten by the installer** | **NO** |
|
||
|
||
And the idempotence guard reads only the marker (`manager.go:285-288`):
|
||
|
||
```go
|
||
if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) {
|
||
m.setStatus(...)
|
||
return // idempotent: this exact descriptor already converged
|
||
}
|
||
```
|
||
|
||
Unchanged descriptor → same hash → **early return → `finishConverged` never runs → the seed never
|
||
happens** — into a config file that no longer has it. The directory even carries
|
||
`agent.json.pre-seed` (14:25) proving the seeder ran exactly once, in the Phase 0 window.
|
||
|
||
**RED-PROOF.** Moving *only* the marker aside and restarting the agent — nothing else — produced the
|
||
seed instantly:
|
||
|
||
```
|
||
17:02:26 pbsdr: seeded escrow.pbs_storage_id (the ceremony one-liner needs no flags) storage_id=felhom-pbs
|
||
17:02:26 pbsdr: converged state=adopted storage_id=felhom-pbs
|
||
grep -c escrow agent.json: 0 → 1 preflight overall ok: false → TRUE
|
||
```
|
||
|
||
**Consequence: on any rebuilt box whose PBS-DR descriptor has not changed, the customer cannot run the
|
||
escrow ceremony at all** — which is precisely when they must, because the rebuild just changed their
|
||
repository key. → **R-221** *(deviation 5: the marker move-aside. It doubles as the red-proof.)*
|
||
|
||
---
|
||
|
||
## ★ Phase 3 step 4 — THE RESULT. R-198's retention WORKS IN PRODUCTION.
|
||
|
||
Second ceremony run through the customer wizard at 15:03:12Z (`restic_pw_sealed: true`,
|
||
`uploaded: true`), its code claimed one-shot into a **separate** `0600` file — **ten words, 90
|
||
characters, verified different from the Phase 0 code** without either being rendered.
|
||
|
||
Hub SQLite with `-wal`, freshness 99 s:
|
||
|
||
```
|
||
=========== host_escrow_superseded ===========
|
||
id host_id k_blob identity_blob_len restic_pw_sha256 superseded_at
|
||
3 demo-hp-bb76ea 383 -1 (NULL) 8e03eddf9ff7fde2 2026-08-04 07:15:36
|
||
4 demo-felhom-8363b5 383 -1 (NULL) 48741892f0ef4d59 2026-08-04 07:20:08
|
||
5 c11-36d660 383 572 626e424670248db3 2026-08-05 15:03:14 ★
|
||
|
||
=========== host_escrow (current) ===========
|
||
c11-36d660 blob 383 identity_blob 572 e11a6c542b73477a 2026-08-05T15:03:14Z
|
||
```
|
||
|
||
| the question | the answer |
|
||
|---|---|
|
||
| does the superseded row carry the previous sealed package? | **YES** |
|
||
| is its length equal to step 1's? | **YES — 572 B, exactly** |
|
||
| is it the OLD key? | **YES — `626e424670248db3`, the Phase 0 key, sealed under the FIRST code** |
|
||
| and rows 3–4, superseded before hub v0.93.0? | **`identity_blob` NULL — the old, lossy shape** |
|
||
|
||
> **This is the first supersession that has ever occurred in production since the fix, and the
|
||
> retention held.** What was unit-proven only is now proven live. The blob that R-198 exists to keep
|
||
> was kept, byte-length exact, on the first real opportunity to lose it.
|
||
|
||
**Step 5 — both signals fired, at the instant of supersession:**
|
||
|
||
```
|
||
15:03:14 offsite_repo_key_changed warning "Offsite repository key CHANGED for host c11-36d660…" (R-197)
|
||
15:03:14 escrow_superseded info "A korábbi helyreállítási csomag megőrizve (új kulcs érkezett)."
|
||
14:57:03 offsite_selfheal_restaged info (R-204 item 4's event, also correct)
|
||
```
|
||
|
||
**Step 6 — the run REFUSED rather than starting a fresh history:**
|
||
|
||
```
|
||
15:03:41 [escrow-confirm] hub-verified: the escrow covers the current repo password (hash e11a6c542b73…)
|
||
— EscrowState auto-confirmed escrowed; offsite runs enabled
|
||
15:04:43 [offbox] offsite repo ORPHANED — remote holds backups written under a previous,
|
||
no-longer-available key; runs will skip until reset
|
||
15:04:43 Event pushed: offbox_repo_orphaned (warning)
|
||
```
|
||
|
||
The Phase 0 history — snapshot `f3d9cd67`, holding all three sentinels — is **intact and untouched** on
|
||
the remote. Nothing was overwritten. The escrow auto-confirm also worked with no manual
|
||
`confirm-escrow`, which the 2026-08-04 drill had needed (R-204 item 3, closed, confirmed live here).
|
||
|
||
### R-222 — step 7: the retained package has no read path, and reaching for it is reported as a wrong code
|
||
|
||
With agent 0.125.0 restored so the answer would not be confounded by R-216, the **first** recovery
|
||
code — the one that genuinely opens the orphaned history — was entered:
|
||
|
||
```
|
||
15:05:59 unlock HTTP 200 in 1.111 s ← a REAL unseal attempt, not a 404
|
||
[WARN] recovery: unlock failed: agentapi: POST /escrow/recover-offsite-password: HTTP 400:
|
||
the recovery code did not open the sealed bundle, or the bundle could not be fetched
|
||
— nothing was written
|
||
```
|
||
|
||
**The engine is correct**: the retrieval serves the CURRENT row, the first code cannot open a bundle
|
||
sealed under the second, it fails closed, and nothing is written. Exactly as the RECON predicted
|
||
(`ListSupersededEscrow`: zero production callers).
|
||
|
||
**The customer-facing message is not:**
|
||
|
||
> *„A megadott helyreállítási kódot nem fogadtuk el. Ellenőrizd, hogy mind a tíz szót pontosan…"*
|
||
|
||
— and the same screen simultaneously states the hub holds a package sealed at **`2026-08-05T15:03:14Z`**
|
||
(the NEW one), with no hint that an older package exists, is deliberately retained, or corresponds to
|
||
the code just typed.
|
||
|
||
> **So the customer holding exactly the right code for their orphaned history is told they mistyped
|
||
> it.** And this is the same sentence R-216 produces for a missing agent capability, and the same
|
||
> sentence a genuinely wrong code produces. **Three different situations — wrong code / the machine
|
||
> cannot ask / the machine asked about the wrong package — one message.**
|
||
|
||
This is also the live proof of **R-202**'s open concern: the orphan card promises the old history *„a
|
||
hozzá tartozó helyreállítási kóddal később visszaállítható lehet"*. Here is a customer **with** that
|
||
code, and the answer is "wrong code". → **R-222.** Per the brief: the gap is recorded, **nothing was
|
||
built.**
|
||
|
||
|
||
---
|
||
|
||
## The fix pass (2026-08-05, evening) — and the golden rebake it turned into
|
||
|
||
Six findings fixed in controller **v0.201.0** + hub **v0.97.0/0.97.1**, each red-proofed. Full account:
|
||
`felhom-controller/REPORT.md`.
|
||
|
||
### R-223 — vouching agent 0.125.0 was NOT a dropdown, and the manifest had been frozen
|
||
|
||
The first vouch attempt was **REFUSED**, correctly, by R-120's gate (`hub/internal/web/configs.go:1162`):
|
||
|
||
```
|
||
[WARN] artifact vouch REFUSED: golden 0.192.0 is older than the newest controller
|
||
the fleet reports (0.201.0) — a fresh install would land on stale application code (R-120)
|
||
```
|
||
|
||
The manifest was **verified unchanged afterwards** rather than inferred from the HTTP 303 (which is a
|
||
redirect carrying the refusal, not a success).
|
||
|
||
**The artifacts form saves as a unit**, so the agent could not be vouched while the golden was stale —
|
||
and the golden had been 0.192.0 since before controller 0.193.0 shipped. **The Day-0 manifest had
|
||
therefore been effectively unvouchable for days, and nobody had cause to notice because nobody tried.**
|
||
That is the finding; the "one dropdown" I first reported was wrong.
|
||
|
||
### The rebake — golden 0.201.0
|
||
|
||
Baked in the drill VM on DooPlex (the sanctioned exception for bakes, `target-selection.md`).
|
||
Preconditions checked first: 6.6 T free, **no qemu held the qcow2 — proved by the revert succeeding**,
|
||
`virgin` intact, Debian template `13.6-1` listed rather than assumed (the point release rots) and
|
||
checksum-verified on download.
|
||
|
||
| | |
|
||
|---|---|
|
||
| artifact | **658 165 766 B**, sha256 `e730d7cab343eb35bcef627829b40c160f83945e0a9bd57da546efc21f007654` |
|
||
| markers | `Result=success` · `ExecMainStatus=0` · **no** `FATAL`/`excluding`/`ERROR` · `docker OK (overlay2; data-root /var/lib/docker)` · `including mount point rootfs ('/')` **and** `mp0 ('/var/lib/felhom')` · `upload OK (HTTP 201)` |
|
||
| round trip | **the bytes downloaded from Gitea hash identically.** A local checksum proves only that the local file is what it is |
|
||
| token hygiene | token read from a 0600 file **inside** the VM, never on a command line; literal-value grep on the saved log → **0 hits** |
|
||
| teardown | guest 9100 `--purge`d · token/script/log shredded **after** the log was copied out for evidence · qemu exited (checked by `ps -eo comm`, not `pgrep -f`, which self-matches) · disk reverted to `virgin` |
|
||
|
||
*Doc drift noticed, not acted on: the runbook's pass markers say `including mount point` for rootfs
|
||
**and mp0 and mp1**. `build-golden.sh` v3.0.0 bakes ONE data volume, so two mount points is correct
|
||
and the runbook describes an older layout.*
|
||
|
||
### The manifest, read back after the save
|
||
|
||
```
|
||
agent_version = 0.125.0 agent_sha256 = f7d8339b53d92a6c…7016b3
|
||
golden_version = 0.201.0 golden_sha256 = e730d7cab343eb35…f007654
|
||
min_agent = 0.125.0
|
||
```
|
||
|
||
**A machine installed from now on lands on current agent AND current controller**, and can open a
|
||
recovery package on day one. R-223 closed.
|
||
|
||
### The held floor — what is and is not established
|
||
|
||
**Established, positively, that the hold WAS in force**: the hub logged it with the corrected reason
|
||
(*"floor 0.200.0 is ABOVE the vouched golden 0.192.0…"*, 22:12:06), and the box observed it from its own
|
||
side — `[offsite-apply] settle-gate: GO — floor still unknown after 1m30s`, which is what a withheld
|
||
directive looks like to a controller.
|
||
|
||
**NOT established by a positive observable: that the hold is now released.** After the vouch, c11's
|
||
next report (22:27:06) produced **no** HELD line and the Hosts dashboard renders **no** held reason —
|
||
but both of those are absences, and this project's own rule is that an absent log line is not evidence.
|
||
The hub logs nothing on the served path by design, and `SetFloor` is DEBUG-only and fires only on a
|
||
change, so the box-side positive would need a controller restart — **which is not taken, because Phase 2
|
||
needs the venue as it is.**
|
||
|
||
What can be said with certainty is the arithmetic the code performs: floor `0.200.0` ≤ golden `0.201.0`
|
||
(so the beyond-golden branch cannot fire) and agent `0.125.0` ≥ min_agent `0.125.0` (so the
|
||
below-MinAgent branch cannot fire) — the exact combination
|
||
`TestResolveManagedFloor_R216_FloorAboveGolden/floor_at_the_golden_→_unchanged_behaviour` pins as
|
||
served. **Deduction plus a pinned test, not a live positive. Recorded as such.**
|