3f4fb3825f
gates / gates (push) Successful in 17s
Capability map: the unaided-recovery row turns FAILED -> PROVEN-LIVE, scoped, with what it still does not claim stated in the row itself: shape (c) did not fire positively (with the mint guard holding there is no local key, so the offer comes from shape (a)); and 'unaided' here means possible-without-a-shell, not obvious, because two obstacles are unsignposted. OPEN-ITEMS: R-201 closed with its evidence. Five new rows R-249..R-253 (the retrieval passphrase in page HTML; the host-key scan ladder vs AAAA settle; the listing's per-tag rows; the two unsignposted restore steps). R-243 annotated rather than re-filed: on a REBUILD offsite_delivery_stuck does not skip, so the row's gap is narrower than it reads. STATUS.md: headline changed, and trimmed 97 -> 92 lines rather than extended, per its own header. Teardown recorded as OWED with its before-measurements, the stop-and-age gate, and the positive controls that must survive.
222 lines
12 KiB
Markdown
222 lines
12 KiB
Markdown
# REPORT — the fifth walk (R-201), 2026-08-07
|
||
|
||
**Runbook-style validation, supervised. One machine destroyed on purpose, with the operator's
|
||
confirmation at the §6 STOP. No product code written.** Venue: `demo-hp` VM **325 `walk5-appliance`**,
|
||
customer `walk5`, host `walk5-4bada5`. Full evidence: `documentation/tests/walk5-r201-2026-08-07/`.
|
||
|
||
---
|
||
|
||
## 1. THE VERDICT — both halves, separately
|
||
|
||
### THE DATA: **PASS**
|
||
|
||
All three sentinels came back **byte-identical** out of snapshot `5b0f20f7`, under the key recovered
|
||
from the sealed package with R:
|
||
|
||
```
|
||
11eb7fb2d1891f62a685e3d9a9fda44e0a342d437d4aebaa15fcb4c0fd79d539 61 WALK5-SENTINEL-A.txt
|
||
6b504d1e83bf4c673334d0b9e701f3e5f07311d954fdc4463a6b0a55b261ae32 66 WALK5-őrszem-ékezetes-árvíztűrő.txt
|
||
0baaf402733e6a85f3d9ce8c4c3a57b05692f5f6733da6a23da76d2cd7635210 12582912 WALK5-SENTINEL-C-12MB.bin
|
||
name hex 57414c4b352d c591 72737a656d2d c3a9 6b657a657465732d c3a1 7276 c3ad 7a74 c5b1 72 c591 2e747874
|
||
```
|
||
|
||
Identical to Phase A in every byte **including the accented filename's name bytes**. Read back with
|
||
`os.listdir` on a **bytes** path, so no decode/encode round trip could launder a `U+FFFD` into looking
|
||
correct — the check that caught this three times before.
|
||
|
||
### THE JOURNEY: **PASS — the first time in five walks**
|
||
|
||
**No step needed a command line inside the guest.** Everything that *progressed* the journey was an
|
||
HTTP request a browser makes. The previous walk needed **three** guest command lines; this needed
|
||
**zero**. The reset-code hatch was used **once, in Phase A**, where §3 permits it.
|
||
|
||
**Named per §3 so the claim is not read wider than it is** — the guest command lines used were the
|
||
`w5watch.log` sampler, `docker logs`, the settings reads, the restic listing and the final sentinel
|
||
verification. **All instrumentation:** none changed state, none was needed to progress, and removing
|
||
them all would have changed nothing but my ability to describe what happened.
|
||
|
||
---
|
||
|
||
## 2. §5's OBSERVATION — the first live exercise of R-241's fix, and it stands on its own
|
||
|
||
**At 14:58:52Z, unaided, before anyone had logged in:**
|
||
|
||
```
|
||
[WARN] [offbox] NOT minting a repository password: the hub holds a sealed recovery package for this
|
||
box, and a fresh key would orphan the history that package protects (R-241). The transport is
|
||
configured; the tier stays down until the customer's recovery code places the escrowed key.
|
||
[INFO] [offbox] apply-offsite: transport configured …, tier HELD awaiting the escrowed key
|
||
```
|
||
|
||
| §5 asks | answer |
|
||
|---|---|
|
||
| does it declare a need, and when is it staged/collected? | **yes** — declared `needs_credential` 14:38:49Z and 14:53:49Z; hub staged **14:56:34Z**; collected + applied **14:58:52Z**. **Zero human actions**, on a box not yet claimed |
|
||
| **is any repository key written?** | **NO** — sampled every ~20 s from 14:24:52Z; `repo_password` absent at every sample. The directory holds `applied_marker`, `known_hosts`, `ssh_key` and nothing else |
|
||
| what state does it report instead? | `enabled=true`, `escrow_state=pending`, no key → the derived **`awaiting_recovery_key`** holding state |
|
||
| the two fingerprints, before login | hub's package seals `eabf427c7274…144f`; **the box holds NONE** |
|
||
|
||
**Positive control, because an absent line is not evidence:** the scheduler logged
|
||
`agent-channel-health` ×5, `stack-scan` ×2, `system-health`, `backup-cache` and
|
||
`offsite-credential-retry` in the same window. The absence of a mint is **explained**, not merely
|
||
observed.
|
||
|
||
> **HONEST SCOPE.** With the mint guard holding there is **no local key**, so the offer fires on
|
||
> **shape (a)**, not shape (c). Shape (c) was measured in **Phase A, in its negative half** — hub hash
|
||
> == local hash, correctly silent. **This walk proves the mint guard positively and the discriminator
|
||
> negatively.** A positive shape-(c) firing needs a box holding a *different* key, which v0.206.0 now
|
||
> prevents from arising by itself.
|
||
|
||
---
|
||
|
||
## 3. THE RTO
|
||
|
||
**71.7 s**, login (15:04:25.417Z) → open store (15:05:37.156Z). Of that, **12.44 s was the unseal
|
||
itself**; ~22 s was **my own harness retry** (I scraped the CSRF token from a `<meta>` tag the recovery
|
||
page does not carry, got a 403, re-read it from the form). **A customer clicking the button sees
|
||
≈50 s.** Both numbers are given because 71.7 s is what was measured.
|
||
|
||
---
|
||
|
||
## 4. DEAD ENDS, in the customer's terms
|
||
|
||
**By §3's definition — something needing a shell inside the guest — there were ZERO.** Two obstacles
|
||
were met, both cleared **from the dashboard**, and **neither is signposted**:
|
||
|
||
| # | what the customer sees | what got past it | known? |
|
||
|---|---|---|---|
|
||
| 1 | „nincs elérhető adatmeghajtó a visszaállításhoz" | Tárhely → Meghajtók → „Meglévő meghajtó csatolása" re-registers both surviving disks | **NEW — R-252** |
|
||
| 2 | „a(z) calibre-web nincs telepítve — előbb állítsd helyre az alkalmazást" — **on a page that says three lines above „Nincs telepítve — a visszaállítás előbb újratelepíti"** | redeploy from the catalogue (~90 s), then re-run the restore | **NEW — R-253** |
|
||
|
||
**So: the machinery works end to end and the data is provably safe. The unaided journey now succeeds,
|
||
and it succeeds through two obstacles a customer must guess their way past.**
|
||
|
||
---
|
||
|
||
## 5. WHAT EACH INSTALL LANDED ON — and delivery is part of the pass
|
||
|
||
| | vouched | first install | after the rebuild |
|
||
|---|---|---|---|
|
||
| agent | 0.127.0 | **0.127.0** | **0.127.0** |
|
||
| controller | golden 0.206.0 | **0.206.0** | **0.206.0** |
|
||
|
||
**No hand upgrade either time, and no downgrade on the reinstall.** This is the first walk of the five
|
||
where the box under test **is the box a customer receives** — R-239's delivery gap, the headline of both
|
||
previous walks, is closed for this run.
|
||
|
||
---
|
||
|
||
## 6. THE RECOVERY SCREEN, QUOTED
|
||
|
||
Appeared **without being sought**: `/` → 302 `/launcher` → 302 **`/recovery`**.
|
||
|
||
> „Ezt a gépet újratelepítették. **A korábbi, házon kívüli mentéseid megvannak** — a Felhom központi
|
||
> rendszere őriz hozzájuk egy lezárt csomagot, amelyet **2026-08-07T12:51:02Z** zártunk le. […]
|
||
> **A helyreállítási kódot senki nem tudja pótolni** — sem a Felhom, sem az ügyfélszolgálat, sem az
|
||
> üzemeltető. […] Ha megadod a kódot, **feloldjuk a mentéseid zárolását és megmutatjuk, mi van
|
||
> bennük**. **Ebben a lépésben semmit nem állítunk vissza és semmi nem változik.**"
|
||
|
||
The sealed-at timestamp **matches `host_escrow.created_at` exactly**. *(Copy wart, recorded not filed:
|
||
it is a raw ISO-8601 string on a Hungarian customer screen where every other date reads `2026-08-07 14:57`.)*
|
||
|
||
---
|
||
|
||
## 7. THE LISTING, against Phase A
|
||
|
||
| | Phase A | the screen |
|
||
|---|---|---|
|
||
| app | `calibre-web` | **`calibre-web`** ✅ |
|
||
| when | `5b0f20f7` @ 12:57:41Z | **2026-08-07 14:57** ✅ (CEST) |
|
||
| size | 12.784 MiB | **12.8 MB** ✅ |
|
||
|
||
A second row `felhom-offbox · 12.8 MB` also appears — the tier's marker tag rendered as an app, and the
|
||
total doubled. **R-251.**
|
||
|
||
---
|
||
|
||
## 8. §4.6's TWO PRE-DESTRUCTION CHECKS — both pass, neither previously exercised on a clean box
|
||
|
||
- **The recovery offer is NOT shown**: `/` → `/launcher`, **zero** recovery mentions on either landing
|
||
page, `/recovery` 302s. And the reason is measured, not assumed — box key, box's ACK-cached hub hash
|
||
and hub `restic_pw_sha256` are all `eabf427c7274…144f`, so **shape (c) compares equal and stays
|
||
silent**.
|
||
- **The restore page lists the app with the future-backup toggle OFF** — identical rendering both ways.
|
||
**R-237's fix, live**; the last walk measured 0 entries and a 302 here.
|
||
|
||
---
|
||
|
||
## 9. PHASE A's SEVEN RECORDS
|
||
|
||
Controller `0.206.0` · agent `0.127.0` · PBS wrapper **matches vouched** · guests 1/1 · DR recipe
|
||
**present** · key escrow **present** · snapshot `5b0f20f7` · 1 snapshot · **12.0 MB** (12 611 563 B) ·
|
||
`host_escrow` blob **383 B**, `identity_blob` **572 B**, `stale_at` **NULL**, 0 superseded rows ·
|
||
escrowed key fingerprint `a6:86:f7:fb:…:4c:f9` · box key == hub hash == `eabf427c7274…144f`.
|
||
|
||
**Sentinels listed BY NAME** out of the snapshot with `restic ls latest --long` — see §1.
|
||
|
||
> **The §4.5 gate earned its place again, and this time it caught MY fault.** The first off-site run
|
||
> reported **`ok` in 28 s with 0 snapshots**: I had sent the per-app toggle as `enabled=1`, and the
|
||
> handler accepts only `on`/`true`, so it recorded *off* and the run correctly backed up nothing.
|
||
> Re-toggled, selection verified in the rendered page, re-run → 1 snapshot, 12.0 MB. **A green tick is
|
||
> not evidence a file is in a snapshot.**
|
||
|
||
---
|
||
|
||
## 10. R — SHREDDED, with a working control
|
||
|
||
One `0600` file on **DooPlex only**, never rendered, never an argument, never a log line. Shape only:
|
||
**82 characters, 10 hyphen-separated tokens**.
|
||
|
||
```
|
||
plant → ~/.config/walk5/R_PLANTED_CONTROL.txt
|
||
sweep → 2 hits (the real file + the planted control) ← the control PROVES the sweep works
|
||
shred → both, then the pattern file itself
|
||
sweep → 0 hits
|
||
```
|
||
|
||
**Every sweep path was asserted to exist first** — a sweep pointed at a missing path returns zero for
|
||
the wrong reason. The appliance's copy was `shred -u`'d mid-walk and its absence verified.
|
||
|
||
---
|
||
|
||
## 11. NEW FINDINGS — the highest register ID moved **R-248 → R-253**
|
||
|
||
| ID | |
|
||
|---|---|
|
||
| **R-249** | **The retrieval passphrase ships in the customer page's HTML** (`data-secret`), so any headless read puts it in a transcript — with no reveal action and **no audit event**, where the break-glass credential emits one. Found by doing it. **MEDIUM** |
|
||
| **R-250** | **A customer create can fail fail-closed** because the host-key scan ladder (~60 s) is shorter than the fresh sub-account's DNS/**AAAA-before-A** settle (~100 s measured). Retry is safe and idempotent; nothing says so. **LOW-MEDIUM** |
|
||
| **R-251** | The recovery listing renders **one row per restic tag**, showing the customer an "app" they never installed and their data counted twice. **Cosmetic** |
|
||
| **R-252** | After a rebuild the restore refuses — **the drives lost their registration** — and nothing on the recovery path says to re-attach them |
|
||
| **R-253** | The restore refuses because the app is not installed, **on a page that says the restore reinstalls it**. Two shipped sentences that contradict each other, in the customer's language, at the last step of a recovery |
|
||
|
||
**Recorded against an existing row rather than minted:** **R-243** claims `offsite_delivery_stuck`
|
||
"skips the applied shape". **On a rebuild it does not skip** — 88 s after the destruction the hub
|
||
emitted the warning and wrote an **operator-channel** `notification_log` row naming a guest rebuild as
|
||
the cause, correctly. The gap is real for the state R-243 describes and **not** for the state a rebuild
|
||
produces; the row is annotated so it is not read wider than it measures.
|
||
|
||
---
|
||
|
||
## 12. TEARDOWN — OWED, not done
|
||
|
||
The machine is the evidence until the verdict is written. Full enumeration, the "before" measurements,
|
||
the stop-and-age gate, and the positive controls that must survive:
|
||
`documentation/tests/walk5-r201-2026-08-07/teardown-owed.md`. **R-244's residue will grow by this
|
||
venue.**
|
||
|
||
---
|
||
|
||
## 13. WHAT DID NOT RUN, AND WHY
|
||
|
||
- **A positive shape-(c) firing.** Structurally unreachable on a healthy v0.206.0 rebuild — see §2.
|
||
- **A soak / scheduled cycle.** The previous walk covered it; §7 does not ask for one and adding it
|
||
would have delayed the destruction past the operator's window.
|
||
- **Any product code.** §0 forbids it: five findings were filed and the walk continued.
|
||
- **Teardown.** §10 defers it deliberately.
|
||
- **`felhom-offbox`'s second listing row and the raw ISO date** were observed, not chased.
|
||
|
||
## Documents updated
|
||
|
||
`00-capability-map.md` (the unaided-recovery row → **PROVEN-LIVE, scoped**), `OPEN-ITEMS.md`
|
||
(**R-201 CLOSED**; R-249…R-253 filed; R-243 annotated), `STATUS.md` (headline changed; trimmed 97 → 92
|
||
lines rather than extended), and the journal + teardown ledger.
|