drill 0.243.0 complete: 0 interventions, the connect e-mail proven, NOT ready for a volunteer
gates / gates (push) Successful in 21s

The automatic connect e-mail is proven with a real mailbox: the host record was
deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second
later, with selfbind_link_sent (host delete) on the timeline. The requirement was
two minutes. The hub refuses to delete an ONLINE host with no override, so the
record had to fall stale first — that wait is part of the proof.

Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box.
The verdict is still no, for a new reason: a one-drive box with no off-site tier
keeps none of the household's own files in any backup, the page says otherwise,
and the restore that should save them makes it worse (R-537, R-538).

Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing
on the off-site server written or removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 14:27:04 +02:00
parent ba33db3108
commit dfd854474e
6 changed files with 196 additions and 49 deletions
+74 -45
View File
@@ -1,52 +1,81 @@
# REPORT — before the volunteer: the big night's P1 fixes, the two rulings, the publish (2026-09-15)
# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16)
Releases: **agent v0.131.0**, **controller v0.243.0** (MinAgent 0.131.0), **hub v0.114.0**, catalog templates (no version),
**installer ISO 1.27.1 published**. Baselines re-verified at start: controller `406755f`, agent `4586f0f`, felhom.eu
`a028a9a`, catalog `882a43e` — all equal to the task table. Architecture read and cited: 03 §4, 05 (new §13/§14), 07 §6,
08 §6.2, 02 (settings after install).
## Claims in the prompt that turned out wrong — first, as asked
## Claims in the prompt that turned out wrong (first)
1. **"The floor carries the agent release to every box"** — false. The hub HOLDS a floor whose declared MinAgent is
above the box's agent; an agent updates only by an operator-signed `agent_update` job (R-530). Delivered to demo-hp
with the operator's keys; demo-felhom and Peti's box stay on 0.130.0.
2. **"`--restart always` does not restart a killed container"** — CORRECT, now measured (Docker 29.8.0, both policies
exited 60 s after `docker kill`).
3. **"Appliance registration knows no customer"** — correct as read (`api/appliance.go`); not a trigger.
4. **"a `node_down` mail would have gone at ~60 min"** — not measured; unchanged as an inference.
5. **"R-110 tag move" for the ISO** — R-110 governs the host installer's `installer-v*` tag, which the ISO does not
change; no tag was moved. **"index page on iso.felhom.eu"** — the bucket serves no index (R-504); the index is the
website page `felhom.eu/letoltes`, published with the ISO.
6. **"`08-alarm-ladder.md` §5 holds the cooldown"** — it is in §6.1/§6.2; the ruling was written there.
7. **"Park note in `RUNBOOK-manual-build.md` §4"** — §4 is the golden image; the park note went to §3.1 (controller).
1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused
exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran,
and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing
written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1).
2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot.
After a successful claim the same URL becomes the password-reset surface with a "request a new code"
button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says
"Túl sok próbálkozás — próbáld újra 15 perc múlva."
3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than
assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's
own bake, and the guest runs controller 0.243.0 with no self-update.
4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box —
there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the
customer's namespace, empty before and after.
5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false
and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was
true.
## Parts
- **A (R-523)** agent supervisor — live on 9201: idle 59 s, parked held + unpark 25 s, swap deferred; crash-loop guard
tripped for real and the hub mailed `controller_crashloop`; after the 30-minute pause the agent restarted it
(09:36:24Z, dashboard 200 at 09:36:30Z) and the hub minted `controller_restarted_by_agent`; mid-deploy timing not measured.
- **B (R-513)** generated FileBrowser password — live on 9202 (generated) and 9201 (operator-set left alone); public 401.
- **C (R-517/R-518)** per-tier page live on 9201; absent-tier skip unit-proven; „0 B" found live and fixed (unreleased).
- **D.1 (R-509)** self-bind auto-send — shipped and red-proofed; **live mail NOT checked** (a throwaway customer's
delete runs the RESET cascade on ep0 — fenced).
- **D.2** `node_*` bypass the quiet hour (operator ruling) — red-proofed; recorded in 08 §6.2 and CONTEXT.md.
- **D.3 (R-511)** re-issue adopts — shipped, red-proofed; not provable on tester-1 (no host). **STOP held**: listing
posted, operator said yes, one snapshot forgotten on ep0 (1 928 820 672 B logical), token kept, nothing else touched.
- **D.4 (R-510)** three GETs → 530 ×3 (no box, no tunnel); stays open; day-0 A.1 names the setting.
- **E (R-512/R-514/R-515)** catalog — live on 9202: stranger 400, 20/20 documents, peak 772 MB; OOM line not proven (R-528).
- **F** ISO 1.27.1 — uploaded; round trip sha256 `25637007…c053`, 1 705 322 496 B; `.sha256` 200; manifest 200;
1.26.1 kept for rollback; `felhom.eu/letoltes` 200. G11 PASS, G12 not measurable with the token.
- **G** docs: 03, 05 §13–14, 07 §6.4, 08 §6.2, 02, CONTEXT, RUNBOOK §3.1, day0 A.1, VOLUNTEER prerequisites.
## What I exercised
## Extra acts, stated
- The felhom.eu push was blocked by the due-checks gate (R-433 due today): the Gmail-read mailbox holds no Hetzner reply;
re-dated to 2026-09-22 with the reason in the row. No `--no-verify` anywhere.
- The operator's three signing keys arrived mode 664; set to 600 (R-533).
A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first
console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps,
ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the
morning-after checks and a three-layer teardown.
## What broke — product, and mine
**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token
lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim),
**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup
page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves
the app listing files it cannot open, after making the app's own wastebasket unreachable).
**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was
restoring caps after the memory test — the second time this exact mistake has happened, repaired
per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the
controller-injected environment and crash-looped it; repaired through the controller's own API. I measured
my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice
before reading it off the dashboard.
## Rows
Register rows **232 → 232**: closed 9 (R-493, R-495, R-496, R-512, R-513, R-514, R-515, R-517, R-523); opened 9
(R-525 … R-533); narrowed R-509, R-510, R-511, R-518; R-433 re-dated.
## Teardown, three layers
Machine: 9202 throwaways removed; 9201 controller running again (09:36:24Z); the killed homebox deploy created no container but left its `app.yaml`
(09:04:55Z, keys only read) — moved aside to `/root/homebox-app.yaml.p1fixes-residue`, noted under R-531.
Host: demo-hp park marker removed; agent 0.131.0 stays (the release). ep0: one tester-1 snapshot removed on yes.
Hub: demo-hp floor override 0.243.0 kept (it delivers the release); no customer created or deleted.
Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533,
plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's
first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538.
## The automatic connect e-mail (R-509) — PASSED
The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host,
with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the
box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host
delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second
later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The
requirement was two minutes.
## Verdict
**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the
supervisor restarts a dead controller, the file manager has its own password, the backup page tells the
truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the
connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are
in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes
things worse.
## Teardown
Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state,
token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers
9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host
record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**.
Nothing on the off-site server was written, removed or pruned — its listing is empty before and after.
## Checks
`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at
**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by
`head_sha`.