drill 0.243.0 complete: 0 interventions, the connect e-mail proven, NOT ready for a volunteer
gates / gates (push) Successful in 21s

The automatic connect e-mail is proven with a real mailbox: the host record was
deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second
later, with selfbind_link_sent (host delete) on the timeline. The requirement was
two minutes. The hub refuses to delete an ONLINE host with no override, so the
record had to fall stale first — that wait is part of the proof.

Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box.
The verdict is still no, for a new reason: a one-drive box with no off-site tier
keeps none of the household's own files in any backup, the page says otherwise,
and the restore that should save them makes it worse (R-537, R-538).

Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing
on the off-site server written or removed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-16 14:27:04 +02:00
parent ba33db3108
commit dfd854474e
6 changed files with 196 additions and 49 deletions
+74 -45
View File
@@ -1,52 +1,81 @@
# REPORT — before the volunteer: the big night's P1 fixes, the two rulings, the publish (2026-09-15)
# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16)
Releases: **agent v0.131.0**, **controller v0.243.0** (MinAgent 0.131.0), **hub v0.114.0**, catalog templates (no version),
**installer ISO 1.27.1 published**. Baselines re-verified at start: controller `406755f`, agent `4586f0f`, felhom.eu
`a028a9a`, catalog `882a43e` — all equal to the task table. Architecture read and cited: 03 §4, 05 (new §13/§14), 07 §6,
08 §6.2, 02 (settings after install).
## Claims in the prompt that turned out wrong — first, as asked
## Claims in the prompt that turned out wrong (first)
1. **"The floor carries the agent release to every box"** — false. The hub HOLDS a floor whose declared MinAgent is
above the box's agent; an agent updates only by an operator-signed `agent_update` job (R-530). Delivered to demo-hp
with the operator's keys; demo-felhom and Peti's box stay on 0.130.0.
2. **"`--restart always` does not restart a killed container"** — CORRECT, now measured (Docker 29.8.0, both policies
exited 60 s after `docker kill`).
3. **"Appliance registration knows no customer"** — correct as read (`api/appliance.go`); not a trigger.
4. **"a `node_down` mail would have gone at ~60 min"** — not measured; unchanged as an inference.
5. **"R-110 tag move" for the ISO** — R-110 governs the host installer's `installer-v*` tag, which the ISO does not
change; no tag was moved. **"index page on iso.felhom.eu"** — the bucket serves no index (R-504); the index is the
website page `felhom.eu/letoltes`, published with the ISO.
6. **"`08-alarm-ladder.md` §5 holds the cooldown"** — it is in §6.1/§6.2; the ruling was written there.
7. **"Park note in `RUNBOOK-manual-build.md` §4"** — §4 is the golden image; the park note went to §3.1 (controller).
1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused
exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran,
and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing
written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1).
2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot.
After a successful claim the same URL becomes the password-reset surface with a "request a new code"
button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says
"Túl sok próbálkozás — próbáld újra 15 perc múlva."
3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than
assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's
own bake, and the guest runs controller 0.243.0 with no self-update.
4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box —
there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the
customer's namespace, empty before and after.
5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false
and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was
true.
## Parts
- **A (R-523)** agent supervisor — live on 9201: idle 59 s, parked held + unpark 25 s, swap deferred; crash-loop guard
tripped for real and the hub mailed `controller_crashloop`; after the 30-minute pause the agent restarted it
(09:36:24Z, dashboard 200 at 09:36:30Z) and the hub minted `controller_restarted_by_agent`; mid-deploy timing not measured.
- **B (R-513)** generated FileBrowser password — live on 9202 (generated) and 9201 (operator-set left alone); public 401.
- **C (R-517/R-518)** per-tier page live on 9201; absent-tier skip unit-proven; „0 B" found live and fixed (unreleased).
- **D.1 (R-509)** self-bind auto-send — shipped and red-proofed; **live mail NOT checked** (a throwaway customer's
delete runs the RESET cascade on ep0 — fenced).
- **D.2** `node_*` bypass the quiet hour (operator ruling) — red-proofed; recorded in 08 §6.2 and CONTEXT.md.
- **D.3 (R-511)** re-issue adopts — shipped, red-proofed; not provable on tester-1 (no host). **STOP held**: listing
posted, operator said yes, one snapshot forgotten on ep0 (1 928 820 672 B logical), token kept, nothing else touched.
- **D.4 (R-510)** three GETs → 530 ×3 (no box, no tunnel); stays open; day-0 A.1 names the setting.
- **E (R-512/R-514/R-515)** catalog — live on 9202: stranger 400, 20/20 documents, peak 772 MB; OOM line not proven (R-528).
- **F** ISO 1.27.1 — uploaded; round trip sha256 `25637007…c053`, 1 705 322 496 B; `.sha256` 200; manifest 200;
1.26.1 kept for rollback; `felhom.eu/letoltes` 200. G11 PASS, G12 not measurable with the token.
- **G** docs: 03, 05 §13–14, 07 §6.4, 08 §6.2, 02, CONTEXT, RUNBOOK §3.1, day0 A.1, VOLUNTEER prerequisites.
## What I exercised
## Extra acts, stated
- The felhom.eu push was blocked by the due-checks gate (R-433 due today): the Gmail-read mailbox holds no Hetzner reply;
re-dated to 2026-09-22 with the reason in the row. No `--no-verify` anywhere.
- The operator's three signing keys arrived mode 664; set to 600 (R-533).
A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first
console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps,
ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the
morning-after checks and a three-layer teardown.
## What broke — product, and mine
**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token
lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim),
**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup
page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves
the app listing files it cannot open, after making the app's own wastebasket unreachable).
**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was
restoring caps after the memory test — the second time this exact mistake has happened, repaired
per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the
controller-injected environment and crash-looped it; repaired through the controller's own API. I measured
my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice
before reading it off the dashboard.
## Rows
Register rows **232 → 232**: closed 9 (R-493, R-495, R-496, R-512, R-513, R-514, R-515, R-517, R-523); opened 9
(R-525 … R-533); narrowed R-509, R-510, R-511, R-518; R-433 re-dated.
## Teardown, three layers
Machine: 9202 throwaways removed; 9201 controller running again (09:36:24Z); the killed homebox deploy created no container but left its `app.yaml`
(09:04:55Z, keys only read) — moved aside to `/root/homebox-app.yaml.p1fixes-residue`, noted under R-531.
Host: demo-hp park marker removed; agent 0.131.0 stays (the release). ep0: one tester-1 snapshot removed on yes.
Hub: demo-hp floor override 0.243.0 kept (it delivers the release); no customer created or deleted.
Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533,
plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's
first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538.
## The automatic connect e-mail (R-509) — PASSED
The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host,
with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the
box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host
delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second
later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The
requirement was two minutes.
## Verdict
**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the
supervisor restarts a dead controller, the file manager has its own password, the backup page tells the
truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the
connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are
in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes
things worse.
## Teardown
Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state,
token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers
9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host
record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**.
Nothing on the off-site server was written, removed or pruned — its listing is empty before and after.
## Checks
`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at
**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by
`head_sha`.
+47
View File
@@ -1,5 +1,51 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-09-16 (drill on a fresh box) — the fixes hold; the backup promise does not.**
> **Ready for a volunteer: NO — one reason, and it is new.** On a brand-new box with one drive, the
> household's own files are in **no backup at all**, and the backup page says they are. I deleted five
> photos the way a child would, restored from the box's own backup, and the folder came back listing all
> five photos — none of which opens. The bytes had never been copied. The app's own wastebasket still held
> them, and the restore made that unreachable too.
**What I proved on a fresh box.** The installer downloads and installs; the box lands on the golden this
drill baked (checked by checksum, not by trust); the connect e-mail and the bind page work; the dashboard
opens through the tunnel from outside; the file manager has its own password and „admin/admin" is refused;
four apps installed and were used; the backup page tells the truth per tier; „Mentés most" stopped the apps
for 26 seconds, inside what the button promises.
**The five faults.** A controller killed during an install: back in 37 seconds. Two reboots a minute apart:
everything back in 124 seconds, and the box did not count the reboots against its own safety brake. Wrong
passwords five times: the app lets you keep trying, the box's own setup code locks for 15 minutes after two
and e-mails you — correctly. Memory pressure: the box still cannot see it (second box, same result).
The deleted photo folder: see above.
**The automatic connect e-mail: it works.** I deleted the box's record on the hub and the „connect your
Felhom box" e-mail reached the customer **one second later**, naming the reason. That was the last thing
waiting to be proven with a real mailbox.
**Decisions I took.** None under the unattended rule.
**Needs you.**
1. **Say whether the backup page may keep promising what it does not hold.** Today, a new box with one drive
backs up its apps' settings and databases — not the household's own files. The page says otherwise, and a
restore then reports success while the files are gone. If you do nothing: the first volunteer can lose
their photos and be told everything is fine. I can fix the wording and the refusal in the controller; the
real protection needs a second drive or the off-site copy switched on.
2. **Grant the off-site server one permission.** The re-issue fails on a missing grant, so a rebuilt or new
box gets no off-site copy at all. If you do nothing: the third backup level stays unavailable for every
new box, and the fix already written stays dead.
3. **Rule on the restart brake.** The box stops retrying after three restarts in fifteen minutes. I measured
that a controller dying every twenty minutes is restarted forever, and the only trace is a note that
e-mails nobody. Options: leave it (the box heals itself and the timeline records it); add a second,
slower counter that raises a warning; or make the fifth restart in a day a warning. My pick: the second
counter — it keeps the healing and ends the silence. If you do nothing: a slowly failing box stays
invisible until someone reads the timeline.
---
## Previous note
**Updated 2026-09-15 (P1 fixes) — the big night's blockers, fixed and shipped.**
> **Ready for a volunteer: almost.** The file manager has a real password, the backup page tells the truth, the
@@ -739,3 +785,4 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
File diff suppressed because one or more lines are too long
@@ -1,8 +1,8 @@
# DRILL — prove the P1 fixes on a fresh box (2026-09-16)
**Interventions: _pending_** (O1/O2 pre-declared, counted apart).
**Ready for a volunteer: _pending_.**
**The automatic connect e-mail: _pending_.**
**Interventions: 0** (O1 the operator's self-bind press and O2 the PBS re-issue press were pre-declared and are counted apart).
**Ready for a volunteer: NO — and the reason is new, not one of the old ones.** Every P1 fix this drill set out to prove did hold on a fresh box. But a one-drive box with no off-site tier — the state every fresh install starts in — keeps **none of the household's own files in any backup**, while the backup page says it does, and a restore then reports success and leaves the app listing photos it cannot open (R-537, R-538).
**The automatic connect e-mail: PASSED.** The host record was deleted at 12:22:59Z and the mail „Kösd össze a Felhom dobozodat” reached `tester1@felhom.eu` at **12:23:00Z — one second later**, with `selfbind_link_sent … (host delete)` on the customer timeline. The requirement was two minutes.
> Baselines at start (re-verified against live Gitea): controller `383a30b3c07b` v0.243.0 (`Unreleased`: the
> „0 B" tile fix), agent `e98b857684f4` v0.131.0, felhom.eu `351296114c4d` hub v0.114.0, catalog `94bc5febaca2`.
@@ -61,3 +61,14 @@
## 12:58 CEST ("Operator email suppressed ... cooldown") and mailed again at 13:18 CEST - one hour after
## the 12:08 mail. So the one-hour operator cooldown both SUPPRESSES and RELEASES correctly, proven from
## the hub log and from the inbox independently.
##
## THE STALENESS ALARM, fired by the teardown itself (and it is TRUE - the box really is gone):
## the box's last report: 13:56:14 CEST (11:56:14Z); the machine destroyed ~13:57:30 CEST.
## 14:22:00 CEST host_stale "tester-1-652049 ok -> stale" -> OPERATOR EMAIL SENT, same second.
## That is 25m46s after the last report, i.e. the 30-minute staleness threshold measured from the
## LAST REPORT, not from the moment of death - exactly as the dead-man's-switch is designed.
## The hub's delete-impact endpoint flipped to {"status":"stale","deletable":true} in the same minute
## (12:21:27Z deletable=False -> 12:22:28Z deletable=True), which is what released the host delete.
## NOTE on R-529 (the widened host_* cooldown bypass): this fired ONCE, so the 5-minute dedupe was still
## not exercised on this box - a second host_* within the hour would be needed, and the box was deleted
## instead. The bypass remains proven only by its red-proofed test and the 2026-09-15 node_* run.
@@ -4,3 +4,63 @@
form fields: NONE
## waiting for the host record to fall stale (delete is refused while ONLINE, by design)
2026-09-16T11:59:21Z status=ok deletable=False
2026-09-16T12:00:21Z status=ok deletable=False
2026-09-16T12:01:21Z status=ok deletable=False
2026-09-16T12:02:21Z status=ok deletable=False
2026-09-16T12:03:22Z status=ok deletable=False
2026-09-16T12:04:22Z status=ok deletable=False
2026-09-16T12:05:22Z status=ok deletable=False
2026-09-16T12:06:23Z status=ok deletable=False
2026-09-16T12:07:23Z status=ok deletable=False
2026-09-16T12:08:23Z status=ok deletable=False
2026-09-16T12:09:24Z status=ok deletable=False
2026-09-16T12:10:24Z status=ok deletable=False
2026-09-16T12:11:24Z status=ok deletable=False
2026-09-16T12:12:25Z status=ok deletable=False
2026-09-16T12:13:25Z status=ok deletable=False
2026-09-16T12:14:25Z status=ok deletable=False
2026-09-16T12:15:25Z status=ok deletable=False
2026-09-16T12:16:26Z status=ok deletable=False
2026-09-16T12:17:26Z status=ok deletable=False
2026-09-16T12:18:26Z status=ok deletable=False
2026-09-16T12:19:27Z status=ok deletable=False
2026-09-16T12:20:27Z status=ok deletable=False
2026-09-16T12:21:27Z status=ok deletable=False
2026-09-16T12:22:28Z status=stale deletable=True
DELETABLE at 2026-09-16T12:22:28Z
## 2026-09-16T12:22:59Z DELETING the host record (customer tester-1 is KEPT; RESET never used)
pre-state: {"deletable":true,"escrow_present":false,"guests":1,"log_bundles":0,"pbs_secret_present":false,"recovery_present":true,"reports":10,"status":"stale","wg_peer_bound":true}
delete POST at 2026-09-16T12:22:59Z
http=303
host record after: 404 (404 = gone)
customer record after: 200 (200 = KEPT)
hub log right after the delete:
2026/09/16 14:22:00 [INFO] Host staleness: tester-1-652049 ok → stale (host_stale)
2026/09/16 14:22:00 [INFO] Operator email sent for tester-1/host_stale
2026/09/16 14:22:59 [INFO] host deleted: tester-1-652049 (escrow deleted: false)
2026/09/16 14:22:59 [INFO] self-bind link emailed to the registered address of tester-1
2026/09/16 14:22:59 [INFO] self-bind link (hash 029595d7…, valid 7 days) emailed to the registered address of tester-1
2026/09/16 14:22:59 [INFO] self-bind link auto-minted for tester-1 on host delete (the console banner's promised email now exists)
customer timeline, newest events after the host delete:
Sep 16 12:22 | info | selfbind_link_sent | Self-bind link e-mailed (host delete) | hub
Sep 16 12:22 | warning | host_stale | Host tester-1-652049: no report for 30m | hub
Sep 16 11:56 | info | controller_started | Controller elindult (0.243.0) | controller
Sep 16 11:39 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
Sep 16 11:37 | info | controller_restarted_by_agent | Host tester-1-652049 guest 9201: the agent restarted the controller (controller container exited on 2 consecutive sweeps) — restart #2 since the agent started | hub
Sep 16 11:34 | info | controller_started | Controller elindult (0.243.0) | controller
host list now: 0 occurrences of the deleted host id (0 = gone)
## RESULT - the automatic connect e-mail after a host delete (R-509) - PASSED
## host delete POST 2026-09-16T12:22:59Z -> 303, host record 404, customer record 200 (KEPT)
## hub log, same second: "host deleted: tester-1-652049 (escrow deleted: false)"
## "self-bind link emailed to the registered address of tester-1"
## "self-bind link (hash 029595d7..., valid 7 days) emailed ..."
## "self-bind link auto-minted for tester-1 on host delete (the console banner's
## promised email now exists)"
## MAILBOX, read independently: a NEW message at 2026-09-16T12:23:00Z - ONE SECOND after the delete -
## to tester1@felhom.eu, subject "[Felhom] Kosd ossze a Felhom dobozodat", body "Elkeszult a Felhom
## dobozod, es keszen all az osszekotesre ... https://hub.felhom.eu/bind/..."
## (The 09:59:56Z message in the same thread is the O1 operator-pressed one; this is a second, new one.)
## TIMELINE EVENT: "selfbind_link_sent | Self-bind link e-mailed (host delete) | hub" - the occasion is
## named, as required.
## Well inside the two-minute requirement: 1 second.