drill 0.243.0 complete: 0 interventions, the connect e-mail proven, NOT ready for a volunteer
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
The automatic connect e-mail is proven with a real mailbox: the host record was deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second later, with selfbind_link_sent (host delete) on the timeline. The requirement was two minutes. The hub refuses to delete an ONLINE host with no override, so the record had to fall stale first — that wait is part of the proof. Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box. The verdict is still no, for a new reason: a one-drive box with no off-site tier keeps none of the household's own files in any backup, the page says otherwise, and the restore that should save them makes it worse (R-537, R-538). Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing on the off-site server written or removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,52 +1,81 @@
|
||||
# REPORT — before the volunteer: the big night's P1 fixes, the two rulings, the publish (2026-09-15)
|
||||
# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16)
|
||||
|
||||
Releases: **agent v0.131.0**, **controller v0.243.0** (MinAgent 0.131.0), **hub v0.114.0**, catalog templates (no version),
|
||||
**installer ISO 1.27.1 published**. Baselines re-verified at start: controller `406755f`, agent `4586f0f`, felhom.eu
|
||||
`a028a9a`, catalog `882a43e` — all equal to the task table. Architecture read and cited: 03 §4, 05 (new §13/§14), 07 §6,
|
||||
08 §6.2, 02 (settings after install).
|
||||
## Claims in the prompt that turned out wrong — first, as asked
|
||||
|
||||
## Claims in the prompt that turned out wrong (first)
|
||||
1. **"The floor carries the agent release to every box"** — false. The hub HOLDS a floor whose declared MinAgent is
|
||||
above the box's agent; an agent updates only by an operator-signed `agent_update` job (R-530). Delivered to demo-hp
|
||||
with the operator's keys; demo-felhom and Peti's box stay on 0.130.0.
|
||||
2. **"`--restart always` does not restart a killed container"** — CORRECT, now measured (Docker 29.8.0, both policies
|
||||
exited 60 s after `docker kill`).
|
||||
3. **"Appliance registration knows no customer"** — correct as read (`api/appliance.go`); not a trigger.
|
||||
4. **"a `node_down` mail would have gone at ~60 min"** — not measured; unchanged as an inference.
|
||||
5. **"R-110 tag move" for the ISO** — R-110 governs the host installer's `installer-v*` tag, which the ISO does not
|
||||
change; no tag was moved. **"index page on iso.felhom.eu"** — the bucket serves no index (R-504); the index is the
|
||||
website page `felhom.eu/letoltes`, published with the ISO.
|
||||
6. **"`08-alarm-ladder.md` §5 holds the cooldown"** — it is in §6.1/§6.2; the ruling was written there.
|
||||
7. **"Park note in `RUNBOOK-manual-build.md` §4"** — §4 is the golden image; the park note went to §3.1 (controller).
|
||||
1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused
|
||||
exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran,
|
||||
and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing
|
||||
written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1).
|
||||
2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot.
|
||||
After a successful claim the same URL becomes the password-reset surface with a "request a new code"
|
||||
button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says
|
||||
"Túl sok próbálkozás — próbáld újra 15 perc múlva."
|
||||
3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than
|
||||
assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's
|
||||
own bake, and the guest runs controller 0.243.0 with no self-update.
|
||||
4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box —
|
||||
there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the
|
||||
customer's namespace, empty before and after.
|
||||
5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false
|
||||
and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was
|
||||
true.
|
||||
|
||||
## Parts
|
||||
- **A (R-523)** agent supervisor — live on 9201: idle 59 s, parked held + unpark 25 s, swap deferred; crash-loop guard
|
||||
tripped for real and the hub mailed `controller_crashloop`; after the 30-minute pause the agent restarted it
|
||||
(09:36:24Z, dashboard 200 at 09:36:30Z) and the hub minted `controller_restarted_by_agent`; mid-deploy timing not measured.
|
||||
- **B (R-513)** generated FileBrowser password — live on 9202 (generated) and 9201 (operator-set left alone); public 401.
|
||||
- **C (R-517/R-518)** per-tier page live on 9201; absent-tier skip unit-proven; „0 B" found live and fixed (unreleased).
|
||||
- **D.1 (R-509)** self-bind auto-send — shipped and red-proofed; **live mail NOT checked** (a throwaway customer's
|
||||
delete runs the RESET cascade on ep0 — fenced).
|
||||
- **D.2** `node_*` bypass the quiet hour (operator ruling) — red-proofed; recorded in 08 §6.2 and CONTEXT.md.
|
||||
- **D.3 (R-511)** re-issue adopts — shipped, red-proofed; not provable on tester-1 (no host). **STOP held**: listing
|
||||
posted, operator said yes, one snapshot forgotten on ep0 (1 928 820 672 B logical), token kept, nothing else touched.
|
||||
- **D.4 (R-510)** three GETs → 530 ×3 (no box, no tunnel); stays open; day-0 A.1 names the setting.
|
||||
- **E (R-512/R-514/R-515)** catalog — live on 9202: stranger 400, 20/20 documents, peak 772 MB; OOM line not proven (R-528).
|
||||
- **F** ISO 1.27.1 — uploaded; round trip sha256 `25637007…c053`, 1 705 322 496 B; `.sha256` 200; manifest 200;
|
||||
1.26.1 kept for rollback; `felhom.eu/letoltes` 200. G11 PASS, G12 not measurable with the token.
|
||||
- **G** docs: 03, 05 §13–14, 07 §6.4, 08 §6.2, 02, CONTEXT, RUNBOOK §3.1, day0 A.1, VOLUNTEER prerequisites.
|
||||
## What I exercised
|
||||
|
||||
## Extra acts, stated
|
||||
- The felhom.eu push was blocked by the due-checks gate (R-433 due today): the Gmail-read mailbox holds no Hetzner reply;
|
||||
re-dated to 2026-09-22 with the reason in the row. No `--no-verify` anywhere.
|
||||
- The operator's three signing keys arrived mode 664; set to 600 (R-533).
|
||||
A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first
|
||||
console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps,
|
||||
ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the
|
||||
morning-after checks and a three-layer teardown.
|
||||
|
||||
## What broke — product, and mine
|
||||
|
||||
**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token
|
||||
lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim),
|
||||
**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup
|
||||
page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves
|
||||
the app listing files it cannot open, after making the app's own wastebasket unreachable).
|
||||
|
||||
**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was
|
||||
restoring caps after the memory test — the second time this exact mistake has happened, repaired
|
||||
per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the
|
||||
controller-injected environment and crash-looped it; repaired through the controller's own API. I measured
|
||||
my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice
|
||||
before reading it off the dashboard.
|
||||
|
||||
## Rows
|
||||
Register rows **232 → 232**: closed 9 (R-493, R-495, R-496, R-512, R-513, R-514, R-515, R-517, R-523); opened 9
|
||||
(R-525 … R-533); narrowed R-509, R-510, R-511, R-518; R-433 re-dated.
|
||||
|
||||
## Teardown, three layers
|
||||
Machine: 9202 throwaways removed; 9201 controller running again (09:36:24Z); the killed homebox deploy created no container but left its `app.yaml`
|
||||
(09:04:55Z, keys only read) — moved aside to `/root/homebox-app.yaml.p1fixes-residue`, noted under R-531.
|
||||
Host: demo-hp park marker removed; agent 0.131.0 stays (the release). ep0: one tester-1 snapshot removed on yes.
|
||||
Hub: demo-hp floor override 0.243.0 kept (it delivers the release); no customer created or deleted.
|
||||
Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533,
|
||||
plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's
|
||||
first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538.
|
||||
|
||||
## The automatic connect e-mail (R-509) — PASSED
|
||||
|
||||
The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host,
|
||||
with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the
|
||||
box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host
|
||||
delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second
|
||||
later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The
|
||||
requirement was two minutes.
|
||||
|
||||
## Verdict
|
||||
|
||||
**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the
|
||||
supervisor restarts a dead controller, the file manager has its own password, the backup page tells the
|
||||
truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the
|
||||
connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are
|
||||
in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes
|
||||
things worse.
|
||||
|
||||
## Teardown
|
||||
|
||||
Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state,
|
||||
token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers
|
||||
9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host
|
||||
record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**.
|
||||
Nothing on the off-site server was written, removed or pruned — its listing is empty before and after.
|
||||
|
||||
## Checks
|
||||
|
||||
`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at
|
||||
**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by
|
||||
`head_sha`.
|
||||
|
||||
Reference in New Issue
Block a user