drill 0.243.0 complete: 0 interventions, the connect e-mail proven, NOT ready for a volunteer
gates / gates (push) Successful in 21s
gates / gates (push) Successful in 21s
The automatic connect e-mail is proven with a real mailbox: the host record was deleted at 12:22:59Z and the mail reached the customer at 12:23:00Z, one second later, with selfbind_link_sent (host delete) on the timeline. The requirement was two minutes. The hub refuses to delete an ONLINE host with no override, so the record had to fall stale first — that wait is part of the proof. Interventions: 0. Every P1 fix this drill set out to prove held on a fresh box. The verdict is still no, for a new reason: a one-drive box with no off-site tier keeps none of the household's own files in any backup, the page says otherwise, and the restore that should save them makes it worse (R-537, R-538). Teardown, three layers, stated. Customer tester-1 kept; RESET never used; nothing on the off-site server written or removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,52 +1,81 @@
|
||||
# REPORT — before the volunteer: the big night's P1 fixes, the two rulings, the publish (2026-09-15)
|
||||
# REPORT — DRILL: prove the P1 fixes on a fresh box (2026-09-16)
|
||||
|
||||
Releases: **agent v0.131.0**, **controller v0.243.0** (MinAgent 0.131.0), **hub v0.114.0**, catalog templates (no version),
|
||||
**installer ISO 1.27.1 published**. Baselines re-verified at start: controller `406755f`, agent `4586f0f`, felhom.eu
|
||||
`a028a9a`, catalog `882a43e` — all equal to the task table. Architecture read and cited: 03 §4, 05 (new §13/§14), 07 §6,
|
||||
08 §6.2, 02 (settings after install).
|
||||
## Claims in the prompt that turned out wrong — first, as asked
|
||||
|
||||
## Claims in the prompt that turned out wrong (first)
|
||||
1. **"The floor carries the agent release to every box"** — false. The hub HOLDS a floor whose declared MinAgent is
|
||||
above the box's agent; an agent updates only by an operator-signed `agent_update` job (R-530). Delivered to demo-hp
|
||||
with the operator's keys; demo-felhom and Peti's box stay on 0.130.0.
|
||||
2. **"`--restart always` does not restart a killed container"** — CORRECT, now measured (Docker 29.8.0, both policies
|
||||
exited 60 s after `docker kill`).
|
||||
3. **"Appliance registration knows no customer"** — correct as read (`api/appliance.go`); not a trigger.
|
||||
4. **"a `node_down` mail would have gone at ~60 min"** — not measured; unchanged as an inference.
|
||||
5. **"R-110 tag move" for the ISO** — R-110 governs the host installer's `installer-v*` tag, which the ISO does not
|
||||
change; no tag was moved. **"index page on iso.felhom.eu"** — the bucket serves no index (R-504); the index is the
|
||||
website page `felhom.eu/letoltes`, published with the ISO.
|
||||
6. **"`08-alarm-ladder.md` §5 holds the cooldown"** — it is in §6.1/§6.2; the ruling was written there.
|
||||
7. **"Park note in `RUNBOOK-manual-build.md` §4"** — §4 is the golden image; the park note went to §3.1 (controller).
|
||||
1. **"The WG hook now adopts a stuck endpoint token."** It does not, on a real box. The hook refused
|
||||
exactly as R-511 describes, the operator pressed the explicit Re-issue, hub v0.114.0's adopt path ran,
|
||||
and the ENDPOINT refused it: `missing Datastore.Modify on /datastore/felhom-offsite` → 502, nothing
|
||||
written. The fix is sound and **inert** until the ep0 grant is given (**R-534**, P1).
|
||||
2. **"The claim is one-shot — measure it, do not assume."** Correct to doubt it: it is NOT one-shot.
|
||||
After a successful claim the same URL becomes the password-reset surface with a "request a new code"
|
||||
button. And the lock-out is not at five wrong codes, it is at **two** — the third try already says
|
||||
"Túl sok próbálkozás — próbáld újra 15 perc múlva."
|
||||
3. **"The fresh box lands on golden 0.243.0."** This one was TRUE, and it is now measured rather than
|
||||
assumed: the golden archive left on the box has sha256 `e2d1843c…c10a`, byte-identical to this drill's
|
||||
own bake, and the guest runs controller 0.243.0 with no self-update.
|
||||
4. **"Restore one DB-backed app from the off-site tier onto 9202."** Not walkable at all on this box —
|
||||
there is no off-site tier to restore from (consequence of item 1), proven by a read-only listing of the
|
||||
customer's namespace, empty before and after.
|
||||
5. **My own wrong reading, recorded:** I first reported "no lock-out" for the claim code. That was false
|
||||
and I corrected it in the same evidence file — the lock-out is real and the alarm for it fired and was
|
||||
true.
|
||||
|
||||
## Parts
|
||||
- **A (R-523)** agent supervisor — live on 9201: idle 59 s, parked held + unpark 25 s, swap deferred; crash-loop guard
|
||||
tripped for real and the hub mailed `controller_crashloop`; after the 30-minute pause the agent restarted it
|
||||
(09:36:24Z, dashboard 200 at 09:36:30Z) and the hub minted `controller_restarted_by_agent`; mid-deploy timing not measured.
|
||||
- **B (R-513)** generated FileBrowser password — live on 9202 (generated) and 9201 (operator-set left alone); public 401.
|
||||
- **C (R-517/R-518)** per-tier page live on 9201; absent-tier skip unit-proven; „0 B" found live and fixed (unreleased).
|
||||
- **D.1 (R-509)** self-bind auto-send — shipped and red-proofed; **live mail NOT checked** (a throwaway customer's
|
||||
delete runs the RESET cascade on ep0 — fenced).
|
||||
- **D.2** `node_*` bypass the quiet hour (operator ruling) — red-proofed; recorded in 08 §6.2 and CONTEXT.md.
|
||||
- **D.3 (R-511)** re-issue adopts — shipped, red-proofed; not provable on tester-1 (no host). **STOP held**: listing
|
||||
posted, operator said yes, one snapshot forgotten on ep0 (1 928 820 672 B logical), token kept, nothing else touched.
|
||||
- **D.4 (R-510)** three GETs → 530 ×3 (no box, no tunnel); stays open; day-0 A.1 names the setting.
|
||||
- **E (R-512/R-514/R-515)** catalog — live on 9202: stranger 400, 20/20 documents, peak 772 MB; OOM line not proven (R-528).
|
||||
- **F** ISO 1.27.1 — uploaded; round trip sha256 `25637007…c053`, 1 705 322 496 B; `.sha256` 200; manifest 200;
|
||||
1.26.1 kept for rollback; `felhom.eu/letoltes` 200. G11 PASS, G12 not measurable with the token.
|
||||
- **G** docs: 03, 05 §13–14, 07 §6.4, 08 §6.2, 02, CONTEXT, RUNBOOK §3.1, day0 A.1, VOLUNTEER prerequisites.
|
||||
## What I exercised
|
||||
|
||||
## Extra acts, stated
|
||||
- The felhom.eu push was blocked by the due-checks gate (R-433 due today): the Gmail-read mailbox holds no Hetzner reply;
|
||||
re-dated to 2026-09-22 with the reason in the row. No `--no-verify` anywhere.
|
||||
- The operator's three signing keys arrived mode 664; set to 600 (R-533).
|
||||
A fresh box from the published ISO 1.27.1, walked as a volunteer: download by checksum, install, first
|
||||
console, connect e-mail and self-bind, claim, tunnel from outside, data drive, file manager, four apps,
|
||||
ten minutes of real use, the backup page and „Mentés most", version labels. Then five faults, then the
|
||||
morning-after checks and a three-layer teardown.
|
||||
|
||||
## What broke — product, and mine
|
||||
|
||||
**Product.** Five new rows: **R-534** (P1, the off-site tier cannot be provisioned — the endpoint token
|
||||
lacks a grant), **R-535** (P2, the console keeps showing the pairing banner after bind and claim),
|
||||
**R-536** (P2, „app installed" is sent when the install is merely accepted), **R-537** (P1, the app-backup
|
||||
page claims it holds the app's data and it does not), **R-538** (P1, a restore reports success and leaves
|
||||
the app listing files it cannot open, after making the app's own wastebasket unreachable).
|
||||
|
||||
**Mine, recorded because they cost real time.** A blunt `sed` rewrote Redis's memory cap while I was
|
||||
restoring caps after the memory test — the second time this exact mistake has happened, repaired
|
||||
per-service. A hand-run `docker compose up -d` inside the guest recreated Paperless without the
|
||||
controller-injected environment and crash-looped it; repaired through the controller's own API. I measured
|
||||
my own CSRF error twice before measuring the claim page. And I guessed the file manager's address twice
|
||||
before reading it off the dashboard.
|
||||
|
||||
## Rows
|
||||
Register rows **232 → 232**: closed 9 (R-493, R-495, R-496, R-512, R-513, R-514, R-515, R-517, R-523); opened 9
|
||||
(R-525 … R-533); narrowed R-509, R-510, R-511, R-518; R-433 re-dated.
|
||||
|
||||
## Teardown, three layers
|
||||
Machine: 9202 throwaways removed; 9201 controller running again (09:36:24Z); the killed homebox deploy created no container but left its `app.yaml`
|
||||
(09:04:55Z, keys only read) — moved aside to `/root/homebox-app.yaml.p1fixes-residue`, noted under R-531.
|
||||
Host: demo-hp park marker removed; agent 0.131.0 stays (the release). ep0: one tester-1 snapshot removed on yes.
|
||||
Hub: demo-hp floor override 0.243.0 kept (it delivers the release); no customer created or deleted.
|
||||
Opened 5 (R-534 … R-538), updated 3 (R-511, R-528, R-531), closed 2 earlier in the day (R-529, R-533,
|
||||
plus R-510 from the walk). **Counted, not asserted:** the register held 236 open rows before this session's
|
||||
first commit today and holds **238** now. Three of the five new rows are P1: R-534, R-537, R-538.
|
||||
|
||||
## The automatic connect e-mail (R-509) — PASSED
|
||||
|
||||
The host record `tester-1-652049` was deleted at **12:22:59Z** (the hub refuses to delete an ONLINE host,
|
||||
with no override by design, so the record had to fall stale first — it did at 12:22:28Z, 25m46s after the
|
||||
box's last report). In the same second the hub logged „self-bind link auto-minted for tester-1 on host
|
||||
delete", and the mailbox received „[Felhom] Kösd össze a Felhom dobozodat" at **12:23:00Z — one second
|
||||
later**. The customer timeline carries `selfbind_link_sent — Self-bind link e-mailed (host delete)`. The
|
||||
requirement was two minutes.
|
||||
|
||||
## Verdict
|
||||
|
||||
**Ready for a volunteer: no.** Every P1 fix this drill set out to prove held on a fresh box — the
|
||||
supervisor restarts a dead controller, the file manager has its own password, the backup page tells the
|
||||
truth per tier and skips an absent tier without stopping apps, the tunnel opens from outside, and the
|
||||
connect e-mail now sends itself. Interventions: **0**. What stops a volunteer is new: their own files are
|
||||
in no backup on a one-drive box, the page says otherwise, and the restore that should save them makes
|
||||
things worse.
|
||||
|
||||
## Teardown
|
||||
|
||||
Three layers, stated. Evidence off the box first (R-320): agent journal, controller log and box state,
|
||||
token-leak control 0. **Machine:** VM 334 purged with its disks; `qm list` empty; demo-hp's own containers
|
||||
9201 and 9202 untouched. **Host:** nothing to remove — the drill box WAS the nested VM. **Hub:** the host
|
||||
record deleted, the customer `tester-1` kept with its e-mail, domain and tunnel; **RESET was never used**.
|
||||
Nothing on the off-site server was written, removed or pruned — its listing is empty before and after.
|
||||
|
||||
## Checks
|
||||
|
||||
`python3 scripts/repo_gates.py --fast` — all 14 gates OK. `scripts/unproven.py --summary` — unchanged at
|
||||
**35 of 55 not walked**. CI for the session's pushes: job **642**, conclusion **success**, matched by
|
||||
`head_sha`.
|
||||
|
||||
@@ -1,5 +1,51 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-09-16 (drill on a fresh box) — the fixes hold; the backup promise does not.**
|
||||
|
||||
> **Ready for a volunteer: NO — one reason, and it is new.** On a brand-new box with one drive, the
|
||||
> household's own files are in **no backup at all**, and the backup page says they are. I deleted five
|
||||
> photos the way a child would, restored from the box's own backup, and the folder came back listing all
|
||||
> five photos — none of which opens. The bytes had never been copied. The app's own wastebasket still held
|
||||
> them, and the restore made that unreachable too.
|
||||
|
||||
**What I proved on a fresh box.** The installer downloads and installs; the box lands on the golden this
|
||||
drill baked (checked by checksum, not by trust); the connect e-mail and the bind page work; the dashboard
|
||||
opens through the tunnel from outside; the file manager has its own password and „admin/admin" is refused;
|
||||
four apps installed and were used; the backup page tells the truth per tier; „Mentés most" stopped the apps
|
||||
for 26 seconds, inside what the button promises.
|
||||
|
||||
**The five faults.** A controller killed during an install: back in 37 seconds. Two reboots a minute apart:
|
||||
everything back in 124 seconds, and the box did not count the reboots against its own safety brake. Wrong
|
||||
passwords five times: the app lets you keep trying, the box's own setup code locks for 15 minutes after two
|
||||
and e-mails you — correctly. Memory pressure: the box still cannot see it (second box, same result).
|
||||
The deleted photo folder: see above.
|
||||
|
||||
**The automatic connect e-mail: it works.** I deleted the box's record on the hub and the „connect your
|
||||
Felhom box" e-mail reached the customer **one second later**, naming the reason. That was the last thing
|
||||
waiting to be proven with a real mailbox.
|
||||
|
||||
**Decisions I took.** None under the unattended rule.
|
||||
|
||||
**Needs you.**
|
||||
1. **Say whether the backup page may keep promising what it does not hold.** Today, a new box with one drive
|
||||
backs up its apps' settings and databases — not the household's own files. The page says otherwise, and a
|
||||
restore then reports success while the files are gone. If you do nothing: the first volunteer can lose
|
||||
their photos and be told everything is fine. I can fix the wording and the refusal in the controller; the
|
||||
real protection needs a second drive or the off-site copy switched on.
|
||||
2. **Grant the off-site server one permission.** The re-issue fails on a missing grant, so a rebuilt or new
|
||||
box gets no off-site copy at all. If you do nothing: the third backup level stays unavailable for every
|
||||
new box, and the fix already written stays dead.
|
||||
3. **Rule on the restart brake.** The box stops retrying after three restarts in fifteen minutes. I measured
|
||||
that a controller dying every twenty minutes is restarted forever, and the only trace is a note that
|
||||
e-mails nobody. Options: leave it (the box heals itself and the timeline records it); add a second,
|
||||
slower counter that raises a warning; or make the fifth restart in a day a warning. My pick: the second
|
||||
counter — it keeps the healing and ends the silence. If you do nothing: a slowly failing box stays
|
||||
invisible until someone reads the timeline.
|
||||
|
||||
---
|
||||
|
||||
## Previous note
|
||||
|
||||
**Updated 2026-09-15 (P1 fixes) — the big night's blockers, fixed and shipped.**
|
||||
|
||||
> **Ready for a volunteer: almost.** The file manager has a real password, the backup page tells the truth, the
|
||||
@@ -739,3 +785,4 @@ off. **`peti-felhom` is a real machine we have not heard from since 15 July** an
|
||||
|
||||
The 2026-08-09 batch (R-279 … R-292), still untriaged; the three remaining R-264 readers; R-317 (one
|
||||
line in the agent); R-327 (decide the naming claim's status); R-359 (nothing reads the off-site store).
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -1,8 +1,8 @@
|
||||
# DRILL — prove the P1 fixes on a fresh box (2026-09-16)
|
||||
|
||||
**Interventions: _pending_** (O1/O2 pre-declared, counted apart).
|
||||
**Ready for a volunteer: _pending_.**
|
||||
**The automatic connect e-mail: _pending_.**
|
||||
**Interventions: 0** (O1 the operator's self-bind press and O2 the PBS re-issue press were pre-declared and are counted apart).
|
||||
**Ready for a volunteer: NO — and the reason is new, not one of the old ones.** Every P1 fix this drill set out to prove did hold on a fresh box. But a one-drive box with no off-site tier — the state every fresh install starts in — keeps **none of the household's own files in any backup**, while the backup page says it does, and a restore then reports success and leaves the app listing photos it cannot open (R-537, R-538).
|
||||
**The automatic connect e-mail: PASSED.** The host record was deleted at 12:22:59Z and the mail „Kösd össze a Felhom dobozodat” reached `tester1@felhom.eu` at **12:23:00Z — one second later**, with `selfbind_link_sent … (host delete)` on the customer timeline. The requirement was two minutes.
|
||||
|
||||
> Baselines at start (re-verified against live Gitea): controller `383a30b3c07b` v0.243.0 (`Unreleased`: the
|
||||
> „0 B" tile fix), agent `e98b857684f4` v0.131.0, felhom.eu `351296114c4d` hub v0.114.0, catalog `94bc5febaca2`.
|
||||
|
||||
@@ -61,3 +61,14 @@
|
||||
## 12:58 CEST ("Operator email suppressed ... cooldown") and mailed again at 13:18 CEST - one hour after
|
||||
## the 12:08 mail. So the one-hour operator cooldown both SUPPRESSES and RELEASES correctly, proven from
|
||||
## the hub log and from the inbox independently.
|
||||
##
|
||||
## THE STALENESS ALARM, fired by the teardown itself (and it is TRUE - the box really is gone):
|
||||
## the box's last report: 13:56:14 CEST (11:56:14Z); the machine destroyed ~13:57:30 CEST.
|
||||
## 14:22:00 CEST host_stale "tester-1-652049 ok -> stale" -> OPERATOR EMAIL SENT, same second.
|
||||
## That is 25m46s after the last report, i.e. the 30-minute staleness threshold measured from the
|
||||
## LAST REPORT, not from the moment of death - exactly as the dead-man's-switch is designed.
|
||||
## The hub's delete-impact endpoint flipped to {"status":"stale","deletable":true} in the same minute
|
||||
## (12:21:27Z deletable=False -> 12:22:28Z deletable=True), which is what released the host delete.
|
||||
## NOTE on R-529 (the widened host_* cooldown bypass): this fired ONCE, so the 5-minute dedupe was still
|
||||
## not exercised on this box - a second host_* within the hour would be needed, and the box was deleted
|
||||
## instead. The bypass remains proven only by its red-proofed test and the 2026-09-15 node_* run.
|
||||
|
||||
@@ -4,3 +4,63 @@
|
||||
form fields: NONE
|
||||
## waiting for the host record to fall stale (delete is refused while ONLINE, by design)
|
||||
2026-09-16T11:59:21Z status=ok deletable=False
|
||||
2026-09-16T12:00:21Z status=ok deletable=False
|
||||
2026-09-16T12:01:21Z status=ok deletable=False
|
||||
2026-09-16T12:02:21Z status=ok deletable=False
|
||||
2026-09-16T12:03:22Z status=ok deletable=False
|
||||
2026-09-16T12:04:22Z status=ok deletable=False
|
||||
2026-09-16T12:05:22Z status=ok deletable=False
|
||||
2026-09-16T12:06:23Z status=ok deletable=False
|
||||
2026-09-16T12:07:23Z status=ok deletable=False
|
||||
2026-09-16T12:08:23Z status=ok deletable=False
|
||||
2026-09-16T12:09:24Z status=ok deletable=False
|
||||
2026-09-16T12:10:24Z status=ok deletable=False
|
||||
2026-09-16T12:11:24Z status=ok deletable=False
|
||||
2026-09-16T12:12:25Z status=ok deletable=False
|
||||
2026-09-16T12:13:25Z status=ok deletable=False
|
||||
2026-09-16T12:14:25Z status=ok deletable=False
|
||||
2026-09-16T12:15:25Z status=ok deletable=False
|
||||
2026-09-16T12:16:26Z status=ok deletable=False
|
||||
2026-09-16T12:17:26Z status=ok deletable=False
|
||||
2026-09-16T12:18:26Z status=ok deletable=False
|
||||
2026-09-16T12:19:27Z status=ok deletable=False
|
||||
2026-09-16T12:20:27Z status=ok deletable=False
|
||||
2026-09-16T12:21:27Z status=ok deletable=False
|
||||
2026-09-16T12:22:28Z status=stale deletable=True
|
||||
DELETABLE at 2026-09-16T12:22:28Z
|
||||
## 2026-09-16T12:22:59Z DELETING the host record (customer tester-1 is KEPT; RESET never used)
|
||||
pre-state: {"deletable":true,"escrow_present":false,"guests":1,"log_bundles":0,"pbs_secret_present":false,"recovery_present":true,"reports":10,"status":"stale","wg_peer_bound":true}
|
||||
delete POST at 2026-09-16T12:22:59Z
|
||||
http=303
|
||||
|
||||
host record after: 404 (404 = gone)
|
||||
customer record after: 200 (200 = KEPT)
|
||||
hub log right after the delete:
|
||||
2026/09/16 14:22:00 [INFO] Host staleness: tester-1-652049 ok → stale (host_stale)
|
||||
2026/09/16 14:22:00 [INFO] Operator email sent for tester-1/host_stale
|
||||
2026/09/16 14:22:59 [INFO] host deleted: tester-1-652049 (escrow deleted: false)
|
||||
2026/09/16 14:22:59 [INFO] self-bind link emailed to the registered address of tester-1
|
||||
2026/09/16 14:22:59 [INFO] self-bind link (hash 029595d7…, valid 7 days) emailed to the registered address of tester-1
|
||||
2026/09/16 14:22:59 [INFO] self-bind link auto-minted for tester-1 on host delete (the console banner's promised email now exists)
|
||||
customer timeline, newest events after the host delete:
|
||||
Sep 16 12:22 | info | selfbind_link_sent | Self-bind link e-mailed (host delete) | hub
|
||||
Sep 16 12:22 | warning | host_stale | Host tester-1-652049: no report for 30m | hub
|
||||
Sep 16 11:56 | info | controller_started | Controller elindult (0.243.0) | controller
|
||||
Sep 16 11:39 | warning | backup_tier_skipped | Whole-guest backup tier felhom-pbs skipped: its storage does not exist on the host (never provisioned or removed). No app was stopped for it. | controller
|
||||
Sep 16 11:37 | info | controller_restarted_by_agent | Host tester-1-652049 guest 9201: the agent restarted the controller (controller container exited on 2 consecutive sweeps) — restart #2 since the agent started | hub
|
||||
Sep 16 11:34 | info | controller_started | Controller elindult (0.243.0) | controller
|
||||
host list now: 0 occurrences of the deleted host id (0 = gone)
|
||||
## RESULT - the automatic connect e-mail after a host delete (R-509) - PASSED
|
||||
## host delete POST 2026-09-16T12:22:59Z -> 303, host record 404, customer record 200 (KEPT)
|
||||
## hub log, same second: "host deleted: tester-1-652049 (escrow deleted: false)"
|
||||
## "self-bind link emailed to the registered address of tester-1"
|
||||
## "self-bind link (hash 029595d7..., valid 7 days) emailed ..."
|
||||
## "self-bind link auto-minted for tester-1 on host delete (the console banner's
|
||||
## promised email now exists)"
|
||||
## MAILBOX, read independently: a NEW message at 2026-09-16T12:23:00Z - ONE SECOND after the delete -
|
||||
## to tester1@felhom.eu, subject "[Felhom] Kosd ossze a Felhom dobozodat", body "Elkeszult a Felhom
|
||||
## dobozod, es keszen all az osszekotesre ... https://hub.felhom.eu/bind/..."
|
||||
## (The 09:59:56Z message in the same thread is the O1 operator-pressed one; this is a second, new one.)
|
||||
## TIMELINE EVENT: "selfbind_link_sent | Self-bind link e-mailed (host delete) | hub" - the occasion is
|
||||
## named, as required.
|
||||
## Well inside the two-minute requirement: 1 second.
|
||||
|
||||
Reference in New Issue
Block a user