diff --git a/REPORT.md b/REPORT.md index c6288f2e..9c9a24ab 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,45 +1,52 @@ -# REPORT — the second night: scratch guest built, controller v0.242.0, rotation restarted (2026-09-13/14) +# REPORT — before the volunteer: the big night's P1 fixes, the two rulings, the publish (2026-09-15) -*Overwritten each session. Evidence: `documentation/audits/nightly-2026-09-13b-bentopdf/` (the guest -build and the bentopdf walk) and `documentation/audits/v0242-2026-09-14/` (red-proofs, floor, three -live passes). Controller detail: `felhom-controller/REPORT.md`.* +Releases: **agent v0.131.0**, **controller v0.243.0** (MinAgent 0.131.0), **hub v0.114.0**, catalog templates (no version), +**installer ISO 1.27.1 published**. Baselines re-verified at start: controller `406755f`, agent `4586f0f`, felhom.eu +`a028a9a`, catalog `882a43e` — all equal to the task table. Architecture read and cited: 03 §4, 05 (new §13/§14), 07 §6, +08 §6.2, 02 (settings after install). -**Rules:** `.claude/rules/unprompted-work.md`. **Architecture named:** `01-topology-and-trust.md` §2 -(the scratch guest is a second guest of the same customer), `07-backup-architecture.md` §6.3 and the -R-237 store-keyed list rule, `09-update-architecture.md` (holds). **Baselines:** felhom.eu `72ee053`, -controller `3e81330` (v0.241.0), catalog `6d6eec3`. +## Claims in the prompt that turned out wrong (first) +1. **"The floor carries the agent release to every box"** — false. The hub HOLDS a floor whose declared MinAgent is + above the box's agent; an agent updates only by an operator-signed `agent_update` job (R-530). Delivered to demo-hp + with the operator's keys; demo-felhom and Peti's box stay on 0.130.0. +2. **"`--restart always` does not restart a killed container"** — CORRECT, now measured (Docker 29.8.0, both policies + exited 60 s after `docker kill`). +3. **"Appliance registration knows no customer"** — correct as read (`api/appliance.go`); not a trigger. +4. **"a `node_down` mail would have gone at ~60 min"** — not measured; unchanged as an inference. +5. **"R-110 tag move" for the ISO** — R-110 governs the host installer's `installer-v*` tag, which the ISO does not + change; no tag was moved. **"index page on iso.felhom.eu"** — the bucket serves no index (R-504); the index is the + website page `felhom.eu/letoltes`, published with the ISO. +6. **"`08-alarm-ladder.md` §5 holds the cooldown"** — it is in §6.1/§6.2; the ruling was written there. +7. **"Park note in `RUNBOOK-manual-build.md` §4"** — §4 is the golden image; the park note went to §3.1 (controller). -## Step 0 — R-481, operator ruling option 1 → LXC 9202 on demo-hp, persists +## Parts +- **A (R-523)** agent supervisor — live on 9201: idle 59 s, parked held + unpark 25 s, swap deferred; crash-loop guard + tripped for real and the hub mailed `controller_crashloop`; after the 30-minute pause the agent restarted it + (09:36:24Z, dashboard 200 at 09:36:30Z) and the hub minted `controller_restarted_by_agent`; mid-deploy timing not measured. +- **B (R-513)** generated FileBrowser password — live on 9202 (generated) and 9201 (operator-set left alone); public 401. +- **C (R-517/R-518)** per-tier page live on 9201; absent-tier skip unit-proven; „0 B" found live and fixed (unreleased). +- **D.1 (R-509)** self-bind auto-send — shipped and red-proofed; **live mail NOT checked** (a throwaway customer's + delete runs the RESET cascade on ep0 — fenced). +- **D.2** `node_*` bypass the quiet hour (operator ruling) — red-proofed; recorded in 08 §6.2 and CONTEXT.md. +- **D.3 (R-511)** re-issue adopts — shipped, red-proofed; not provable on tester-1 (no host). **STOP held**: listing + posted, operator said yes, one snapshot forgotten on ep0 (1 928 820 672 B logical), token kept, nothing else touched. +- **D.4 (R-510)** three GETs → 530 ×3 (no box, no tunnel); stays open; day-0 A.1 names the setting. +- **E (R-512/R-514/R-515)** catalog — live on 9202: stranger 400, 20/20 documents, peak 772 MB; OOM line not proven (R-528). +- **F** ISO 1.27.1 — uploaded; round trip sha256 `25637007…c053`, 1 705 322 496 B; `.sha256` 200; manifest 200; + 1.26.1 kept for rollback; `felhom.eu/letoltes` 200. G11 PASS, G12 not measurable with the token. +- **G** docs: 03, 05 §13–14, 07 §6.4, 08 §6.2, 02, CONTEXT, RUNBOOK §3.1, day0 A.1, VOLUNTEER prerequisites. -Restored from the vouched golden 0.236.0 onto a `dir` storage re-added at `/mnt/hdd_1` -(`nvme-scratch`), sized like 9201, unprivileged; the demo-hp customer seeded with hub OFF, tunnel OFF, -agent OFF, off-site OFF, self-update OFF; image set by hand (the one place allowed). Disposition in -the guest, on the host and on the hub side (`operations/nodes.md`, the register). Two CC decisions, -tagged in `CONTEXT.md` and `01-topology-and-trust.md`: no real-data-drive bind, no cloudflared. +## Extra acts, stated +- The felhom.eu push was blocked by the due-checks gate (R-433 due today): the Gmail-read mailbox holds no Hetzner reply; + re-dated to 2026-09-22 with the reason in the row. No `--no-verify` anywhere. +- The operator's three signing keys arrived mode 664; set to 600 (R-533). -## Step 1–3 — controller v0.242.0 (`d698ce3`, docs `406755f`), one release, floor-delivered - -R-487 (removed app listed with its restore; the lists are keyed on the drives), R-491 (removal -clears the update hold), R-490 (`/api/system/info` reachable + fallback), R-489 (`volumes_removed` -difference — **half: a restore-recreated volume has no compose label and is missed; row kept open, -measured**), R-476 (Tier-2 copy dated by its data), R-456 (rule pinned). Red-proofs ×11, full gate -green, floor 0.242.0 with MinAgent 0.129.0: demo-hp +16 s, demo-felhom +17 s. Live on 9202 with an -opengist throwaway through the endpoints the UI invokes (no browser on DooPlex); the first pass was -refused 409 (a running app must be stopped before removal) and repeated. - -## Step 2 — rotation reset to bentopdf on 9202: clean - -Front door, use (browser-side tool, no data route — recorded, not faked), backup + second copy, -remove-with-data keep-backups, Tier-2 unit restore, guarded update (same pin; backed up first because -the restore rewrote `deployed_at`, R-478), remove-everything. No new rows from the walk. - -## Register - -210 open / 194 closed → **205 / 200**. Closed: R-481, R-487, R-491, R-490, R-476, R-456. Re-scoped: -R-489. Opened: R-492 (delete `cfg.Paths.HDDPath`). Security note: demo-hp's retrieval passphrase -appeared once in a tool output on DooPlex during the guest build (redaction list missed the key). +## Rows +Register rows **232 → 232**: closed 9 (R-493, R-495, R-496, R-512, R-513, R-514, R-515, R-517, R-523); opened 9 +(R-525 … R-533); narrowed R-509, R-510, R-511, R-518; R-433 re-dated. ## Teardown, three layers - -Machine: throwaways (opengist, bentopdf) removed with data and backups; 9202 persists on purpose. -Host: the `nvme-scratch` storage and 9202 stay (disposition recorded). Hub: floor raised; nothing else. +Machine: 9202 throwaways removed; 9201 controller running again (09:36:24Z); the killed homebox deploy created no container but left its `app.yaml` +(09:04:55Z, keys only read) — moved aside to `/root/homebox-app.yaml.p1fixes-residue`, noted under R-531. +Host: demo-hp park marker removed; agent 0.131.0 stays (the release). ep0: one tester-1 snapshot removed on yes. +Hub: demo-hp floor override 0.243.0 kept (it delivers the release); no customer created or deleted. diff --git a/STATUS.md b/STATUS.md index ae00aff9..b003050d 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,5 +1,42 @@ # STATUS — what works, what's broken, what's next +**Updated 2026-09-15 (P1 fixes) — the big night's blockers, fixed and shipped.** + +> **Ready for a volunteer: almost.** The file manager has a real password, the backup page tells the truth, the +> installer is published with its download page, and a dead controller now comes back by itself — proven on the HP. +> Two things still stand in the way: a volunteer's own tunnel is unproven until their box exists, and every box except +> the HP still runs the old agent until you sign its update. + +**Decisions I took.** None under the unattended rule. Three things I did not do, each with its reason: I did not send a +real connect e-mail to test it, because deleting a test customer would touch the off-site server; I did not build a +„release PBS token" button, because the off-site server has no way to remove only the token; the memory warning does not +say „restarted", because nothing restarts the app. + +**What I exercised, on the HP.** Killed the controller: back in 59 seconds. Parked it: it stayed off. Killed it during an +update: the update rolled itself back. After three kills in 13 minutes the box stopped retrying for 30 minutes and you were +e-mailed — that is the safety brake, and it worked. When the brake ended, it started the controller again by itself, and told you that too. +File manager: a new box gets its own password; the HP, where you set one, was left alone. Backup page: correct on the HP. +Vaultwarden refuses strangers. Paperless took 20 documents at once without losing one. + +**What broke, and whether I fixed it.** The backup tile printed „0 B" after an agent restart — fixed, ships next release. +The memory-warning check sees nothing inside our guests, because Docker reports nothing there — not fixed, filed. + +**Rows.** Closed 9, opened 9. Register: 232 before, 232 after. + +**Needs you.** +1. **Sign the agent update for the N100 (and later Peti's box).** If you do nothing: the N100 keeps the old agent, its + dead controller does not restart, and its controller update stays held. +2. **Say where your signing keys live between sessions.** They were readable by other users on DooPlex; I tightened them. + If you do nothing: the keys stay on the build server. +3. **When the next box for „Tester 1" exists, check the tunnel opens from outside.** If you do nothing: the „No TLS Verify" + fix stays unproven. +4. **Allow one test customer to be deleted (it resets on the off-site server), or accept the unit test.** If you do + nothing: the automatic connect e-mail is untested with a real mailbox. + +--- + +## Previous note + **Updated 2026-09-15 (morning note) — the big night: a household's first month on one fresh box.** > **Ready for a volunteer: not yet — three things stop them.** The dashboard link still does not open through diff --git a/documentation/audits/evidence-p1fixes-2026-09-15/A4-crashloop-resume-9201.txt b/documentation/audits/evidence-p1fixes-2026-09-15/A4-crashloop-resume-9201.txt new file mode 100644 index 00000000..3bbf0f65 --- /dev/null +++ b/documentation/audits/evidence-p1fixes-2026-09-15/A4-crashloop-resume-9201.txt @@ -0,0 +1,7 @@ +## 2026-09-15T09:36:30Z dashboard health: 200 +Sep 15 11:34:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:34:52.719+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s +Sep 15 11:35:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:35:22.727+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s +Sep 15 11:35:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:35:52.706+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s +Sep 15 11:36:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:36:22.781+02:00 level=WARN msg="controller-supervisor: controller is NOT running — restarting the bootstrap unit" vmid=9201 status=exited unit=felhom-controller-bootstrap.service +Sep 15 11:36:24 demo-hp felhom-agent[1526161]: time=2026-09-15T11:36:24.022+02:00 level=WARN msg="controller-supervisor: RESTARTED the controller" vmid=9201 reason="controller container exited on 2 consecutive sweeps" +status=running started=2026-09-15T09:36:23.860962402Z diff --git a/documentation/audits/evidence-p1fixes-2026-09-15/A4-hub-events.txt b/documentation/audits/evidence-p1fixes-2026-09-15/A4-hub-events.txt index 9ab87862..f473def4 100644 --- a/documentation/audits/evidence-p1fixes-2026-09-15/A4-hub-events.txt +++ b/documentation/audits/evidence-p1fixes-2026-09-15/A4-hub-events.txt @@ -2,3 +2,6 @@ 2026/09/15 11:14:37 [INFO] Controller supervisor: controller_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps) 2026/09/15 11:14:37 [INFO] Controller supervisor: controller_crashloop demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps) 2026/09/15 11:14:38 [INFO] Operator email sent for demo-hp/controller_crashloop +## 2026-09-15T09:44:52Z after the 09:36:24Z resume restart: +2026/09/15 11:44:37 [INFO] Controller supervisor: controller_restarted_by_agent demo-hp-bb76ea/9201 (controller container exited on 2 consecutive sweeps) +## NOTE: the crash-loop line appears twice above only because the watch script printed it twice; the single hub pod logged it ONCE (checked per pod). diff --git a/documentation/audits/evidence-p1fixes-2026-09-15/A4-kill-middeploy-9201.txt b/documentation/audits/evidence-p1fixes-2026-09-15/A4-kill-middeploy-9201.txt index cef385b3..a43a674b 100644 --- a/documentation/audits/evidence-p1fixes-2026-09-15/A4-kill-middeploy-9201.txt +++ b/documentation/audits/evidence-p1fixes-2026-09-15/A4-kill-middeploy-9201.txt @@ -9,3 +9,10 @@ Sep 15 11:09:52 demo-hp felhom-agent[1526161]: time=2026-09-15T11:09:52.697+02:0 Sep 15 11:10:22 demo-hp felhom-agent[1526161]: time=2026-09-15T11:10:22.739+02:00 level=WARN msg="controller-supervisor: crash-loop pause in force — not restarting" vmid=9201 since=2026-09-15T09:05:51Z resume_after=30m0s ## teardown: remove http=502 at 2026-09-15T09:10:40Z ## CORRECTION 2026-09-15T09:11:57Z: the 'health 200 at 09:09:08Z — 248 s after the kill' line is WRONG. docker inspect: felhom-controller exited 09:05:02Z (exit 137, the kill) and was NOT restarted. What answered 200 once is unexplained. What happened instead, and it is the guard working as designed: the agent had restarted this controller at 08:54:24Z, 08:56:53Z and 08:58:24Z (my idle, unpark and failed-swap tests; the swap's own rollback at 09:01:57Z was correctly NOT counted), so on the second not-running sweep after this kill it logged 'CRASH-LOOP … restarts_in_window=3' at 09:05:52Z and paused restarts for 30 min. Mid-deploy restart timing is therefore NOT measured here; the resume after the pause is captured separately. +## homebox after resume: {'state': 'not_deployed', 'deployed': False, 'deploying': False} +homebox not deployed — nothing to remove +## after cleanup: 0 +## after cleanup: /opt/docker/stacks/homebox/app.yaml +## homebox app.yaml: 2026-09-15 09:04:55.994631593 +0000 451 +## homebox app.yaml: deployed deployed_at env locked_fields desired_state +## homebox app.yaml: moved-aside