From 32a1520833a6fbc31277daa1bfccaabfccf4d436 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 6 Oct 2026 05:08:10 +0200 Subject: [PATCH] burn-down night: morning note, STATUS top (164), CONTEXT, report (199 -> 164; 1 opened, 36 closed) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 12 +++ REPORT-burndown3-2026-10-06.md | 14 +++- STATUS.md | 25 +++++- .../night-burndown-2026-10-05/MORNING-NOTE.md | 80 +++++++++++++++++++ 4 files changed, 125 insertions(+), 6 deletions(-) create mode 100644 documentation/audits/night-burndown-2026-10-05/MORNING-NOTE.md diff --git a/CONTEXT.md b/CONTEXT.md index 662f5a5b..5bdd0fcd 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,18 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-06 (night) — the burn-down night (releases).** Register 199 → 164 (1 opened: R-889 disk percent vs `df`; +> 36 closed). Releases in window 1 (00:30–00:58): hub v0.138.0 (manifest `edde9e13`; R-879 seal — 16 legacy values sealed, +> raw DB copy 0 plaintext; roll-back = `felhom-hub -unseal-box-secrets` in the pod first), agent v0.148.0 (tag = `861d32a`, +> sha `3e68a087…`, bundle `a6fa4f58…`; signed jobs ×3 + bundle ×3), controller v0.298.0 (`84d6ef7`, MinAgent 0.131.0), +> golden 0.298.0 (`a2e730e9…`, PINNED Docker set — a first unpinned bake was never vouched), vouch agent 0.148.0 / +> golden 0.298.0 / min_agent 0.131.0, floors 0.298.0 for demo-hp, demo-felhom, tester-1; catalog `c265b37`. On `main` +> unreleased: installer 1.32.0 (tag `installer-v1.32.0` not cut), controller (R-585 rest, R-621 panel, R-516 bundle), +> hub (R-872 early-return log lines). **Decisions taken by CC unattended — operator may reverse: `09` §3 131–136** +> (R-130 recommendation, R-879 no-key refusal, R-729/R-545 password kept when needed, R-682 finish-and-keep, R-426 +> one-register exemption, R-621 link). Decoy exemptions 20 → 0 (R-426). Night watches: R-872 not judged (Tester 2 +> 34.8 h old < 48 h) — re-dated 2026-10-07; R-887 2 lost jobs. Report: `REPORT-burndown3-2026-10-06.md`. + > **2026-10-05 (late night) — burn-down round 2 (releases).** Register 292 → 199 (1 opened: R-888; 94 closed: 43 > accepted by the operator 18:23, 51 fixed). Releases: hub v0.137.0 (`557629d`, deployed), agent v0.147.0 (tag, sha > `642c4d19…`, bundle `326527d0…`, signed jobs to 3 boxes), controller v0.297.0 (`1453cfc` + `6f1ba1f`), golden 0.297.0 diff --git a/REPORT-burndown3-2026-10-06.md b/REPORT-burndown3-2026-10-06.md index e21aed45..7d40e616 100644 --- a/REPORT-burndown3-2026-10-06.md +++ b/REPORT-burndown3-2026-10-06.md @@ -78,9 +78,19 @@ repository password is kept whenever anything could depend on it · 134 R-682 a green first. One helper amended a commit before its result line (allowed by the brief). - The hub deploy re-sent one true operator mail: „Tester 2 down" (the restart re-checks a down box). -## Night watches (§5) — FILLED IN AFTER 05:00 +## Night watches (§5) -(see below) +- **R-872 at 05:00:** `Deadline check: 4 customers, 0 backup missed … 1 skipped (down)`, no per-box line. From a hub.db + copy with -wal (a different channel): Tester 2 first reported 2026-10-04 16:13:44Z → 34.8 h at 05:00, under the 48 h + line → not judged, by design — but silently. No alarm owed, none fired. Fixed without a row: the early returns log why + (hub, unreleased, red-proved). Dated check re-dated to 2026-10-07 (stated in the row). Not closed. +- **R-887:** 2 lost jobs of ~30 runs (felhom.eu 1384, felhom-controller 1401); each re-run once → success. Added to the row. + +## CI, last commit of every repo + +felhom.eu `307ecdf0` (run at 05:06, see the night log), felhom-controller `c67b26be` → run 1401 (lost, re-run success), +felhom-agent `37e98f45` → 1399 success, app-catalog `d955df1f` → 1400 success. `unproven.py --summary`: unchanged +(NOT WALKED 35 of 55). ## Teardown diff --git a/STATUS.md b/STATUS.md index a43e89ab..28ceff1e 100644 --- a/STATUS.md +++ b/STATUS.md @@ -3,10 +3,27 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was sent to it.** -**Updated 2026-10-05 (late night, burn-down round 2): every box of ours healthy. The open-items list is at 199 (was 292 -at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10-05.md`.** +**Updated 2026-10-06 06:00 (the burn-down night): every box of ours healthy on hub 0.138.0, agent 0.148.0, controller +0.298.0. The open-items list is at 164 (was 199). Morning note: `documentation/audits/night-burndown-2026-10-05/MORNING-NOTE.md`; +report: `REPORT-burndown3-2026-10-06.md`.** -## Tonight, last (2026-10-05): the list at 199 +## Tonight (2026-10-06): the list at 164 + +**What happened:** 36 rows closed, 1 opened. Released and delivered the normal way to demo-hp, demo-felhom and Tester 1: +hub 0.138.0 (secrets in the hub database are sealed; checked on a copy: none readable), agent 0.148.0 (+ its root +files), controller 0.298.0 and a new install image 0.298.0. The app catalog was updated. Six small decisions were taken +by CC unattended (`09` §3 decisions 131–136); each can be reversed. + +**Needs you (none urgent; if you do nothing, each stays as it is):** +1. **Publish installer 1.32.0?** Nine uninstall/pre-flight fixes wait on `main`. If nothing: new installs keep the old + installer. Pick: yes. +2. **R-889 — the disk percentage** reads ~5 points low (not `df`'s formula), so fill alarms come late. Pick: use `df`'s + number and keep the alarm levels. +3. **Ten rows need your answer** and **seventeen need a design** — listed in the morning note, section 6. +4. Still open from before: R-469 (one paragraph in the catalog's instruction file), R-887 (CI runner fetch timeout: + 2 more lost jobs tonight, both re-run green), R-831/R-870 (printed tokens), R-616's Gitea admin token rotation. + +## Before that (2026-10-05, round 2): the list at 199 **What happened:** - **Your answer is recorded** (below): 43 rows closed as accepted. @@ -22,7 +39,7 @@ at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10 **The numbers:** 292 before → **199 after**; 1 opened; 94 closed. -**Needs you (none urgent; if you do nothing, each stays open as it is):** +**Needs you (none urgent; if you do nothing, each stays open as it is):** *(items 2, 3 and 5 were answered by the reviewer's picks at ~21:00 — `09` §3 decisions 128–130 — and built or closed the same night)* 1. **R-469** — a one-paragraph rewording in the app catalog's instruction file; my permission check refused editing instruction files. Say "go" and it is done in a minute. 2. **R-126** — should a network share be offered as an export destination? Two options in the row; pick one. diff --git a/documentation/audits/night-burndown-2026-10-05/MORNING-NOTE.md b/documentation/audits/night-burndown-2026-10-05/MORNING-NOTE.md new file mode 100644 index 00000000..8ee4d7d3 --- /dev/null +++ b/documentation/audits/night-burndown-2026-10-05/MORNING-NOTE.md @@ -0,0 +1,80 @@ +# Morning note — the burn-down night, 2026-10-06 + +## 1. Is anything broken right now? + +**No.** The hub, demo-hp, demo-felhom and Tester 1 are healthy on the new versions. Tester 2 was off all night, as usual, +and nothing was sent to it. CI is green on every repository's last pushed commit (the felhom.eu one was still running at +05:10; the full report says which run). + +**You do not need to do anything this morning.** + +One small thing: when the hub restarted at 00:49 it mailed you „Tester 2 is down". That is true (it is off). The hub +sends this once after every restart for a box that is already down. That is how it was built. + +## 2. The four numbers + +Rows before: **199**. Rows after: **164**. Opened: **1**. Closed: **36**. + +## 3. Decisions I took myself (you may reverse any of them) + +1. The installer's „120 GB minimum" is now called a recommendation. It always only warned; now the words say so. +2. If the hub ever loses its sealing key, it refuses to store a new box secret. It never stores one unsealed. +3. When a household removes its own off-site backup target, the box deletes the login key. It keeps the backup password + when any backup could still need it. +4. If „Remove app" is cut off by a restart, the box finishes it at the next start. It keeps the app's data and backups; + the household can delete them with a second press. +5. One check in our register tools lost its exemption, because you closed the reason for it yesterday. +6. When an app update is held, the app page now links to the app's saved log. It does not show the raw log. + +## 4. What was fixed and released + +- **Security:** the hub database no longer holds box keys, owner passphrases or backup tokens in readable form. I read a + copy of the real database: 0 readable values. Every box still logs in. The hub's login cookie can no longer be planted + from another web address. A box no longer stores the catalog password in its files. +- **What a household sees:** app pages show the real web address, not „wiki.DOMAIN". A household can remove its own + off-site backup target. A cut-off app removal finishes by itself. More warnings and mails follow the household's + language, and the pages use the friendly „te" form. The dashboard says when the box cannot be reached from outside. + An app export to a network drive needs a password (your ruling). +- **Backups and monitoring:** a lost drive sends one mail, not five. A broken agent link now reports when it recovers. +- **Tools:** every check we run now has a test that proves it can fail. Building those tests found several checks that + could be fooled; all are fixed. The slowest test group now runs in 2 seconds instead of 7 minutes. +- **Released:** hub 0.138.0, agent 0.148.0, controller 0.298.0 and a new install image. All three of our boxes run them. + The app catalog got honest memory figures, and wger was closed to strangers (wger stays hidden). +- **Waiting on main for the next release:** nine installer fixes (the uninstall leaves no old keys and no open tunnel), + and more language fixes for the dashboard. + +## 5. What failed, and why + +- **The scratch-box tests did not run.** Their helper stalled and I stopped it. Nothing changed on the scratch box. + Four catalog fixes wait for that test: visitor addresses for four apps, and a nextcloud health check. +- **My first install-image bake used the wrong Docker version.** I saw its warning, never approved that image, and baked + it again correctly. The instructions now include the missing step. +- **One agent test was red here for 4½ hours**, because an installer change confused it. The released agent was not + affected. It is fixed. +- **Two CI jobs were lost** (the known Gitea fault). Each passed when I ran it again. + +## 6. Rows moved to you or to a design + +- **Need your answer (10):** a weekly disk trim on customer boxes; deleting broken backup leftovers; letting app updates + trust Docker's own health check; what lifting a held update should do; a longer quiet time after a crash (your + ruling from last night was already true in the code, so it changed nothing); three catalog questions; a check that + would run Docker on this server; one privacy sentence for an app's phone client. +- **Need a design (17):** among them an operator way into a running box (three rows wait for it), keeping the household + logged in when a setting changes, and a household timeline. I wrote three one-page proposals: pausing apps once per + backup copy, restoring over a newer database, and memory kills that nobody reports. + +## 7. The two night watches + +1. **The 05:00 missed-backup check** worked as designed, but you could not see it. Tester 2 has been off since Sunday + evening, but it first reported only 35 hours ago, and the check waits 48 hours before it expects anything. So no + alarm was owed, and none came. The check did not say why it skipped; now it does (in the next hub release). It + will judge Tester 2 tomorrow at 05:00, if Tester 2 is still off. +2. **Lost CI jobs:** 2 tonight, both fine on a second run. + +## 8. Decisions for you + +1. **Publish the installer now?** Nine fixes are ready. Pick: **yes, publish it** — I can do it in a few minutes. + If you do nothing: new installs keep the old installer, and an uninstall still leaves old keys and the tunnel behind. +2. **The disk percentage.** The box shows disk use about 5 points lower than Linux's own `df`, so disk alarms come late. + Pick: **use the `df` number and keep the alarm levels**; full boxes then alarm a little earlier. If you do nothing: + the number stays a little optimistic.