burn-down night: morning note, STATUS top (164), CONTEXT, report (199 -> 164; 1 opened, 36 closed)
gates / gates (push) Successful in 2m18s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-06 05:08:10 +02:00
parent 307ecdf005
commit 32a1520833
4 changed files with 125 additions and 6 deletions
+12
View File
@@ -16,6 +16,18 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-06 (night) — the burn-down night (releases).** Register 199 → 164 (1 opened: R-889 disk percent vs `df`;
> 36 closed). Releases in window 1 (00:30–00:58): hub v0.138.0 (manifest `edde9e13`; R-879 seal — 16 legacy values sealed,
> raw DB copy 0 plaintext; roll-back = `felhom-hub -unseal-box-secrets` in the pod first), agent v0.148.0 (tag = `861d32a`,
> sha `3e68a087…`, bundle `a6fa4f58…`; signed jobs ×3 + bundle ×3), controller v0.298.0 (`84d6ef7`, MinAgent 0.131.0),
> golden 0.298.0 (`a2e730e9…`, PINNED Docker set — a first unpinned bake was never vouched), vouch agent 0.148.0 /
> golden 0.298.0 / min_agent 0.131.0, floors 0.298.0 for demo-hp, demo-felhom, tester-1; catalog `c265b37`. On `main`
> unreleased: installer 1.32.0 (tag `installer-v1.32.0` not cut), controller (R-585 rest, R-621 panel, R-516 bundle),
> hub (R-872 early-return log lines). **Decisions taken by CC unattended — operator may reverse: `09` §3 131–136**
> (R-130 recommendation, R-879 no-key refusal, R-729/R-545 password kept when needed, R-682 finish-and-keep, R-426
> one-register exemption, R-621 link). Decoy exemptions 20 → 0 (R-426). Night watches: R-872 not judged (Tester 2
> 34.8 h old < 48 h) — re-dated 2026-10-07; R-887 2 lost jobs. Report: `REPORT-burndown3-2026-10-06.md`.
> **2026-10-05 (late night) — burn-down round 2 (releases).** Register 292 → 199 (1 opened: R-888; 94 closed: 43
> accepted by the operator 18:23, 51 fixed). Releases: hub v0.137.0 (`557629d`, deployed), agent v0.147.0 (tag, sha
> `642c4d19…`, bundle `326527d0…`, signed jobs to 3 boxes), controller v0.297.0 (`1453cfc` + `6f1ba1f`), golden 0.297.0
+12 -2
View File
@@ -78,9 +78,19 @@ repository password is kept whenever anything could depend on it · 134 R-682 a
green first. One helper amended a commit before its result line (allowed by the brief).
- The hub deploy re-sent one true operator mail: „Tester 2 down" (the restart re-checks a down box).
## Night watches (§5) — FILLED IN AFTER 05:00
## Night watches (§5)
(see below)
- **R-872 at 05:00:** `Deadline check: 4 customers, 0 backup missed … 1 skipped (down)`, no per-box line. From a hub.db
copy with -wal (a different channel): Tester 2 first reported 2026-10-04 16:13:44Z → 34.8 h at 05:00, under the 48 h
line → not judged, by design — but silently. No alarm owed, none fired. Fixed without a row: the early returns log why
(hub, unreleased, red-proved). Dated check re-dated to 2026-10-07 (stated in the row). Not closed.
- **R-887:** 2 lost jobs of ~30 runs (felhom.eu 1384, felhom-controller 1401); each re-run once → success. Added to the row.
## CI, last commit of every repo
felhom.eu `307ecdf0` (run at 05:06, see the night log), felhom-controller `c67b26be` → run 1401 (lost, re-run success),
felhom-agent `37e98f45` → 1399 success, app-catalog `d955df1f` → 1400 success. `unproven.py --summary`: unchanged
(NOT WALKED 35 of 55).
## Teardown
+21 -4
View File
@@ -3,10 +3,27 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
sent to it.**
**Updated 2026-10-05 (late night, burn-down round 2): every box of ours healthy. The open-items list is at 199 (was 292
at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10-05.md`.**
**Updated 2026-10-06 06:00 (the burn-down night): every box of ours healthy on hub 0.138.0, agent 0.148.0, controller
0.298.0. The open-items list is at 164 (was 199). Morning note: `documentation/audits/night-burndown-2026-10-05/MORNING-NOTE.md`;
report: `REPORT-burndown3-2026-10-06.md`.**
## Tonight, last (2026-10-05): the list at 199
## Tonight (2026-10-06): the list at 164
**What happened:** 36 rows closed, 1 opened. Released and delivered the normal way to demo-hp, demo-felhom and Tester 1:
hub 0.138.0 (secrets in the hub database are sealed; checked on a copy: none readable), agent 0.148.0 (+ its root
files), controller 0.298.0 and a new install image 0.298.0. The app catalog was updated. Six small decisions were taken
by CC unattended (`09` §3 decisions 131–136); each can be reversed.
**Needs you (none urgent; if you do nothing, each stays as it is):**
1. **Publish installer 1.32.0?** Nine uninstall/pre-flight fixes wait on `main`. If nothing: new installs keep the old
installer. Pick: yes.
2. **R-889 — the disk percentage** reads ~5 points low (not `df`'s formula), so fill alarms come late. Pick: use `df`'s
number and keep the alarm levels.
3. **Ten rows need your answer** and **seventeen need a design** — listed in the morning note, section 6.
4. Still open from before: R-469 (one paragraph in the catalog's instruction file), R-887 (CI runner fetch timeout:
2 more lost jobs tonight, both re-run green), R-831/R-870 (printed tokens), R-616's Gitea admin token rotation.
## Before that (2026-10-05, round 2): the list at 199
**What happened:**
- **Your answer is recorded** (below): 43 rows closed as accepted.
@@ -22,7 +39,7 @@ at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10
**The numbers:** 292 before → **199 after**; 1 opened; 94 closed.
**Needs you (none urgent; if you do nothing, each stays open as it is):**
**Needs you (none urgent; if you do nothing, each stays open as it is):** *(items 2, 3 and 5 were answered by the reviewer's picks at ~21:00 — `09` §3 decisions 128–130 — and built or closed the same night)*
1. **R-469** — a one-paragraph rewording in the app catalog's instruction file; my permission check refused editing
instruction files. Say "go" and it is done in a minute.
2. **R-126** — should a network share be offered as an export destination? Two options in the row; pick one.
@@ -0,0 +1,80 @@
# Morning note — the burn-down night, 2026-10-06
## 1. Is anything broken right now?
**No.** The hub, demo-hp, demo-felhom and Tester 1 are healthy on the new versions. Tester 2 was off all night, as usual,
and nothing was sent to it. CI is green on every repository's last pushed commit (the felhom.eu one was still running at
05:10; the full report says which run).
**You do not need to do anything this morning.**
One small thing: when the hub restarted at 00:49 it mailed you „Tester 2 is down". That is true (it is off). The hub
sends this once after every restart for a box that is already down. That is how it was built.
## 2. The four numbers
Rows before: **199**. Rows after: **164**. Opened: **1**. Closed: **36**.
## 3. Decisions I took myself (you may reverse any of them)
1. The installer's „120 GB minimum" is now called a recommendation. It always only warned; now the words say so.
2. If the hub ever loses its sealing key, it refuses to store a new box secret. It never stores one unsealed.
3. When a household removes its own off-site backup target, the box deletes the login key. It keeps the backup password
when any backup could still need it.
4. If „Remove app" is cut off by a restart, the box finishes it at the next start. It keeps the app's data and backups;
the household can delete them with a second press.
5. One check in our register tools lost its exemption, because you closed the reason for it yesterday.
6. When an app update is held, the app page now links to the app's saved log. It does not show the raw log.
## 4. What was fixed and released
- **Security:** the hub database no longer holds box keys, owner passphrases or backup tokens in readable form. I read a
copy of the real database: 0 readable values. Every box still logs in. The hub's login cookie can no longer be planted
from another web address. A box no longer stores the catalog password in its files.
- **What a household sees:** app pages show the real web address, not „wiki.DOMAIN". A household can remove its own
off-site backup target. A cut-off app removal finishes by itself. More warnings and mails follow the household's
language, and the pages use the friendly „te" form. The dashboard says when the box cannot be reached from outside.
An app export to a network drive needs a password (your ruling).
- **Backups and monitoring:** a lost drive sends one mail, not five. A broken agent link now reports when it recovers.
- **Tools:** every check we run now has a test that proves it can fail. Building those tests found several checks that
could be fooled; all are fixed. The slowest test group now runs in 2 seconds instead of 7 minutes.
- **Released:** hub 0.138.0, agent 0.148.0, controller 0.298.0 and a new install image. All three of our boxes run them.
The app catalog got honest memory figures, and wger was closed to strangers (wger stays hidden).
- **Waiting on main for the next release:** nine installer fixes (the uninstall leaves no old keys and no open tunnel),
and more language fixes for the dashboard.
## 5. What failed, and why
- **The scratch-box tests did not run.** Their helper stalled and I stopped it. Nothing changed on the scratch box.
Four catalog fixes wait for that test: visitor addresses for four apps, and a nextcloud health check.
- **My first install-image bake used the wrong Docker version.** I saw its warning, never approved that image, and baked
it again correctly. The instructions now include the missing step.
- **One agent test was red here for 4½ hours**, because an installer change confused it. The released agent was not
affected. It is fixed.
- **Two CI jobs were lost** (the known Gitea fault). Each passed when I ran it again.
## 6. Rows moved to you or to a design
- **Need your answer (10):** a weekly disk trim on customer boxes; deleting broken backup leftovers; letting app updates
trust Docker's own health check; what lifting a held update should do; a longer quiet time after a crash (your
ruling from last night was already true in the code, so it changed nothing); three catalog questions; a check that
would run Docker on this server; one privacy sentence for an app's phone client.
- **Need a design (17):** among them an operator way into a running box (three rows wait for it), keeping the household
logged in when a setting changes, and a household timeline. I wrote three one-page proposals: pausing apps once per
backup copy, restoring over a newer database, and memory kills that nobody reports.
## 7. The two night watches
1. **The 05:00 missed-backup check** worked as designed, but you could not see it. Tester 2 has been off since Sunday
evening, but it first reported only 35 hours ago, and the check waits 48 hours before it expects anything. So no
alarm was owed, and none came. The check did not say why it skipped; now it does (in the next hub release). It
will judge Tester 2 tomorrow at 05:00, if Tester 2 is still off.
2. **Lost CI jobs:** 2 tonight, both fine on a second run.
## 8. Decisions for you
1. **Publish the installer now?** Nine fixes are ready. Pick: **yes, publish it** — I can do it in a few minutes.
If you do nothing: new installs keep the old installer, and an uninstall still leaves old keys and the tunnel behind.
2. **The disk percentage.** The box shows disk use about 5 points lower than Linux's own `df`, so disk alarms come late.
Pick: **use the `df` number and keep the alarm levels**; full boxes then alarm a little earlier. If you do nothing:
the number stays a little optimistic.