burn-down night: morning note, STATUS top (164), CONTEXT, report (199 -> 164; 1 opened, 36 closed)
gates / gates (push) Successful in 2m18s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-06 05:08:10 +02:00
parent 307ecdf005
commit 32a1520833
4 changed files with 125 additions and 6 deletions
@@ -0,0 +1,80 @@
# Morning note — the burn-down night, 2026-10-06
## 1. Is anything broken right now?
**No.** The hub, demo-hp, demo-felhom and Tester 1 are healthy on the new versions. Tester 2 was off all night, as usual,
and nothing was sent to it. CI is green on every repository's last pushed commit (the felhom.eu one was still running at
05:10; the full report says which run).
**You do not need to do anything this morning.**
One small thing: when the hub restarted at 00:49 it mailed you „Tester 2 is down". That is true (it is off). The hub
sends this once after every restart for a box that is already down. That is how it was built.
## 2. The four numbers
Rows before: **199**. Rows after: **164**. Opened: **1**. Closed: **36**.
## 3. Decisions I took myself (you may reverse any of them)
1. The installer's „120 GB minimum" is now called a recommendation. It always only warned; now the words say so.
2. If the hub ever loses its sealing key, it refuses to store a new box secret. It never stores one unsealed.
3. When a household removes its own off-site backup target, the box deletes the login key. It keeps the backup password
when any backup could still need it.
4. If „Remove app" is cut off by a restart, the box finishes it at the next start. It keeps the app's data and backups;
the household can delete them with a second press.
5. One check in our register tools lost its exemption, because you closed the reason for it yesterday.
6. When an app update is held, the app page now links to the app's saved log. It does not show the raw log.
## 4. What was fixed and released
- **Security:** the hub database no longer holds box keys, owner passphrases or backup tokens in readable form. I read a
copy of the real database: 0 readable values. Every box still logs in. The hub's login cookie can no longer be planted
from another web address. A box no longer stores the catalog password in its files.
- **What a household sees:** app pages show the real web address, not „wiki.DOMAIN". A household can remove its own
off-site backup target. A cut-off app removal finishes by itself. More warnings and mails follow the household's
language, and the pages use the friendly „te" form. The dashboard says when the box cannot be reached from outside.
An app export to a network drive needs a password (your ruling).
- **Backups and monitoring:** a lost drive sends one mail, not five. A broken agent link now reports when it recovers.
- **Tools:** every check we run now has a test that proves it can fail. Building those tests found several checks that
could be fooled; all are fixed. The slowest test group now runs in 2 seconds instead of 7 minutes.
- **Released:** hub 0.138.0, agent 0.148.0, controller 0.298.0 and a new install image. All three of our boxes run them.
The app catalog got honest memory figures, and wger was closed to strangers (wger stays hidden).
- **Waiting on main for the next release:** nine installer fixes (the uninstall leaves no old keys and no open tunnel),
and more language fixes for the dashboard.
## 5. What failed, and why
- **The scratch-box tests did not run.** Their helper stalled and I stopped it. Nothing changed on the scratch box.
Four catalog fixes wait for that test: visitor addresses for four apps, and a nextcloud health check.
- **My first install-image bake used the wrong Docker version.** I saw its warning, never approved that image, and baked
it again correctly. The instructions now include the missing step.
- **One agent test was red here for 4½ hours**, because an installer change confused it. The released agent was not
affected. It is fixed.
- **Two CI jobs were lost** (the known Gitea fault). Each passed when I ran it again.
## 6. Rows moved to you or to a design
- **Need your answer (10):** a weekly disk trim on customer boxes; deleting broken backup leftovers; letting app updates
trust Docker's own health check; what lifting a held update should do; a longer quiet time after a crash (your
ruling from last night was already true in the code, so it changed nothing); three catalog questions; a check that
would run Docker on this server; one privacy sentence for an app's phone client.
- **Need a design (17):** among them an operator way into a running box (three rows wait for it), keeping the household
logged in when a setting changes, and a household timeline. I wrote three one-page proposals: pausing apps once per
backup copy, restoring over a newer database, and memory kills that nobody reports.
## 7. The two night watches
1. **The 05:00 missed-backup check** worked as designed, but you could not see it. Tester 2 has been off since Sunday
evening, but it first reported only 35 hours ago, and the check waits 48 hours before it expects anything. So no
alarm was owed, and none came. The check did not say why it skipped; now it does (in the next hub release). It
will judge Tester 2 tomorrow at 05:00, if Tester 2 is still off.
2. **Lost CI jobs:** 2 tonight, both fine on a second run.
## 8. Decisions for you
1. **Publish the installer now?** Nine fixes are ready. Pick: **yes, publish it** — I can do it in a few minutes.
If you do nothing: new installs keep the old installer, and an uninstall still leaves old keys and the tunnel behind.
2. **The disk percentage.** The box shows disk use about 5 points lower than Linux's own `df`, so disk alarms come late.
Pick: **use the `df` number and keep the alarm levels**; full boxes then alarm a little earlier. If you do nothing:
the number stays a little optimistic.