burn-down night: morning note, STATUS top (164), CONTEXT, report (199 -> 164; 1 opened, 36 closed)
gates / gates (push) Successful in 2m18s
gates / gates (push) Successful in 2m18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,80 @@
|
||||
# Morning note — the burn-down night, 2026-10-06
|
||||
|
||||
## 1. Is anything broken right now?
|
||||
|
||||
**No.** The hub, demo-hp, demo-felhom and Tester 1 are healthy on the new versions. Tester 2 was off all night, as usual,
|
||||
and nothing was sent to it. CI is green on every repository's last pushed commit (the felhom.eu one was still running at
|
||||
05:10; the full report says which run).
|
||||
|
||||
**You do not need to do anything this morning.**
|
||||
|
||||
One small thing: when the hub restarted at 00:49 it mailed you „Tester 2 is down". That is true (it is off). The hub
|
||||
sends this once after every restart for a box that is already down. That is how it was built.
|
||||
|
||||
## 2. The four numbers
|
||||
|
||||
Rows before: **199**. Rows after: **164**. Opened: **1**. Closed: **36**.
|
||||
|
||||
## 3. Decisions I took myself (you may reverse any of them)
|
||||
|
||||
1. The installer's „120 GB minimum" is now called a recommendation. It always only warned; now the words say so.
|
||||
2. If the hub ever loses its sealing key, it refuses to store a new box secret. It never stores one unsealed.
|
||||
3. When a household removes its own off-site backup target, the box deletes the login key. It keeps the backup password
|
||||
when any backup could still need it.
|
||||
4. If „Remove app" is cut off by a restart, the box finishes it at the next start. It keeps the app's data and backups;
|
||||
the household can delete them with a second press.
|
||||
5. One check in our register tools lost its exemption, because you closed the reason for it yesterday.
|
||||
6. When an app update is held, the app page now links to the app's saved log. It does not show the raw log.
|
||||
|
||||
## 4. What was fixed and released
|
||||
|
||||
- **Security:** the hub database no longer holds box keys, owner passphrases or backup tokens in readable form. I read a
|
||||
copy of the real database: 0 readable values. Every box still logs in. The hub's login cookie can no longer be planted
|
||||
from another web address. A box no longer stores the catalog password in its files.
|
||||
- **What a household sees:** app pages show the real web address, not „wiki.DOMAIN". A household can remove its own
|
||||
off-site backup target. A cut-off app removal finishes by itself. More warnings and mails follow the household's
|
||||
language, and the pages use the friendly „te" form. The dashboard says when the box cannot be reached from outside.
|
||||
An app export to a network drive needs a password (your ruling).
|
||||
- **Backups and monitoring:** a lost drive sends one mail, not five. A broken agent link now reports when it recovers.
|
||||
- **Tools:** every check we run now has a test that proves it can fail. Building those tests found several checks that
|
||||
could be fooled; all are fixed. The slowest test group now runs in 2 seconds instead of 7 minutes.
|
||||
- **Released:** hub 0.138.0, agent 0.148.0, controller 0.298.0 and a new install image. All three of our boxes run them.
|
||||
The app catalog got honest memory figures, and wger was closed to strangers (wger stays hidden).
|
||||
- **Waiting on main for the next release:** nine installer fixes (the uninstall leaves no old keys and no open tunnel),
|
||||
and more language fixes for the dashboard.
|
||||
|
||||
## 5. What failed, and why
|
||||
|
||||
- **The scratch-box tests did not run.** Their helper stalled and I stopped it. Nothing changed on the scratch box.
|
||||
Four catalog fixes wait for that test: visitor addresses for four apps, and a nextcloud health check.
|
||||
- **My first install-image bake used the wrong Docker version.** I saw its warning, never approved that image, and baked
|
||||
it again correctly. The instructions now include the missing step.
|
||||
- **One agent test was red here for 4½ hours**, because an installer change confused it. The released agent was not
|
||||
affected. It is fixed.
|
||||
- **Two CI jobs were lost** (the known Gitea fault). Each passed when I ran it again.
|
||||
|
||||
## 6. Rows moved to you or to a design
|
||||
|
||||
- **Need your answer (10):** a weekly disk trim on customer boxes; deleting broken backup leftovers; letting app updates
|
||||
trust Docker's own health check; what lifting a held update should do; a longer quiet time after a crash (your
|
||||
ruling from last night was already true in the code, so it changed nothing); three catalog questions; a check that
|
||||
would run Docker on this server; one privacy sentence for an app's phone client.
|
||||
- **Need a design (17):** among them an operator way into a running box (three rows wait for it), keeping the household
|
||||
logged in when a setting changes, and a household timeline. I wrote three one-page proposals: pausing apps once per
|
||||
backup copy, restoring over a newer database, and memory kills that nobody reports.
|
||||
|
||||
## 7. The two night watches
|
||||
|
||||
1. **The 05:00 missed-backup check** worked as designed, but you could not see it. Tester 2 has been off since Sunday
|
||||
evening, but it first reported only 35 hours ago, and the check waits 48 hours before it expects anything. So no
|
||||
alarm was owed, and none came. The check did not say why it skipped; now it does (in the next hub release). It
|
||||
will judge Tester 2 tomorrow at 05:00, if Tester 2 is still off.
|
||||
2. **Lost CI jobs:** 2 tonight, both fine on a second run.
|
||||
|
||||
## 8. Decisions for you
|
||||
|
||||
1. **Publish the installer now?** Nine fixes are ready. Pick: **yes, publish it** — I can do it in a few minutes.
|
||||
If you do nothing: new installs keep the old installer, and an uninstall still leaves old keys and the tunnel behind.
|
||||
2. **The disk percentage.** The box shows disk use about 5 points lower than Linux's own `df`, so disk alarms come late.
|
||||
Pick: **use the `df` number and keep the alarm levels**; full boxes then alarm a little earlier. If you do nothing:
|
||||
the number stays a little optimistic.
|
||||
Reference in New Issue
Block a user