Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
18 KiB
STATUS — what works, what's broken, what's next
Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).
Updated 2026-10-04 (evening): the HOST's security fixes install themselves too, the hub has a fleet view and four OS alarms, and the tunnel status is true. Both demo boxes run controller 0.292.0 and host agent 0.141.1. Hub 0.131.1. New installs get golden 0.292.0 with agent 0.141.1 (vouched).
Today (2026-10-04, evening): host fixes, the fleet view, a true tunnel status
Decisions I took myself (you may reverse each):
- Only a box the installer itself recorded as "appliance" gets host updates (a record the agent cannot change).
- The four alarm times: no update run for 7 days; "reboot needed" for 14 days; the demo boxes approve nothing for 7 days; a box has fixes nobody approved for 14 days. All four are settings.
- I released the host agent and the hub a second time today. The first release said "reboot needed" wrongly; the new 14-day alarm would then have mailed you about boxes you had already rebooted.
One decision for you (a safe default if you say nothing) — the Docker engine updates (designed, not built):
- Turn on Docker's "live-restore" on every box. With it, a Docker engine update restarts no app (measured: 0
restarts). Without it, every app stops for about 30 seconds per engine update.
- A (my pick): turn it on — in the new-install image and once on existing boxes. It goes on without restarting anything. Cost: it must never be turned off by a plain restart again (that stops every app and starts none).
- B: leave it off. Every Docker update then means ~30 seconds of every app being down, at night.
- If you say nothing: nothing changes; Docker updates stay unbuilt.
What I did:
- The host's Debian fixes now install themselves, after the guest's, on the same night run, only on appliances, never a kernel or boot package, never a reboot. Proven: demo-felhom installed 108 host fixes and stayed healthy; all 108 came from Debian. A box made "ring 1" for a test installed exactly the one approved host version it lacked.
- A way to put one host package back by hand is written and proven on demo-hp (and the test taught it two fixes).
- The fleet view in the hub: one line per box with its updates, "reboot needed since", and the tunnel.
- The tunnel status is now true: running, not running, or unknown. I blocked demo-hp's tunnel: after two reports (about 30 minutes) you got a "tunnel down" mail; when I unblocked it, "recovered". A simply stopped tunnel heals itself within 5 minutes, before the hub can even see it.
- The night run is fast: 23–32 seconds when there is nothing to install (target was under 60).
- The kernel test on demo-hp (your two reboots): Secure Boot works with it, but GRUB's "boot once" does not work on our boxes: the second plain reboot came up on the NEW kernel again. So a new kernel that hangs would stay. That must be solved before kernel updates. demo-hp now runs the newer kernel, healthy. A watchdog chip exists on demo-hp.
- Found and fixed during the live test: "reboot needed" was wrong in two ways (fixed in the second release).
- New-install image baked and vouched (0.292.0 with host agent 0.141.1). The golden waiver is gone: not needed.
- Rows: 2 closed, 2 opened-and-closed the same day, 3 opened. The list went from 332 to 333.
Needs you later (nothing breaks if you wait):
- A host that crashes does not restart by itself (Linux's "panic" setting is off). Changing it changes how every box behaves, so it is your call, together with kernel updates.
Today (2026-10-04, afternoon): the guest's security fixes install themselves
One decision for you (a safe default if you say nothing):
- Undo for a guest update that goes wrong. I tested it first, as you asked: Proxmox cannot take a snapshot of a
customer box at all (the box is linked to the household's drives, and Proxmox refuses). So there is no automatic undo.
Today a failed check stops, mails you, and last night's whole-box backup (minutes old) is the undo, by hand.
- A (my pick): keep it so. Nothing new to build. A restore takes about 1–3 minutes plus losing what apps wrote since the backup.
- B: build our own disk snapshot under Proxmox. Automatic, but a new mechanism nobody has tested, with a risk to the disk pool.
- If you say nothing: A stays.
What I did:
- Your two choices are built. A returning household's new box sets the old off-site copy aside on night one and starts a new one (nothing deleted). An approved version that Debian already replaced comes from Debian's dated archive.
- Guest security fixes now install themselves. Each night, after the whole-box backup, the demo boxes install
Debian's fixes. When both demo boxes run the same versions healthy for 24 hours and one night, the hub approves that
set; every other box then installs exactly those versions. Docker, the host and the kernel are not touched.
- Proven live: each demo box installed 53 fixes and stayed healthy. With a 2-minute test wait the hub approved the set; then I put back 24 hours. demo-felhom, made "ring 1" for the test, installed exactly the 3 approved versions it lacked and nothing newer. A run where I stopped an app failed its check and mailed you (you got that mail).
- You can switch OS updates off per box in the hub. They are ON by default. The household sees one line in its timeline: "System security fixes installed".
- The tunnel and two other built-in programs are current (cloudflared was 4 months old). A new monthly check catches them falling behind. The tunnel came back by itself on both demo boxes within 20 seconds.
- Three bugs found and fixed during the live test, before release.
- Two gaps found, now rows: existing boxes cannot receive the new update tool through the product (only new installs, or by hand — the demo boxes got it by hand); and the hub's "tunnel status" for every box always reads "inactive" because it checks the wrong place.
- Rows: 4 closed (one opened and closed the same day), 5 opened. The list went from 331 to 333.
Today (2026-10-04, day): off-site closed, operating-system updates measured
Two decisions for you — each has a safe default if you say nothing:
- A returning household's first night. A new box for someone who had a box before makes no off-site copy, because
the old copy was made with a key the new box does not have.
- A (my pick): the new box sets the old copy aside by itself on its first night, then starts a new one. An un-claimed box already does exactly this. Nothing is deleted; you can put the old copy back. Cost: a small box change. A household that wanted to continue the old copy with its recovery code must do that before night one.
- B: ask the household on the recovery-code evening. Cost: new screens in two languages. If they skip it, the gap stays.
- If you say nothing: such a household has no off-site copy until someone presses "start a new off-site backup", and you get a mail each time. Nobody is blocked today (no returning household is waiting).
- Approved OS updates, when Debian has already replaced the version. Debian keeps only two versions of a package.
- A (my pick): the box then fetches the exact approved version from Debian's own dated archive (snapshot.debian.org, signed by Debian). Measured: 2–3 seconds. Cost: one more outside service we rely on — but in 3 months it would have been needed zero times for the packages a box has.
- B: the box waits for the next approval. Cost: no new service; a box can stay unpatched for one cycle.
- If you say nothing: nothing is blocked now; the first build step can start with B and switch later.
What I did:
- A restored box is safe by default. The automatic restore-test was already safe (measured). The agent's disaster restore now refuses to run next to a live original. Restores by hand use a new safe script: no start on boot, no link to the real drives, network off. All proven live on demo-hp. New agent 0.139.0 on both demo boxes.
- The weekly clean-up cannot get stuck any more. You can allow one bigger clean-up for one box (hub 0.129.0). Proven on a lab copy: 98 backups, the normal limit refused, the bigger one cleaned to 13, and the next week was normal again.
- A dated check for 12 October looks at the first clean-up that really deletes old backups.
- OS updates, measured on the demo boxes (nothing built for customers):
- The boxes are far behind: 188 updates per host, about 59 per guest. Every guest runs a different Docker version.
- A Debian update of the guest took 24 seconds. No app stopped. The undo (restore the guest) took 73 seconds.
- A Debian update of demo-hp's host took 60 seconds. The guests kept running.
- A broken (killed) update does not fix itself. Two repair commands fix it in 5 seconds.
- A Docker update restarts every app (about 30 seconds silent). A Docker setting ("live-restore") avoids that completely — but switching it off again stopped every app and started none. Recorded as a trap.
- The new kernel booted fine, and the box fell back to the old kernel on the next restart. But if a new kernel hangs, the box keeps trying it. Someone must then switch it off and on. No hardware watchdog is used.
- About 40 packages from Proxmox have ordinary names (for example the disk system ZFS and the Secure Boot loader). So "which lane" must follow where a package comes from, not its name.
- The internet tunnel program on every box (cloudflared) is 4 months old. Nothing updates it. New row.
- Thank you for being near the box. demo-hp restarted twice. It now runs the old kernel and starts the new one at its next restart.
- Rows: 2 closed, 5 opened. The list went from 328 to 331.
Today (2026-10-04): off-site safety finished
- The weekly clean-up is ON and no longer stops itself. The fake-backup check now skips young copies that a manual backup replaced the same day, instead of refusing. I ran one window on demo-hp: no refusal, nothing removed (127 → 127), because every candidate was still young. The first real removals come when those copies are older than 8 days — around 11 October. demo-felhom gets its first window at its next night run.
- DooPlex keeps 8 weekly copies of ep0 (your choice). Old copies go weekly; anything ep0 deletes still never reaches the copy by itself. Today's nightly copy ran fine.
- tester-1's 3 old keys are gone. The daily alarm for tester-1 stopped. (You got one last alarm mail this morning, sent just before I removed them.)
- A restore from the DooPlex copy works. I restored demo-hp's whole box onto a scratch machine in about 3 minutes, read its data, then deleted it. One trap found: a restored box starts automatically and points at the real drives. Run beside the original, that would be two copies of the same box. I switched it off in time; the runbook now warns, and a row asks to make it safe by default.
- "Delete my old set-aside backups" works again. The hub does it, 7 days after the household's request; the household or you can cancel in that time. Tested on tester-1 with a planted test folder.
- Your answers are recorded, including that you keep the current Hetzner key for now (a row holds the 3 steps).
- Rows: 6 closed, 4 opened. The list went from 330 to 328.
Today (2026-10-03, evening): your choices A and A — built
- A box can no longer delete its off-site backups. Both demo boxes now use a key that can only add. I tried a delete from each box: refused. Old backups all stayed (demo-felhom 11 → 13, demo-hp 91 → 100).
- No box gets the storage password any more. The box gives the hub only its public key; the hub puts it in the storage account. I asked for the password with demo-hp's own login: refused.
- The hub keeps the passwords locked (encrypted). A copy of the hub database no longer reveals them.
- Every day the hub checks each storage account's key file. It found old unlocked keys from earlier boxes: 4 on demo-felhom's account and 5 on demo-hp's — now removed. tester-1's account still has 3 (its box is gone); you get a daily alarm for it until they go.
- The weekly clean-up window works, but it is switched OFF. I opened one window on demo-hp by hand: it opened, the fake-backup check refused, and it closed in 3 seconds. The check refused because I had made a manual backup today — it is too strict. I fix that next; until then nothing deletes old backups (there is plenty of room).
- ep0's whole-box backups are copied to DooPlex every night. First copy: 12 GB in 3 minutes, all 4 backups. DooPlex cannot read them (encrypted per household). A failed copy mails you; the test mail arrived.
- One bug, found and fixed live: the first new box version misread its backup count as 0, and you got one false alarm mail ("demo-felhom: fell from 11 to 0"). Ignore that mail. Fixed 15 minutes later (0.289.1).
- Rows: 3 closed, 6 opened, 1 opened and closed the same day. The list went from 327 to 330.
Today (2026-10-03, later): off-site backup safety, step 1 — measured, nothing built
- The "add only" lock works. I tested it on tester-1's storage account (your choice; no box uses it now). New backups go in. Restore works. Every delete is refused. I put the account back exactly as it was.
- But the lock alone does not protect us yet. A broken-into box can ask the hub for the storage password. With that password it can log in and remove the lock. First the box must stop getting the password.
- A second trap: with "add only", an attacker can add fake backups dated in the future. The normal clean-up rule then deletes all the real backups. Any clean-up must check for this.
- The hub keeps every storage password in plain form. Anyone who reads the hub database can delete every household's off-site backups. New row.
- ep0: Hetzner cannot snapshot the extra disk at all. One of the three old ideas does not exist.
- Rows: 2 closed, 3 opened. The list went from 326 to 327 rows. The dated check for 6 October is done.
Today (2026-10-03): the to-do list is in order — paperwork only, no machine touched
- Finished items left the open list. It went from 442 rows to 326. Nothing was deleted; each moved row names where its full text is.
- Every open row now has one category and one severity (P1 now · P2 before the first paying customer · P3 during the first customers · P4 later). No row is P1. 27 rows are P2.
- Two new automatic checks refuse a finished row left in the open list, and a new row without a category or a severity.
- Your four new items are on the roadmap: security updates for the box's own system; legal pages and business papers; "what if the household leaves Felhom"; a second login step for the dashboard. Two of them are also real findings today: a box never receives system security updates, and the website has no privacy notice, terms or imprint.
- The ranked list and my reasoning: the triage recommendation in the audits folder.
Tester-2 — read only, from the hub. The customer record exists. Tester-2's box has not registered yet.
Today (afternoon): every app checked again
- No data-loss fault. I checked all 58 apps with the fixed check. It made each app really save something, then looked where the data landed. No app saves data where the backup does not copy it.
- 40 apps: proven correct. 18 apps: not proven either way. In those 18 the check could not fill every folder (for example an upload folder stays empty because my test saves no upload). Nothing was found outside a backed-up folder. I list them and work on them later.
- papra was broken for new installs, now fixed. It needs more memory than its limit, so a fresh install crashed in a loop. I measured it and raised the limit. No box runs papra.
- plant-it's program image is gone from Docker Hub. It is already hidden from new installs. No box runs it.
Removing an app now tells the truth (your choice A)
- When the household removes an app, its own files (books, videos) stay. The dialog and the result now say so, and name the folder. Proven on the scratch box: the video was still there after the remove.
Before the first paying customer
Everything here must be done before the first customer who pays:
- SparkyFitness: written permission from the author (the e-mail draft is ready). If none → hidden from new installs.
- Tandoor: written permission from the authors. If none → hidden from new installs.
- A lawyer reviews the licence list (Tandoor, SparkyFitness, Emby, n8n, Plex, the paid "enterprise" parts, Redis).
- Agent updates: I may sign agent updates only until the first paying customer; after that you sign them.
Your licence decisions are recorded: Emby, Plex and n8n stay. recipe-importer needs nothing unless we share it.
Also today
- New version 0.288.0 and a new golden (0.288.0).
- Rows. 7 closed, 6 opened. The list went from 436 to 442 rows.
What needs you
- The undo choice at the top of today's section (A: keep the backup as the undo; B: build our own snapshot). If you say nothing, A stays. Earlier today's two decisions are built. The Hetzner key change stays your call.
- plant-it: keep the hidden template as it is, or remove it entirely (its image no longer exists). If you say nothing: it stays hidden; nothing runs it.
- Send the SparkyFitness request, and ask the Tandoor authors (the "Before the first paying customer" list).
- Phone test (2 minutes), only if you want it: say so, and I put MeTube back on demo-hp with a family login.
Standing steps
- Monthly security re-test: last run 2026-10-01, next due ~2026-11-01. (You start it with the standing brief.)
- Weekly: the golden bake (next around 9 October; today's bake was 0.288.0).
- Registry clean-up: only when the registry disk fills; "show me" mode first, then a person decides.